Intelligent resource allocation method and device for eMBB / URLLC coexistence scene, terminal and storage medium
Through deep reinforcement learning method and environmental interaction, the problem of poor resource allocation effect in the coexistence scenario of eMBB/URLLC is solved, and the optimal resource allocation strategy is realized that while meeting the URLLC reliability and low latency requirements, it can maximize the service satisfaction of eMBB services.
Patent Information
- Application Number
- CN202510616119.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-14
AI Technical Summary
In the eMBB/URLLC coexistence scenario, it is difficult for the prior art to control the throughput and service satisfaction of eMBB while meeting the reliability of URLLC, which affects the resource allocation effect.
The deep reinforcement learning method is adopted to continuously interact with the environment to deal with the random and occasional characteristics of URLLC services, while meeting the QoS needs of URLLC users, maximize the service satisfaction of eMBB services and solve the optimal resource allocation strategy. The specific steps include obtaining the maximum transmission data rate of the eMBB user, calculating the transmission data rate loss of the eMBB user after the URLLC perforation, calculating the transmission data rate of the URLLC user based on the finite block long encoding formula, and establishing a resource allocation model and a perforation model using a proportional fair algorithm, modeling through Markov process and realizing collaborative training of the agent based on the PF-TSIPRA algorithm.
While meeting the QoS needs of URLLC users, it can maximize the service satisfaction of eMBB services, realize the optimal resource allocation strategy, and reduce the impact of URLLC perforation on eMBB services.
Smart Images

Figure CN120129080A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an intelligent resource allocation method, device, terminal and storage medium for eMBB / URLLC coexistence scenarios, and belongs to the field of wireless communication technology. Background Art
[0002] With the continuous advancement towards 6G wireless communication, the service types and scopes of cellular communication are also increasing day by day, such as industrial automation, intelligent transportation, tactile Internet, etc. Among them, enhanced mobile broadband (eMBB) and ultra-reliable low-latency communication (URLLC) are two main types of communication services. Among them, eMBB services such as high-definition video streaming, virtual reality (VR), augmented reality (AR), etc., focus on higher transmission data rates; URLLC services include autonomous driving, industrial Internet of Things, etc. Such services are extremely sensitive to real-time performance and pay more attention to the characteristics of low latency and high reliability.
[0003] In most scenarios, the two services coexist. Since the eMBB service mainly focuses on high throughput in dense application scenarios, while the URLLC service focuses on the strict latency and reliability of critical tasks, the QoS (quality of service) requirements of the two services are very different. Therefore, how to effectively allocate resources for the two coexisting services in limited bandwidth and power resources is a key issue. The puncturing scheme proposed by 3GPP (3rd Generation Partnership Project) gives an answer. By combining the mini-slot technology, the URLLC service will preempt the resources of the eMBB service when it arrives to meet its QoS requirements. However, puncturing will cause the transmission data rate of the eMBB service to be impaired. Therefore, it is necessary to minimize the impact of URLLC puncturing on the eMBB service while meeting the QoS requirements of the URLLC service.
[0004] In actual scenarios, the URLLC service has the characteristics of randomness and occasionality, and it is difficult to obtain an optimal resource allocation scheme through traditional scheduling methods. Considering the channel differences between eMBB users and the impact of URLLC puncturing, simply pursuing the maximization of throughput may cause some eMBB users to fail to reach the target transmission data rate, and the service satisfaction is affected.
[0005] In the prior art, due to the significantly different service quality requirements of the two services, it is usually difficult to control the throughput and service satisfaction of eMBB to the maximum while meeting the reliability of URLLC, which affects the resource allocation effect when eMBB / URLLC coexists. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies in the prior art and provide an intelligent resource allocation method, device, terminal, and storage medium for the eMBB / URLLC coexistence scenario. By using the deep reinforcement learning method and continuously interacting with the environment to cope with the random and occasional characteristics of the URLLC service, while meeting the QoS requirements of URLLC users, the service satisfaction of the eMBB service is maximized, and the optimal resource allocation strategy is solved.
[0007] To solve the above technical problems, the present invention is implemented by the following technical solutions:
[0008] In the first aspect, the present invention provides an intelligent resource allocation method for the eMBB / URLLC coexistence scenario, including the following steps:
[0009] Obtain the maximum transmission data rate of eMBB users without URLLC puncturing.
[0010] Adopt the puncturing method to achieve the coexistence of eMBB / URLLC, and calculate the transmission data rate loss of eMBB users after being punctured by URLLC according to the maximum transmission data rate of eMBB users.
[0011] Based on the finite block length coding formula, calculate the transmission data rate of URLLC users and achieve the reliability constraint of URLLC.
[0012] According to the maximum transmission data rate of eMBB users, the transmission data rate loss of eMBB users, and the reliability constraint of URLLC, use the proportional fairness algorithm to establish an eMBB resource allocation model on the time slot scale and a URLLC puncturing model on the mini-time slot scale respectively.
[0013] Model the eMBB resource allocation model on the time slot scale and the URLLC puncturing model on the mini-time slot scale as Markov processes respectively.
[0014] For the Markov process of the eMBB resource allocation model on the time slot scale, set an eMBB agent, and for the Markov process of the URLLC puncturing model on the mini-time slot scale, set a URLLC agent, and realize the collaborative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm.
[0015] Obtain the first input parameter and the second input parameter, input the first input parameter into the trained eMBB agent, and output the resource block allocation matrix and the power allocation matrix; input the second input parameter into the trained URLLC agent, and output the puncturing matrix, the puncturing bandwidth ratio matrix and the puncturing power ratio matrix to achieve resource allocation.
[0016] Further, the maximum transmission data rate of the eMBB user without URLLC puncturing is obtained, and the specific expression is as follows:
[0017] ;
[0018] In the formula: is the time slot; is the resource block; is the eMBB user; is the maximum transmission data rate that the eMBB user can achieve on the resource block k in the time slot t without puncturing; is the eMBB user downlink transmission power on the resource block k in the time slot t; is the small-scale channel fading factor; is the bandwidth of the resource block k; is the noise power spectral density; is the path loss factor; is the eMBB user distance between and the next-generation base station gNodeB; is the path loss exponent.
[0019] Further, calculating the transmission data rate loss of the eMBB user after being punctured by URLLC according to the maximum transmission data rate of the eMBB user specifically includes:
[0020] Obtain the bandwidth and transmission power obtained when URLLC punctures eMBB, and the specific expressions are as follows:
[0021] ;
[0022] ;
[0023] In the formula: is the URLLC user; the time slot includes mini time slots, is an integer, and the set of mini time slots is specifically represented as ; is the URLLC user In a time slot Resource block The mini - time slot of The perforated bandwidth; For URLLC users In a time slot Resource block The mini - time slot of The perforated transmission power; For the time slot Resource block The mini - time slot of The proportion of the perforated bandwidth, ; For the time slot Resource block The mini - time slot of The proportion of the perforated power, ;
[0024] According to the bandwidth and transmission power obtained when URLLC perforates eMBB, calculate the data rate of the eMBB user On the th mini - time slot of resource block k in time slot t, and the specific expression is as follows:
[0025] ;
[0026] In the formula: Is the data rate of the eMBB user after being perforated by URLLC On the th mini - time slot of resource block k in time slot t;
[0027] According to the maximum transmission data rate of the eMBB user without being perforated by URLLC and the data rate of the eMBB user after being perforated by URLLC In a time slot Resource block The mini - time slot of Calculate the loss of the transmission data rate of the eMBB user after being perforated by URLLC In a time slot Resource block The mini - time slot of And the specific expression is as follows:
[0028] ;
[0029] In the formula: Is the data rate of the eMBB user after being perforated by URLLC In a time slot Resource block The mini - time slot of The data rate loss on the transmission line.
[0030] Further, the calculation of the URLLC user transmission data rate based on the finite block length coding formula specifically includes:
[0031] Get time slot Resource Block Mini time slots The specific expression of perforation is as follows:
[0032] ;
[0033] Where: For time slot Resource Block Mini time slots perforation condition;
[0034] According to time slot Resource Block Mini time slots Based on the finite block length coding formula, the URLLC user In time slot No. The data rate of a mini-time slot is expressed as follows:
[0035] ;
[0036] ;
[0037] ;
[0038] Where: For URLLC users signal-to-noise ratio; represents channel dispersion; is the inverse Gaussian function; Mini-slot Length; is the maximum tolerable transmission packet loss probability of URLLC; , Resource Block A collection of For resource blocks in the collection the number of For time slot Resource Block Mini time slots perforation condition; For URLLC users In time slot No. Data rate of a mini-slot;
[0039] Let represent the URLLC user in the time slot at the th mini-slot, then the specific expression of the transmission delay constraint of the URLLC user is as follows:
[0040] ;
[0041] In the formula: is the tolerable transmission delay of URLLC;
[0042] The reliability constraint for implementing URLLC is specifically expressed as follows:
[0043] ;
[0044] In the formula: is the packet loss rate of URLLC per unit time; represents the URLLC user in the time slot at the th mini-slot not meeting the transmission delay constraint.
[0045] Furthermore, the establishment of the eMBB resource allocation model at the time slot scale and the URLLC puncturing model at the mini-slot scale includes:
[0046] Establishing the eMBB resource allocation model at the time slot scale specifically includes:
[0047] Obtaining the weight of the eMBB user in the time slot The specific expression is as follows:
[0048] ;
[0049] In the formula: is the weight of the eMBB user in the time slot , is the historical average transmission data rate of the eMBB user The specific expression is as follows:
[0050] ;
[0051] In the formula: is the time window length for calculating the historical average transmission data rate; is the eMBB user in the time slot The total transmission data rate; is the eMBB user calculated based on the time slot as a reference historical average transmission data rate;
[0052] Obtain the time slot resource block allocation situation, according to the time slot resource block allocation situation and the maximum transmission data rate of the eMBB user without being perforated by URLLC, obtain the eMBB user in the time slot total transmission data rate, the specific expression is as follows:
[0053] ;
[0054] ;
[0055] In the formula: is the time slot resource block allocation situation, is the eMBB user in the time slot total transmission data rate;
[0056] According to the weight of the eMBB user in the time slot weight, time slot resource block allocation situation and eMBB user in the time slot total transmission data rate, establish an eMBB resource allocation model on the time slot scale, the specific expression is as follows:
[0057] ;
[0058] ;
[0059] ;
[0060] ;
[0061] ;
[0062] In the formula: is the maximum transmission power of the next-generation base station gNodeB; , is the set of eMBB users e, is the number of eMBB users e in the set; Indicates resource block allocation matrix of order Indicates power allocation matrix of order
[0063] Build a URLLC puncturing model on the mini-slot scale, specifically including:
[0064] According to the eMBB user after being punctured by URLLC in the time slot resource block in the mini-slot on the transmission data rate loss, calculate the total transmission data rate loss of the eMBB user in the mini-slot on the total transmission data rate loss, the specific expression is as follows:
[0065] ;
[0066] In the formula: is the total transmission data rate loss of the eMBB user in the mini-slot on;
[0067] According to the weight of the eMBB user in the time slot the puncturing situation of the time slot resource block in the mini-slot and the total transmission data rate loss of the eMBB user in the mini-slot on, build a URLLC puncturing model on the mini-slot scale, the specific expression is as follows:
[0068] ;
[0069] ;
[0070] ;
[0071] ;
[0072] ;
[0073] ;
[0074] ;
[0075] In the formula: , is the set of URLLC users u, is the number of URLLC users u in the set; denotes the order puncturing matrix of the puncturing position of URLLC user u; denotes order puncturing bandwidth ratio matrix; denotes order puncturing power ratio matrix.
[0076] Furthermore, modeling the eMBB resource allocation model at the time slot scale and the URLLC puncturing model at the mini-time slot scale as Markov processes respectively includes:
[0077] Modeling the eMBB resource allocation model at the time slot scale as a Markov process, the expression of the state space is as follows:
[0078] ;
[0079] In the formula: is the signal-to-noise ratio of eMBB user at time slot ; is the total transmission data rate of eMBB user at time slot ; is the state space of modeling the eMBB resource allocation model at the time slot scale as a Markov process; is used to represent eMBB;
[0080] The specific expression of the action space of modeling the eMBB resource allocation model at the time slot scale as a Markov process is as follows:
[0081] ;
[0082] In the formula: is the action space of modeling the eMBB resource allocation model at the time slot scale as a Markov process;
[0083] Formulate the reward function of modeling the eMBB resource allocation model at the time slot scale as a Markov process, the specific expression is as follows:
[0084] ;
[0085] In the formula: is the reward function of modeling the eMBB resource allocation model at the time slot scale as a Markov process;
[0086] Modeling the URLLC puncturing model at the mini-time slot scale as a Markov process, the expression of the state space is as follows:
[0087] ;
[0088] In the formula: is the state space for modeling the URLLC puncturing model on the mini-slot scale as a Markov process; represents an eMBB user The total transmission data rate loss on the mini-slot ; is the mini-slot The signal-to-noise ratio of the URLLC user u within; is used to represent URLLC;
[0089] The specific expression of the action space for modeling the URLLC puncturing model on the mini-slot scale as a Markov process is as follows:
[0090] ;
[0091] In the formula: is the action space for modeling the URLLC puncturing model on the mini-slot scale as a Markov process;
[0092] Formulate the reward function for modeling the URLLC puncturing model on the mini-slot scale as a Markov process, and the specific expression is as follows:
[0093] ;
[0094] In the formula: is the reward function for modeling the URLLC puncturing model on the mini-slot scale as a Markov process; is the penalty term, indicating the total number of URLLC users u that do not meet the transmission delay constraint within the mini-slot ; is the penalty coefficient.
[0095] Furthermore, the collaborative training of the eMBB agent and the URLLC agent is implemented based on the PF-TSIPRA algorithm, specifically including:
[0096] The specific training steps of the PF-TSIPRA algorithm are as follows:
[0097] Step 1: Randomly initialize the training network parameters of the eMBB agent , the training network parameters of the URLLC agent and copy them to the corresponding target networks , Initialize the experience replay pool of the eMBB agent and the experience replay pool of the URLLC agent ;
[0098] In the formula: Actor network parameters for the eMBB agent; Critic network parameters for the eMBB agent; Actor network parameters for the URLLC agent; Critic network parameters for the URLLC agent; Target Actor network parameters for the eMBB agent; Target Critic network parameters for the eMBB agent; Target Actor network parameters for the URLLC agent; Target Critic network parameters for the URLLC agent;
[0099] Step 2: At each time slot t, the eMBB agent observes the state , and based on selects the first action ;
[0100] Where: is the exploration noise; is the corresponding Actor network of the eMBB agent;
[0101] Step 3: At each mini time slot n, the URLLC agent observes the state , and based on selects the second action ;
[0102] Where: is the corresponding Actor network of the URLLC agent;
[0103] Step 4: The URLLC agent executes the second action , obtains the second reward and the observation state of the URLLC agent at the next moment ; Stores in the URLLC agent experience replay pool and ends the mini time slot n;
[0104] Step 5: The eMBB agent executes the first action , obtains the first reward and the observation state of the eMBB agent at the next moment ; Stores in the eMBB agent experience replay pool and ends the time slot t;
[0105] Step 6: From the eMBB agent experience replay pool and the URLLC agent experience replay pool randomly sample batch data respectively 、 Based on 、 calculate the target Q value;
[0106] In the formula: = 1, 2,..., C; represents the index of the th sample; C represents the number of sampled data; is the reward discount factor; 、 are the evaluations of the target Critic network 、 for the state 、 and the action 、 respectively; 、 are the decisions made by the target Actor network 、 according to the state 、 respectively; ; ; is the target Critic network of the eMBB agent corresponding to ; is the target Critic network of the URLLC agent corresponding to ; is the target Actor network of the eMBB agent corresponding to ; is the target Actor network of the URLLC agent corresponding to ;
[0107] Step 7: Update the Critic network by minimizing the loss function based on 、 ;
[0108] In the formula: 、 are the evaluations of the Critic network 、 for the current state 、 and the action 、 respectively; is the Critic network of the eMBB agent corresponding to ; is The Critic network of the corresponding URLLC agent;
[0109] Step 8: Use the gradient ascent method based on and Update the Actor network;
[0110] Where: is the action generated by the Actor network of the eMBB agent according to the current policy; is the action generated by the Actor network of the URLLC agent according to the current policy;
[0111] Step 9: Based on and and and Update and and and respectively; where is the soft update coefficient;
[0112] Step 10: Repeat Step 2 to Step 9 until the maximum number of training epochs is reached to complete the training work.
[0113] In a second aspect, the present invention provides an intelligent resource allocation device for an eMBB / URLLC coexistence scenario, and the device includes:
[0114] Acquisition module: used to obtain the maximum transmission data rate of eMBB users without URLLC puncturing;
[0115] First calculation module: used to achieve the coexistence of eMBB / URLLC by means of puncturing, and calculate the transmission data rate loss of eMBB users after being punctured by URLLC according to the maximum transmission data rate of eMBB users;
[0116] Second calculation module: used to calculate the URLLC user transmission data rate based on the finite block length coding formula and implement the reliability constraint of URLLC;
[0117] Model establishment module: used to establish an eMBB resource allocation model on the time slot scale and a URLLC puncturing model on the mini-time slot scale respectively using the proportional fairness algorithm according to the maximum transmission data rate of eMBB users, the transmission data rate loss of eMBB users, and the reliability constraint of URLLC;
[0118] Conversion module: used to model the eMBB resource allocation model on the time slot scale and the URLLC puncturing model on the mini-time slot scale as Markov processes respectively;
[0119] Training module: used to set up an eMBB agent for the Markov process of the eMBB resource allocation model at the time slot scale, set up a URLLC agent for the Markov process of the URLLC puncturing model at the mini-time slot scale, and implement the collaborative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm;
[0120] Output module: used to obtain the first input parameter and the second input parameter, input the first input parameter into the trained eMBB agent, and output the resource block allocation matrix and the power allocation matrix; input the second input parameter into the trained URLLC agent, and output the puncturing matrix, the puncturing bandwidth ratio matrix, and the puncturing power ratio matrix to achieve resource allocation.
[0121] In a third aspect, the present invention provides a terminal, including a processor and a storage medium;
[0122] The storage medium is used to store instructions;
[0123] The processor is used to operate according to the instructions to execute the steps of the method according to the first aspect.
[0124] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method according to the first aspect.
[0125] Compared with the prior art, the beneficial effects achieved by the present invention:
[0126] The present invention proposes an intelligent resource allocation method for the eMBB / URLLC coexistence scenario aiming at the eMBB and URLLC resource allocation problems in the downlink scenario; this solution decomposes the scheduling problem into two sub-problems, allocates bandwidth and power resources for eMBB users at the time slot scale; performs URLLC puncturing at the mini-time slot scale to solve a reasonable bandwidth and power preemption scheme; this solution takes into account the channel differences between eMBB users and the impact of URLLC puncturing, constructs puncturing preferences using the proportional fairness method, and maintains fairness among eMBB users; by continuously interacting with the environment to cope with the random and occasional characteristics of URLLC services, it maximizes the service satisfaction of eMBB services while meeting the QoS requirements of URLLC users and solves the optimal resource allocation strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0127] Figure 1 is a schematic flowchart of an intelligent resource allocation method for an eMBB / URLLC coexistence scenario provided by an embodiment of the present invention;
[0128] Figure 2 is a schematic diagram of puncturing in an eMBB / URLLC coexistence scenario provided by an embodiment of the present invention;
[0129] Figure 3 It is the execution flowchart of PF - TSIPRA provided according to the embodiments of the present invention. Specific embodiments
[0130] The technical solution of the present invention will be described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0131] The term "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " generally represents an "or" relationship between the associated objects before and after.
[0132] Embodiment 1:
[0133] As Figure 1 shown, the present invention provides an intelligent resource allocation method for eMBB / URLLC co - existence scenarios, including the following steps:
[0134] Step 1: Obtain the maximum transmission data rate of eMBB users without URLLC puncturing.
[0135] Step 2: Achieve the co - existence of eMBB / URLLC by means of puncturing, and calculate the transmission data rate loss of eMBB users after being punctured by URLLC according to the maximum transmission data rate of eMBB users.
[0136] Step 3: Calculate the transmission data rate of URLLC users based on the finite - block - length coding formula and achieve the reliability constraint of URLLC.
[0137] Step 4: According to the maximum transmission data rate of eMBB users, the transmission data rate loss of eMBB users, and the reliability constraint of URLLC, use the proportional fairness algorithm to establish an eMBB resource allocation model on the time - slot scale and a URLLC puncturing model on the mini - time - slot scale respectively.
[0138] Step 5: Model the eMBB resource allocation model on the time - slot scale and the URLLC puncturing model on the mini - time - slot scale as Markov processes respectively.
[0139] Step 6: Set up the eMBB agent for the Markov process of the eMBB resource allocation model at the time slot scale, and set up the URLLC agent for the Markov process of the URLLC puncturing model at the mini-time slot scale. Based on the PF-TSIPRA algorithm, implement the collaborative training of the eMBB agent and the URLLC agent;
[0140] Step 7: Obtain the first input parameter and the second input parameter. Input the first input parameter into the trained eMBB agent to output the resource block allocation matrix and the power allocation matrix; input the second input parameter into the trained URLLC agent to output the puncturing matrix, the puncturing bandwidth ratio matrix, and the puncturing power ratio matrix, so as to achieve resource allocation.
[0141] Specifically, this solution considers the downlink scenario supported by the next generation node B (gNodeB), adopts orthogonal frequency division multiple access, and its coverage area contains several eMBB users and URLLC users, which are represented by the sets , respectively; the available bandwidth is divided into the form of resource blocks, denoted as ; in the time dimension, the time slot is used as the basic unit, denoted as , where the duration of each time slot is 1 ms. On this basis, the concept of mini-time slot is proposed in 5G NR (a global 5G standard based on the brand new air interface design of OFDM), and each time slot is further divided into several mini-time slots, denoted as ; in order to meet the QoS requirements of high reliability and low latency for URLLC users, 3GPP proposed the concept of puncturing scheduling. When the URLLC service arrives, it will puncture the ongoing eMBB service on the mini-time slot to ensure the low latency and reliability of the URLLC service.
[0142] Different from the eMBB service, the URLLC service has higher requirements for latency, so it can be transmitted across resource blocks; in this solution, it is assumed that there is always eMBB data transmission, and the URLLC service can puncture multiple eMBB resource blocks, and the puncturing method is as shown in Figure 2 . The left half of the figure is a simple scenario diagram containing several eMBB users and several URLLC users, and the right half is a URLLC puncturing diagram. It can be seen from the figure that after the URLLC service arrives, it will preempt the bandwidth resources of the eMBB service on the mini-time slot scale; the gNodeB allocates bandwidth and power resources for the eMBB service at the beginning of each time slot, and the URLLC service punctures the eMBB service on the mini-time slot scale after it arrives.
[0143] An embodiment, obtaining the maximum transmission data rate of eMBB users without URLLC puncturing, specifically includes:
[0144] Obtaining the allocation situation of time slots Resource blocks :
[0145] ;
[0146] In the formula: is the allocation situation of time slot Resource blocks ;
[0147] In the case of no URLLC service puncturing, the maximum transmission data rate that eMBB user e can achieve on time slot Resource blocks is:
[0148] ;
[0149] In the formula: is the time slot; is the resource block; is the eMBB user; is the maximum transmission data rate that the eMBB user can achieve on time slot t and resource block k in the case of no URLLC service puncturing; is the eMBB user 's downlink transmission power on time slot t and resource block k; is the small-scale channel fading factor; is the bandwidth of resource block k; is the noise power spectral density; is the path loss factor; is the eMBB user 's distance from the next-generation base station gNodeB; is the path loss exponent;
[0150] Then the total transmission data rate of eMBB user on time slot is:
[0151] ;
[0152] In the formula: is the total transmission data rate of eMBB user on time slot .
[0153] An embodiment realizes the coexistence of eMBB / URLLC in a puncturing manner, and calculates the transmission data rate loss of eMBB users after being punctured by URLLC according to the maximum transmission data rate of eMBB users, specifically including:
[0154] To meet the reliability and low-latency requirements of URLLC services, on the mini-slot scale, URLLC services will puncture the resources of ongoing eMBB services, resulting in the loss of eMBB service transmission power and bandwidth resources; obtain the mini-slots resource blocks of The puncturing situation is as follows:
[0155] ;
[0156] In the formula: is the mini-slot resource block of The puncturing situation;
[0157] Obtain the bandwidth and transmission power obtained when URLLC punctures eMBB, and the specific expressions are as follows:
[0158] ;
[0159] ;
[0160] In the formula: is the URLLC user; the time slot includes mini-slots, is an integer, , The set of mini-slots is ; is the URLLC user in the time slot resource block of The bandwidth obtained by puncturing the mini-slot; is the URLLC user in the time slot resource block of The transmission power obtained by puncturing the mini-slot; is the time slot resource block of The puncturing bandwidth ratio of the mini-slot, ; is the time slot Resource block mini-slot of puncturing power ratio of, ;
[0161] Assume that each eMBB service may be punctured by multiple URLLC services, and the URLLC service may also cross resource blocks to puncture the eMBB service. According to the bandwidth and transmission power obtained when the URLLC punctures the eMBB, calculate the eMBB user after being punctured by the URLLC at the th mini-slot of resource block k in time slot t, the specific expression is as follows and can be approximated as:
[0162] ;
[0163] In the formula: is the data rate of the eMBB user after being punctured by the URLLC at the th mini-slot of resource block k in time slot t;
[0164] According to the maximum transmission data rate of the eMBB user without being punctured by the URLLC and the data rate of the eMBB user after being punctured by the URLLC at the time slot resource block mini-slot of calculate the transmission data rate loss of the eMBB user after being punctured by the URLLC at the time slot resource block mini-slot of The specific expression is as follows:
[0165] ;
[0166] In the formula: is the transmission data rate loss of the eMBB user after being punctured by the URLLC at the time slot resource block mini-slot of ;
[0167] According to the transmission data rate loss of the eMBB user after being punctured by the URLLC at the time slot resource block mini-slot of calculate the total transmission data rate loss of the eMBB user at the mini-slot The specific expression is as follows:
[0168] ;
[0169] where: is an eMBB user The total transmission data rate loss on the mini-slot is
[0170] An embodiment calculates the transmission data rate of URLLC users based on the finite block length coding formula and realizes the reliability constraint of URLLC, specifically including:
[0171] The tolerable transmission delay of URLLC services is , and its maximum tolerable transmission packet loss probability is ; Since the packet length of URLLC is usually short, the Shannon formula based on infinite block length is no longer applicable, so the finite block length channel coding formula is used to approximate the transmission rate of URLLC services.
[0172] Based on the finite block length coding formula, calculate the data rate of URLLC users in the th mini-slot of the time slot, and the specific expression is as follows:
[0173] ;
[0174] ;
[0175] ;
[0176] where: is the signal-to-noise ratio of URLLC user ; represents the channel dispersion, and when the signal-to-noise ratio is high can be approximated as 1; is the Gaussian inverse function, ; is the length of the mini-slot ; , is the set of resource blocks , is the number of resource blocks in the set; is the time slot The resource block The puncturing situation of the mini-slot; is the data rate of URLLC user in the th mini-slot of the time slot;
[0177] Let Indicates a URLLC user In a time slot The th mini - time - slot's transmission traffic demand, then the specific expression of the transmission delay constraint for the URLLC user is as follows:
[0178] ;
[0179] In the formula: is the tolerable transmission delay of URLLC;
[0180] Meanwhile, the reliability constraint of the URLLC service also needs to be considered. This constraint is described as the packet loss rate of the URLLC service per unit time being less than or equal to the maximum tolerable packet loss rate of the URLLC service. To achieve the reliability constraint of URLLC, the specific expression is as follows:
[0181] ;
[0182] In the formula: is the packet loss rate of URLLC per unit time; Indicates the URLLC user In the time slot The th mini - time - slot does not meet the transmission delay constraint. The specific calculation method is as follows:
[0183] ;
[0184] Among them, As a binary indicator, it indicates whether the user In the time slot The th mini - time - slot meets the transmission delay constraint. Taking 1 means not meeting the transmission delay constraint, otherwise taking 0. The calculation method is as follows:
[0185] .
[0186] An embodiment is to establish an eMBB resource allocation model on the time - slot scale and a URLLC puncturing model on the mini - time - slot scale respectively using the proportional fairness algorithm according to the maximum transmission data rate of eMBB, the transmission data rate loss of eMBB users, and the URLLC reliability constraint. Specifically, it includes:
[0187] The gNodeB allocates bandwidth resources and power resources for eMBB services at the beginning of each time slot, enabling as many eMBB users as possible to reach the target transmission data rate; when URLLC services arrive, it preempts the bandwidth and power resources of eMBB services in the mini-slot dimension so that URLLC services can be served promptly; the optimization problem proposed in this scheme includes two sub-problems: eMBB resource allocation and URLLC puncturing, which are respectively established on two time scales: time slots and mini-slots.
[0188] The service satisfaction of eMBB users is an important QoS metric for eMBB services (the service satisfaction of eMBB users is defined as the percentage of eMBB users whose transmission data rate can reach the target transmission data rate), which means that all users need to be taken into account during resource scheduling to maximize fairness among users as much as possible; because different eMBB users have different channel conditions, simply using the maximization of the total transmission data rate of users as the objective function may cause the gNodeB to ignore those eMBB users with poor channel conditions, and random URLLC puncturing may also cause the transmission data rate of these users to be lower than the target transmission data rate; therefore, the proportional fairness algorithm is adopted for eMBB resource allocation and URLLC puncturing, introducing the historical throughput of eMBB users as a weight to schedule to each user as much as possible while maximizing the total throughput.
[0189] Establish an eMBB resource allocation model on the time slot scale, specifically including:
[0190] Obtain eMBB users In the time slot The weight is as follows:
[0191] ;
[0192] In the formula: Is the weight of eMBB user In the time slot The weight, Is the historical average transmission data rate of eMBB user The specific expression is as follows:
[0193] ;
[0194] In the formula: Is the time window length for calculating the historical average transmission data rate; Is the eMBB user In the time slot The total transmission data rate; Is based on the time slot Is the eMBB user calculated based on Historical average transmission data rate;
[0195] On the time slot scale, the gNodeB allocates bandwidth and frequency resources for eMBB services at the beginning of the time slot, and according to the eMBB users In the time slot weight, time slot resource block allocation situation and eMBB users In the time slot total transmission data rate, establish an eMBB resource allocation model on the time slot scale, and the specific expression is as follows:
[0196] ;
[0197] (1);
[0198] (2);
[0199] (3);
[0200] (4);
[0201] In the formula: is the maximum transmission power of the next-generation base station gNodeB; , is the set of eMBB users e, is the number of eMBB users e in the set; represents order resource block allocation matrix; represents order power allocation matrix; Constraints 1 and 2 mean that the resource block can only be allocated to one user within the time slot; Constraints 3 and 4 mean that the sum of the powers allocated to all resource blocks should be lower than the maximum transmission power of the base station; The goal of the optimization problem P1 is to find the optimal solutions of resource block allocation and power allocation, and maximize the eMBB user throughput while taking into account the fairness of eMBB users;
[0202] Establish a URLLC puncturing model on the mini time slot scale, specifically including:
[0203] The URLLC service is scheduled on the mini time slot scale. By allocating reasonable puncturing bandwidth ratio and power ratio for the URLLC service to meet the delay and reliability requirements of the URLLC service, and at the same time minimizing the transmission data rate loss caused by puncturing to the eMBB service, according to the eMBB users In the time slot weight, time slot resource block mini time slot The perforation situation of and eMBB users In the mini-slot The total transmission data rate loss on, establish a URLLC perforation model on the mini-slot scale, and the specific expression is as follows:
[0204] ;
[0205] (5);
[0206] (6);
[0207] (7);
[0208] (8);
[0209] (9);
[0210] (10);
[0211] In the formula: , is the set of URLLC users u, is the number of URLLC users u in the set; represents the order perforation matrix of the perforation position of URLLC user u; represents order perforation bandwidth occupancy ratio matrix; represents order perforation power occupancy ratio matrix; Constraint 5 ensures the low-latency and reliability requirements of URLLC; Constraint 6 is a binary perforation indicator; Constraints 7 and 9 represent the bandwidth resources and power resources allocated to URLLC services; Constraints 8 and 10 represent that the bandwidth resources and power resources allocated to URLLC services do not exceed the bandwidth resources and power resources of the eMBB service in the current resource block.
[0212] An embodiment models the eMBB resource allocation model on the slot scale and the URLLC perforation model on the mini-slot scale as Markov processes respectively, specifically including:
[0213] In the previous embodiment, the optimization problem P1 requires making resource block allocation and power allocation decisions, which contain both discrete variables and continuous variables, and its objective function is non-linear; the puncturing model of the optimization problem P2 also contains integer decisions and discrete solution variables, and its objective function is also non-linear; therefore, these two problems are mixed integer non-linear programming problems. The introduction of integer decision variables will expand the solution space and make it complex to find the optimal solution; at the same time, considering the randomness of URLLC services, the two sub-problems are respectively modeled as Markov processes (markov decision process, MDP), and reinforcement learning is used to handle the uncertainty of the environment and find the optimal solution.
[0214] Among them, MDP usually includes a state space , an action space and a reward function . When in state , the agent executes action , obtains reward and the state at the next moment.
[0215] MDP parameters for eMBB services: The state space of the eMBB agent is as follows:
[0216] ;
[0217] In the formula: is the signal-to-noise ratio of eMBB user at time slot ; is the total transmission data rate of eMBB user at time slot ; is the state space of the eMBB resource allocation model modeled as a Markov process on the time slot scale; is used to represent eMBB;
[0218] The specific expression of the action space of the eMBB resource allocation model modeled as a Markov process on the time slot scale is as follows:
[0219] ;
[0220] In the formula: is the action space of the eMBB resource allocation model modeled as a Markov process on the time slot scale;
[0221] Based on the optimization objective of P1, the reward function of the eMBB resource allocation model modeled as a Markov process on the time slot scale is formulated, and the specific expression is as follows:
[0222] ;
[0223] Wherein: is the reward function for modeling the eMBB resource allocation model as a Markov process on the time slot scale.
[0224] MDP parameters of URLLC services: The state space of the URLLC agent is as follows:
[0225] ;
[0226] Wherein: is the state space for modeling the URLLC puncturing model as a Markov process on the mini time slot scale; represents the eMBB user on the mini time slot total transmission data rate loss; is the mini time slot signal-to-noise ratio of URLLC user u within; is used to represent URLLC;
[0227] The specific expression of the action space for modeling the URLLC puncturing model as a Markov process on the mini time slot scale is as follows:
[0228] ;
[0229] Wherein: is the action space for modeling the URLLC puncturing model as a Markov process on the mini time slot scale;
[0230] Based on the goal of minimizing the transmission data rate loss of eMBB users, the reward function for modeling the URLLC puncturing model as a Markov process on the mini time slot scale is formulated, and the specific expression is as follows:
[0231] ;
[0232] Wherein: is the reward function for modeling the URLLC puncturing model as a Markov process on the mini time slot scale; is the penalty term, indicating the total number of URLLC users u that do not meet the transmission delay constraint within the mini time slot ; is the penalty coefficient.
[0233] As Figure 3 shown, an embodiment realizes the collaborative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm, specifically including:
[0234] The specific training steps of the PF-TSIPRA algorithm are as follows:
[0235] Step 1: Randomly initialize the training network parameters of the eMBB agent and the training network parameters of the URLLC agent and copy them to the corresponding target networks , Initialize the experience replay pool of the eMBB agent and the experience replay pool of the URLLC agent ;
[0236] Where: is the Actor network parameter of the eMBB agent; is the Critic network parameter of the eMBB agent; is the Actor network parameter of the URLLC agent; is the Critic network parameter of the URLLC agent; is the target Actor network parameter of the eMBB agent; is the target Critic network parameter of the eMBB agent; is the target Actor network parameter of the URLLC agent; is the target Critic network parameter of the URLLC agent;
[0237] Step 2: At each time slot t, the eMBB agent observes the state and selects the first action based on ;
[0238] Where: is the exploration noise; is the corresponding Actor network of the eMBB agent;
[0239] Step 3: At each mini time slot n, the URLLC agent observes the state and selects the second action based on ;
[0240] Where: is the corresponding Actor network of the URLLC agent;
[0241] Step 4: The URLLC agent executes the second action , obtains the second reward and the observation state of the URLLC agent at the next moment ; Store in the experience replay pool of the URLLC agent In it, end mini-slot n;
[0242] Step 5: The eMBB agent executes the first action , and obtains the first reward and the eMBB agent's observation state at the next moment ; Store in the eMBB agent's experience replay pool and end slot t;
[0243] Step 6: Randomly sample batch data from the eMBB agent's experience replay pool and the URLLC agent's experience replay pool respectively 、 , and calculate the target Q value based on 、 ;
[0244] In the formula: = 1, 2,..., C; represents the index of the th sample; C represents the number of sampled data; is the reward discount factor; 、 are the evaluations of the target Critic network 、 for the state 、 and the actions 、 respectively; 、 are the decisions made by the target Actor network 、 according to the states 、 respectively; ; ; is the target Critic network of the eMBB agent corresponding to ; is the target Critic network of the URLLC agent corresponding to ; is the target Actor network of the eMBB agent corresponding to ; is the target Actor network of the URLLC agent corresponding to ;
[0245] Step 7: Update the Critic network by minimizing the loss function based on 、 ;
[0246] In the formula: , are the evaluations of the Critic network , for the current state , of the action , respectively; is the Critic network of the corresponding eMBB agent; is the Critic network of the corresponding URLLC agent;
[0247] Step 8: Update the Actor network based on , using the gradient ascent method;
[0248] In the formula: is the action generated by the Actor network of the eMBB agent according to the current policy; is the action generated by the Actor network of the URLLC agent according to the current policy;
[0249] Step 9: Update , , , respectively based on , , and ; where is the soft update coefficient;
[0250] Step 10: Repeat Steps 2 to 9 until the maximum number of training epochs is reached to complete the training work.
[0251] Specifically, this solution proposes a two-stage intelligent puncture resource allocation algorithm based on proportional fairness (PF-TSIPRA), which separately sets the eMBB agent and the URLLC agent to solve the optimization problems P1 and P2; the eMBB agent executes at the time slot scale, makes decisions at the beginning of each time slot, and at the same time, the URLLC agent makes decisions in a loop on the mini time slots within each time slot, as specifically shown in Figure 3 shown.
[0252] The eMBB agent includes the Actor network , the Critic network , the target Actor network , the target Critic network , and their corresponding network parameters are respectively 、 、 and ; The URLLC agent includes an Actor network , a Critic network , the target Actor network , the target Critic network , and their corresponding network parameters are respectively 、 、 and ; and are used to distinguish eMBB and URLLC;
[0253] When in the state 、 , the Actor networks 、 generate actions 、 respectively according to the current policy, and add exploration noise , and the specific expressions are as follows:
[0254] ;
[0255] ;
[0256] After the eMBB agent executes the first action , it obtains the first reward , generates the state at the next moment , and stores in the eMBB agent experience replay pool ; The URLLC agent executes the second action , obtains the second reward and the observed state of the URLLC agent at the next moment ; Stores in the URLLC agent experience replay pool ; Subsequently, the eMBB agent and the URLLC agent respectively randomly sample several groups of sample data from the eMBB agent experience replay pool and the URLLC agent experience replay pool , denoted as 、 , for updating their respective Actor networks and Critic networks, and calculating the target Q-values of the eMBB agent's Critic network using the Bellman equation and the target Q-values of the URLLC agent's Critic network , the specific expressions are as follows:
[0257] ;
[0258] ;
[0259] In the formula: = 1, 2,..., C; represents the index of the -th sample; C represents the number of sampled data; is the reward discount factor; , are the evaluations of the target Critic network , for the actions , in the states , respectively; , are the decisions made by the target Actor network and according to the states , respectively; where, ; .
[0260] The Critic network parameters , are updated by minimizing the loss function:
[0261] ;
[0262] ;
[0263] In the formula: is the learning rate of the Critic network; is the -corresponding Critic network loss function; is the -corresponding Critic network loss function, calculated as follows:
[0264] ;
[0265] ;
[0266] Among them, , are the Critic networks respectively , evaluate the current state , and the actions , therein;
[0267] Using the gradient ascent method, update the Actor network parameters by maximizing the Q value of the Critic network:
[0268] ;
[0269] ;
[0270] In the formula: is the learning rate of the Actor network; and are calculated as follows:
[0271] ;
[0272] ;
[0273] In the formula: is the action generated by the Actor network of the eMBB agent according to the current policy; is the action generated by the Actor network of the URLLC agent according to the current policy;
[0274] Finally, use soft update to update the target network parameters , , and :
[0275] ;
[0276] ;
[0277] ;
[0278] ;
[0279] In the formula: is the soft update coefficient;
[0280] Repeat the above steps until the maximum number of training epochs is reached to complete the training work.
[0281] After training, in each time slot, the eMBB agent outputs the resource block allocation matrix according to the input parameter and a power allocation matrix ; in each mini-slot, the URLLC agent outputs a puncturing matrix , a puncturing bandwidth ratio matrix , and a puncturing power ratio matrix according to the input parameters .
[0282] The present invention aims at the eMBB and URLLC resource allocation problems in the downlink scenario, and proposes an intelligent resource allocation method for the eMBB / URLLC coexistence scenario; this solution decomposes the scheduling problem into two sub-problems, allocates bandwidth and power resources for eMBB users on the time-slot scale; performs URLLC puncturing on the mini-slot scale to solve a reasonable bandwidth and power preemption scheme; this solution takes into account the channel differences among eMBB users and the impact of URLLC puncturing, constructs a puncturing preference using the proportional fairness method to maintain fairness among eMBB users; responds to the random and occasional characteristics of URLLC services by continuously interacting with the environment, maximizes the service satisfaction of eMBB services while meeting the QoS requirements of URLLC users, and solves the optimal resource allocation strategy.
[0283] Embodiment 2:
[0284] The present invention provides an intelligent resource allocation device for the eMBB / URLLC coexistence scenario, and the device includes:
[0285] An acquisition module: used to obtain the maximum transmission data rate of eMBB users without URLLC puncturing;
[0286] A first calculation module: used to achieve the coexistence of eMBB / URLLC by puncturing, and calculate the loss of the transmission data rate of eMBB users after being punctured by URLLC according to the maximum transmission data rate of eMBB users;
[0287] A second calculation module: used to calculate the transmission data rate of URLLC users based on the finite block length coding formula and achieve the reliability constraint of URLLC;
[0288] A model establishment module: used to respectively establish an eMBB resource allocation model on the time-slot scale and a URLLC puncturing model on the mini-slot scale using the proportional fairness algorithm according to the maximum transmission data rate of eMBB users, the loss of the transmission data rate of eMBB users, and the reliability constraint of URLLC;
[0289] A conversion module: used to respectively model the eMBB resource allocation model on the time-slot scale and the URLLC puncturing model on the mini-slot scale as Markov processes;
[0290] Training module: It is used to set up the eMBB agent for the Markov process of the eMBB resource allocation model at the time slot scale, set up the URLLC agent for the Markov process of the URLLC puncturing model at the mini-slot scale, and implement the collaborative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm;
[0291] Output module: It is used to obtain the first input parameter and the second input parameter, input the first input parameter into the trained eMBB agent, and output the resource block allocation matrix and the power allocation matrix; input the second input parameter into the trained URLLC agent, and output the puncturing matrix, the puncturing bandwidth ratio matrix and the puncturing power ratio matrix to achieve resource allocation.
[0292] Embodiment 3:
[0293] An embodiment of the present invention further provides a terminal, including a processor and a storage medium;
[0294] The storage medium is used to store instructions;
[0295] The processor is used to operate according to the instructions to execute the steps of the method described in Embodiment 1.
[0296] Embodiment 4:
[0297] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method described in Embodiment 1.
[0298] Since the storage medium provided by the embodiment of the present invention can execute the method provided by the first embodiment of the present invention, it has the corresponding functional modules and beneficial effects of the execution method.
[0299] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0300] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0301] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0302] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0303] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principles of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. An intelligent resource allocation method for eMBB / URLLC coexistence scenario, characterized in that: The following steps are involved: Get the maximum transmission data rate of eMBB users without URLLC puncturing; The coexistence of eMBB / URLLC is achieved by using the puncturing method. The transmission data rate loss of the eMBB user after URLLC puncturing is calculated based on the maximum transmission data rate of the eMBB user. Calculate the URLLC user transmission data rate based on the finite block length coding formula and implement the URLLC reliability constraint; According to the maximum transmission data rate of eMBB users, the transmission data rate loss of eMBB users and the reliability constraint of URLLC, the proportional fairness algorithm is used to establish the eMBB resource allocation model on the time slot scale and the URLLC perforation model on the mini time slot scale. The eMBB resource allocation model at the time slot scale and the URLLC perforation model at the mini-time slot scale are modeled as Markov processes respectively; The eMBB agent is set for the Markov process of the eMBB resource allocation model at the time slot scale, and the URLLC agent is set for the Markov process of the URLLC perforation model at the mini time slot scale. The collaborative training of the eMBB agent and the URLLC agent is implemented based on the PF-TSIPRA algorithm. Obtain a first input parameter and a second input parameter, input the first input parameter to a trained eMBB agent, and output a resource block allocation matrix and a power allocation matrix; The second input parameter is input into the trained URLLC agent, and the perforation matrix, perforation bandwidth ratio matrix and perforation power ratio matrix are output to realize resource allocation.
2. The intelligent resource allocation method for the eMBB / URLLC coexistence scenario according to claim 1, characterized in that: The maximum transmission data rate of the eMBB user without URLLC puncturing is obtained, and the specific expression is as follows: ; Where: is the time slot; is a resource block; For eMBB users; For eMBB users without perforation The maximum transmission data rate achievable in resource block k in time slot t; For eMBB users Downlink transmission power in resource block k in time slot t; is the small-scale channel fading factor; is the bandwidth of resource block k; is the noise power spectral density; is the path loss factor; For eMBB users The distance to the next generation base station gNodeB; is the path loss exponent.
3. The intelligent resource allocation method for the eMBB / URLLC coexistence scenario according to claim 2, characterized in that: The calculating, according to the maximum transmission data rate of the eMBB user, the transmission data rate loss of the eMBB user after URLLC puncturing specifically includes: The bandwidth and transmission power obtained when URLLC punctures eMBB are obtained. The specific expressions are as follows: ; ; Where: For URLLC users; time slot include Mini time slots, is an integer, , The set of mini-slots is , specifically expressed as ; For URLLC users In time slot Resource Block Mini time slots Bandwidth obtained by perforation; For URLLC users In time slot Resource Block Mini time slots Transmission power obtained by perforation; For time slot Resource Block Mini time slots The proportion of perforation bandwidth, ; For time slot Resource Block Mini time slots The proportion of perforation power, ; According to the bandwidth and transmission power obtained when URLLC punctures eMBB, the eMBB user after URLLC puncture is calculated. In time slot t, resource block k The data rate on a mini-time slot is expressed as follows: ; Where: For eMBB users after URLLC perforation In time slot t, resource block k The data rate per mini-slot; According to the maximum transmission data rate of the eMBB user without URLLC puncturing and the maximum transmission data rate of the eMBB user after URLLC puncturing In time slot Resource Block Mini time slots The data rate of the eMBB user after URLLC puncturing is calculated. In time slot Resource Block Mini time slots The transmission data rate loss on the , the specific expression is as follows: ; Where: For eMBB users after URLLC perforation In time slot Resource Block Mini time slots The data rate loss on the transmission line.
4. The intelligent resource allocation method for the eMBB / URLLC coexistence scenario according to claim 3, characterized in that: The calculation of the URLLC user transmission data rate based on the finite block length coding formula specifically includes: Get time slot Resource Block Mini time slots The specific expression of perforation is as follows: ; Where: For time slot Resource Block Mini time slots perforation condition; According to time slot Resource Block Mini time slots Based on the finite block length coding formula, the URLLC user In time slot No. The data rate of a mini-time slot is expressed as follows: ; ; ; Where: For URLLC users signal-to-noise ratio; represents channel dispersion; is the inverse Gaussian function; Mini-slot Length; is the maximum tolerable transmission packet loss probability of URLLC; , Resource Block A collection of For resource blocks in the collection the number of For time slot Resource Block Mini time slots perforation condition; For URLLC users In time slot No. The data rate per mini-slot; make Indicates URLLC user In time slot No. The transmission traffic demand of mini-time slots is The specific expression of the transmission delay constraint is as follows: ; Where: is the tolerable transmission delay of URLLC; The reliability constraint for implementing URLLC is specifically expressed as follows: ; Where: is the packet loss rate of URLLC per unit time; Indicates URLLC user In time slot No. The transmission delay constraint is not met within the mini-time slot.
5. The intelligent resource allocation method for the eMBB / URLLC coexistence scenario according to claim 4, characterized in that: The method of respectively establishing an eMBB resource allocation model on a time slot scale and a URLLC perforation model on a mini time slot scale includes: Establish an eMBB resource allocation model on a time slot scale, including: Get eMBB users In time slot The specific expression is as follows: ; Where: For eMBB users In time slot The weight of For eMBB users The historical average transmission data rate is expressed as follows: ; Where: The length of the time window for calculating the historical average transmission data rate; For eMBB users In time slot The total transmission data rate; Time slot eMBB users calculated for the baseline The historical average transmission data rate; Get time slot Resource Block The allocation of time slots Resource Block The allocation of eMBB users and the maximum transmission data rate of eMBB users without URLLC perforation are obtained. In time slot The total transmission data rate is expressed as follows: ; ; Where: For time slot Resource Block The distribution of For eMBB users In time slot The total transmission data rate; According to eMBB users In time slot The weight and time slot Resource Block Distribution and eMBB users In time slot The total transmission data rate is used to establish the eMBB resource allocation model on the time slot scale. The specific expression is as follows: ; ; ; ; ; Where: The maximum transmission power of the next generation base station gNodeB; , is the set of eMBB users e, is the number of eMBB users e in the set; express A resource block allocation matrix of order; express The power allocation matrix of order; Establish a URLLC perforation model on a mini-timeslot scale, including: According to the eMBB user after URLLC puncturing In time slot Resource Block Mini time slots The transmission data rate loss on the eMBB user In the mini time slot The total transmission data rate loss on the , the specific expression is as follows: ; Where: For eMBB users In the mini time slot Total transmission data rate loss on According to eMBB users In time slot The weight and time slot Resource Block Mini time slots Perforation Situation and eMBB Users In the mini time slot The total transmission data rate loss on the mini-time slot scale is calculated to establish the URLLC perforation model. The specific expression is as follows: ; ; ; ; ; ; ; Where: , is the set of URLLC users u, is the number of URLLC users u in the set; Indicates the URLLC user u puncture position order perforated matrix; express The perforation bandwidth ratio matrix of order; express The perforation power ratio matrix of each order.
6. The intelligent resource allocation method for the eMBB / URLLC coexistence scenario according to claim 5, characterized in that: The eMBB resource allocation model on the time slot scale and the URLLC perforation model on the mini time slot scale are modeled as Markov processes respectively, including: The eMBB resource allocation model at the time slot scale is modeled as a Markov process, and the expression of the state space is as follows: ; Where: For eMBB users In time slot signal-to-noise ratio; For eMBB users In time slot The total transmission data rate; Model the eMBB resource allocation model at the time slot scale as the state space of a Markov process; Used to represent eMBB; The eMBB resource allocation model at the time slot scale is modeled as the action space of a Markov process. The specific expression is as follows: ; Where: Model the eMBB resource allocation model at the time slot scale as the action space of a Markov process; The eMBB resource allocation model on the time slot scale is modeled as a reward function of a Markov process. The specific expression is as follows: ; Where: Model the eMBB resource allocation model at the time slot scale as a reward function of a Markov process; The URLLC perforation model on the mini-slot scale is modeled as a Markov process, and the expression of the state space is as follows: ; Where: The URLLC perforation model at the mini-slot scale is modeled as the state space of a Markov process; Indicates eMBB users In the mini time slot Total transmission data rate loss on It's a mini time slot The signal-to-noise ratio of the inner URLLC user u; Used to represent URLLC; The URLLC perforation model on the mini-slot scale is modeled as the action space of a Markov process. The specific expression is as follows: ; Where: The URLLC perforation model at the mini-slot scale is modeled as the action space of a Markov process; The URLLC perforation model on the mini-slot scale is formulated as a reward function of a Markov process, and the specific expression is as follows: ; Where: Model the URLLC perforation model on the mini-slot scale as a reward function of a Markov process; is the penalty term, representing the mini-slot The total number of URLLC users u that do not meet the transmission delay constraint; is the penalty coefficient.
7. The intelligent resource allocation method for the eMBB / URLLC coexistence scenario according to claim 6, characterized in that: The collaborative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm specifically includes: The specific training steps of the PF-TSIPRA algorithm are as follows: Step 1: Randomly initialize the training network parameters of the eMBB agent , URLLC agent training network parameters And copy it to the corresponding target network , Initialize the eMBB agent experience replay pool and URLLC agent experience replay pool ; Where: Actor network parameters of the eMBB agent; is the Critic network parameter of the eMBB agent; Actor network parameters of the URLLC agent; are the critic network parameters of the URLLC agent; The target Actor network parameters of the eMBB agent; is the target critic network parameter of the eMBB agent; The target Actor network parameters of the URLLC agent; is the target critic network parameter of the URLLC agent; Step 2: At each time slot t, the eMBB agent observes the state , and based on Select first action ; Where: To explore noise; for The corresponding Actor network of eMBB agents; Step 3: At each mini-slot n, the URLLC agent observes the state , and based on Select the second action ; Where: for The corresponding URLLC agent's Actor network; Step 4: URLLC agent performs the second action , get the second reward and the URLLC agent observation state at the next moment ;Will Deposit into URLLC agent experience replay pool In, end mini-slot n; Step 5: eMBB agent performs the first action , get the first reward and the observed state of the eMBB agent at the next moment ;Will Stored in the eMBB agent experience replay pool In, end time slot t; Step 6: Replay experience from the eMBB agent pool and URLLC agent experience replay pool Randomly sample batch data in , ,based on , Calculate the target Q value; Where: =1, 2, ..., C; Indicates The index of the sample; C represents the number of sample data extracted; is the reward discount factor; , The target critic network is , Status , Next action , Evaluation; , The target Actor network , According to the status , Decisions made; ; ; for The corresponding target critic network of the eMBB agent; for The corresponding URLLC agent’s target critic network; for The corresponding target Actor network of the eMBB agent; for The target Actor network of the corresponding URLLC agent; Step 7: Based on , Minimize the loss function to update the Critic network; Where: , Critic network , Current status , Next action , Evaluation; for The corresponding Critic network of the eMBB agent; for The corresponding URLLC agent’s Critic network; Step 8: Use the gradient ascent method based on , Update the Actor network; Where: Actions generated by the Actor network of the eMBB agent according to the current strategy; The actions generated by the URLLC agent's Actor network according to the current strategy; Step 9: Based on , , , Update separately , , and ;in, is the soft update coefficient; Step 10: Repeat steps 2 to 9 until the maximum number of training cycles is reached and the training is completed.
8. An intelligent resource allocation device for eMBB / URLLC coexistence scenario, characterized in that: The device comprises: Acquisition module: used to obtain the maximum transmission data rate of eMBB users without URLLC puncturing; The first calculation module is used to achieve the coexistence of eMBB / URLLC by using a puncturing method, and calculate the transmission data rate loss of the eMBB user after the URLLC puncturing according to the maximum transmission data rate of the eMBB user; The second calculation module is used to calculate the URLLC user transmission data rate based on the finite block length coding formula and implement the reliability constraint of the URLLC; Model building module: used to build an eMBB resource allocation model on a time slot scale and a URLLC perforation model on a mini time slot scale using a proportional fairness algorithm according to the maximum transmission data rate of the eMBB user, the transmission data rate loss of the eMBB user, and the reliability constraint of the URLLC; Conversion module: used to model the eMBB resource allocation model at the time slot scale and the URLLC perforation model at the mini time slot scale as Markov processes respectively; Training module: used to set the eMBB agent for the Markov process of the eMBB resource allocation model on the time slot scale, set the URLLC agent for the Markov process of the URLLC perforation model on the mini time slot scale, and implement the collaborative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm; Output module: used to obtain the first input parameter and the second input parameter, input the first input parameter to the trained eMBB intelligent agent, output the resource block allocation matrix and the power allocation matrix; input the second input parameter to the trained URLLC intelligent agent, output the perforation matrix, the perforation bandwidth ratio matrix and the perforation power ratio matrix to realize resource allocation.
9. A terminal, characterized in that: including processor and storage medium; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
EMBB and URLLC combined scheduling method and system
CN113691350A
Flow scheduling method for coexistence of eMBB and uRLLC equipment
CN114222371A
EMBB and URLLC service resource allocation method and system based on deep reinforcement learning
CN119110417A
Allocating radio resources using artificial intelligence
US20240223344A1