Intelligent Resource Allocation Method, Device, Terminal and Storage Medium for eMBB / URLLC Coexistence Scenario

In the eMBB/URLLC coexistence scenario, using the intelligent resource allocation method of deep reinforcement learning and Markov process, and collaborative training of eMBB and URLLC agents is solved, the impact of URLLC perforation on eMBB services is achieved, and the throughput and service satisfaction of eMBB services are maximized while satisfying URLLC reliability.

CN120129080BActive Publication Date: 2025-07-11NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510616119.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-11
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

In the eMBB/URLLC coexistence scenario, it is difficult for the prior art to maximize the throughput and service satisfaction of the eMBB service while meeting the URLLC reliability requirements, especially because the impact of URLLC perforation on the eMBB service is difficult to control.

Method used

Deep reinforcement learning method is adopted to establish an eMBB resource allocation model on the slot scale and a URLLC perforation model on the mini slot scale. The Markov process and PF-TSIPRA algorithm are used to realize collaborative training of eMBB agents and URLLC agents, optimize resource allocation strategies, meet the QoS needs of URLLC and maximize service satisfaction of eMBB services.

Benefits of technology

While meeting the QoS needs of URLLC users, it maximizes the service satisfaction of eMBB services, optimizes resource allocation through collaborative training agents, reduces the impact of URLLC perforation on eMBB services, and realizes the optimal resource allocation strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120129080B_ABST
    Figure CN120129080B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent resource allocation method, device, terminal and storage medium for the eMBB / URLLC coexistence scenario in the field of wireless communication technology, aiming to solve the problem that it is difficult to control the throughput and service satisfaction of eMBB to be maximized under the premise of meeting the reliability of URLLC in the prior art. The method includes realizing the collaborative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm according to the Markov process of the eMBB resource allocation model at the time slot scale and the Markov process of the URLLC puncturing model at the mini time slot scale; the present invention constructs the puncturing preference by using the proportional fairness method to maintain the fairness of eMBB users; by continuously interacting with the environment to cope with the random and occasional characteristics of URLLC services, while meeting the QoS requirements of URLLC users, maximizing the service satisfaction of eMBB services, and solving the optimal resource allocation strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an intelligent resource allocation method, device, terminal and storage medium for eMBB / URLLC coexistence scenarios, belonging to the field of wireless communication technology. Background Art

[0002] With the continuous advancement towards 6G wireless communication, the service types and scope of cellular communication are also increasing day by day, such as industrial automation, intelligent transportation, tactile Internet, etc. Among them, enhanced mobile broadband (eMBB) and ultra-reliable low-latency communication (URLLC) are two main types of communication services. Among them, eMBB services such as high-definition video streaming, virtual reality (VR), augmented reality (AR), etc., focus on higher transmission data rates; URLLC services include autonomous driving, industrial Internet of Things, etc. Such services are extremely sensitive to real-time performance and pay more attention to the characteristics of low latency and high reliability.

[0003] In most scenarios, the two services coexist. Since the eMBB service mainly focuses on high throughput in dense application scenarios, while the URLLC service focuses on the strict latency and reliability of critical tasks, the QoS (quality of service) requirements of the two services are very different. Therefore, how to effectively allocate resources for the two coexisting services in limited bandwidth and power resources is a key issue. The puncturing scheme proposed by 3GPP (3rd Generation Partnership Project) gives an answer. By combining the mini-slot technology, the URLLC service will preempt the resources of the eMBB service when it arrives to meet its QoS requirements. However, puncturing will cause the transmission data rate of the eMBB service to be damaged. Therefore, it is necessary to minimize the impact of URLLC puncturing on the eMBB service while meeting the QoS requirements of the URLLC service.

[0004] In actual scenarios, the URLLC service has the characteristics of randomness and occasionality, and it is difficult to obtain an optimal resource allocation scheme through traditional scheduling methods. Considering the channel differences among eMBB users and the impact of URLLC puncturing, simply pursuing the maximization of throughput may cause some eMBB users to fail to reach the target transmission data rate, and the service satisfaction is affected.

[0005] In the prior art, due to the significantly different quality of service requirements of the two services, it is usually difficult to control the throughput and service satisfaction of eMBB to the maximum while meeting the reliability of URLLC, which affects the resource allocation effect when eMBB / URLLC coexists. Summary of the Invention

[0006] The purpose of the present invention is to overcome the deficiencies in the prior art and provide an intelligent resource allocation method, device, terminal, and storage medium for the eMBB / URLLC coexistence scenario. By using the deep reinforcement learning method and continuously interacting with the environment to cope with the random and occasional characteristics of the URLLC service, while meeting the QoS requirements of URLLC users, the service satisfaction of the eMBB service is maximized, and the optimal resource allocation strategy is solved.

[0007] To solve the above technical problems, the present invention is implemented by the following technical solutions:

[0008] In a first aspect, the present invention provides an intelligent resource allocation method for an eMBB / URLLC coexistence scenario, including the following steps:

[0009] Obtain the maximum transmission data rate of eMBB users without being punctured by URLLC.

[0010] Implement the coexistence of eMBB / URLLC by using the puncturing method, and calculate the transmission data rate loss of eMBB users after being punctured by URLLC according to the maximum transmission data rate of eMBB users.

[0011] Calculate the transmission data rate of URLLC users based on the finite block length coding formula and implement the reliability constraint of URLLC.

[0012] According to the maximum transmission data rate of eMBB users, the transmission data rate loss of eMBB users, and the reliability constraint of URLLC, use the proportional fairness algorithm to establish an eMBB resource allocation model on the time slot scale and a URLLC puncturing model on the mini time slot scale respectively.

[0013] Model the eMBB resource allocation model on the time slot scale and the URLLC puncturing model on the mini time slot scale as Markov processes respectively.

[0014] Set an eMBB agent for the Markov process of the eMBB resource allocation model on the time slot scale, and set a URLLC agent for the Markov process of the URLLC puncturing model on the mini time slot scale, and implement the collaborative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm.

[0015] Obtain the first input parameter and the second input parameter, input the first input parameter into the trained eMBB agent, and output the resource block allocation matrix and the power allocation matrix; input the second input parameter into the trained URLLC agent, and output the puncturing matrix, the puncturing bandwidth ratio matrix and the puncturing power ratio matrix to achieve resource allocation.

[0016] Further, the maximum transmission data rate of the eMBB user without URLLC puncturing is obtained, and the specific expression is as follows:

[0017] ;

[0018] In the formula: is the time slot; is the resource block; is the eMBB user; is the maximum transmission data rate that the eMBB user can achieve on the resource block k in the time slot t without puncturing; is the eMBB user downlink transmission power on the resource block k in the time slot t; is the small-scale channel fading factor; is the bandwidth of the resource block k; is the noise power spectral density; is the path loss factor; is the eMBB user distance between and the next-generation base station gNodeB; is the path loss exponent.

[0019] Further, calculating the transmission data rate loss of the eMBB user after being punctured by URLLC according to the maximum transmission data rate of the eMBB user specifically includes:

[0020] Obtain the bandwidth and transmission power obtained when URLLC punctures eMBB, and the specific expressions are as follows:

[0021] ;

[0022] ;

[0023] In the formula: is the URLLC user; the time slot includes mini time slots, is an integer, ,, the set of mini time slots is specifically expressed as ; is the URLLC user In a time slot resource block of the mini - slot the perforated bandwidth; For URLLC users In a time slot resource block of the mini - slot the perforated transmission power; For the time slot resource block of the mini - slot the proportion of the perforated bandwidth, ; For the time slot resource block of the mini - slot the proportion of the perforated power, ;

[0024] According to the bandwidth and transmission power obtained when URLLC perforates eMBB, calculate the data rate of the eMBB user in the th mini - slot of time slot t and resource block k. The specific expression is as follows:

[0025] ;

[0026] In the formula: is the data rate of the eMBB user in the th mini - slot of time slot t and resource block k;

[0027] According to the maximum transmission data rate of the eMBB user without being perforated by URLLC and the data rate of the eMBB user in the time slot resource block of the mini - slot calculate the loss of the transmission data rate of the eMBB user in the time slot resource block of the mini - slot . The specific expression is as follows:

[0028] ;

[0029] In the formula: is the data rate of the eMBB user in the time slot resource block of the mini - slot Transmission data rate loss on

[0030] Furthermore, calculating the transmission data rate of URLLC users based on the finite block length coding formula specifically includes:

[0031] Obtain the time slot Resource block of the mini-slot The puncturing situation, and the specific expression is as follows:

[0032] ;

[0033] In the formula: is the time slot Resource block of the mini-slot The puncturing situation;

[0034] According to the puncturing situation of the time slot Resource block of the mini-slot Based on the finite block length coding formula, calculate the data rate of URLLC users in the time slot at the th mini-slot, and the specific expression is as follows:

[0035] ;

[0036] ;

[0037] ;

[0038] In the formula: is the signal-to-noise ratio of URLLC users ; represents the channel dispersion; is the Gaussian inverse function; is the length of the mini-slot ; is the maximum tolerable transmission packet loss probability of URLLC; , is the set of resource blocks , is the number of resource blocks in the set; is the time slot Resource block of the mini-slot The puncturing situation; is the URLLC user in the time slot at the Data rate of a mini-slot;

[0039] Let represent the URLLC user in the time slot at the th mini-slot, then the specific expression of the transmission delay constraint of the URLLC user is as follows:

[0040] ;

[0041] In the formula: is the tolerable transmission delay of URLLC;

[0042] The reliability constraint for implementing URLLC is specifically expressed as follows:

[0043] ;

[0044] In the formula: is the packet loss rate of URLLC per unit time; represents the URLLC user in the time slot at the th mini-slot not meeting the transmission delay constraint.

[0045] Furthermore, the establishment of the eMBB resource allocation model at the time slot scale and the URLLC puncturing model at the mini-slot scale includes:

[0046] Establishing the eMBB resource allocation model at the time slot scale specifically includes:

[0047] Obtaining the weight of the eMBB user in the time slot The specific expression is as follows:

[0048] ;

[0049] In the formula: is the weight of the eMBB user in the time slot , is the historical average transmission data rate of the eMBB user The specific expression is as follows:

[0050] ;

[0051] In the formula: is the time window length for calculating the historical average transmission data rate; is the eMBB user in the time slot The total transmission data rate; is the eMBB user calculated based on the time slot as a reference historical average transmission data rate;

[0052] Obtain the allocation of the time slot resource block allocation, and according to the allocation of the time slot resource block allocation and the maximum transmission data rate of the eMBB user without being punctured by URLLC, obtain the eMBB user in the time slot total transmission data rate, and the specific expression is as follows:

[0053] ;

[0054] ;

[0055] In the formula: is the allocation of the time slot resource block allocation, is the eMBB user in the time slot total transmission data rate;

[0056] According to the weight of the eMBB user in the time slot weight, time slot resource block allocation and eMBB user in the time slot total transmission data rate, establish an eMBB resource allocation model on the time slot scale, and the specific expression is as follows:

[0057] ;

[0058] ;

[0059] ;

[0060] ;

[0061] ;

[0062] In the formula: is the maximum transmission power of the next-generation base station gNodeB; , is the set of eMBB users e, is the number of eMBB users e in the set; Indicates resource block allocation matrix of order Indicates power allocation matrix of order

[0063] Establish a URLLC puncturing model on the mini-slot scale, specifically including:

[0064] According to the eMBB user after being punctured by URLLC in the time slot resource block in the mini-slot transmission data rate loss on, calculate the total transmission data rate loss of the eMBB user in the mini-slot on, the specific expression is as follows:

[0065] ;

[0066] In the formula: is the total transmission data rate loss of the eMBB user in the mini-slot on;

[0067] According to the weight of the eMBB user in the time slot of, the puncturing situation of the time slot resource block in the mini-slot of and the total transmission data rate loss of the eMBB user in the mini-slot on, establish a URLLC puncturing model on the mini-slot scale, the specific expression is as follows:

[0068] ;

[0069] ;

[0070] ;

[0071] ;

[0072] ;

[0073] ;

[0074] ;

[0075] In the formula: , is the set of URLLC users u, is the number of URLLC users u in the set; denotes the order puncturing matrix of the URLLC user u's puncturing position; denotes order puncturing bandwidth ratio matrix; denotes order puncturing power ratio matrix.

[0076] Furthermore, modeling the eMBB resource allocation model at the time slot scale and the URLLC puncturing model at the mini time slot scale as Markov processes respectively includes:

[0077] Modeling the eMBB resource allocation model at the time slot scale as a Markov process, the expression of the state space is as follows:

[0078] ;

[0079] In the formula: is the signal-to-noise ratio of the eMBB user at the time slot ; is the total transmission data rate of the eMBB user at the time slot ; is the state space of modeling the eMBB resource allocation model at the time slot scale as a Markov process; is used to represent eMBB;

[0080] The specific expression of the action space of modeling the eMBB resource allocation model at the time slot scale as a Markov process is as follows:

[0081] ;

[0082] In the formula: is the action space of modeling the eMBB resource allocation model at the time slot scale as a Markov process;

[0083] Formulate the reward function of modeling the eMBB resource allocation model at the time slot scale as a Markov process, the specific expression is as follows:

[0084] ;

[0085] In the formula: is the reward function of modeling the eMBB resource allocation model at the time slot scale as a Markov process;

[0086] Modeling the URLLC puncturing model at the mini time slot scale as a Markov process, the expression of the state space is as follows:

[0087] ;

[0088] In the formula: is the state space of modeling the URLLC puncturing model as a Markov process on the mini-slot scale; represents the eMBB user on the mini-slot total transmission data rate loss; is the mini-slot signal-to-noise ratio of the URLLC user u within; is used to represent URLLC;

[0089] The specific expression of the action space of modeling the URLLC puncturing model as a Markov process on the mini-slot scale is as follows:

[0090] ;

[0091] In the formula: is the action space of modeling the URLLC puncturing model as a Markov process on the mini-slot scale;

[0092] Formulate the reward function of modeling the URLLC puncturing model as a Markov process on the mini-slot scale, and the specific expression is as follows:

[0093] ;

[0094] In the formula: is the reward function of modeling the URLLC puncturing model as a Markov process on the mini-slot scale; is the penalty term, indicating the total number of URLLC users u that do not meet the transmission delay constraint within the mini-slot ; is the penalty coefficient.

[0095] Furthermore, the collaborative training of the eMBB agent and the URLLC agent is realized based on the PF-TSIPRA algorithm, specifically including:

[0096] The specific training steps of the PF-TSIPRA algorithm are as follows:

[0097] Step 1: Randomly initialize the training network parameters of the eMBB agent , the training network parameters of the URLLC agent and copy them to the corresponding target networks , Initialize the experience replay pool of the eMBB agent and the experience replay pool of the URLLC agent ;

[0098] In the formula: Actor network parameters for the eMBB agent; Critic network parameters for the eMBB agent; Actor network parameters for the URLLC agent; Critic network parameters for the URLLC agent; Target Actor network parameters for the eMBB agent; Target Critic network parameters for the eMBB agent; Target Actor network parameters for the URLLC agent; Target Critic network parameters for the URLLC agent;

[0099] Step 2: At each time slot t, the eMBB agent observes the state , and based on selects the first action ;

[0100] Where: is the exploration noise; is the corresponding Actor network of the eMBB agent;

[0101] Step 3: At each mini time slot n, the URLLC agent observes the state , and based on selects the second action ;

[0102] Where: is the corresponding Actor network of the URLLC agent;

[0103] Step 4: The URLLC agent executes the second action , obtains the second reward and the observation state of the URLLC agent at the next moment ; stores in the URLLC agent's experience replay pool and ends mini time slot n;

[0104] Step 5: The eMBB agent executes the first action , obtains the first reward and the observation state of the eMBB agent at the next moment ; stores in the eMBB agent's experience replay pool and ends time slot t;

[0105] Step 6: From the eMBB agent's experience replay pool and the URLLC agent experience replay pool randomly sample batch data respectively 、 Based on 、 calculate the target Q value;

[0106] In the formula: = 1, 2,..., C; represents the index of the th sample; C represents the number of sampled data; is the reward discount factor; 、 are the evaluations of the target Critic network 、 for the state 、 and the action 、 respectively; 、 are the decisions made by the target Actor network 、 according to the state 、 respectively; ; ; is the target Critic network of the eMBB agent corresponding to ; is the target Critic network of the URLLC agent corresponding to ; is the target Actor network of the eMBB agent corresponding to ; is the target Actor network of the URLLC agent corresponding to ;

[0107] Step 7: Update the Critic network by minimizing the loss function based on 、 ;

[0108] In the formula: 、 are the evaluations of the Critic network 、 for the current state 、 and the action 、 respectively; is the Critic network of the eMBB agent corresponding to ; is The Critic network of the corresponding URLLC agent;

[0109] Step 8: Using the gradient ascent method based on , update the Actor network;

[0110] Where: is the action generated by the Actor network of the eMBB agent according to the current policy; is the action generated by the Actor network of the URLLC agent according to the current policy;

[0111] Step 9: Based on , , , respectively update , , and ; where is the soft update coefficient;

[0112] Step 10: Repeat Steps 2 to 9 until the maximum number of training epochs is reached to complete the training work.

[0113] In a second aspect, the present invention provides an intelligent resource allocation device for an eMBB / URLLC coexistence scenario, and the device includes:

[0114] An acquisition module: used to obtain the maximum transmission data rate of eMBB users without URLLC puncturing;

[0115] A first calculation module: used to achieve the coexistence of eMBB / URLLC by means of puncturing, and calculate the loss of the transmission data rate of eMBB users after being punctured by URLLC according to the maximum transmission data rate of eMBB users;

[0116] A second calculation module: used to calculate the transmission data rate of URLLC users based on the finite block length coding formula and achieve the reliability constraint of URLLC;

[0117] A model establishment module: used to establish an eMBB resource allocation model on the time slot scale and a URLLC puncturing model on the mini-time slot scale respectively using the proportional fairness algorithm according to the maximum transmission data rate of eMBB users, the loss of the transmission data rate of eMBB users, and the reliability constraint of URLLC;

[0118] Conversion module: used to model the eMBB resource allocation model at the time slot scale and the URLLC puncturing model at the mini time slot scale as Markov processes respectively;

[0119] Training module: used to set an eMBB agent for the Markov process of the eMBB resource allocation model at the time slot scale, set a URLLC agent for the Markov process of the URLLC puncturing model at the mini time slot scale, and implement the collaborative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm;

[0120] Output module: used to obtain the first input parameter and the second input parameter, input the first input parameter into the trained eMBB agent, and output the resource block allocation matrix and the power allocation matrix; input the second input parameter into the trained URLLC agent, and output the puncturing matrix, the puncturing bandwidth ratio matrix and the puncturing power ratio matrix to achieve resource allocation.

[0121] In a third aspect, the present invention provides a terminal, including a processor and a storage medium;

[0122] The storage medium is used to store instructions;

[0123] The processor is used to operate according to the instructions to execute the steps of the method according to the first aspect.

[0124] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to the first aspect are implemented.

[0125] Compared with the prior art, the beneficial effects achieved by the present invention:

[0126] Aiming at the eMBB and URLLC resource allocation problems in the downlink scenario, the present invention proposes an intelligent resource allocation method for the eMBB / URLLC coexistence scenario; this solution decomposes the scheduling problem into two sub-problems, allocating bandwidth and power resources for eMBB users at the time slot scale; performing URLLC puncturing at the mini time slot scale to solve a reasonable bandwidth and power preemption scheme; this solution takes into account the channel differences between eMBB users and the impact of URLLC puncturing, constructs a puncturing preference using the proportional fairness method, and maintains fairness among eMBB users; by continuously interacting with the environment to cope with the random and occasional characteristics of URLLC services, while meeting the QoS requirements of URLLC users, maximizing the service satisfaction of eMBB services, and solving the optimal resource allocation strategy. Description of the drawings

[0127] Figure 1It is a schematic flowchart of an intelligent resource allocation method for an eMBB / URLLC coexistence scenario provided according to an embodiment of the present invention;

[0128] Figure 2 It is a perforation schematic diagram of an eMBB / URLLC coexistence scenario provided according to an embodiment of the present invention;

[0129] Figure 3 It is an execution flowchart of PF-TSIPRA provided according to an embodiment of the present invention. Detailed implementation manners

[0130] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.

[0131] The term "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " generally represents an "or" relationship between the associated objects before and after.

[0132] Embodiment 1:

[0133] As Figure 1 shown, the present invention provides an intelligent resource allocation method for an eMBB / URLLC coexistence scenario, including the following steps:

[0134] Step 1: Obtain the maximum transmission data rate of eMBB users without URLLC perforation;

[0135] Step 2: Adopt a perforation method to achieve the coexistence of eMBB / URLLC, and calculate the transmission data rate loss of eMBB users after being perforated by URLLC according to the maximum transmission data rate of eMBB users;

[0136] Step 3: Calculate the transmission data rate of URLLC users based on the finite block length coding formula and achieve the reliability constraint of URLLC;

[0137] Step 4: According to the maximum transmission data rate of eMBB users, the transmission data rate loss of eMBB users, and the reliability constraint of URLLC, use the proportional fairness algorithm to establish an eMBB resource allocation model on the time slot scale and a URLLC perforation model on the mini time slot scale respectively;

[0138] Step 5: Model the eMBB resource allocation model at the time slot scale and the URLLC puncturing model at the mini-time slot scale as Markov processes respectively;

[0139] Step 6: Set up an eMBB agent for the Markov process of the eMBB resource allocation model at the time slot scale and a URLLC agent for the Markov process of the URLLC puncturing model at the mini-time slot scale, and implement the collaborative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm;

[0140] Step 7: Obtain the first input parameter and the second input parameter, input the first input parameter into the trained eMBB agent to output the resource block allocation matrix and the power allocation matrix; input the second input parameter into the trained URLLC agent to output the puncturing matrix, the puncturing bandwidth ratio matrix and the puncturing power ratio matrix, so as to achieve resource allocation.

[0141] Specifically, this solution considers the downlink scenario supported by the next generation node B (gNodeB), adopts orthogonal frequency division multiple access, and its coverage area contains a number of eMBB users and URLLC users, which are represented by the sets , respectively; the available bandwidth is divided into the form of resource blocks, denoted as ; in the time dimension, the time slot is used as the basic unit, denoted as , where the duration of each time slot is 1 ms. On this basis, the concept of mini-time slot is proposed in 5G NR (a global 5G standard based on the brand-new air interface design of OFDM), and each time slot is further divided into several mini-time slots, denoted as ; in order to meet the QoS requirements of high reliability and low latency of URLLC users, 3GPP has proposed the concept of puncturing scheduling. When the URLLC service arrives, it will puncture the eMBB service being transmitted on the mini-time slot to ensure the low latency and reliability of the URLLC service.

[0142] Different from the eMBB service, the URLLC service has higher requirements for latency, so it can be transmitted across resource blocks; in this solution, it is assumed that there is always eMBB data transmission, and the URLLC service can puncture multiple eMBB resource blocks, and the puncturing method is as Figure 2As shown in the figure, the left half of the figure is a simple scenario diagram containing several eMBB users and several URLLC users, and the right half is a URLLC puncturing diagram. It can be seen from the figure that after the arrival of the URLLC service, it will preempt the bandwidth resources of the eMBB service on the mini-slot scale; the gNodeB allocates bandwidth and power resources for the eMBB service at the beginning of each time slot, and the URLLC service perforates the eMBB service on the mini-slot scale after its arrival.

[0143] An embodiment, obtaining the maximum transmission data rate of eMBB users without being punctured by URLLC, specifically includes:

[0144] Obtaining the time slot resource block allocation situation:

[0145] ;

[0146] In the formula: is the time slot resource block allocation situation;

[0147] In the case of not being punctured by the URLLC service, the maximum transmission data rate that eMBB user e can achieve on the time slot resource block is:

[0148] ;

[0149] In the formula: is the time slot; is the resource block; is the eMBB user; is the eMBB user in the case of not being punctured by the URLLC service the maximum transmission data rate that can be achieved on the time slot t resource block k; is the eMBB user downlink transmission power on the time slot t resource block k; is the small-scale channel fading factor; is the bandwidth of resource block k; is the noise power spectral density; is the path loss factor; is the eMBB user distance between and the next-generation base station gNodeB; is the path loss exponent;

[0150] Then the total transmission data rate of eMBB user on the time slot is:

[0151] ;

[0152] Wherein: is an eMBB user in the time slot total transmission data rate.

[0153] An embodiment realizes the coexistence of eMBB / URLLC by means of puncturing, and calculates the transmission data rate loss of the eMBB user after being punctured by URLLC according to the maximum transmission data rate of the eMBB user, specifically including:

[0154] To meet the reliability and low latency requirements of URLLC services, on the mini-slot scale, the URLLC service will puncture the eMBB service resources that are being transmitted, resulting in the loss of eMBB service transmission power and bandwidth resources; obtain the time slot resource block mini-slot puncturing situation, the specific expression is as follows:

[0155] ;

[0156] Wherein: is the time slot resource block mini-slot puncturing situation;

[0157] Obtain the bandwidth and transmission power obtained when URLLC punctures eMBB, the specific expression is as follows:

[0158] ;

[0159] ;

[0160] Wherein: is a URLLC user; the time slot includes mini-slots, is an integer, , the set of mini-slots is specifically expressed as ; is the URLLC user in the time slot resource block mini-slot bandwidth obtained by puncturing; is the URLLC user in the time slot resource block Mini - time slot The transmission power obtained by perforation; For the time slot Resource block Mini - time slot The perforation bandwidth ratio of; ; For the time slot Resource block Mini - time slot The perforation power ratio of; ;

[0161] Assume that each eMBB service may be perforated by multiple URLLC services, and the URLLC service may also perforate the eMBB service across resource blocks. According to the bandwidth and transmission power obtained when the URLLC perforates the eMBB, calculate the eMBB user On the th mini - time slot of resource block k in time slot t, the data rate is as follows and can be approximated as:

[0162] ;

[0163] In the formula: Is the data rate of the eMBB user after being perforated by URLLC On the th mini - time slot of resource block k in time slot t;

[0164] According to the maximum transmission data rate of the eMBB user without being perforated by URLLC and the data rate of the eMBB user In the time slot Resource block Mini - time slot Calculate the loss of the transmission data rate of the eMBB user after being perforated by URLLC In the time slot Resource block Mini - time slot The specific expression is as follows:

[0165] ;

[0166] In the formula: Is the loss of the transmission data rate of the eMBB user after being perforated by URLLC In the time slot Resource block Mini - time slot ;

[0167] According to the eMBB user after being perforated by URLLC In the time slot Resource block Mini-slot of Transmission data rate loss on, calculate eMBB users In the mini-slot Total transmission data rate loss on, the specific expression is as follows:

[0168] ;

[0169] In the formula: Is the eMBB user In the mini-slot Total transmission data rate loss on.

[0170] An embodiment, calculate the URLLC user transmission data rate based on the finite block length coding formula, and implement the reliability constraint of URLLC, specifically including:

[0171] The tolerable transmission delay of the URLLC service is , and its maximum tolerable transmission packet loss probability is ; Since the URLLC packet length is usually short, the Shannon formula based on infinite block length is no longer applicable, so the finite block length channel coding formula is used to approximate the transmission rate of the URLLC service.

[0172] Based on the finite block length coding formula, calculate the URLLC user In the time slot Of the Data rate of the nth mini-slot, the specific expression is as follows:

[0173] ;

[0174] ;

[0175] ;

[0176] In the formula: Is the signal-to-noise ratio of the URLLC user ; Represents the channel dispersion, when the signal-to-noise ratio is high Can be approximated as 1; Is the Gaussian inverse function, ; Is the length of the mini-slot ; , Is the resource block Set of Is the number of resource blocks in the set ; Is the time slot Resource block mini - time - slot puncturing situation; for URLLC users in the time - slot the data rate of the

[0177] Let represent the traffic demand of URLLC users in the time - slot the th mini - time - slot, then the specific expression of the transmission delay constraint for URLLC users is as follows:

[0178] ;

[0179] In the formula: is the tolerable transmission delay of URLLC;

[0180] Meanwhile, the reliability constraint of URLLC services needs to be considered. This constraint is described as the packet loss rate of URLLC services per unit time being less than or equal to the maximum tolerable packet loss rate of URLLC services. To achieve the reliability constraint of URLLC, the specific expression is as follows:

[0181] ;

[0182] In the formula: is the packet loss rate of URLLC per unit time; represents that URLLC users in the time - slot the th mini - time - slot do not meet the transmission delay constraint. The specific calculation method is as follows:

[0183] ;

[0184] Among them, as a binary indicator, represents whether the user in the time - slot the th mini - time - slot meets the transmission delay constraint. Taking 1 means not meeting the transmission delay constraint, otherwise taking 0. The calculation method is as follows:

[0185] .

[0186] An embodiment is to establish an eMBB resource allocation model at the time - slot scale and a URLLC puncturing model at the mini - time - slot scale respectively using the proportional fairness algorithm according to the maximum transmission data rate of eMBB, the transmission data rate loss of eMBB users, and the URLLC reliability constraint. Specifically, it includes:

[0187] At the beginning of each time slot, the gNodeB allocates bandwidth resources and power resources for eMBB services, enabling as many eMBB users as possible to reach the target transmission data rate; when a URLLC service arrives, it preempts the bandwidth and power resources of the eMBB service in the mini-slot dimension, allowing the URLLC service to be served promptly. The optimization problem proposed in this solution includes two sub-problems: eMBB resource allocation and URLLC puncturing, which are respectively established on two time scales: time slots and mini-slots.

[0188] The service satisfaction of eMBB users is an important QoS metric for eMBB services (the service satisfaction of eMBB users is defined as the percentage of eMBB users whose transmission data rate can reach the target transmission data rate), which means that all users need to be considered during resource scheduling to maximize fairness among users as much as possible; since different eMBB users have different channel conditions, simply using the maximization of the total transmission data rate of users as the objective function may cause the gNodeB to ignore those eMBB users with poor channel conditions, and random URLLC puncturing may also cause the transmission data rate of these users to be lower than the target transmission data rate; therefore, a proportional fairness algorithm is adopted for eMBB resource allocation and URLLC puncturing, introducing the historical throughput of eMBB users as a weight to schedule to each user as much as possible while maximizing the total throughput.

[0189] Establish an eMBB resource allocation model on the time slot scale, specifically including:

[0190] Obtain the weight of eMBB user in time slot , and the specific expression is as follows:

[0191] ;

[0192] In the formula: is the weight of eMBB user in time slot , is the historical average transmission data rate of eMBB user , and the specific expression is as follows:

[0193] ;

[0194] In the formula: is the time window length for calculating the historical average transmission data rate; is the total transmission data rate of eMBB user in time slot ; is the time slot eMBB users calculated based on the benchmark Historical average transmission data rate;

[0195] On the time slot scale, the gNodeB allocates bandwidth and frequency resources for eMBB services at the beginning of the time slot. According to the eMBB users In the time slot Weight, time slot Resource block Allocation situation and eMBB users In the time slot Total transmission data rate, establish an eMBB resource allocation model on the time slot scale, and the specific expression is as follows:

[0196] ;

[0197] (1);

[0198] (2);

[0199] (3);

[0200] (4);

[0201] In the formula: Is the maximum transmission power of the next-generation base station gNodeB; , Is the set of eMBB users e, Is the number of eMBB users e in the set; Indicates Order resource block allocation matrix; Indicates Order power allocation matrix; Constraints 1 and 2 mean that the resource block can only be allocated to one user within the time slot; Constraints 3 and 4 mean that the sum of the powers allocated to all resource blocks should be lower than the maximum transmission power of the base station; The goal of the optimization problem P1 is to find the optimal solutions for resource block allocation and power allocation, maximizing the eMBB user throughput while taking into account the fairness of eMBB users;

[0202] Establish a URLLC puncturing model on the mini time slot scale, specifically including:

[0203] The URLLC service is scheduled on the mini time slot scale. By allocating reasonable puncturing bandwidth ratios and power ratios for the URLLC service to meet the delay and reliability requirements of the URLLC service, while minimizing the transmission data rate loss caused by puncturing to the eMBB service. According to the eMBB users In the time slot Weight, time slot Resource block mini-slot of puncturing situation and eMBB users of in the mini-slot total transmission data rate loss on, establish a URLLC puncturing model on the mini-slot scale, and the specific expression is as follows:

[0204] ;

[0205] (5);

[0206] (6);

[0207] (7);

[0208] (8);

[0209] (9);

[0210] (10);

[0211] In the formula: , is the set of URLLC users u, is the number of URLLC users u in the set; represents the -order puncturing matrix of the puncturing position of URLLC user u; represents -order puncturing bandwidth ratio matrix; represents -order puncturing power ratio matrix; Constraint 5 ensures the low-latency and reliability requirements of URLLC; Constraint 6 is a binary puncturing indicator; Constraints 7 and 9 represent the bandwidth resources and power resources allocated to URLLC services; Constraints 8 and 10 represent that the bandwidth resources and power resources allocated to URLLC services do not exceed the bandwidth resources and power resources of the current resource block eMBB service.

[0212] An embodiment models the eMBB resource allocation model on the time-slot scale and the URLLC puncturing model on the mini-slot scale as Markov processes respectively, specifically including:

[0213] In the previous embodiment, the optimization problem P1 needs to make resource block allocation and power allocation decisions, which include both discrete variables and continuous variables, and its objective function is non-linear; the puncturing model of the optimization problem P2 also includes integer decisions and discrete solution variables, and its objective function is also non-linear; therefore, these two problems are mixed integer non-linear programming problems. The introduction of integer decision variables will expand the solution space and make it complicated to find the optimal solution; at the same time, considering the randomness of URLLC services, the two sub-problems are respectively modeled as Markov processes (Markov decision process, MDP), and reinforcement learning is used to handle the uncertainty of the environment and find the optimal solution.

[0214] Among them, MDP usually includes a state space , an action space and a reward function . When in state , the agent executes action , obtains reward and the state at the next moment.

[0215] MDP parameters for eMBB services: The state space of the eMBB agent is as follows:

[0216] ;

[0217] In the formula: is the signal-to-noise ratio of eMBB user in time slot ; is the total transmission data rate of eMBB user in time slot ; is the state space for modeling the eMBB resource allocation model as a Markov process on the time slot scale; is used to represent eMBB;

[0218] The specific expression of the action space for modeling the eMBB resource allocation model as a Markov process on the time slot scale is as follows:

[0219] ;

[0220] In the formula: is the action space for modeling the eMBB resource allocation model as a Markov process on the time slot scale;

[0221] Based on the optimization objective of P1, the reward function for modeling the eMBB resource allocation model as a Markov process on the time slot scale is formulated, and the specific expression is as follows:

[0222] ;

[0223] In the formula: is the reward function for modeling the eMBB resource allocation model as a Markov process on the time slot scale.

[0224] MDP parameters of URLLC services: The state space of the URLLC agent is as follows:

[0225] ;

[0226] In the formula: is the state space for modeling the URLLC puncturing model as a Markov process on the mini - time - slot scale; represents the eMBB user in the mini - time - slot on the total transmission data rate loss; is the mini - time - slot the signal - to - noise ratio of the URLLC user u within; is used to represent URLLC;

[0227] The specific expression of the action space for modeling the URLLC puncturing model as a Markov process on the mini - time - slot scale is as follows:

[0228] ;

[0229] In the formula: is the action space for modeling the URLLC puncturing model as a Markov process on the mini - time - slot scale;

[0230] Based on the goal of minimizing the transmission data rate loss of eMBB users, the reward function for modeling the URLLC puncturing model as a Markov process on the mini - time - slot scale is formulated, and the specific expression is as follows:

[0231] ;

[0232] In the formula: is the reward function for modeling the URLLC puncturing model as a Markov process on the mini - time - slot scale; is the penalty term, representing the total number of URLLC users u that do not meet the transmission delay constraint within the mini - time - slot ; is the penalty coefficient.

[0233] As Figure 3 shown, in an embodiment, the coordinated training of the eMBB agent and the URLLC agent is implemented based on the PF - TSIPRA algorithm, specifically including:

[0234] The specific training steps of the PF - TSIPRA algorithm are as follows:

[0235] Step 1: Randomly initialize the training network parameters of the eMBB agent and the URLLC agent, and copy them to the corresponding target networks ; initialize the experience replay pool of the eMBB agent and the experience replay pool of the URLLC agent ;

[0236] Wherein: is the Actor network parameter of the eMBB agent; is the Critic network parameter of the eMBB agent; is the Actor network parameter of the URLLC agent; is the Critic network parameter of the URLLC agent; is the target Actor network parameter of the eMBB agent; is the target Critic network parameter of the eMBB agent; is the target Actor network parameter of the URLLC agent; is the target Critic network parameter of the URLLC agent;

[0237] Step 2: At each time slot t, the eMBB agent observes the state and selects the first action based on ;

[0238] Wherein: is the exploration noise; is the Actor network of the corresponding eMBB agent;

[0239] Step 3: At each mini time slot n, the URLLC agent observes the state and selects the second action based on ;

[0240] Wherein: is the Actor network of the corresponding URLLC agent;

[0241] Step 4: The URLLC agent executes the second action , obtains the second reward and the observation state of the URLLC agent at the next moment ; store in the experience replay pool of the URLLC agent ​End the mini-slot n;

[0242] Step 5: The eMBB agent executes the first action to obtain the first reward and the eMBB agent's observed state at the next moment ; Store in the eMBB agent's experience replay pool to end the time slot t;

[0243] Step 6: Randomly sample batch data from the eMBB agent's experience replay pool and the URLLC agent's experience replay pool respectively 、 , and calculate the target Q value based on 、 ;

[0244] In the formula: = 1, 2,..., C; represents the index of the th sample; C represents the number of sampled data; is the reward discount factor; 、 are the evaluations of the target Critic network 、 for the states 、 and the actions 、 respectively; 、 are the decisions made by the target Actor network 、 according to the states 、 respectively; ; ; is the target Critic network of the eMBB agent corresponding to ; is the target Critic network of the URLLC agent corresponding to ; is the target Actor network of the eMBB agent corresponding to ; is the target Actor network of the URLLC agent corresponding to ;

[0245] Step 7: Update the Critic network by minimizing the loss function based on 、 ;

[0246] In the formula: and are the evaluations of the Critic network and for the current state and the actions and respectively; is the Critic network of the corresponding eMBB agent; is the Critic network of the corresponding URLLC agent;

[0247] Step 8: Use the gradient ascent method to update the Actor network based on and ;

[0248] In the formula: is the action generated by the Actor network of the eMBB agent according to the current policy; is the action generated by the Actor network of the URLLC agent according to the current policy;

[0249] Step 9: Based on and and and update and and and respectively; where is the soft update coefficient;

[0250] Step 10: Repeat Steps 2 to 9 until the maximum number of training cycles is reached to complete the training work.

[0251] Specifically, this solution proposes a two-stage intelligent puncture resource allocation algorithm based on proportional fairness (PF-TSIPRA), which separately sets the eMBB agent and the URLLC agent to solve the optimization problems P1 and P2; the eMBB agent executes on the time slot scale, makes decisions at the beginning of each time slot, and at the same time, the URLLC agent executes decisions cyclically on the mini time slots within each time slot, as specifically shown in Figure 3 as follows.

[0252] The eMBB agent includes an Actor network , a Critic network , a target Actor network , and a target Critic network , and the corresponding network parameters are respectively 、 、 and ; The URLLC agent includes an Actor network , a Critic network , a target Actor network , and a target Critic network , and the corresponding network parameters are respectively 、 、 and ; and are used to distinguish eMBB and URLLC;

[0253] When in state 、 , the Actor networks 、 generate actions 、 respectively according to the current policy, and add exploration noise , and the specific expression is as follows:

[0254] ;

[0255] ;

[0256] After the eMBB agent executes the first action , it obtains the first reward , generates the state at the next moment , and stores in the eMBB agent experience replay pool ; The URLLC agent executes the second action , obtains the second reward and the observed state of the URLLC agent at the next moment ; Stores in the URLLC agent experience replay pool ; Subsequently, the eMBB agent and the URLLC agent respectively randomly sample several groups of sample data from the eMBB agent experience replay pool and the URLLC agent experience replay pool , denoted as 、 , which is used to update the respective Actor network and Critic network, and calculates the target Q value of the eMBB agent's Critic network using the Bellman equation and the target Q value of the URLLC agent's Critic network , and the specific expressions are as follows:

[0257] ;

[0258] ;

[0259] In the formula: = 1, 2,..., C; represents the index of the th sample; C represents the number of sampled data; is the reward discount factor; , are the evaluations of the target Critic network , for the states , and the actions , respectively; , are the decisions made by the target Actor network and according to the states , respectively; among them, ; .

[0260] The Critic network parameters , are updated by minimizing the loss function:

[0261] ;

[0262] ;

[0263] In the formula: is the learning rate of the Critic network; is the corresponding Critic network loss function; is the corresponding Critic network loss function, and the calculation is as follows:

[0264] ;

[0265] ;

[0266] Among them, , are the Critic networks respectively , for the current state , the actions , under it, and the evaluation;

[0267] Using the gradient ascent method, update the Actor network parameters by maximizing the Q value of the Critic network:

[0268] ;

[0269] ;

[0270] In the formula: is the learning rate of the Actor network; and are calculated as follows:

[0271] ;

[0272] ;

[0273] In the formula: is the action generated by the Actor network of the eMBB agent according to the current policy; is the action generated by the Actor network of the URLLC agent according to the current policy;

[0274] Finally, use the soft update method to update the target network parameters , , and :

[0275] ;

[0276] ;

[0277] ;

[0278] ;

[0279] In the formula: is the soft update coefficient;

[0280] Repeat the above steps until the maximum number of training cycles is reached to complete the training work.

[0281] After training, in each time slot, the eMBB agent outputs the resource block allocation matrix according to the input parameters ​ and a power allocation matrix ; in each mini-slot, the URLLC agent outputs a puncturing matrix , a puncturing bandwidth ratio matrix , and a puncturing power ratio matrix according to the input parameters .

[0282] The present invention aims at the eMBB and URLLC resource allocation problems in the downlink scenario, and proposes an intelligent resource allocation method for the eMBB / URLLC coexistence scenario; this solution decomposes the scheduling problem into two sub-problems, allocates bandwidth and power resources for eMBB users at the time slot scale; performs URLLC puncturing at the mini-slot scale to solve a reasonable bandwidth and power preemption scheme; this solution takes into account the channel differences among eMBB users and the impact of URLLC puncturing, constructs puncturing preferences using the proportional fairness method to maintain fairness among eMBB users; continuously interacts with the environment to cope with the random and occasional characteristics of URLLC services, maximizes the service satisfaction of eMBB services while meeting the QoS requirements of URLLC users, and solves the optimal resource allocation strategy.

[0283] Embodiment 2:

[0284] The present invention provides an intelligent resource allocation device for the eMBB / URLLC coexistence scenario, and the device includes:

[0285] An acquisition module: used to obtain the maximum transmission data rate of eMBB users without URLLC puncturing;

[0286] A first calculation module: used to achieve the coexistence of eMBB / URLLC in a puncturing manner, and calculate the loss of the transmission data rate of eMBB users after being punctured by URLLC according to the maximum transmission data rate of eMBB users;

[0287] A second calculation module: used to calculate the transmission data rate of URLLC users based on the finite block length coding formula and achieve the reliability constraint of URLLC;

[0288] A model establishment module: used to establish an eMBB resource allocation model at the time slot scale and a URLLC puncturing model at the mini-slot scale respectively using the proportional fairness algorithm according to the maximum transmission data rate of eMBB users, the loss of the transmission data rate of eMBB users, and the reliability constraint of URLLC;

[0289] A conversion module: used to model the eMBB resource allocation model at the time slot scale and the URLLC puncturing model at the mini-slot scale as Markov processes respectively;

[0290] Training module: used to set up the eMBB agent for the Markov process of the eMBB resource allocation model at the time slot scale, set up the URLLC agent for the Markov process of the URLLC puncturing model at the mini time slot scale, and implement the collaborative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm;

[0291] Output module: used to obtain the first input parameter and the second input parameter, input the first input parameter into the trained eMBB agent, and output the resource block allocation matrix and the power allocation matrix; input the second input parameter into the trained URLLC agent, and output the puncturing matrix, the puncturing bandwidth ratio matrix, and the puncturing power ratio matrix to achieve resource allocation.

[0292] Embodiment 3:

[0293] The embodiment of the present invention also provides a terminal, including a processor and a storage medium;

[0294] The storage medium is used to store instructions;

[0295] The processor is used to operate according to the instructions to execute the steps of the method described in Embodiment 1.

[0296] Embodiment 4:

[0297] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method described in Embodiment 1.

[0298] Since the storage medium provided by the embodiment of the present invention can execute the method provided by the first embodiment of the present invention, therefore, it has the corresponding functional modules and beneficial effects for executing the method.

[0299] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0300] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0301] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0302] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0303] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. An intelligent resource allocation method for eMBB / URLLC coexistence scenarios, characterized in that, Including the following steps: Obtain the maximum transmission data rate of eMBB users without URLLC puncturing; Adopt a puncturing method to achieve the coexistence of eMBB / URLLC, and calculate the transmission data rate loss of eMBB users after being punctured by URLLC according to the maximum transmission data rate of eMBB users; Calculate the transmission data rate of URLLC users based on the finite block length coding formula and achieve the reliability constraint of URLLC; According to the maximum transmission data rate of eMBB users, the transmission data rate loss of eMBB users, and the reliability constraint of URLLC, use the proportional fairness algorithm to establish an eMBB resource allocation model on the time slot scale and a URLLC puncturing model on the mini-time slot scale respectively; Model the eMBB resource allocation model on the time slot scale and the URLLC puncturing model on the mini-time slot scale as Markov processes respectively; Set an eMBB agent for the Markov process of the eMBB resource allocation model on the time slot scale, and set a URLLC agent for the Markov process of the URLLC puncturing model on the mini-time slot scale, and realize the cooperative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm; Obtain the first input parameter and the second input parameter, input the first input parameter into the trained eMBB agent, and output the resource block allocation matrix and the power allocation matrix; Input the second input parameter into the trained URLLC agent, output the puncturing matrix, the puncturing bandwidth ratio matrix, and the puncturing power ratio matrix, and realize resource allocation; Among them, the realization of the cooperative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm includes: Step a: Randomly initialize the training network parameters of the eMBB agent and the URLLC agent and copy them to the corresponding target networks, and initialize the experience replay pools of the eMBB agent and the URLLC agent; Step b: At each time slot, the eMBB agent observes the state and generates a first action through the Actor network; at each mini-time slot, the URLLC agent observes the state and generates a second action through the Actor network. The URLLC agent executes the second action, obtains the second reward and the observation state of the URLLC agent at the next moment, and stores the corresponding experience data in the URLLC agent experience replay pool; the eMBB agent executes the first action, obtains the first reward and the observation state of the eMBB agent at the next moment, and stores the corresponding experience data in the eMBB agent experience replay pool; the eMBB agent and the URLLC agent respectively randomly sample batch data from their own experience replay pools for updating their respective target networks; Step c: Repeat step b until the maximum number of training cycles is reached.

2. The intelligent resource allocation method for the eMBB / URLLC coexistence scenario according to claim 1, wherein, The obtaining of the maximum transmission data rate of eMBB users without URLLC puncturing is specifically expressed as follows: ; Wherein: is a time slot; is a resource block; is an eMBB user; In the case of no puncturing, the maximum transmission data rate that an eMBB user can achieve on resource block k in time slot t; is the downlink transmission power of an eMBB user on resource block k in time slot t; is the small-scale channel fading factor; is the bandwidth of resource block k; is the noise power spectral density; is the path loss factor; is the distance between an eMBB user and the next-generation base station gNodeB; is the path loss exponent.

3. The intelligent resource allocation method for the eMBB / URLLC coexistence scenario according to claim 2, characterized in that, The calculation of the transmission data rate loss of eMBB users after being punctured by URLLC according to the maximum transmission data rate of eMBB users specifically includes: Obtain the bandwidth and transmission power obtained when URLLC perforates eMBB, and the specific expressions are as follows: ; ; Wherein: is a URLLC user; the time slot includes mini time slots, is an integer, , the set of mini time slots is ; specifically expressed as is the URLLC user in the time slot resource block the bandwidth obtained by perforating the mini time slot ; is the URLLC user in the time slot resource block the transmission power obtained by perforating the mini time slot ; is the ratio of the perforated bandwidth of the mini time slot resource block in the time slot , ; is the ratio of the perforated power of the mini time slot resource block in the time slot , ;​​​​​​​​​​​​​​​​​ Calculate the data rate of the eMBB user after being punctured by URLLC according to the bandwidth and transmission power obtained when eMBB is punctured by URLLC At the th mini-slot of resource block k in time slot t, the specific expression is as follows: ; In the formula: is the data rate of the eMBB user after being perforated by URLLC in the th mini-slot of resource block k in time slot t; According to the maximum transmission data rate of eMBB users without URLLC puncturing and the eMBB users after being punctured by URLLC In the time slot Resource block of the mini-slot The data rate on, calculate the eMBB users after being punctured by URLLC In the time slot Resource block of the mini-slot The loss of the transmission data rate on is as follows: ; Wherein: is the eMBB user after being perforated by URLLC in the time slot resource block of the mini-slot transmission data rate loss on.

4. The intelligent resource allocation method for the eMBB / URLLC coexistence scenario according to claim 3, wherein Calculate the URLLC user transmission data rate based on the finite block length coding formula, specifically including: Obtain time slot Resource block Mini time slot Puncturing situation, and the specific expression is as follows: ; In the formula: is a time slot resource block mini time slot puncturing situation; According to time slots Resource block Mini-slot Based on the puncturing situation of the mini-slot and the finite blocklength coding formula, calculate the data rate of the URLLC user In time slots The th mini-slot, and the specific expression is as follows: ; ; ; Wherein: is the signal-to-noise ratio of the URLLC user ; represents the channel dispersion; is the Gaussian inverse function; is the mini-slot length; is the maximum tolerable transmission packet loss probability of URLLC; , is the set of resource blocks ; is the number of resource blocks in the set; is the time slot resource block mini-slot puncturing situation; is the URLLC user in the time slot at the th mini-slot data rate; Let represent the URLLC user in the time slot at the th mini - time - slot's transmission traffic demand. Then the specific expression of the transmission delay constraint for the URLLC user is as follows: ; Wherein: is the tolerable transmission delay of URLLC; Implement the reliability constraint of URLLC, and the specific expression is as follows: ; Wherein: is the packet loss rate of URLLC per unit time; represents a URLLC user in the time slot at the th mini-slot fails to meet the transmission delay constraint.

5. The intelligent resource allocation method for the eMBB / URLLC coexistence scenario according to claim 4, wherein, Respectively establish an eMBB resource allocation model at the time slot scale and a URLLC perforation model at the mini-time slot scale, including: Establish an eMBB resource allocation model at the time slot scale, specifically including: Obtain eMBB users In the time slot The weight is as follows: ; Wherein: is the eMBB user in the time slot weight, is the eMBB user historical average transmission data rate, and the specific expression is as follows: ; Wherein: is the time window length for calculating the historical average transmission data rate; is the eMBB user in the time slot total transmission data rate; is based on the time slot is the historical average transmission data rate of the eMBB user calculated based on ; Obtain time slot Resource block 's allocation situation, according to the time slot Resource block 's allocation situation and the maximum transmission data rate of eMBB users without being perforated by URLLC, obtain the total transmission data rate of eMBB users in the time slot as follows: ; ; Wherein: is the time slot resource block allocation situation of is the eMBB user in the time slot total transmission data rate of; According to eMBB users In the time slot weight, time slot resource block allocation situation and eMBB users In the time slot total transmission data rate, establish an eMBB resource allocation model on the time slot scale, and the specific expression is as follows: ; ; ; ; ; Wherein: is the maximum transmission power of the next-generation base station gNodeB; , is the set of eMBB users e, is the number of eMBB users e in the set; represents an RB allocation matrix of order represents a power allocation matrix of order Establish a URLLC perforation model at the mini-time slot scale, specifically including: According to the eMBB user after being perforated by URLLC In the time slot Resource block Of the mini-slot Calculate the transmission data rate loss of the eMBB user In the mini-slot The total transmission data rate loss on is as follows: ; Wherein: is the eMBB user in the mini-slot total transmission data rate loss; According to eMBB users In the time slot Weight, time slot Resource block Mini-slot Puncturing situation and eMBB users On the mini-slot Based on the total transmission data rate loss of eMBB users on the mini-slot, a URLLC puncturing model on the mini-slot scale is established, and the specific expression is as follows: ; ; ; ; ; ; ; In the formula: , is the set of URLLC users u, is the number of URLLC users u in the set; represents the -order puncturing matrix of the puncturing position of URLLC user u; represents -order puncturing bandwidth ratio matrix; represents -order puncturing power ratio matrix.

6. The intelligent resource allocation method for the eMBB / URLLC coexistence scenario according to claim 5, characterized in that, Model the eMBB resource allocation model at the time slot scale and the URLLC perforation model at the mini-time slot scale as Markov processes respectively, including: Model the eMBB resource allocation model at the time slot scale as a Markov process, and the expression of the state space is as follows: ; Wherein: is the eMBB user in the time slot signal-to-noise ratio; is the eMBB user in the time slot total transmission data rate; is the state space of the eMBB resource allocation model modeled as a Markov process on the time slot scale; used to represent eMBB; The specific expression of the action space of the eMBB resource allocation model at the time slot scale modeled as a Markov process is as follows: ; In the formula: is the action space for modeling the eMBB resource allocation model as a Markov process on the time slot scale; Formulate the reward function of the eMBB resource allocation model at the time slot scale modeled as a Markov process, and the specific expression is as follows: ; In the formula: is the reward function for modeling the eMBB resource allocation model as a Markov process on the time slot scale; Model the URLLC perforation model at the mini-time slot scale as a Markov process, and the expression of the state space is as follows: ; Wherein: The state space for modeling the URLLC puncturing model on the mini-slot scale as a Markov process; Denotes an eMBB user On the mini-slot Total transmission data rate loss; Is the mini-slot Signal-to-noise ratio of the URLLC user u within; Used to represent URLLC; The specific expression of the action space of the URLLC perforation model at the mini-time slot scale modeled as a Markov process is as follows: ; Wherein: is the action space for modeling the URLLC puncturing model on the mini-slot scale as a Markov process; Formulate the reward function of the URLLC perforation model at the mini-time slot scale modeled as a Markov process, and the specific expression is as follows: ; Wherein: is the reward function for modeling the URLLC puncturing model on the mini-slot scale as a Markov process; is the penalty term, indicating the total number of URLLC users u that do not satisfy the transmission delay constraint within the mini-slot ; is the penalty coefficient.

7. The intelligent resource allocation method for the eMBB / URLLC coexistence scenario according to claim 6, wherein, Implement the collaborative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm, specifically including: The specific training steps of the PF-TSIPRA algorithm are as follows: Step 1: Randomly initialize the training network parameters of the eMBB agent and the training network parameters of the URLLC agent and copy them to the corresponding target networks and Initialize the experience replay pool of the eMBB agent and the experience replay pool of the URLLC agent ; Wherein: are the Actor network parameters of the eMBB agent; are the Critic network parameters of the eMBB agent; are the Actor network parameters of the URLLC agent; are the Critic network parameters of the URLLC agent; are the target Actor network parameters of the eMBB agent; are the target Critic network parameters of the eMBB agent; are the target Actor network parameters of the URLLC agent; are the target Critic network parameters of the URLLC agent; Step 2: At each time slot t, the eMBB agent observes the state , and based on selects the first action ; In the formula: is the exploration noise; is the Actor network of the corresponding eMBB agent; Step 3: At each mini-slot n, the URLLC agent observes the state , and based on selects the second action ; In the formula: is the Actor network of the corresponding URLLC agent; Step 4: The URLLC agent executes the second action , and obtains the second reward and the URLLC agent observation state at the next moment ; Store the experience data into the URLLC agent experience replay pool , and end the mini-slot n; Step 5: The eMBB agent executes the first action , obtains the first reward and the observation state of the eMBB agent at the next moment ; stores the experience data in the eMBB agent's experience replay pool , and ends time slot t; Step 6: Randomly sample a batch of data from the eMBB agent experience replay pool and the URLLC agent experience replay pool respectively, and calculate the target Q value based on , , , ; where: = 1, 2, …, C; represents the index of the th sample; C represents the number of sampled data; is the reward discount factor; and are the evaluations of the target Critic network and for the states and and the actions and respectively; and are the decisions made by the target Actor network and according to the states and respectively; ; ; is the target Critic network of the eMBB agent corresponding to ; is the target Critic network of the URLLC agent corresponding to ; is the target Actor network of the eMBB agent corresponding to ; is the target Actor network of the URLLC agent corresponding to ; Step 7: Based on , minimize the loss function to update the Critic network; Where: , are the evaluations of the current state , by the Critic networks , for the actions , respectively; is the Critic network of the eMBB agent corresponding to ; is the Critic network of the URLLC agent corresponding to ; Step 8: Use the gradient ascent method based on 、 Update the Actor network; In the formula: is the action generated by the Actor network of the eMBB agent according to the current policy; is the action generated by the Actor network of the URLLC agent according to the current policy; Step 9: Based on , , , update , , and respectively; where is the soft update coefficient. Step 10: Repeat steps 2 to 9 until the maximum number of training cycles is reached to complete the training work.

8. An intelligent resource allocation device for eMBB / URLLC coexistence scenarios, characterized in that, The device includes: Acquisition module: used to obtain the maximum transmission data rate of eMBB users when not perforated by URLLC; First calculation module: used to implement the coexistence of eMBB / URLLC in a perforation manner, and calculate the loss of the transmission data rate of eMBB users after being perforated by URLLC according to the maximum transmission data rate of eMBB users; Second calculation module: used to calculate the URLLC user transmission data rate based on the finite block length coding formula and implement the reliability constraint of URLLC; Model establishment module: used to respectively establish an eMBB resource allocation model at the time slot scale and a URLLC perforation model at the mini-time slot scale using the proportional fairness algorithm according to the maximum transmission data rate of eMBB users, the loss of the transmission data rate of eMBB users, and the reliability constraint of URLLC; Conversion module: used to model the eMBB resource allocation model at the time slot scale and the URLLC perforation model at the mini-time slot scale as Markov processes respectively; Training module: used to set up an eMBB agent for the Markov process of the eMBB resource allocation model at the time slot scale, set up a URLLC agent for the Markov process of the URLLC puncturing model at the mini time slot scale, and implement the collaborative training of the eMBB agent and the URLLC agent based on the PF-TSIPRA algorithm; Output module: used to obtain the first input parameter and the second input parameter, input the first input parameter into the trained eMBB agent, and output the resource block allocation matrix and the power allocation matrix; input the second input parameter into the trained URLLC agent, and output the puncturing matrix, the puncturing bandwidth ratio matrix and the puncturing power ratio matrix to achieve resource allocation.

9. A terminal, characterized in that, It includes a processor and a storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it realizes the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • EMBB and URLLC combined scheduling method and system

    CN113691350A

  • EMBB and URLLC service resource allocation method and system based on deep reinforcement learning

    CN119110417A