Multiplexing Method for eMBB and URLLC in Wireless Power Supply Communication Network
By dividing time slots and micro-time slots in a 5G wireless network, combining preemptive puncture technology and a hybrid depth deterministic strategy gradient method, the allocation of subcarriers, time and energy resources is optimized, and the problem of multiplexing of eMBB and URLLC services in wireless energy-supply communication networks is solved, achieving efficient service multiplexing and higher total eMBB speed.
Patent Information
- Application Number
- CN202211373715.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-03
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-11-03
AI Technical Summary
In 5G wireless networks, how to efficiently multiplex eMBB and URLLC services so that they share the same resources in the same network, especially in wireless powered communication networks, considering constraints such as user battery capacity and RF/DC circuit sensitivity.
By dividing time slots and micro-time slots in the wireless energy-supply communication network, and retrieving and redistributing subcarriers at the beginning of each time slot, combined with preemptive puncture technology, URLLC traffic preempts resources for eMBB transmission, and uses a hybrid depth deterministic strategy gradient method to optimize the allocation of subcarriers, time and energy resources to maximize the total eMBB rate.
It realizes efficient multiplexing of eMBB and URLLC services in WPCN, and has a higher expected implementation rate than static or semi-static spectrum resource allocation, which improves the total eMBB rate and ensures the stability of uplink transmission.
Smart Images

Figure CN115915206B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communications, and particularly relates to a multiplexing method for eMBB and URLLC in a wireless power supply communication network. Background Art
[0002] In order to further improve our cities, living environments and industries, the next-generation wireless network needs to support services and applications with various different requirements. To achieve this goal, the fifth-generation wireless network (5G) is expected to support three major usage scenarios, namely eMBB (enhanced mobile broadband), URLLC (ultra-reliable low-latency communications), and mMTC (massive machine-type communications). Among them, eMBB services require higher data rates to further improve current mobile services such as high-definition video and virtual reality; URLLC services focus on supporting low-latency transmission of small data packets with high reliability, which covers applications such as autonomous vehicles, industrial automation, and vehicle communications; mMTC supports services that connect a large number of devices, and each device intermittently transmits small data packets, covering applications such as smart cities. Among the above three services, the two main services that the 5G wireless network will support are eMBB and URLLC. However, how to accommodate these wireless services with different requirements in the same network so that they can share the same resources has become a challenging problem, and dynamic multiplexing technology has become the key to solving this problem.
[0003] Currently, efficiently multiplexing eMBB and URLLC on shared channels has become one of the main goals of 5G wireless networks. In existing research, static and semi-static resource partitioning that allocates resources for different services in advance has been proven to be spectrally inefficient. According to the concept of allocating resources on demand for URLLC communications, the 3GPP standard recommends using superposition / preemption techniques to multiplex URLLC and eMBB services in 5G networks. Since time is divided into time slots (slots), and each slot consists of several mini-slots (minislots), the main idea of the superposition / puncturing framework is to transmit URLLC packets on a mini-slot basis over the resources occupied by the ongoing transmission of the service type when the URLLC packets arrive. Specifically, eMBB communications share the time-frequency resources within each mini-slot, and these resources can be allocated based on the channel state information of each user at the beginning of each slot. To accommodate URLLC traffic under strict latency constraints, the arriving URLLC packets are transmitted by preempting the resources of the ongoing eMBB transmission and scheduling them into the next minislot. If the base station (BS) allocates transmission power for both eMBB and URLLC traffic simultaneously, it is called superposition. If the BS selects zero transmission power for eMBB traffic, then this is called puncturing. On this basis, some existing research has studied the joint optimization of eMBB and URLLC traffic scheduling by adopting the method of preemptive puncturing, maximizing the eMBB rate under the strict latency requirements of URLLC, where URLLC preempts the radio resources allocated for eMBB transmission.
[0004] In addition, with the development of communication technologies, more and more devices are deployed in places with limited energy acquisition. To enable a large number of intelligent Internet of Things devices to operate for a long time in harsh environments and hard-to-reach locations, wireless-powered communication networks (WPCN) technology is an attractive solution. WPCN is used to solve the problem of energy scarcity in wireless networks, where wireless nodes are usually powered by batteries that are either non-replaceable or require very high replacement costs, which limits the lifespan of the nodes. While WPCN can enable wireless nodes to perform active transmissions by consuming the energy obtained from dedicated radio frequency sources.
[0005] In the research of WPCN, most of the existing studies assume that the battery capacity of users is infinite and ignore the received power threshold of the energy harvesting circuit at the user side, that is, the sensitivity of RF / DC (radio frequency to direct current). In addition, most of the existing work is to allocate the energy in all batteries for uplink information transmission. However, due to channel changes, it is impossible to ensure that enough energy can be harvested in each time slot, so it may lead to insufficient energy for uplink information transmission. Moreover, there is currently a lack of a scheme in WPCN that simultaneously considers the dynamic multiplexing of eMBB and URLLC services. Summary of the Invention
[0006] To solve the above problems existing in the prior art, an embodiment of the present invention provides a multiplexing method for eMBB and URLLC in a wireless powered communication network. The specific technical solution is as follows:
[0007] A multiplexing method for eMBB and URLLC in a wireless powered communication network, which is applied to a WPCN under preset scenario conditions; the WPCN under the preset scenario conditions includes a hybrid access point HAP and multiple users; each user is equipped with a rechargeable battery and has eMBB and URLLC service requirements; the time domain is divided into multiple time slots of equal duration, and each time slot is divided into multiple micro-slots of equal duration; subcarriers are retrieved and re-allocated to users at the beginning of each time slot for eMBB transmission; each arriving URLLC packet of a user preempts the subcarriers allocated to the user in the next micro-slot for URLLC transmission; the user adopts a harvest-then-transmit protocol in each time slot; the method includes:
[0008] Under the constraints of considering URLLC delay, the sensitivity of the RF / DC circuit, the user's battery capacity, and subcarrier availability, with the goal of maximizing the total eMBB rate of the uplink in the WPCN, an optimization problem is constructed to solve the optimal allocation scheme of subcarriers, time, and energy resources for all users;
[0009] Use the preset hybrid deep deterministic policy gradient method to solve the optimization problem to obtain the optimal solution.
[0010] In an embodiment of the present invention, the construction process of the optimization problem includes:
[0011] Determine the expression of the received power of the energy signal sent by the user to the HAP within the time slot t when the user performs transmission;
[0012] According to the expression of the received power, the comparison relationship between the received power and the sensitivity of the RF / DC circuit, the time allocation ratio of the downlink WET phase in time slot t, and the duration of time slot t, an expression for the received energy of the user within time slot t is obtained; wherein, time slot t is divided into a downlink WET phase and an uplink WIT phase, and the uplink WIT phase corresponds to the WPCN uplink
[0013] According to the expression of the received energy of the user within time slot t, the maximum battery capacity, the battery capacity of the user at the end of the downlink WET phase in time slot t-1, and the energy reservation ratio for the uplink WIT phase in time slot t-1, an expression for the updated battery capacity of the user within time slot t is obtained
[0014] According to the expression of the updated battery capacity of the user within time slot t, the energy reservation ratio for the uplink WIT phase of the user within time slot t, the time allocation ratio of the uplink WIT phase in time slot t, and the duration of time slot t, an expression for the uplink transmission power of the user in the uplink WIT phase within time slot t is obtained
[0015] Based on the expression of the uplink transmission power of the user in the uplink WIT phase within time slot t and the bandwidth of the subcarrier, an expression for the uplink received signal-to-noise ratio of the user on any subcarrier b within time slot t is obtained
[0016] According to the binary indication of the subcarriers allocated for eMBB transmission and the subcarriers preempted by URLLC packets, the bandwidth of each subcarrier b, and the expression for the uplink received signal-to-noise ratio of the user on any subcarrier b within time slot t, an expression for the eMBB transmission rate of the user on micro-slot m within time slot t is obtained
[0017] According to the binary indication of the subcarriers preempted by URLLC packets, the bandwidth of each subcarrier b, the expression for the uplink received signal-to-noise ratio of the user on any subcarrier b within time slot t, the codeword block length of the user, the decoding error probability, the Gaussian cumulative distribution function, and the channel dispersion, an expression for the URLLC transmission rate of the user on micro-slot m within time slot t is obtained
[0018] According to all the obtained expressions, the optimization objective is set to maximize the total eMBB rate of the uplink of all users in the WPCN, the solution items are set to the subcarriers, time, and energy resources of all users, and the corresponding constraint conditions are set to construct an optimization problem
[0019] In an embodiment of the present invention, the expression for the received energy of the user within time slot t is:
[0020]
[0021] wherein, E u,tDenote the received energy of user \(u\) in time slot \(t\). Denote the received power of user \(u\) for the energy signal transmitted by the HAP in time slot \(t\); \(\tau\) u,t Denote the time allocation ratio of the downlink WET phase in time slot \(t\); \(t\) 0 Denote the duration of one time slot; \(\varphi\) denote the sensitivity of the RF / DC circuit of user \(u\); \(1(\cdot)\) denote the binary indicator function, which characterizes the comparison relationship between the received power and the sensitivity of the RF / DC circuit. When the content in \((\cdot)\) is true, \(1(\cdot)=1\); otherwise, \(1(\cdot)=0\).
[0022] In an embodiment of the present invention, the expression of the updated battery capacity of the user in time slot \(t\) is:
[0023] \(Q\) u,t \(=\min\{\rho\) u,t-1 \(Q\) u,t-1 \(+E\) u,t , \(Q\) max \(\}\)
[0024] where, \(Q\) u,t Denote the updated battery capacity of user \(u\) in time slot \(t\); \(Q\) u,t-1 Denote the battery capacity of user \(u\) at the end of the downlink WET phase in time slot \(t - 1\); \(Q\) max Denote the maximum battery capacity; \(\rho\) u,t-1 \(\in[0,1]\), denote the energy reservation ratio of user \(u\) for the uplink WIT phase in time slot \(t - 1\); \(\min\) denotes taking the minimum value.
[0025] In an embodiment of the present invention, the expression of the uplink transmission power of the user in the uplink WIT phase in time slot \(t\) is:
[0026]
[0027] where, Denote the uplink transmission power of user \(u\) in the uplink WIT phase in time slot \(t\); \(\rho\) u,t \(\in[0,1]\), denote the energy reservation ratio of user \(u\) for the uplink WIT phase in time slot \(t\), \(Q\) u,t Denote the battery energy queue of user \(u\) at the end of WET in time slot \(t\).
[0028] In an embodiment of the present invention, the expression of the eMBB transmission rate of the user on micro-slot \(m\) in time slot \(t\) is:
[0029]
[0030] where, Denote the eMBB transmission rate of user \(u\) on micro-slot \(m\) in time slot \(t\); \(B\) denote the set of all available subcarriers; \(f\)b represents the bandwidth of sub - carrier b; γ u,b,t represents the uplink received signal - to - noise ratio of user u on any sub - carrier b within time slot t; x u,b,t,m , y u,b,t,m ∈{0, 1}, respectively representing the binary indicators of the sub - carriers allocated for eMBB transmission and the sub - carriers preempted by URLLC packets.
[0031] In an embodiment of the present invention, the expression of the URLLC transmission rate of the user on the mini - slot m within the time slot t is:
[0032]
[0033] where, represents the URLLC transmission rate of user u on the mini - slot m within the time slot t; B represents the set of sub - carriers; C u,b,t represents the channel dispersion; n u represents the codeword block length of user u; ε is the decoding error probability; Q -1 (ε) represents the inverse function of the Gaussian cumulative distribution function.
[0034] In an embodiment of the present invention, the expression of the optimization problem is:
[0035]
[0036] s.t.
[0037]
[0038]
[0039]
[0040]
[0041]
[0042]
[0043]
[0044]
[0045] where, (P) represents the optimization problem; the content after s.t. represents the constraint conditions of the optimization problem; X = {x u,b,t,m} u∈U,b∈B,t∈T,m∈M ; Y = {y u,b,t,m} u∈U,b∈B,t∈T,m∈M, where X and Y represent the subcarrier allocation of all users in the solution items; u represents a user, u ∈ user set U = {1, 2, …, U}, and U represents the number of users; b represents a subcarrier, b ∈ subcarrier set B = {1, 2, …, B}, and B represents the number of subcarriers; t represents a time slot, t ∈ T = {1, 2, …, T}, T represents the set of time slots divided in the time domain, and the duration of each time slot is t 0 ; m represents a mini - time slot, m ∈ M = {1, 2, …, M}, M represents the number of mini - time slots divided in a time domain, and Μ represents the set of mini - time slots divided in a time slot; x u,b,t,m , y u,b,t,m ∈ {0, 1}, representing the binary indicators of the subcarriers allocated to eMBB transmissions and the subcarriers preempted by URLLC packets respectively; x u,b,t,m = 1 indicates that subcarrier b on mini - time slot m within time slot t is allocated to user u for eMBB transmission, otherwise x u,b,t,m = 0; y u,b,t,m = 1 indicates that subcarrier b on mini - time slot m within time slot t is preempted by user u for URLLC transmission, otherwise y u,b,t,m = 0; τ u,t represents the time allocation ratio of the downlink WET phase in time slot t, representing the time allocation in the solution items, 0 ≤ τ u,t ≤ 1; ρ u,t ∈ [0, 1], representing the energy ratio reserved by user u for the uplink WIT phase within time slot t, representing the energy resource allocation in the solution items; γ t ∈ (0, 1), representing the discount factor of time slot t; represents the eMBB transmission rate of user u on mini - time slot m within time slot t; E u,t represents the received energy of user u within time slot t; Q max represents the maximum battery capacity; Q u,t-1 represents the battery capacity of user u at the end of the downlink WET phase within time slot t - 1; ω u represents the minimum data rate requirement for the eMBB transmission of user u; F u represents the URLLC packet length of user u; represents the URLLC transmission rate of user u on mini - time slot m within time slot t; ψ represents the maximum allowable delay of URLLC packets; The first and second items in the constraint conditions are to ensure that a subcarrier can only be allocated to one user at the same time; The third item is to ensure that URLLC traffic can only preempt subcarriers that have been allocated to eMBB, that is, the preemption rationality constraint; The fourth item is the user battery capacity constraint to avoid energy overflow; The fifth and sixth items are the system eMBB minimum data rate requirement constraint and the URLLC transmission delay requirement constraint respectively.
[0046] In one embodiment of the present invention, solving the optimization problem using the preset hybrid deep deterministic policy gradient method to obtain the optimal solution includes:
[0047] Using the preset hybrid deep deterministic policy gradient method to decouple the optimization problem into a discrete sub-problem with subcarrier allocation optimization and a continuous sub-problem with time and energy resource allocation optimization, and alternately solving the two sub-problems to obtain the optimal solution.
[0048] In one embodiment of the present invention, the solving process of the discrete sub-problem with subcarrier allocation optimization includes:
[0049] After the time allocation ratio in the downlink WET phase and the energy ratio reserved for the uplink WIT phase {τ u,t-1 , ρ u,t-1} in a given time slot t - 1, rewrite the optimization problem as the first optimization sub-problem:
[0050] (P1)
[0051] s.t.
[0052]
[0053]
[0054]
[0055]
[0056]
[0057] Using an optimization tool to solve the first optimization sub-problem to obtain the discrete subcarrier allocation and preemption results {X * , Y *}; where X t and Y t are the X and Y corresponding to time slot t respectively; X * is the optimal solution of X t ; Y * is the optimal solution of X t ;
[0058] Correspondingly, the solving process of the continuous sub-problem with time and energy resource allocation optimization includes:
[0059] After the discrete subcarrier allocation and preemption results {X * , Y *} obtained from the first optimization sub-problem are given, rewrite the optimization problem as the second optimization sub-problem:
[0060] (p2)
[0061] s.t.
[0062]
[0063]
[0064]
[0065]
[0066]
[0067] Solve the second optimization sub - problem using the DDPG algorithm to obtain the optimal time - allocation ratio τ of the downlink WET phase within the optimal continuous time slot t u,t and the energy ratio ρ reserved for the uplink WIT phase u,t .
[0068] Advantages of the present invention:
[0069] In the solution provided by the embodiments of the present invention, first, under the constraints of considering URLLC latency, the sensitivity of the RF / DC circuit, the user battery capacity, and sub - carrier availability, an optimization problem for solving the optimal allocation scheme of sub - carriers, time, and energy resources for all users is constructed with the goal of maximizing the total eMBB rate of the uplink in WPCN; then the preset hybrid deep deterministic policy gradient method is used to solve the optimization problem to obtain the optimal solution. The embodiments of the present invention study the multiplexing of eMBB and URLLC transmissions on the shared channel in WPCN, where the HAP powers multiple users through WET and multiplexes its traffic onto the eMBB transmission by means of a preemptive puncturing method, realizing the service multiplexing of eMBB and URLLC in WPCN. Compared with static or semi - static spectrum resource allocation, it has a higher expected implementation rate and achieves a higher total eMBB rate.
[0070] Furthermore, when constructing the optimization problem, the embodiments of the present invention consider the energy reservation in each user's battery to avoid the interruption of uplink information transmission due to the lack of received energy in subsequent time slots, which can ensure the stability of uplink transmission. When solving this non - convex optimization problem with mixed optimization variables of discrete sub - carrier allocation and continuous time and energy resource allocation, the hybrid deep deterministic policy gradient method proposed by the embodiments of the present invention decouples it into a discrete sub - problem and a continuous sub - problem with the same objective as the original problem and solves it through alternating optimization, which can obtain the solution result simply and quickly. Description of the Drawings
[0071] Figure 1Schematic diagram of the multiplexing method for eMBB and URLLC in a wireless power supply communication network provided by an embodiment of the present invention;
[0072] Figure 2 Graph showing the comparison results of rewards of four algorithms in different episodes in the experiment of an embodiment of the present invention;
[0073] Figure 3 For the different battery capacities Q after the four algorithms converge in the experiment of an embodiment of the present invention max Graph showing the comparison results of the total eMBB rate;
[0074] Figure 4 Graph showing the comparison results of the total eMBB rate of four algorithms under different RF / DC circuit sensitivities after convergence in the experiment of an embodiment of the present invention. Specific implementation manners
[0075] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0076] For the multiplexing method for eMBB and URLLC in a wireless power supply communication network provided by an embodiment of the present invention, please refer to Figure 1 as shown below.
[0077] The multiplexing method for eMBB and URLLC in a wireless power supply communication network provided by an embodiment of the present invention is applied to a WPCN under preset scenario conditions; the WPCN under the preset scenario conditions includes a hybrid access point HAP and multiple users; each user is equipped with a rechargeable battery and has eMBB and URLLC service requirements; the time domain is divided into multiple time slots with equal durations, and each time slot is divided into multiple micro - time slots with equal durations; sub - carriers are retrieved and re - allocated to users at the beginning of each time slot for eMBB transmission; each URLLC packet arriving at a user preempts the sub - carriers assigned to the user in the next micro - time slot for URLLC transmission; users adopt a harvest - then - transmit protocol in each time slot; the method includes:
[0078] S1. Under the constraints of considering URLLC latency, the sensitivity of the RF / DC circuit, the user battery capacity, and sub - carrier availability, with the goal of maximizing the total eMBB rate of the uplink in the WPCN, an optimization problem is constructed to solve the optimal allocation scheme of sub - carriers, time, and energy resources for all users;
[0079] S2. Solve the optimization problem using a preset hybrid deep deterministic policy gradient method to obtain an optimal solution.
[0080] To facilitate the understanding of the solution of the embodiments of the present invention, the WPCN under preset scenario conditions is first described.
[0081] The embodiments of the present invention consider a WPCN under such preset scenario conditions, which includes a HAP (hybrid access point) and a group of users U = {1, 2,..., U}, where U represents the user set. Each user has eMBB and URLLC service requirements, and each user is equipped with a rechargeable battery. For easier analysis, it is assumed that the HAP and each user are each equipped with a single antenna for information transmission and reception. Let B = {1, 2,..., B} represent the set of all available subcarriers, and the bandwidth of each subcarrier can be expressed as f b Hz, where b ∈ B, but the bandwidth f b of each subcarrier is not the same, and the total system bandwidth is The time domain is divided into T time slots of equal duration, and the duration of each time slot is t 0 , and the set of all time slots is represented by T. On each subcarrier, it is assumed that the channel has reciprocity and the channel fading coefficient remains unchanged within a time slot, but changes in adjacent time slots. All subcarriers are retrieved at the beginning of each time slot and reallocated and scheduled to users based on CSI (channel state information) for eMBB transmission of users. Due to the strict latency requirements of URLLC transmission, each time slot is further divided into multiple micro time slots of equal duration, and the set of micro time slots within each time slot is represented by M = {1, 2,..., M}. Each URLLC packet arriving at a user will preempt the subcarriers allocated to the user in the next micro time slot for URLLC transmission, that is, it will be immediately scheduled to the next micro time slot to occupy the eMBB subcarriers for URLLC transmission without waiting for the completion of eMBB transmission on these subcarriers. Without loss of generality, the embodiments of the present invention assume that there is always a URLLC packet arriving in each micro time slot, but it arrives randomly at a certain user. Therefore, for a certain user, there may not necessarily be a URLLC packet arriving in a micro time slot.
[0082] It is assumed that all users adopt a harvest-then-transmit protocol in each time slot, that is, the users first harvest energy from the energy signal broadcast by the HAP, and then use the harvested energy to transmit information to the HAP. The process of the users harvesting energy from the energy signal broadcast by the HAP corresponds to the downlink transmission process; the process of the users using the harvested energy to transmit information to the HAP corresponds to the uplink transmission process.
[0083] For example, if user u is scheduled to transmit in time slot t, then time slot t is divided into a downlink WET (wireless energy transfer) phase with a duration of τ u,t t 0 and an uplink WIT (wireless information transfer) phase with a duration of (1 - τ u,t )t 0 , where 0 ≤ τ u,t ≤ 1, and τ u,t represents the time allocation ratio of the downlink WET phase in time slot t.
[0084] The following separately describes S1 and S2.
[0085] Specifically, for S1, in the embodiments of the present invention, the construction process of the optimization problem may include:
[0086] 1) Determine the expression of the received power of the user for the energy signal sent by the HAP within time slot t when the user performs transmission;
[0087] Specifically, assuming that the downlink transmission power of the HAP is equal on each sub - carrier, denoted as P DL , then the received power of user u within time slot t can be expressed as:
[0088]
[0089] where η c represents the energy conversion efficiency of the RF / DC circuit of user u; d u represents the distance between user u and the HAP; α represents the path loss exponent; represents the Rayleigh fading coefficient from the HAP to user u on sub - carrier b within time slot t; x u,b,t,m represents the binary indicator of the sub - carrier allocated for eMBB transmission; x u,b,t,m ∈ {0, 1}. Specifically, x u,b,t,m = 1 indicates that sub - carrier b on mini - slot m within time slot t is allocated to user u for eMBB transmission, otherwise x u,b,t,m = 0.
[0090] 2) According to the expression of the received power, the comparison relationship between the received power and the sensitivity of the RF / DC circuit, the time allocation ratio of the downlink WET phase in time slot t, and the duration of time slot t, obtain the expression of the received energy of the user within time slot t;
[0091] Among them, as described above, time slot \(t\) is divided into a downlink WET phase and an uplink WIT phase, and the uplink WIT phase corresponds to the WPCN uplink.
[0092] In the embodiment of the present invention, it is considered that when the received power of the user is less than the sensitivity \(\varphi\) of its RF / DC circuit, the user cannot collect the energy from the HAP.
[0093] Therefore, the expression of the received energy of the user within time slot \(t\) is:
[0094]
[0095] where \(E\) u,t represents the received energy of user \(u\) within time slot \(t\); represents the received power of user \(u\) for the energy signal transmitted by the HAP within time slot \(t\); \(\tau\) u,t represents the time allocation ratio of the downlink WET phase in time slot \(t\); \(t\) 0 represents the duration of one time slot; \(\varphi\) represents the sensitivity of the RF / DC circuit of user \(u\); \(1(\cdot)\) represents the binary indicator function, which characterizes the comparison relationship between the received power and the sensitivity of the RF / DC circuit. When the content in \((\cdot)\) is true, \(1(\cdot)=1\); otherwise, \(1(\cdot)=0\).
[0096] It can be seen that in the process of constructing the optimization problem in the embodiment of the present invention, the sensitivity of the user's RF / DC circuit is considered as a constraint.
[0097] 3) According to the expression of the received energy of the user within time slot \(t\), the maximum battery capacity, the battery capacity of the user at the end of the downlink WET phase within time slot \(t - 1\), and the energy ratio reserved for the uplink WIT phase within time slot \(t - 1\), obtain the expression of the updated battery capacity of the user within time slot \(t\);
[0098] Based on the above, the update equation of the battery capacity of user \(u\) within time slot \(t\) can be obtained, that is, the expression of the updated battery capacity of the user within time slot \(t\) is:
[0099] \(Q\) u,t \(=\min\{\rho\) u,t-1 \(Q\) u,t-1 \(+E\) u,t , \(Q\) max \}\)
[0100] where \(Q\) u,t represents the updated battery capacity of user \(u\) within time slot \(t\); \(Q\) u,t-1 represents the battery capacity of user \(u\) at the end of the downlink WET phase within time slot \(t - 1\); \(Q\) max represents the maximum battery capacity; \(\rho\) u,t-1∈[0,1] represents the proportion of energy reserved by user u for the uplink WIT phase in time slot t - 1; min represents finding the minimum value.
[0101] It can be seen that in the process of constructing the optimization problem in the embodiments of the present invention, the battery capacity of the user is considered as a constraint.
[0102] 4) According to the expression of the updated battery capacity of the user in time slot t, the proportion of energy reserved by the user for the uplink WIT phase in time slot t, the time allocation proportion of the uplink WIT phase in time slot t, and the duration of time slot t, obtain the expression of the uplink transmission power of the user in the uplink WIT phase in time slot t;
[0103] Based on the foregoing, the expression of the uplink transmission power of the user in the uplink WIT phase in time slot t is:
[0104]
[0105] Among them, represents the uplink transmission power of user u in the uplink WIT phase in time slot t; ρ u,t ∈[0,1], represents the proportion of energy reserved by user u for the uplink WIT phase in time slot t, Q u,t represents the battery energy queue of user u at the end of WET in time slot t. In the embodiments of the present invention, the energy reservation ρ u,t of each user's battery is used as an optimization variable, and a part of the received energy is reserved for subsequent time slots to avoid insufficient power supply for uplink transmission.
[0106] 5) Based on the expression of the uplink transmission power of the user in the uplink WIT phase in time slot t and the bandwidth of the subcarrier, obtain the expression of the uplink received signal - to - noise ratio of the user on any subcarrier b in time slot t;
[0107] Based on the foregoing, the uplink received SNR (signal to noise ratio) γ u,b,t of user u on subcarrier b in time slot t is expressed as:
[0108]
[0109] Among them, σ 2 represents the power spectral density of additive noise; the bandwidth of each subcarrier is f b Hz.
[0110] 6) According to the binary indication of the subcarriers allocated for eMBB transmission and the subcarriers preempted by URLLC packets, the bandwidth of each subcarrier b, and the expression of the uplink received signal-to-noise ratio of the user on any subcarrier b within the time slot t, obtain the expression of the eMBB transmission rate of the user on the mini-slot m within the time slot t;
[0111] Based on the above, based on the Shannon formula, the expression of the eMBB transmission rate of the user on the mini-slot m within the time slot t is:
[0112]
[0113] Among them, represents the eMBB transmission rate of user u on the mini-slot m within the time slot t; B represents the set of all available subcarriers; f b represents the bandwidth of subcarrier b; γ u,b,t represents the uplink received signal-to-noise ratio of user u on any subcarrier b within the time slot t; ; x u,b,t,m , y u,b,t,m ∈{0, 1}, respectively representing the binary indication of the subcarriers allocated for eMBB transmission and the subcarriers preempted by URLLC packets.
[0114] And, in order to ensure that the URLLC packets of user u can only preempt the subcarriers that have been allocated to this user for eMBB transmission, the embodiment of the present invention makes a necessary regulation:
[0115]
[0116] 7) According to the binary indication of the subcarriers preempted by URLLC packets, the bandwidth of each subcarrier b, the expression of the uplink received signal-to-noise ratio of the user on any subcarrier b within the time slot t, the codeword block length of the user, the decoding error probability, the Gaussian cumulative distribution function, and the channel dispersion, obtain the expression of the URLLC transmission rate of the user on the mini-slot m within the time slot t;
[0117] Since the packet length of URLLC is much smaller than that of eMBB, using the Shannon formula model will significantly overestimate the delay performance of URLLC transmission, which means that the Shannon formula is no longer applicable to URLLC transmission. Therefore, based on the finite block length theory, the embodiment of the present invention obtains that the expression of the URLLC transmission rate of the user on the mini-slot m within the time slot t is:
[0118]
[0119] Among them, represents the URLLC transmission rate of user u on the mini-slot m within the time slot t; B represents the set of subcarriers; C u,b,t represents the channel dispersion; nu denotes the codeword block length of user u (unit: symbol); ε is the decoding error probability; Q -1 (ε) represents the inverse function of the Gaussian cumulative distribution function.
[0120] 8) According to all the obtained expressions, set the optimization objective to maximize the total eMBB rate of the uplink of all users in the WPCN, set the solution items to the subcarriers, time, and energy resources of all users, and set the corresponding constraint conditions to construct an optimization problem.
[0121] Specifically, the expression of the optimization problem is:
[0122] (P):
[0123] s.t.
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131]
[0132] To facilitate the overall understanding of the parameters in this optimization problem, the meanings of the parameters involved are explained centrally here.
[0133] Among them, (P) represents the optimization problem; the content after s.t. represents the constraint conditions of the optimization problem; X = {x u,b,t,m} u∈U,b∈B,t∈T,m∈M ; Y = {y u,b,t,m} u∈U,b∈B,t∈T,m∈M , X and Y represent the subcarrier allocation of all users in the solution items; τ = {τ u,t} u∈U,t∈T ; ρ = {ρ u,t} u∈U,t∈T, τ and ρ respectively represent the time allocation ratio for the downlink WET phase in the solution item and the energy ratio reserved for the uplink WIT phase; u represents a user, u ∈ user set U = {1, 2, …, U}, where U represents the number of users; b represents a subcarrier, b ∈ subcarrier set B = {1, 2, …, B}, where B represents the number of subcarriers; t represents a time slot, t ∈ T = {1, 2, …, T}, where T represents the number of time slots divided in the time domain, and T represents the set of time slots divided in the time domain, and the duration of each time slot is t 0 ; m represents a mini - slot, m ∈ M = {1, 2, …, M}, where M represents the number of mini - slots divided in a time domain, and Μ represents the set of mini - slots divided in a time slot; x u,b,t,m , y u,b,t,m ∈ {0, 1}, respectively representing the binary indicators of the subcarriers allocated to eMBB transmission and the subcarriers preempted by URLLC packets; x u,b,t,m = 1 indicates that subcarrier b on mini - slot m within time slot t is allocated to user u for eMBB transmission, otherwise x u,b,t,m = 0; y u,b,t,m = 1 indicates that subcarrier b on mini - slot m within time slot t is preempted by user u for URLLC transmission, otherwise y u,b,t,m = 0; τ u,t represents the time allocation ratio of the downlink WET phase in time slot t, representing the time allocation in the solution item, 0 ≤ τ u,t ≤ 1; ρ u,t ∈ [0, 1], representing the energy ratio reserved by user u for the uplink WIT phase within time slot t, representing the energy resource allocation in the solution item; represents the eMBB transmission rate of user u on mini - slot m within time slot t; E u,t represents the received energy of user u within time slot t; Q max represents the maximum battery capacity; Q u,t-1 represents the battery capacity of user u at the end of the downlink WET phase within time slot t - 1; ω u represents the minimum data rate requirement for the eMBB transmission of user u; F u represents the length of the URLLC packet of user u; represents the URLLC transmission rate of user u on mini - slot m within time slot t; ψ represents the maximum allowable delay of the URLLC packet; The first and second items in the constraint conditions are to ensure that a subcarrier can only be allocated to one user at the same time; The third item is to ensure that URLLC traffic can only preempt the subcarriers that have been allocated to eMBB, that is, the preemption rationality constraint; The fourth item is the user battery capacity constraint to avoid energy overflow; The fifth and sixth items are the system eMBB minimum data rate requirement constraint and the URLLC transmission delay requirement constraint respectively.
[0134] It can be understood that the URLLC latency corresponds to the sixth item in the corresponding constraints; the user battery capacity corresponds to the fourth item in the corresponding constraints; the constraints on subcarrier availability correspond to the first and second items in the corresponding constraints, and the sensitivity of the RF / DC circuit has been considered in the process of constructing the optimization problem.
[0135] It can be seen that in the WPCN with the preset scenario conditions of the present invention, the HAP wirelessly transmits energy to the user and receives information from the user. Among them, URLLC adopts a preemptive shielding method to multiplex its traffic onto the eMBB transmission to achieve the multiplexing of eMBB and URLLC services. In order to maximize the total uplink eMBB rate, the embodiments of the present invention formulate an optimization problem, and jointly allocate the user's subcarriers, time, and energy resources under the constraints of URLLC latency, RF / DC sensitivity, user battery capacity, and subcarrier availability.
[0136] Moreover, different from the existing work, the embodiments of the present invention consider the sensitivity of the RF / DC circuit and the limited battery capacity of each user. To avoid insufficient power supply for uplink transmission, the energy reserved in each user's battery is used as an optimization variable, and a part of the received energy is reserved for subsequent time slots to avoid interrupting the uplink information transmission due to no energy received in subsequent time slots.
[0137] Specifically, for S2, in the embodiments of the present invention, solving the optimization problem by using the preset hybrid deep deterministic policy gradient method to obtain the optimal solution includes:
[0138] Using the preset hybrid deep deterministic policy gradient method to decouple the optimization problem into a discrete sub-problem with subcarrier allocation optimization and a continuous sub-problem with time and energy resource allocation optimization, and alternately solving the two sub-problems to obtain the optimal solution.
[0139] Since the optimization problem (P) includes discrete subcarrier allocation and continuous time and energy allocation, and this problem is non-convex, the embodiments of the present invention propose a new alternating algorithm called the hybrid deep deterministic policy gradient method, abbreviated as the Mixed-DDPG method, to solve this optimization problem (P). Specifically, the original problem (P) is decoupled into a discrete sub-problem and a continuous sub-problem, and the optimal solution is obtained by alternately solving these two sub-problems until convergence. The proposed Mixed-DDPG method is proven to converge to a stable state and can achieve a higher total eMBB rate compared with the existing solutions. The following is a specific description.
[0140] (1) The solution process of the discrete sub-problem with subcarrier allocation optimization includes:
[0141] After the time allocation ratio of the given downlink WET phase and the energy ratio reserved for the uplink WIT phase {τ u,t}, u∈U and {ρ u,t}, u∈U the optimization problem is rewritten as the first optimization sub-problem:
[0142] (P1):
[0143] s.t.
[0144]
[0145]
[0146]
[0147]
[0148]
[0149] Using an optimization tool to solve the first optimization sub-problem, the discrete sub-carrier allocation and preemption results {X * , Y *} are obtained; where X t = {x u,b,t,m} u∈U,b∈B,m∈M and Y t = {y u,b,t,m} u∈U,b∈B,m∈M are the X and Y corresponding to time slot t respectively; X * is the optimal solution of X t ; Y * is the optimal solution of X t .
[0150] Specifically, by fixing the time allocation ratio {τ u,t} u∈U of the continuous downlink WET phase and the energy ratio {ρ u,t} u∈U reserved for the uplink WIT phase in the optimization problem (P), the embodiments of the present invention can obtain a discrete sub-problem, which optimizes the binary indication of sub-carrier allocation for eMBB transmission at the start of time slot t + 1 and the sub-carrier preemption of URLLC packets in each time slot m ∈ M of time slot t + 1. In addition, based on the formula of the foregoing optimization problem construction process, for fixed {τ u,t} u∈U and {ρ u,t} u∈U , Independent for different time slots t, and the eMBB sum rate of all users can be maximized separately in each time slot. Since this sub-problem is formulated and solved separately for each time slot t, the objective function of the first optimization sub-problem (P1) is no longer summed over time slot t. Since the optimal solution needs to be obtained for each mini-slot, given {τ u,t} u∈U and {ρ u,t} u∈U the optimization problem (P) is rewritten as the first optimization sub-problem (P1). Since the first optimization sub-problem (P1) is convex, it can be efficiently solved by an optimization tool. Among them, any existing optimization tool can be used, such as Cvx, gurobi, etc. Among them, Cvx is a kind of MATLAB convex optimization toolbox; gurobi is a new generation of large-scale mathematical programming optimizer developed by Gurobi Company in the United States.
[0151] (2) The solution process of the continuous sub-problem with time and energy resource allocation optimization includes:
[0152] Given the discrete subcarrier allocation and preemption results {X * , Y *} obtained from the first optimization sub-problem, the optimization problem is rewritten as a second optimization sub-problem:
[0153] (p2)
[0154] s.t.
[0155]
[0156]
[0157]
[0158]
[0159]
[0160] Use the DDPG algorithm to solve the second optimization sub-problem to obtain the optimal time allocation ratio τ u,t for the downlink WET phase within the continuous time slot t and the energy ratio ρ u,t reserved for the uplink WIT phase.
[0161] Specifically, since the sub - problem (P2) has a large input space, including X, Y, CSI, and the battery states of all users in different time slots, it is difficult to solve using traditional optimization methods. However, reinforcement learning can be used. An embodiment of the present invention adopts a model - free neural network, that is, a DDPG algorithm with a continuous action space and capable of solving continuous variables τ u,t and ρ u,t is used to solve it.
[0162] Among them, the state, action, and reward of DDPG are defined as follows:
[0163] ① State: s t ={H t ,Q t ,X t ,Y t}, where the channel state information is H t ={h u,b,t} u∈U,b∈B , and the battery state is Q t ={Q u,t} u∈U .
[0164] ② Action: a t ={τ t ,ρ t}, where the time - allocation ratio of the downlink WET phase is τ t ={τ u,t} u∈U , and the energy ratio reserved for the uplink WIT phase is ρ t ={ρ u,t} u∈U .
[0165] ③ Reward: When an action a t is selected, the reward r t is immediately obtained by the following formula.
[0166]
[0167] Among them, is the penalty term; δ is the penalty factor; The penalty term represents the penalty imposed on the agent by the system when any constraint is violated, which can help prevent over - fitting.
[0168] DDPG contains an actor network and a critic network that respectively generate the policy and evaluate the policy. Based on the input s t , the actor network μ(s t |θ μ ) selects a deterministic action by the following formula.
[0169]
[0170] where θ μ is the actor network parameter, is the additive Gaussian noise introduced for action exploration, which follows a normal distribution with mean μ 1 and variance σ 1 2 . When given s t , a t and r t obtained from the environment, to reduce the correlation between training sampled data, the critic network will use N mini-batch data randomly sampled from the replay experience pool {(s j , a j , r j , s j+1 )} j=1,…,N to generate the Q-value Q(s i , a i |θ Q ) to evaluate the selected action a t , and update the parameter θ Q by minimizing the loss.
[0171]
[0172] where y j = r j + γQ'(s j+1 , μ'(s j+1 |θ μ' )|θ Q' ), γ is the discount factor, θ μ' and θ Q′ are the parameters of the target actor network μ'(s|θ μ' ) and the target critic network Q′(s,a|θ Q′ ) introduced to ensure the stability of the DDPG learning process. The target critic network is used to output the target Q-value instead of using the same network to output the Q-value and update the parameters simultaneously.
[0173] The actor network is updated using the deterministic sampling gradient policy.
[0174]
[0175] The parameters of the target actor network and the target critic network are softly updated as follows:
[0176] θ μ' ← ζθ μ + (1 - ζ)θ μ'
[0177] θQ' ←ζθ Q +(1 - ζ)θ Q'
[0178] Among them, 0 ≤ ζ = 1 is the soft update parameter.
[0179] For the specific processing process of DDPG, please refer to the relevant technology, and no detailed description will be given here.
[0180] For the complete solution process of the constructed optimization problem, the execution process of the Mixed-DDPG method described in the following steps can be referred to for understanding.
[0181] Step 1: Initialize U, B, {f b} b∈B , η c , P DL , {d u} u∈U , σ 2 , ψ, ε, φ, {ω u} u∈U , Q max ; Set episode max = 200, step max = 200; Let episode = 0;
[0182] Step 2: Let t = 0, update episode + 1, randomly initialize {Q u,t} u∈U , X t , Y t and the replay experience pool Generate channel sampling {H t} t∈T ; Let step = 0;
[0183] Step 3: Update t + 1, update step + 1, obtain H t and Q t , and the state s t = {H t , Q t , X t , Y t};
[0184] Step 4: Select the action a t according to the state s t = {τ u,t , ρ u,t} u∈U ;
[0185] Step 5: Obtain the system reward r t from the environment, and the new battery energy queue Q t+1, obtain the next state value s t+1 = {H t+1 , Q t+1 , X t+1 , Y t+1}; Store (s t , a t , r t , s t+1 ) in the replay experience pool ;
[0186] Step 6, randomly sample N mini-batch data {(s , a j , r j , s j , s j+1 )} j=1,…,N as training data;
[0187] Step 7, update the parameters of the critic network, actor network, and target network respectively;
[0188] Step 8, return to Step 3 until step = step max .
[0189] Step 9, return to Step 2 until episode = episode max .
[0190] Regarding the proposed Mixed-DDPG method, specifically, the optimization problem (P) is decoupled into the two aforementioned sub-problems and the optimal solution is obtained by alternately solving the two sub-problems. First, assume that the discrete variables X t and Y t are known at the t-th time slot, and solve the discrete sub-problem (P2) using the DDPG algorithm. Given the continuous variables {τ u,t} u∈U and {ρ u,t} u∈U , obtain the optimal discrete variables X t+1 and Y t+1 by solving the discrete sub-problem (P1) using an optimization tool. Iteratively execute these steps until convergence.
[0191] To calculate the complexity of the proposed Mixed-DDPG method, first solve the computational complexities of the two sub-problems (P1) and (P2) respectively. For the convex problem (P1), since it can be solved by an optimization tool and has 2UBTM decision variables, the computational complexity of the sub-problem (P1) is For the sub-problem (P2), define j i and k i as the sizes of the input and output of layer i, then the complexity of the sub-problem (P2) is Therefore, the total computational complexity of the proposed Mixed-DDPG method is where L is the number of iterations of the Mixed-DDPG method.
[0192] It can be understood that the optimization problem needs to be solved in each time slot, and the obtained optimal solution is allocated to each user.
[0193] In the solution provided by the embodiments of the present invention, first, under the constraints of considering the URLLC delay, the sensitivity of the RF / DC circuit, the user battery capacity, and the subcarrier availability, with the goal of maximizing the total eMBB uplink rate in the WPCN, an optimization problem for solving the optimal allocation scheme of subcarriers, time, and energy resources for all users is constructed; then, the preset hybrid deep deterministic policy gradient method is used to solve the optimization problem to obtain the optimal solution. The embodiments of the present invention study the multiplexing of eMBB and URLLC transmissions on the shared channel in the WPCN, where the HAP powers multiple users through WET and multiplexes its traffic onto the eMBB transmission by means of a preemptive puncturing method, realizing the service multiplexing of eMBB and URLLC in the WPCN, having a higher expected implementation rate compared with static or semi-static spectrum resource allocation, and achieving a higher total eMBB rate.
[0194] Moreover, when constructing the optimization problem, the embodiments of the present invention consider the energy reservation in each user's battery to avoid the interruption of uplink information transmission due to not receiving energy in subsequent time slots, which can ensure the stability of uplink transmission. When solving this non-convex optimization problem with mixed optimization variables of discrete subcarrier allocation and continuous time and energy resource allocation, the hybrid deep deterministic policy gradient method proposed by the embodiments of the present invention decouples it into a discrete subproblem and a continuous subproblem with the same objective as the original problem, and solves it through alternating optimization, and can obtain the solution result simply and quickly.
[0195] To verify and illustrate the effect of the method provided by the embodiments of the present invention, the following simulation experiments are given.
[0196] The simulation environment includes U = 10 users and B = 20 subcarriers, and the bandwidth of each subcarrier is f b = 1 MHz. The power spectral density of additive white Gaussian noise is σ 2 = -174 dBm / Hz, and t 0 is normalized to 1. For each user, η c is set to 0.5, F u = 20 bytes, ψ = 1 ms, and the distance d u is uniformly distributed within the range (10 m, 15 m). Each battery has a maximum capacity Q max∈ {5, 10, 15, 20, 25, 30} μJ, the circuit sensitivity is set to φ ∈ {3, 9, 15, 21, 27} μW. Additionally, assume α = 3, ε = 0.01, ω u = 10 Mbps, γ = 0.9, M = 8. The critic network and the actor network have three hidden layers, with 128 neurons in each layer. The Adam optimizer is used to train the network, and the initial learning rates λ a = 0.001 and λ c = 0.002 are set for the actor network and the critic network respectively, and the batch size is set to 128.
[0197] To compare with the Mixed-DDPG method proposed in the embodiments of the present invention, the following algorithms are introduced: ZER-DDPG (zero energy reservation DDPG) algorithm, i.e., ρ = 0 (ρ is the proportion of energy reserved in the uplink WIT phase); FERP-DDPG (fixed energy reservation proportional DDPG) algorithm, i.e., ρ = 0.5; FTTP-DDPG (fixed transmission time proportional DDPG) algorithm, i.e., τ = 0.5 (τ is the time allocation proportion in the downlink WET phase).
[0198] Figure 2 Shows the reward values rewards of the four algorithms under different episodes (Chinese meaning: rounds), where φ = 20 μW, Q max = 25 μJ. It can be clearly seen that the Mixed-DDPG method proposed in the embodiments of the present invention has better performance than the other three algorithms, because the Mixed-DDPG method can dynamically adjust the time allocation proportion τ in the downlink WET phase and the proportion of energy reserved for the uplink WIT phase ρ according to the user's battery capacity and channel state. For example, when the four algorithms converge (i.e., when episode reaches 75), the reward value of the Mixed-DDPG algorithm is about 10%, 13%, and 900% higher than that of ZER-DDPG, FERP-DDPG, and FTTP-DDPG respectively.
[0199] In Figure 3 gives the total eMBB rate under different battery capacities Q max after the four algorithms converge, where φ = 20 μW. It can be seen that as Q maxWith the increase of [specific parameter], the performance of the four algorithms also increases and tends to be stable. This is because a larger battery capacity can store more energy for high rates. However, when the battery capacity is larger than the received energy, the battery capacity no longer affects the data rate. When Q max = 5 μJ, the total rate of the proposed Mixed-DDPG method is approximately 4%, 8%, and 95% higher than that of ZER-DDPG, FERP-DDPG, and FTTP-DDPG, respectively.
[0200] Figure 4 Describes the eMBB total rate of the four algorithms under different sensitivities of the RF / DC circuit after convergence, where the battery capacity Q max = 25 μJ. It can be seen that all four algorithms show a downward trend with the increase of φ. This is because a higher φ will result in less received energy, thus increasing the possibility of energy shortage. It can be seen that the proposed Mixed-DDPG method still has the highest eMBB total rate, which is approximately 4%, 12%, and 80% higher than that of ZER-DDPG, FERP-DDPG, and FTTP-DDPG, respectively, when φ = 15 μW. It can be seen that the ZER-DDPG algorithm can have better performance than the FERP-DDPG algorithm because flexibly changing τ can simultaneously affect the total received energy value in the downlink WET phase and the power transmission in the uplink WIT phase, thereby affecting the system rate.
[0201] In summary, the embodiments of the present invention solve the multiplexing problem of eMBB and URLLC services in the WPCN network. Specifically, the embodiments of the present invention allocate time, frequency bands, and energy resources for the two services with different requirements of eMBB and URLLC on the premise of considering limited battery capacity and a preemption mechanism, so as to maximize the total eMBB rate of the system. The embodiments of the present invention also consider the maximum battery capacity and the sensitivity of the RF / DC circuit in the WPCN, making the model more practical. The goal of the optimization problem is to maximize the total eMBB rate of all users over time while satisfying a series of constraints. To solve this problem, the embodiments of the present invention decouple the original problem into two sub-problems and propose the Mixed-DDPG method to solve them alternately. Numerical simulations show that the performance of the method proposed in the embodiments of the present invention is better than that of other comparative algorithms.
[0202] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of these features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0203] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. A multiplexing method for eMBB and URLLC in a wireless power supply communication network, characterized in that, it is applied to a WPCN under preset scenario conditions; the WPCN under the preset scenario conditions includes a hybrid access point HAP and multiple users; each user is equipped with a rechargeable battery and has eMBB and URLLC service requirements; The time domain is divided into multiple time slots of equal duration, and each time slot is divided into multiple micro - time slots of equal duration; sub - carriers are retrieved and re - allocated to users at the beginning of each time slot for eMBB transmission; each arriving URLLC packet preempts the sub - carriers allocated to the user in the next micro - time slot for URLLC transmission; Users adopt a harvest - then - transmit protocol in each time slot; the method includes: Under the constraints of considering URLLC latency, the sensitivity of the RF / DC circuit, user battery capacity, and sub - carrier availability, with the goal of maximizing the total eMBB rate of the uplink in the WPCN, an optimization problem is constructed to solve the optimal allocation scheme of sub - carriers, time, and energy resources for all users; Use the preset hybrid deep deterministic policy gradient method to solve the optimization problem to obtain the optimal solution; Among them, the construction process of the optimization problem includes: Determine the expression of the received power of the energy signal sent by the user to the HAP within the time slot t when the user performs transmission; According to the expression of the received power, the comparison relationship between the received power and the sensitivity of the RF / DC circuit, the time allocation ratio of the downlink WET stage in the time slot t, and the duration of the time slot t, obtain the expression of the received energy of the user within the time slot t; where the time slot t is divided into a downlink WET stage and an uplink WIT stage, and the uplink WIT stage corresponds to the uplink of the WPCN; According to the expression of the received energy of the user within the time slot t, the maximum battery capacity, the battery capacity of the user at the end of the downlink WET stage in the time slot t - 1, and the energy ratio reserved for the uplink WIT stage in the time slot t - 1, obtain the expression of the updated battery capacity of the user within the time slot t; According to the expression of the updated battery capacity of the user within the time slot t, the energy ratio reserved by the user for the uplink WIT stage within the time slot t, the time allocation ratio of the uplink WIT stage in the time slot t, and the duration of the time slot t, obtain the expression of the uplink transmission power of the user in the uplink WIT stage within the time slot t; Based on the expression of the uplink transmission power of the user in the uplink WIT stage within the time slot t and the bandwidth of the sub - carrier, obtain the expression of the uplink received signal - to - noise ratio of the user on any sub - carrier b within the time slot t; According to the binary indication of the sub - carriers allocated for eMBB transmission and the sub - carriers preempted by URLLC packets, the bandwidth of each sub - carrier b, and the expression of the uplink received signal - to - noise ratio of the user on any sub - carrier b within the time slot t, obtain the expression of the eMBB transmission rate of the user on the micro - time slot m within the time slot t; Based on the binary indication of the subcarriers preempted by the URLLC packet, the bandwidth of each subcarrier b, the expression of the uplink received signal-to-noise ratio of the user on any subcarrier b within the time slot t, the codeword block length of the user, the decoding error probability, the Gaussian cumulative distribution function, and the channel dispersion, obtain the expression of the URLLC transmission rate of the user on the mini-slot m within the time slot t; Based on all the obtained expressions, set the optimization objective to maximize the total eMBB rate of the uplink of all users in the WPCN, set the solution items to the subcarriers, time, and energy resources of all users, and set the corresponding constraint conditions to construct an optimization problem; Among them, solving the optimization problem using the preset hybrid deep deterministic policy gradient method to obtain the optimal solution includes: using the preset hybrid deep deterministic policy gradient method to decouple the optimization problem into a discrete subproblem with subcarrier allocation optimization and a continuous subproblem with time and energy resource allocation optimization, and alternately solving the two subproblems to obtain the optimal solution.
2. The method for multiplexing eMBB and URLLC in a wireless power supply communication network according to claim 1, characterized in that the expression of the received energy of the user within the time slot t is: Among them, E u,t represents the received energy of user u in time slot t; represents the received power of user u for the energy signal transmitted by the HAP in time slot t; τ u,t represents the time allocation ratio of the downlink WET phase in time slot t; t 0 represents the duration of a time slot; φ represents the sensitivity of the RF / DC circuit of user u; 1(·) represents the binary indicator function, characterizing the comparison relationship between the received power and the sensitivity of the RF / DC circuit. When the content in (·) is true, 1(·) = 1, otherwise, 1(·) = 0.
3. The method for multiplexing eMBB and URLLC in a wireless power supply communication network according to claim 2, characterized in that the expression of the updated battery capacity of the user within the time slot t is: Q u,t = min{ρ u,t-1 Q u,t-1 + E u,t , Q max} Among them, Q u,t represents the updated battery capacity of user u within time slot t; Q u,t-1 represents the battery capacity of user u at the end of the downlink WET phase within time slot t - 1; Q max represents the maximum battery capacity; ρ u,t-1 ∈[0, 1], represents the proportion of energy reserved by user u for the uplink WIT phase within time slot t - 1; min represents finding the minimum value.
4. The method for multiplexing eMBB and URLLC in a wireless power supply communication network according to claim 3, characterized in that the expression of the uplink transmission power of the user in the uplink WIT phase within the time slot t is: where, represents the uplink transmission power of user u in the uplink WIT phase within time slot t; ρ u,t ∈ [0, 1], represents the energy ratio reserved by user u for the uplink WIT phase within time slot t, and Q u,t represents the battery energy queue of user u at the end of WET within time slot t.
5. The method for multiplexing eMBB and URLLC in a wireless power supply communication network according to claim 4, characterized in that the expression of the eMBB transmission rate of the user on the mini-slot m within the time slot t is: Among them, represents the eMBB transmission rate of user u on micro-slot m within time slot t; B represents the set of all available sub-carriers; f b represents the bandwidth of sub-carrier b; γ u,b,t represents the uplink received signal-to-noise ratio of user u on any sub-carrier b within time slot t; x u,b,t,m , y u,b,t,m ∈ {0, 1} respectively represent the binary indicators of the sub-carriers allocated for eMBB transmission and the sub-carriers preempted by URLLC packets.
6. The method for multiplexing eMBB and URLLC in a wireless power supply communication network according to claim 5, characterized in that the expression of the URLLC transmission rate of the user on the mini-slot m within the time slot t is: Among them, represents the URLLC transmission rate of user u on micro-slot m within time slot t; B represents the sub-carrier set; C u,b,t represents the channel dispersion; n u represents the codeword block length of user u; ε is the decoding error probability; Q -1 (ε) represents the inverse function of the Gaussian cumulative distribution function.
7. The method for multiplexing eMBB and URLLC in a wireless power supply communication network according to claim 1 or 6, characterized in that the expression of the optimization problem is: (P): s.t. Among them, (P) represents the optimization problem; the content after s.t. represents the constraint conditions of the optimization problem; X = {x u,b,t,m} u∈U,b∈B,t∈T,m∈M ; Y = {y u,b,t,m} u∈U,b∈B,t∈T,m∈M , where X and Y represent the subcarrier allocation of all users in the solution items; u represents a user, u ∈ user set U = {1, 2,..., U}, and U represents the number of users; b represents a subcarrier, b ∈ subcarrier set B = {1, 2,..., B}, and B represents the number of subcarriers; t represents a time slot, t ∈ T = {1, 2,..., T}, T represents the set of time slots divided in the time domain, and the duration of each time slot is t 0 ; m represents a mini - time slot, m ∈ M = {1, 2,..., M}, M represents the number of mini - time slots divided in a time domain, and Μ represents the set of mini - time slots divided in a time slot; x u,b,t,m , y u,b,t,m ∈ {0, 1}, representing the binary indicators of the subcarriers allocated for eMBB transmission and the subcarriers preempted by URLLC packets respectively; x u,b,t,m = 1 means that subcarrier b in mini - time slot m within time slot t is allocated to user u for eMBB transmission, otherwise x u,b,t,m = 0; y u,b,t,m = 1 means that subcarrier b in mini - time slot m within time slot t is preempted by user u for URLLC transmission, otherwise y u,b,t,m = 0; τ u,t represents the time allocation ratio of the downlink WET phase in time slot t, representing the time allocation in the solution item, 0 ≤ τ u,t ≤ 1; ρ u,t ∈ [0, 1], representing the energy ratio reserved by user u for the uplink WIT phase within time slot t, representing the energy resource allocation in the solution item; γ t ∈ (0, 1), representing the discount factor of time slot t; represents the eMBB transmission rate of user u in mini - time slot m within time slot t; E u,t represents the received energy of user u within time slot t; Q max represents the maximum battery capacity; Q u,t-1 represents the battery capacity of user u at the end of the downlink WET phase within time slot t - 1; ω u represents the minimum data rate requirement for the eMBB transmission of user u; F u represents the URLLC packet length of user u; Denote the URLLC transmission rate of user \(u\) on micro-slot \(m\) within time slot \(t\); \(\psi\) represents the maximum allowable delay of the URLLC packet; the first and second terms in the constraint conditions are to ensure that a sub-carrier can only be allocated to one user at the same time; the third term is to ensure that the URLLC traffic can only preempt the sub-carriers already allocated to eMBB, that is, the preemption rationality constraint; the fourth term is the user battery capacity constraint to avoid energy overflow; the fifth and sixth terms are the minimum data rate requirement constraint of the system eMBB and the delay requirement constraint of the URLLC transmission, respectively.
8. The method for multiplexing eMBB and URLLC in a wireless power supply communication network according to claim 7, characterized in that the solution process of the discrete subproblem with subcarrier allocation optimization includes: The time allocation ratio of the downlink WET phase and the energy ratio reserved for the uplink WIT phase {τ u,t-1 , ρ u,t-1} within a given time slot t - 1, the optimization problem is rewritten as a first optimization sub - problem: (P1) s.t. Solve the first optimization sub-problem using an optimization tool to obtain the discrete sub-carrier allocation and preemption results {X * , Y *}; where X t and Y t are the X and Y corresponding to time slot t respectively; X * is the optimal solution of X t ; Y * is the optimal solution of X t . Correspondingly, the solution process of the continuous subproblem with time and energy resource allocation optimization includes: After the discrete sub - carrier allocation and pre - emption result {X * , Y *} given by the first optimization sub - problem, rewrite the optimization problem as a second optimization sub - problem: (P2): s.t. The DDPG algorithm is used to solve the second optimization sub-problem to obtain the optimal time allocation ratio τ of the downlink WET phase within consecutive time slots t u,t and the energy ratio ρ reserved for the uplink WIT phase u,t .
Citation Information
Patent Citations
A resource allocation method based on energy accumulation in cognitive wireless power supply network
CN109275149A
Multiplexing ul transmissions with different reliabilities
CN111903175A