A cognitive wireless non-orthogonal multiple access network resource allocation method

CN116723580BActive Publication Date: 2026-09-25CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310565664.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-18
Publication Date
2026-09-25
Estimated Expiration
2043-05-18

AI Technical Summary

Technical Problem

[0005]将CRN技术和EH技术结合起来具有重要的意义,提高频谱效率的同时还延长了网络设备的生命周期,EH-CRN也是近年来的一个研究热点,然而,在EH-CRN通信网络中资源分配问题是非线性的非凸优化问题,需要在不确定和复杂的通信环境中对网络的一系列参数进行调整,对于传统的基于凸优化或启发式的方法来说计算复杂且难以求解

Benefits of technology

[0033]本发明的有益效果为:本发明基于深度强化学习(DRL)方法,对多个次用户下的EH-CRN-NOMA通信系统中的资源分配进行了研究,简化了目标函数的优化变量,从而减少智能体的动作空间的维度来加速算法的收敛,同时设计合适的奖励函数,让收敛后的算法执行最优的策略来适应动态变化的信道环境和逼近网络场景下的优化目标,从而克服无线通信网络中的频谱资源短缺和能量受限问题;与现有技术相比,本发明显著提高了系统吞吐量以及收敛速度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116723580B_ABST
    Figure CN116723580B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of wireless communication, and particularly relates to a cognitive wireless non-orthogonal multiple access network resource allocation method; the method comprises the following steps: constructing an EH-CRN-NOMA system model; constructing a maximum secondary user total throughput optimization problem according to the EH-CRN-NOMA system model; introducing an auxiliary variable to simplify optimization variables of the maximum secondary user total throughput optimization problem, and obtaining simplified optimization variables; calculating a reward function, and solving the optimization problem by using a deep reinforcement learning algorithm according to the reward function and the simplified optimization variables, and obtaining a resource allocation scheme; simulation results show that the system throughput and the convergence speed of the algorithm can be superior to those of a comparative algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication technology, specifically relating to a method for resource allocation in cognitive wireless non-orthogonal multiple access networks. Background Technology

[0002] With the continuous upgrading and iteration of information and communication technologies, the development of wireless local area networks (WLANs), the Internet of Things (IoT), and 5G mobile communications has been promoted, and more and more communication devices are accessing the Internet through wireless technology. However, the demand for radio spectrum from various wireless communication devices is also increasing dramatically. Currently, most radio systems adopt the traditional fixed spectrum allocation scheme, that is, fixed spectrum resources are allocated to licensed users, but the spectrum usage time of these users is intermittent rather than continuous, resulting in some spectrum resources not being fully utilized. Against this background, Cognitive Radio Network (CRN) technology has been proposed to solve the problem of spectrum resource shortage. CRN spectrum allocation is not static; users without spectrum licenses can also access the channel. In CRN, users are divided into primary users (PUs) and secondary users (SUs) based on whether they have obtained spectrum usage rights. Based on the access method of SUs during spectrum sharing, they can be divided into overlay mode and underlay mode. Overlay mode refers to the SU and PU not transmitting at the same time. The SU accesses the idle spectrum of the PU, reducing the idle time of the spectrum and thus improving the spectrum utilization. The core idea of ​​Underlay mode is: the SU and PU transmit at the same time without causing harmful interference to the transmission of the PU.

[0003] Currently, with the increasing demand for high-speed and low-latency wireless communication applications and ubiquitous mobile services, the power requirements for wireless devices are rising, making power supply a bottleneck in the lifecycle of wireless network equipment. Traditional battery-powered systems have many problems. For example, disposable batteries, while inexpensive, have limited lifespans and require periodic replacement, and they can cause environmental pollution, especially for equipment deployed in remote or harsh environments where battery replacement or charging may be difficult or even impossible. In large-scale networks, regularly replacing or charging batteries for wireless devices consumes significant time and effort and contradicts green communication and carbon neutrality initiatives. More seriously, untimely power supply can cause communication interruptions, impacting user experience. Therefore, finding a new power supply method is essential. Furthermore, because the SUs in a CRN (Content Regulator) continuously perform tasks such as spectrum sensing, CRN systems face more severe energy consumption problems compared to traditional wireless communication systems.

[0004] In recent years, energy harvesting (EH) technology has received widespread attention in the communications field. EH offers a new solution for energy-constrained or high-energy-consuming communication systems. This technology can harvest energy not only from environmental energy sources such as solar, thermal, and wind power, but also from radio frequency (RF) signals emitted by other devices, and then provide the extracted energy to energy-constrained devices. Because environmental energy sources are greatly affected by external conditions and exhibit strong randomness, while RF signals are relatively more reliable and stable, energy harvesting from RF has received more attention and research.

[0005] Combining CRN and EH technologies is of great significance, improving spectral efficiency while extending the lifespan of network devices. EH-CRN has also been a research hotspot in recent years. However, resource allocation in EH-CRN communication networks is a nonlinear, nonconvex optimization problem, requiring adjustments to a series of network parameters in uncertain and complex communication environments. Traditional methods based on convex optimization or heuristics are computationally complex and difficult to solve. Therefore, a resource allocation method that can overcome the problems of spectrum resource scarcity and energy constraints in wireless communication networks is urgently needed. Summary of the Invention

[0006] Many existing studies have investigated EH-CRN-NOMA scenarios. However, unlike this invention, these studies involve multiple user units (PUs) but a single number of service units (SUs). PUs and SUs are cyclically scheduled to communicate within different time slots. In contrast, the proposed system only has a single PU, SU, and service base (BS) communicating within any given time slot, equivalent to a scenario with drastically changing channel coefficients for a single PU, SU, and BS. The application of NOMA technology is not fully realized in these studies. To accelerate algorithm solving, many studies have used distributed multi-agent reinforcement learning frameworks in multi-user NOMA scenarios to improve system performance. This approach reduces the burden on a single agent through cooperation among multiple agents, thus adapting to complex networks. However, this method requires communication between agents, increasing communication overhead and potentially leading to network congestion and channel interference.

[0007] To address the shortcomings of existing technologies, this invention proposes a resource allocation method for cognitive wireless non-orthogonal multiple access networks, which includes:

[0008] S1: Construct the EH-CRN-NOMA system model;

[0009] S2: Construct an optimization problem to maximize the total throughput of secondary users based on the EH-CRN-NOMA system model;

[0010] S3: Introduce auxiliary variables to simplify the optimization variables of the problem of maximizing the total throughput of secondary users, and obtain the simplified optimization variables;

[0011] S4: Calculate the reward function, and use a deep reinforcement learning algorithm to solve the optimization problem based on the reward function and the simplified optimization variables to obtain the resource allocation scheme.

[0012] Preferably, the EH-CRN-NOMA system model includes one primary user (PU), M secondary users (SU), and one base station (BS); the m-th secondary user is represented as SU. m (1≤m≤M), its energy is limited and it accesses the PU's spectrum resources in Underlay mode; within time slot t, SU m First in α m (t) Data is transmitted to the BS within time T, where α m (t) represents the time allocation coefficient for the m-th user, 0 ≤ α m (t)≤1, where T represents the length of each time slot; the remaining (1-α) m During time interval (t)T, SU m It charges its own battery by absorbing the radio frequency signal of the PU.

[0013] Preferably, the process of constructing the optimization problem to maximize total throughput of secondary users includes:

[0014] S21: Set the decoding order and calculate the interference experienced by the secondary user during decoding and the interference experienced by the primary user during decoding; calculate the transmission rate of the secondary user based on the interference experienced by the secondary user, and calculate the transmission rate of the primary user based on the interference experienced by the primary user; sum the transmission rates of all secondary users to obtain the objective function for maximizing the total throughput of the secondary users in the optimization problem.

[0015] S22: Set time allocation coefficient constraints, calculate power constraints for secondary users, and set transmission power constraints based on time allocation coefficients and power constraints for secondary users; set minimum transmission rate constraints for primary users and minimum transmission rate constraints for secondary users based on the transmission rates of secondary users and primary users, respectively.

[0016] Furthermore, the decoding order is as follows: first, the secondary user signals are decoded according to signal strength among all secondary users; after decoding the secondary user signals, the primary user signals are decoded last.

[0017] Furthermore, the formula for calculating the transmission rate of the secondary user is:

[0018]

[0019] in, This represents the transmission rate of the m-th secondary user within time slot t. g represents the transmit power of the m-th secondary user within time slot t. m (t) represents the channel gain coefficient from the m-th secondary user to the base station within time slot t. σ represents the interference experienced by the m-th secondary user during decoding within time slot t. 2 This indicates environmental noise.

[0020] The preferred optimization problem for maximizing the total throughput of secondary users is expressed as:

[0021]

[0022] stC1: 0 ≤ α m (t)≤1

[0023]

[0024]

[0025]

[0026] Where, α m (t) represents the time allocation coefficient for the m-th user within time slot t. This represents the transmission power of the m-th secondary user within time slot t. P represents the transmission rate of the m-th secondary user within time slot t, where M represents the number of secondary users. max E represents the maximum transmit power of the secondary user. m (t-1) represents the electricity consumption of the m-th user at the end of time slot t-1, where T represents the time slot length, and E represents the electricity consumption of the m-th user. m (t) represents the electricity consumption of the m-th user at the end of time slot t, E max This indicates the maximum battery level for the next user. R represents the energy collected by the m-th secondary user within time slot t. p (t) represents the primary user's transmission rate within time slot t, R sth R represents the minimum transmission rate for the secondary user. pth This indicates the minimum transmission rate for the primary user.

[0027] The preferred, simplified expression for the optimization variables is as follows:

[0028] E m,c (t)=β m (t)min{E max -E m (t-1), T η P p |H m (t)| 2}-(1-β m (t))min{E m(t-1), TP max}

[0029] Among them, E m,c (t) represents the simplified optimization variable, β m (t) represents an auxiliary variable, E max E represents the maximum battery capacity of the secondary user. m (t-1) represents the energy consumption of the m-th secondary user at the end of time slot t-1, T represents the time slot length, and P represents the energy harvesting efficiency. p H represents the primary user's transmit power. m (t) represents the channel gain coefficient from the primary user to the secondary user, P max This indicates the maximum transmit power for the secondary user.

[0030] Preferably, the formula for calculating the reward function is:

[0031]

[0032] Where, r t r represents the reward value within time slot t. t1 r represents the throughput of successful transmissions by secondary users. t2 The throughput of secondary user transmission failures is represented by B, bandwidth is represented by M, number of secondary users is represented by ω, and indicator factor is represented by R. p (t) represents the transmission rate of the primary user within time slot t.

[0033] The beneficial effects of this invention are as follows: Based on the Deep Reinforcement Learning (DRL) method, this invention studies resource allocation in an EH-CRN-NOMA communication system with multiple sub-users, simplifies the optimization variables of the objective function, thereby reducing the dimension of the agent's action space to accelerate algorithm convergence. At the same time, a suitable reward function is designed to allow the converged algorithm to execute the optimal strategy to adapt to the dynamically changing channel environment and approximate the optimization objective in the network scenario, thereby overcoming the problems of spectrum resource shortage and energy limitation in wireless communication networks. Compared with the prior art, this invention significantly improves system throughput and convergence speed. Attached Figure Description

[0034] Figure 1 This is a flowchart of the cognitive wireless non-orthogonal multiple access network resource allocation method in this invention;

[0035] Figure 2 This is a schematic diagram of the EH-CRN-NOMA system model in this invention;

[0036] Figure 3 This is a comparison chart of the rewards obtained by the PPO(β) algorithm in this invention when M = (2,10);

[0037] Figure 4 To compare the algorithm PPO(α,P) s Comparison chart of rewards obtained when M=(8,10);

[0038] Figure 5 This is a comparison chart of the throughput of successful SUs transmissions in the PPO(β) algorithm of this invention when M = (2,6,10);

[0039] Figure 6 To compare the algorithm PPO(α,P) s A comparison of the throughput of successful transmissions of SUS when M = (2,6,10). Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] This invention proposes a resource allocation method for cognitive wireless non-orthogonal multiple access networks, such as... Figure 1 As shown, the method includes the following:

[0042] S1: Construct the EH-CRN-NOMA system model.

[0043] like Figure 2 As shown, an EH-CRN-NOMA (Energy Harvesting-Cognitive Radio-Non-Orthogonal Multiple Access Network) system model is constructed, including one primary user PU, M secondary users SU, and one base station BS. Each device in the network is equipped with a single antenna, and the m-th secondary user SU is denoted as SU. m (1≤m≤M), the secondary user's spectrum resources are limited by energy and access the PU in Underlay mode. The PU continuously transmits data to the BS in each time slot. m The PU is charged by collecting radio frequency energy using EH technology, and then the collected energy is used to transmit data to the BS. In order to better perform EH and avoid damaging interference to the PU, this invention uses SU m It is deployed in a location that is close to the PU but far from the BS.

[0044] The system includes two types of links: one is the energy harvesting link from the PU to the SUs, and the other includes the link from the PU to the BS and the SUs. m Data transmission link to BS. In this invention, PU is marked as SU. m The channel gain coefficient is H m(t), the channel gain coefficient from PU to BS is h(t), SU m The channel gain coefficient to BS is g m (t). Both of the above channels take into account large-scale path loss and small-scale multipath fading, and assume that they follow block fading, remain constant within each time slot, and change between different time slots.

[0045] At the beginning of each time slot, it is assumed that each SU m Knowing the channel state information of the PU and itself, this condition can be satisfied by broadcasting one or more pilot signals in the channel. SUs all adopt a protocol of transmitting first and then collecting, that is, within time slot t, the SU... m First in α m (t) Data is transmitted to the BS within time T, where α m (t) represents the time allocation coefficient for the m-th user, 0 ≤ α m (t)≤1, where T represents the length of each time slot; the remaining (1-α) m During time interval (t)T, SU m It charges its own battery by absorbing the radio frequency signal of the PU.

[0046] S2: Construct an optimization problem to maximize the total throughput of secondary users based on the EH-CRN-NOMA system model.

[0047] S21: Set the decoding order and calculate the interference experienced by the secondary user during decoding and the interference experienced by the primary user during decoding; calculate the transmission rate of the secondary user based on the interference experienced by the secondary user, and calculate the transmission rate of the primary user based on the interference experienced by the primary user; sum the transmission rates of all secondary users to obtain the objective function for maximizing the total throughput of the secondary users.

[0048] Let the initial energy of SUs be E0, and the upper limit of the battery be E. max At the end of time slot t-1, SU will m The charge level is marked as E m (t-1). Within time slot t, SU m The energy consumed is And meet the restrictions in, SU m The transmission power, in the remaining (1-α) m During time interval (t)T, SU m The energy collected is:

[0049]

[0050] Among them, P pη represents the transmission power of PU; η represents the energy harvesting efficiency. Therefore, at the end of time slot t, SU m The amount of electricity E m (t) can be represented as:

[0051]

[0052] Since all SUs and PUs access the same uplink channel via CRN-NOMA technology, the signal received at BS at time t can be expressed as:

[0053]

[0054] Where, x p (t) and Let P represent the message sent by PU and the m-th SU, respectively. p h represents the transmission power of PU. t denoted by PU to BS; z represents additive white Gaussian noise in the environment.

[0055] The combination of CRN and NOMA technologies allows the interference threshold problem of SU to PU in the original Underlay mode to be solved by Continuous Interference Cancellation (SIC). In this system model, the decoding order of SIC can be determined based on QoS order, that is, the signals of unlicensed SUs are decoded first, and the signals of licensed PUs are decoded last. In this case, if all SUs are decoded successfully, the PU will not be interfered with when decoding last; otherwise, the PU will be interfered with by the signals of SUs that failed to decode. The decoding order can also be divided according to the signal strength at the signal receiver, that is, the BS determines the decoding order based on the product of the received transmit power and the channel power gain. The stronger the signal, the stronger the interference to other signals during decoding, so it is decoded first. This invention takes into account the advantages of the above two decoding methods and combines the two decoding orders, that is, first decode the SUs according to signal strength among all SUs. m After decoding the SUs signal, the PU signal is decoded last. The above decoding process takes into account the failure case. When the m-th SU is decoded, the interference it receives is:

[0056]

[0057] in, This represents the interference experienced by the m-th secondary user during decoding within time slot t; Indicates the decoding order, if SU n The signal strength is greater than the SU currently being decoded. m If the strength is high, then otherwise This refers to the secondary user decoding coefficients, specifically SU. n Was decoding successful? If successful. otherwise

[0058] When decoding the signal of the PU, it is subjected to interference I. p for:

[0059]

[0060] in, This refers to the primary user decoding coefficients, specifically SU. m Was decoding successful? If successful. otherwise

[0061] The transmission rate of the secondary user is calculated based on the interference experienced by the secondary user, and is expressed as follows:

[0062]

[0063] The transmission rate of the primary user is calculated based on the interference experienced by the primary user, and is expressed as:

[0064]

[0065] The objective function for the optimization problem of maximizing total throughput of secondary users is:

[0066]

[0067] S22: Set time allocation coefficient constraints, calculate power constraints for secondary users, and set transmission power constraints based on time allocation coefficients and power constraints for secondary users; set minimum transmission rate constraints for primary users and minimum transmission rate constraints for secondary users based on the transmission rates of secondary users and primary users, respectively.

[0068] The time allocation coefficient constraint is: 0 ≤ α m (t)≤1

[0069] The power limit for secondary users is:

[0070] The transmit power constraint is:

[0071] The QoS requirements of SUs prevent PUs from being subjected to harmful interference from SUs. This invention provides all SUs with... m Set a minimum rate threshold for PU, denoted as R. sth and R pth If SU m R was not satisfied sth The requirement is that it is considered to have failed to decode during the SIC process, causing... If all SUs can be successfully decoded, then the PU will not experience interference from SUs during the final transmission, and the PU will have its own dedicated channel during decoding. pth The introduction of this constraint is to ensure that the PU is not subjected to harmful interference in the worst-case scenario, i.e., the failure of all SUs to decode is considered interference; the minimum transmission rate constraints for the primary user and the secondary user are expressed as follows:

[0072]

[0073] R p (t)≥R pth #(10)

[0074] Therefore, the optimization problem of maximizing the total throughput of secondary users can be expressed as:

[0075]

[0076] Among them, constraint C2 guarantees SU m Transmission power It will not exceed the maximum power P of SUs max This also ensures SU m The energy used to transmit data will not exceed the currently available energy. When energy is insufficient, the solution is to maintain the transmission time while adjusting the SU (Supply Unit). m C3 is to ensure that the transmitted power meets the energy constraint requirements; C4 is to ensure that the collected power does not exceed the battery limit; C5 is the rate requirement for SUs and PU.

[0077] S3: Introduce auxiliary variables to simplify the optimization variables of the problem of maximizing the total throughput of secondary users, and obtain the simplified optimization variables.

[0078] As can be seen from formula (11), the optimization problem becomes increasingly complex with the increase of the number of M variables. Considering its non-convexity and complexity, this invention adopts a centralized DRL algorithm to seek the optimal strategy. First, in order to reduce the burden on the agent under centralized training, the optimization variables of the objective function are simplified, thereby simplifying the action space of the DRL algorithm; specifically:

[0079] The NOMA system in this invention has multiple SUs. Solving for the optimal control strategy for all SUs using a centralized approach places high demands on the individual agent. Therefore, this invention simplifies the optimization variables in formula (8) to reduce their dimensionality, thereby alleviating the burden on the agent. In formula (8), two variables need to be optimized, namely... Furthermore, the ranges of their values ​​may differ; the former is [0, 1], while the latter is [0, P]. max Therefore, a new parameter E is introduced. m,c (t) to replace αm (t) and E m,c (t) represents SU within a time slot m The change in battery energy, i.e., the collected portion minus the consumed portion, is defined as follows:

[0080]

[0081] Among them, E m,c (t) can be positive or negative, depending on the relative magnitude of energy collection and consumption within the time slot. Within time slot t, its value range is as follows:

[0082] -min{E m (t-1), TP max}≤E m,c (t)≤min{E max -E m (t-1), TηP p |H m (t)| 2}#(13)

[0083] It can be seen that its value varies over a large range: if SU m Transmitting data without collecting energy consumes the minimum energy on the left side of the inequality; if SU m It only collects energy and does not transmit data; the collected energy is the minimum value on the right side of the inequality. If the action output space is defined as E... m,c While reducing the dimensionality of (t), its dynamic range may decrease the stability of the DRL algorithm. Therefore, to limit it to a smaller range, an auxiliary variable β is introduced. m (t)(0≤β m (t)≤1), use β m (t) and 1-β m (t) represent the proportions of the left and right sides in inequality (10), respectively. m,c (t) is derived from β m The representation of (t) is as follows:

[0084]

[0085] By introducing auxiliary variables, the action space dimension of the DRL algorithm is increased from the original... That is, 1×2M decreased to [β] m (t), ...] that is, 1×M, and its range is between [0, 1], which is beneficial to the convergence and stability of the DRL algorithm.

[0086] S4: Calculate the reward function, and use a deep reinforcement learning algorithm to solve the optimization problem based on the reward function and the simplified optimization variables to obtain the resource allocation scheme.

[0087] Having simplified the action space as described above, we will now discuss the Markov Decision Process (MDP) during the training of the system model. The quintuple of an MDP consists of state S, action A, state transition probability r, and discount factor. The following section mainly introduces the settings of S, A, and the reward function r in this algorithm, specifically:

[0088] (1) State Space S: The algorithm proposed in this invention needs to control all SUs. The agent should have global information to make a globally optimal strategy. At time slot t, the state space S should contain the channel gain coefficient H of all SUs to PU links. m (t)(i.e., H) m (t)=[H1(t),...,H m (t)]), the channel gain coefficient g of all SUs to BS links m (t), the channel gain h(t) from PU to BS, and the charge E of all SUs at the end of the previous time slot. m (t-1), i.e., s t =[H m (t), g m (t), h(t), E m (t-1)].

[0089] (2) Action space A: Since a new variable E is introduced in formulas (10) and (11) m,c (t) and auxiliary variable β m (t), the action space has changed from the original data transmission time ratio and transmission power magnitude to the current auxiliary variable, namely a. t = [β1(t), ..., β m (t)].

[0090] (3) Reward r: When designing the reward function of this algorithm, the first consideration was maximizing the throughput of SUs, so This needs to be included in the reward. Additionally, a rate threshold is set for both SUs and PU to ensure the QoS requirements of SUs and the normal transmission of PU in the worst-case scenario. The reward should also consider whether the threshold rate requirement is met. Therefore, the reward function in time slot t mainly consists of three parts:

[0091]

[0092]

[0093]

[0094] Wherein, formula (15) represents SU m transmission rate Meet threshold requirement R sth The throughput of a successful transmission; Formula (16) represents SU m The throughput of transmission failure due to the rate not meeting the rate requirement, for the m-th user, it belongs to only one of formulas (15) and (16); ω in formula (17) is an indicator factor, when the rate of PU meets the threshold requirement R pth The value is 0 when the condition is met and -1 when the condition is not met. This means that successful PU transmission receives neither a reward nor a penalty, while a penalty is applied for failure. Considering all three factors simultaneously, the final reward function of this invention is formed:

[0095] r t =fun1(r t1 )-fun2(r t2 )+ωfun3(R p (t)) #(18)

[0096] For the reward function, fun1, fun2, and fun3 are set as follows:

[0097]

[0098]

[0099] fun3(R p (t))=log 10 (1+R p (t)) #(21)

[0100] In the system of this invention, since there are multiple SUs, the total throughput calculated after each interaction with the environment varies greatly. If throughput is directly used as the reward, the algorithm may fail to converge due to excessive fluctuations, or the agent may stagnate due to insufficient rewards. Therefore, three functions fun1, fun2, and fun3 are introduced to adjust the throughput of the system. t1 r t2 and R p The algorithm processes (t) to ensure that the reward fluctuates within a reasonable range, eventually leading to convergence.

[0101] It is worth noting that during the operation of the NOMA system proposed in this invention, SUs accessing the same channel m The data transmission times may vary, making it complex to calculate the total throughput of all SUs based on actual conditions, because SUs... m The transmission times are different; the later the transmission end time, the better. mIts signal-to-noise ratio not only changes frequently but also varies greatly. Considering extreme cases, the signal strength of SU is the highest. m Since the data transmission time is the longest, when calculating its throughput, it is necessary to consider the situation where the signal-to-noise ratio changes due to the termination of transmission and exit of other M-1 SUs, making the throughput calculation segmented. In addition, when calculating the reward, whether the SUs meet the minimum rate requirement is also dynamically changing. The signal-to-noise ratio of the SU that has failed to transmit but has not yet finished will decrease as the transmission of other SUs ends, thereby increasing the rate. This may allow the previously failed SUs to meet the rate requirement again.

[0102] For the two complex solution scenarios mentioned above, and considering the short data transmission intervals, the following simplifications were made when calculating throughput and rewards: Within one time slot length, SU m The signal-to-noise ratio will not be affected by other SUs m The effect of ending the transmission is always equal to the value at the time of SIC decoding.

[0103] The optimization problem is solved using a deep reinforcement learning algorithm based on the reward function and the simplified optimization variables. This yields the resource allocation scheme, which includes the time allocation coefficient and transmission power for each user.

[0104] Evaluation of the present invention:

[0105] The algorithm of this invention is denoted as PPO(β). An algorithm that directly solves for the optimization objective without introducing auxiliary variables is used as a comparison algorithm and denoted as PPO(α, P...). s Apart from the change in action space, all other parameters remain unchanged in both algorithms.

[0106] Figure 3 This indicates the convergence of PPO(β) when M=2 and M=10. Figure 4 PPO(α, P) represents the values ​​of M=8 and M=10. s The convergence of ). From Figure 3 As can be seen, when M=2, the PPO(β) algorithm converges in about 30 rounds, while when M=10, it converges in about 25 rounds, even faster than when M=2. This indicates that the convergence speed of the PPO(β) algorithm is not affected by the number of M. Figure 4 It can be seen that the algorithm converges faster when M=8 than when M=10, with the former reaching convergence in around 60 rounds and the latter around 90 rounds. This indicates that PPO(α, P s The convergence speed of the algorithm is affected by the number of SUs. When M is smaller, the algorithm needs to control fewer parameters and can find the optimal control strategy for the current environment more quickly. (Comparison) Figure 3 and Figure 4 It can be seen that when M=10, the convergence speed of the PPO(β) algorithm is relatively faster than that of the PPO(α, P) algorithm. s The former algorithm can converge very quickly.

[0107] Figure 5 and Figure 6 The PPO(β) algorithm and PPO(α, P) under different M values. s The throughput change graph obtained by the algorithm in each round is worth noting. It is worth noting that the total throughput statistics in this graph only include the successfully transmitted portion. In this part, the throughput of decoding failures is not included. As can be seen, Figure 5 and Figure 6 The throughput in both algorithms increases with increasing M, but the rate of increase differs, with the former showing a significantly greater increase than the latter. When M = 2, the total throughput of SUs in both algorithms is roughly the same. When the number of SUs M = 10, the PPO(β) algorithm achieves a throughput of 45 (bps / Hz), while the PPO(α, P) algorithm achieves a throughput of 45 (bps / Hz). s The algorithm can only achieve a throughput of 25 (bps / Hz), while the former is almost twice that of the latter.

[0108] In summary, simulation results show that the algorithm proposed in this invention achieves better system throughput and convergence speed than the comparative algorithms.

[0109] The above-described embodiments further illustrate the purpose, technical solution, and advantages of the present invention. It should be understood that the above-described embodiments are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made to the present invention within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for resource allocation in a cognitive wireless non-orthogonal multiple access network, characterized in that, include: S1: Construct the EH-CRN-NOMA system model; the EH-CRN-NOMA system model includes one primary user (PU), M secondary users (SU), and one base station (BS); the m-th secondary user is represented as SU. Its energy is limited and it accesses the PU's spectrum resources in Underlay mode; in time slots Inside, SU First Data is transmitted to the BS within a time limit, where, This represents the time allocation coefficient for the m-th user. T represents the length of each time slot; the remaining Within a time period, SU It charges its own battery by absorbing the radio frequency signal of the PU; S2: Construct an optimization problem to maximize the total throughput of secondary users based on the EH-CRN-NOMA system model; the process of constructing the optimization problem to maximize the total throughput of secondary users includes: S21: Set the decoding order and calculate the interference experienced by the secondary user during decoding and the interference experienced by the primary user during decoding; calculate the transmission rate of the secondary user based on the interference experienced by the secondary user, and calculate the transmission rate of the primary user based on the interference experienced by the primary user; sum the transmission rates of all secondary users to obtain the objective function for maximizing the total throughput of the secondary users in the optimization problem. S22: Set time allocation coefficient constraints, calculate power constraints for secondary users, set transmission power constraints based on time allocation coefficients and power constraints for secondary users; set minimum transmission rate constraints for primary users and minimum transmission rate constraints for secondary users based on transmission rates for secondary users and primary users, respectively. S3: Introduce auxiliary variables to simplify the optimization variables in the problem of maximizing the total throughput of secondary users, resulting in simplified optimization variables; the simplified variables are expressed as: ; in, Represents the simplified optimization variables. Represents auxiliary variables. This indicates the maximum battery level for the next user. This represents the battery level of the m-th user at the end of time slot t-1. Indicates the time slot length. Indicates the efficiency of energy harvesting. Indicates the primary user's transmit power. This represents the channel gain coefficient from the primary user to the secondary user. This indicates the maximum transmit power of the secondary user; S4: Calculate the reward function. Based on the reward function and the simplified optimization variables, use a deep reinforcement learning algorithm to solve the optimization problem and obtain the resource allocation scheme. The formula for calculating the reward function is: ; in, This represents the reward value within time slot t. This represents the throughput of successful transmissions by the secondary user. This represents the throughput of secondary user transmission failures. Indicates bandwidth. Indicates the number of secondary users. Indicates the indicator factor. This represents the transmission rate of the primary user within time slot t.

2. The method for resource allocation in a cognitive wireless non-orthogonal multiple access network according to claim 1, characterized in that, The decoding order is as follows: first, decode the secondary user signals according to their signal strength among all secondary users; after decoding the secondary user signals, finally decode the primary user signals.

3. The method for resource allocation in a cognitive wireless non-orthogonal multiple access network according to claim 1, characterized in that, The formula for calculating the transmission rate of a secondary user is: ; in, This represents the transmission rate of the m-th secondary user within time slot t. This represents the transmission power of the m-th secondary user within time slot t. This represents the channel gain coefficient from the m-th secondary user to the base station within time slot t. This represents the interference experienced by the m-th secondary user during decoding within time slot t. This indicates environmental noise.

4. A method for resource allocation in a cognitive wireless non-orthogonal multiple access network according to claim 1, characterized in that, The optimization problem of maximizing total throughput for secondary users can be expressed as: ; in, This represents the time allocation coefficient for the m-th user within time slot t. This represents the transmission power of the m-th secondary user within time slot t. This represents the transmission rate of the m-th secondary user within time slot t. Indicates the number of secondary users. This indicates the maximum transmit power of the secondary user. This represents the battery level of the m-th user at the end of time slot t-1. Indicates the time slot length. This represents the battery level of the m-th user at the end of time slot t. This indicates the maximum battery level for the next user. This represents the energy collected by the m-th user within time slot t. This represents the transmission rate of the primary user within time slot t. This indicates the minimum transmission rate for the secondary user. This indicates the minimum transmission rate for the primary user.

Citation Information

Patent Citations

  • Cognitive radio energy efficiency resource allocation method based on non-orthogonal multiple access

    CN110337148A

  • Deep reinforcement learning user resource allocation method in downlink multi-carrier NOMA scene

    CN115589637A