Energy harvesting cognitive IoT resource allocation method based on deep reinforcement learning

Through the deep reinforcement learning method, the shortage of spectrum resources and communication security problems are solved, and the optimal allocation of IoT resources and communication security enhancement of energy acquisition cognitive is achieved.

CN115712497BActive Publication Date: 2025-08-08FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211278767.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-19
Publication Date
2025-08-08
Estimated Expiration
2042-10-19

AI Technical Summary

Technical Problem

Spectral resources shortage and communication security are important factors that restrict the development of the Internet of Things. Traditional battery power supply increases maintenance costs and is not environmentally friendly, and physical layer security technology is difficult to cope with the security challenges brought about by the increase in computing power.

Method used

The energy acquisition cognitive IoT resource allocation method based on deep reinforcement learning is adopted. By building an energy acquisition cognitive IoT system model, a reinforcement learning model is built, and the links of the sub-transmitter, collaborative jammer and eavesdropping nodes are modeled as agents to perform the optimal allocation of joint energy acquisition time and transmission power.

Benefits of technology

The optimal allocation of IoT resources for energy acquisition cognitive is achieved, communication security performance is enhanced, network changes are adapted to network changes, and confidentiality rate is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115712497B_ABST
    Figure CN115712497B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for allocating energy harvesting cognitive Internet of Things resources based on deep reinforcement learning. The method comprises the following steps: building an energy harvesting cognitive Internet of Things system model and deriving a mathematical model for resource allocation; building a reinforcement learning model to model 2m sub-channels of two links, from a secondary transmitter to a secondary receiver and from a collaborative jammer to an eavesdropping node, and an energy harvesting time allocation network, as 2m+1 reinforcement learning agents, with the remaining parts of the energy harvesting cognitive Internet of Things serving as a reinforcement learning environment, with the agents and the environment continuously interacting; constructing an energy harvesting cognitive Internet of Things resource allocation model based on deep reinforcement learning and training the model; and optimally allocating the energy harvesting time and transmission power for the cognitive Internet of Things using the trained resource allocation model. This method is advantageous for optimally allocating resources for the energy harvesting cognitive Internet of Things.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of wireless communications, and in particular relates to an energy harvesting cognitive Internet of Things resource allocation method based on deep reinforcement learning. Background Art

[0002] With the development of modern wireless communication technology, the Internet of Things (IoT) has become a new paradigm that connects a large number of IoT devices to meet daily service needs. These IoT devices require a large amount of spectrum resources, but the shortage of spectrum resources has become a major factor restricting the future development of the IoT.

[0003] Cognitive radio technology has emerged to address spectrum resource shortages. By sensing the external environment, intelligently adjusting internal system parameters, and discovering and utilizing spectrum holes, cognitive radio achieves efficient spectrum utilization, thereby improving daily service needs. Therefore, applying cognitive radio technology to the Internet of Things (IoT) is a promising solution. The cognitive Internet of Things (CIoT) combines cognitive radio technology with the IoT to address spectrum resource shortages.

[0004] Despite the aforementioned advantages of the cognitive IoT, IoT devices are often powered by traditional batteries. The increasing number of batteries not only increases maintenance costs, but also pollutes the environment as discarded batteries. For large-scale cognitive IoT deployments, frequent battery replacement is not only time-consuming and labor-intensive, but also reduces network communication efficiency. To address this challenge, energy harvesting technology has been applied to the cognitive IoT. Energy harvesting converts natural energy sources such as wind, thermal, solar, and radio frequency energy into electrical energy to power IoT devices, eliminating the need for external battery capacity limitations. Natural energy sources are unstable, and the harvesting process requires large and complex power equipment. However, radio frequency energy signals can be simply received by an antenna. Therefore, considering using radio frequency energy harvesting technology to power IoT devices not only addresses energy efficiency issues but also meets the requirements of green communications.

[0005] While increasing scientific and technological productivity has driven the development of various wireless communication networks, it has also increased the security complexity of these networks. In the current information age, communication between IoT devices has become increasingly frequent within the cognitive Internet of Things (IoT), posing a significant challenge to the security of these networks. The broadcast nature of wireless channels provides opportunities for eavesdropping nodes to access the network and steal confidential information. Traditionally, cryptographic encryption methods have been employed at the upper layers of the network protocol stack to enhance network communication security. However, increased computing power has made these encryption methods vulnerable to cracking by unauthorized users, posing a significant security challenge for the IoT, which stores vast amounts of communication resources. To further enhance network communication security, physical layer security technologies have emerged as a complementary approach to traditional cryptographic encryption. These technologies, including beamforming, artificial noise, and cooperative jamming, primarily exploit the inherent characteristics of wireless channels, such as fading, noise, and interference, to prevent eavesdropping. These methods can guarantee absolute information-theoretical security without requiring the computing power of eavesdropping nodes. Therefore, applying physical layer security technologies to the cognitive IoT is a viable solution for enhancing communication security. Summary of the Invention

[0006] The purpose of the present invention is to provide an energy harvesting cognitive Internet of Things resource allocation method based on deep reinforcement learning, which is conducive to the optimal allocation of energy harvesting cognitive Internet of Things resources.

[0007] To achieve the above objectives, the present invention adopts a technical solution: a method for allocating energy harvesting cognitive IoT resources based on deep reinforcement learning, comprising:

[0008] Build an energy harvesting cognitive IoT system model and derive a mathematical model for resource allocation;

[0009] A reinforcement learning model is built, modeling the 2m sub-channels of the two links (from the secondary transmitter to the secondary receiver and from the cooperative jammer to the eavesdropping node) and an energy harvesting time allocation network t0 as 2m+1 reinforcement learning agents. The rest of the energy harvesting cognitive Internet of Things is the reinforcement learning environment, and the agents continuously interact with the environment.

[0010] Build and train a deep reinforcement learning-based energy harvesting cognitive IoT resource allocation model;

[0011] The trained resource allocation model is used to optimally allocate the energy collection time and transmission power for the cognitive Internet of Things.

[0012] Furthermore, the energy harvesting cognitive IoT resource allocation model based on deep reinforcement learning is trained, which specifically includes the following steps:

[0013] S1. Generate the topology of the cognitive Internet of Things, initialize the channel gain of each link, the number of rounds of training N, and the experience buffer pool D k The maximum capacity N k , and the decision network and target network weight parameters θ k 、 in

[0014] S2. At the beginning of each training round, the positions of all nodes in the cognitive IoT are randomly initialized, the channel gain of each link is updated, and the initial state of the environment is set to S0;

[0015] S3, in each training round t = 0, 1, 2, ..., T max -1 time step, based on the current environment state S t , each agent k obtains local observations of the environment And take action according to the ε-greedy algorithm Updated battery capacity for ST and J and Update channel gain, current environment state S t Transition to the next state S t+1 , agent k obtains the next local observation and reward r t ;

[0016] S4. Update the neural network parameters θ k and That is, from D k Randomly extract a set batch of samples and send them to the decision network to calculate the loss function L(θ k ), and perform gradient descent to minimize L(θ k ) Update the parameters θ k ; Every M consecutive time steps, θ k Copy to the target network weight parameters

[0017] Furthermore, in step S2, at the beginning of each training round, a randomization scheme is used to update the positions of all nodes in the cognitive Internet of Things. The update of the channel gain of each link follows the Rayleigh channel fading model, and the initial state of the environment for this training round is set to:

[0018] S0=S r | t=0 ={G t , SINR t , B t-1} t=0 ={G0, SINR0, {B ST,max , BJ,max}}

[0019] Where, The signal-to-interference-noise ratio set at PR, SR, and E ST and J battery capacity collection And there is B -1 ={B ST,max , B J,max},in, is a set of subchannel numbers, is the subchannel gain, are the transmission powers of the secondary transmitter ST and the cooperative interferer J in the kth subchannel, B ST,max and B J,max The maximum battery capacities of ST and J respectively.

[0020] Furthermore, in step S3, at t=0, 1, 2, ..., T in each training round, max -1 time step, based on the current environment state S t , each agent Obtaining local observations of the environment And take action according to the ε-greedy algorithm:

[0021]

[0022] In the formula, the action space of the agent is c t is the energy harvesting time coefficient, L1 is the discrete time level, L2 is the discrete power level, Q is the estimated state-action value function, p is the randomly generated probability, and ε∈(0,1) is the given probability threshold; the battery capacity of ST and J is updated according to the following formula

[0023] in, are the energy collected by ST and J respectively, The available battery capacities of ST and J, B ST,max 、B ST,max are the maximum battery capacities of ST and J respectively, is the correlation time coefficient, T is the length of the transmission block; then the channel gain is updated based on the Rayleigh channel fading model, and the current state of the environment S t Transition to the next state S t+1 ={G t , SINR t , B t-1}, agent k obtains the next local observation And rewards:

[0024]

[0025] Where,

[0026]

[0027]

[0028] μ1+μ2+μ3=1,

[0029] η1+η2+η3=1,

[0030] 0≤μ1, μ2, μ3, η1, η2, η3≤1,

[0031] And transfer the state Stored in experience buffer pool D k In the middle, set S t+1 is the current state S t .

[0032] Furthermore, in step S4, the neural network parameter θ is updated as follows: k and From D k Randomly extract a set batch of samples and send them into the decision network to calculate the loss function:

[0033]

[0034] Where, is the target state-action value function, a′=argmax a′∈A Q(S t+1 , a′;θ k ), Adam optimizer is used to minimize L(θ k ), thus realizing the parameter θ k Update; at the same time, every M consecutive time steps will θ k Copy to the target network weight parameters

[0035] Furthermore, the learning rate of the Adam optimizer is δ=0.001.

[0036] Compared with the existing technology, the present invention has the following beneficial effects: it provides an energy harvesting cognitive Internet of Things resource allocation method based on deep reinforcement learning, which models the resource allocation problem of the cognitive Internet of Things as a reinforcement learning model, and combines the LSTM network with the classic deep reinforcement learning algorithm D3QN, which accelerates the convergence speed and can effectively enhance the confidentiality performance of the system; the intelligent agent adopts centralized training and distributed decision-making, can quickly adapt to and learn the regular changes of the cognitive Internet of Things, realize the optimal allocation of energy harvesting cognitive Internet of Things resources, maximize the confidentiality rate of the cognitive Internet of Things, and thus enhance the communication security of the energy harvesting cognitive Internet of Things. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a cognitive Internet of Things system model in underlay mode in an embodiment of the present invention;

[0038] Figure 2 A flowchart of a method implementation of an embodiment of the present invention;

[0039] Figure 3 The interaction process between the reinforcement learning agent and the cognitive communication environment in the embodiment of the present invention;

[0040] Figure 4 This is a structural diagram of the energy harvesting cognitive Internet of Things resource allocation model based on deep reinforcement learning in an embodiment of the present invention;

[0041] Figure 5 The change of the system confidentiality rate under different resource allocation strategies in the embodiment of the present invention;

[0042] Figure 6 The impact of different maximum transmission powers of the secondary transmitter on the system security rate in the embodiment of the present invention;

[0043] Figure 7 The impact of different maximum transmission powers of cooperative jammers on the system confidentiality rate in an embodiment of the present invention;

[0044] Figure 8 This is a curve diagram of the system confidentiality rate under different reward discount factors in an embodiment of the present invention. DETAILED DESCRIPTION

[0045] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0046] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs.

[0047] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0048] like Figure 1 The energy-harvesting cognitive IoT shown in this paper is a multi-carrier communication system consisting of a primary user (PU), a secondary user (SU), a collaborative jammer (J), and an eavesdropper (E). The PU consists of a primary transmitter (PT) and a primary receiver (PR), while the SU consists of a secondary transmitter (ST) and a secondary receiver (SR). Both ST and J are energy-harvesting nodes equipped with a battery and a radio frequency energy harvester, and ST accesses the PU's spectrum in an underlay manner. Each of the above nodes is equipped with a single antenna. The radio frequency signal is broadcast by the PT and can be collected and stored by the ST and J. It is assumed that the interference signal transmitted by J can be canceled at the SR but cannot be canceled by the eavesdropper E. It is assumed that the channel state information of the eavesdropper node is known to ST and J.

[0049] Each link consists of m sub-channels. represents the set of m subchannel numbers. The system adopts a quasi-static block fading channel model, that is, each subchannel k obeys a Rayleigh distribution with a mean of 0 and a variance of 1. As the channel gain set of the transmission link where PT-PR, ST-SR, ST-PR, PT-SR, ST-E, PT-ST, PT-J, JE, and J-SR are located. On each subchannel k, the transmission powers of PT, S, and J are The RF signal, the confidential information transmitted by the ST, and the interference signal of the cooperative jammer are They are all independent and identically distributed cyclically symmetric complex Gaussian random variables, and have in

[0050] The length of each transmission block is T, and it is divided into two stages: energy collection and information transmission. and It represents the length ratio of the two stages of ST on the tth transmission block, so there is

[0051]

[0052] Similarly, J has

[0053]

[0054] In the energy harvesting phase, ST and J harvest energy from the received RF signal and store it in the battery. The RF power received by ST and J on subchannel k is

[0055]

[0056]

[0057] Using a nonlinear energy harvesting model to reflect a more realistic energy harvesting scenario, there are

[0058]

[0059]

[0060]

[0061] Where: Is with power The relevant logic function, a l and b l is the energy harvesting circuit parameter, A l is the maximum harvesting power when the energy harvesting process reaches saturation. The energy causal constraints of ST and J on the tth transmission block are:

[0062]

[0063]

[0064]

[0065]

[0066] Among them, B ST,max and B J,max are the maximum battery capacities of ST and J respectively, and are the initial battery capacities of ST and J at the tth transmission block. The maximum transmission power constraints of ST and J are:

[0067]

[0068]

[0069] In the information transmission phase, ST transmits confidential information to SR, and E starts to eavesdrop on the information. In order to achieve secure communication, J transmits interference signals to E to reduce the quality of its eavesdropping service. The signals received by PR, SR, and E on each subchannel k are

[0070]

[0071]

[0072]

[0073] in, are the noise signals received by PR, SR, and E respectively, and there is The signal to interference plus noise ratio (SINR) at PR, SR and E is expressed as

[0074]

[0075]

[0076] Among them, λ1 and λ2 are the SINR thresholds of PR and SR respectively.

[0077] The system confidentiality rate at the tth transmission block can be expressed as

[0078]

[0079] Where,

[0080]

[0081]

[0082]

[0083]

[0084]

[0085] Based on the cognitive IoT system model, under the above constraints, the optimization problem is to maximize the achievable confidentiality rate of the system over L transmission blocks:

[0086]

[0087] st (1), (2), (6), (7), (8), (9) (10), (11), (15), (16), (17)

[0088] The present invention transforms the above optimization problem into a Markov decision process (MDP) problem and solves it through deep reinforcement learning. Figure 2 As shown, this embodiment provides an energy harvesting cognitive IoT resource allocation method based on deep reinforcement learning, including the following steps:

[0089] 1) Build an energy harvesting cognitive IoT system model and derive a mathematical model for resource allocation.

[0090] 2) Build a reinforcement learning model, model the 2m sub-channels of the two links from the secondary transmitter to the secondary receiver and from the cooperative jammer to the eavesdropping node, and an energy harvesting time allocation network t0 as 2m+1 reinforcement learning agents. The rest of the energy harvesting cognitive Internet of Things is the reinforcement learning environment, and the agents interact with the environment continuously, such as Figure 3 shown.

[0091] 3) Build an energy harvesting cognitive IoT resource allocation model based on deep reinforcement learning and train it.

[0092] 4) Through the trained resource allocation model, the optimal allocation of joint energy collection time and transmission power of the cognitive Internet of Things is performed to maximize the system's achievable confidentiality rate, thereby enhancing the communication security of the cognitive Internet of Things.

[0093] The structure of the energy harvesting cognitive IoT resource allocation model based on deep reinforcement learning is as follows: Figure 4 As shown in Figure 2, training the energy harvesting cognitive IoT resource allocation model based on deep reinforcement learning includes the following steps:

[0094] S1. Generate the topology of the cognitive Internet of Things, initialize the channel gain of each link, the number of rounds of training N, and the experience buffer pool D k The maximum capacity N k , and the decision network and target network weight parameters θ k 、 in

[0095] S2. At the beginning of each training round, a randomization scheme is used to update the positions of all nodes in the cognitive Internet of Things. The update of the channel gain of each link follows the Rayleigh channel fading model. The initial state of the environment for this training round is set to:

[0096] S0=S t | t=0 ={G t , SINR t , B t-1} t=0 ={G0, SINR0, {B ST,max , B J,max}},#(19)

[0097] Where, The signal-to-interference-noise ratio set at PR, SR, and E ST and J battery capacity collection And there is B -1 ={B ST,max , B J,max},in, is a set of subchannel numbers, is the subchannel gain, are the transmission powers of the secondary transmitter ST and the cooperative interferer J in the kth subchannel, B ST,max and B J,max The maximum battery capacities of ST and J respectively.

[0098] S3, in each training round t = 0, 1, 2, ..., T max -1 time step, based on the current environment state S t , each agent Obtaining local observations of the environment And take action according to the ε-greedy algorithm:

[0099]

[0100] In the formula, the action space of the agent is c t is the energy harvesting time coefficient, L1 is the discrete time level, L2 is the discrete power level, Q is the estimated state-action value function, p is the randomly generated probability, and ε∈(0,1) is the given probability threshold. The battery capacities of ST and J are updated according to the following formula:

[0101]

[0102]

[0103] in, are the energy collected by ST and J respectively, The available battery capacities of ST and J, B ST,max 、B ST,max are the maximum battery capacities of ST and J respectively, is the correlation time coefficient, T is the length of the transmission block; then the channel gain is updated based on the Rayleigh channel fading model, and the current state of the environment S t Transition to the next state S t+1 ={G t , SINR t , B t-1}, agent k obtains the next local observation And rewards:

[0104]

[0105] Where,

[0106]

[0107]

[0108] μ1+μ2+μ3=1,#(23c)

[0109] η1+η2+η3=1,#(23d)

[0110] 0≤μ1, μ2, μ3, η1, η2, η3≤1, #(23e)

[0111] And transfer the state Stored in experience buffer pool D k In the middle, set S t+1 is the current state S t .

[0112] S4. Update the neural network parameters θ as follows k and From D k Randomly extract a set batch of samples and send them into the decision network to calculate the loss function:

[0113]

[0114] Where, is the target state-action value function, a′=argmax a′∈A Q(S t+1 , a′;θ k ), Adam optimizer with learning rate δ = 0.001 is used to minimize L(θ k ), thus realizing the parameter θ k Update; at the same time, every M consecutive time steps will θ k Copy to the target network weight parameters

[0115] The algorithm implementation for training the deep reinforcement learning-based energy harvesting cognitive IoT resource allocation model is as follows:

[0116]

[0117]

[0118] The feasibility and effectiveness of the method of the present invention are further illustrated by the following simulation.

[0119] Figure 5 The figure shows the change in the confidentiality rate of the cognitive IoT system model during each training round. Compared with other algorithms, the method provided by this invention achieves the best performance in improving the confidentiality rate, demonstrating that this method can effectively enhance the communication security performance of the energy-harvesting cognitive IoT.

[0120] Figure 6 、 Figure 7 The effects of different transmit powers for ST and J on the system's confidentiality rate in a cognitive IoT system model are shown. Compared to other baseline algorithms, this method, which combines an LSTM network with the classic reinforcement learning algorithm D3QN, achieves a higher confidentiality rate even at low transmit power. This demonstrates that this method can adapt to the dynamically changing cognitive IoT environment and effectively enhance its security performance.

[0121] Figure 8 The figure shows the impact of different reward discount factors on the system reward in the method proposed by this invention. As can be seen from the figure, the system reward changes differently under different discount factors. When γ = 0.95, the system reward value is the largest, indicating that the system confidentiality performance under this parameter is optimal.

[0122] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other manner. Any person skilled in the art may utilize the above-disclosed technical content to modify or modify the present invention into equivalent embodiments. However, any simple modifications, equivalent variations, and modifications to the above embodiments that do not depart from the technical content of the present invention and are based on the technical essence of the present invention remain within the scope of protection of the present invention.

Claims

1. A method for allocating energy harvesting cognitive IoT resources based on deep reinforcement learning, characterized in that: include: Build a resource allocation model for energy harvesting cognitive IoT and derive a mathematical model for resource allocation; A deep reinforcement learning-based energy harvesting cognitive IoT resource allocation model was constructed. The 2m subchannels of the links from the secondary transmitter to the secondary receiver and from the collaborative jammer to the eavesdropping node, as well as an energy harvesting time allocation network t0, were modeled as 2m+1 reinforcement learning agents. The remaining components of the energy harvesting cognitive IoT resource allocation model, excluding the reinforcement learning agents, served as the reinforcement learning environment, with which the agents continuously interacted. The deep reinforcement learning-based energy harvesting cognitive IoT resource allocation model was then trained. The trained energy harvesting cognitive IoT resource allocation model is used to optimally allocate the energy harvesting time and transmission power for the cognitive IoT. Training a deep reinforcement learning-based energy harvesting cognitive IoT resource allocation model involves the following steps: S1. Generate the topology of the cognitive Internet of Things, initialize the channel gain of each link, the number of rounds of training N, and the experience buffer pool D k The maximum capacity N k , and the decision network weight parameter θ k and target network weight parameters in S2. At the beginning of each training round, the positions of all nodes in the cognitive IoT are randomly initialized, the channel gain of each link is updated, and the initial state of the environment is set to S0; S3, in each training round t=0,1,2,…,T max -1 time step, based on the current environment state S t , each agent k obtains local observations of the environment And take action according to the ε-greedy algorithm Updated battery capacity for ST and J and Update channel gain, current environment state S t Transition to the next state S t+1 , agent k obtains the next local observation and reward r t ; S4. Update the neural network parameters θ k and That is, from D k Randomly extract a set batch of samples and send them to the decision network to calculate the loss function L(θ k ), and perform gradient descent to minimize L(θ k ) Update the parameters θ k ; Every M consecutive time steps, θ k Copy to the target network weight parameters In step S3, in each training round, at t=0, 1, 2, ..., T max -1 time step, based on the current environment state S t , each agent Obtaining local observations of the environment And take action according to the ε-greedy algorithm: In the formula, the action space of the agent is c t is the energy harvesting time coefficient, L1 is the discrete time level, L2 is the discrete power level, Q is the estimated state-action value function, p is the randomly generated probability, and ε∈(0,1) is the given probability threshold. The battery capacities of ST and J are updated according to the following formula: in, are the energy collected by ST and J respectively, The available battery capacities of ST and J, B sT,max 、B ST,max are the maximum battery capacities of ST and J respectively, is the correlation time coefficient, T is the length of the transmission block; then the channel gain is updated based on the Rayleigh channel fading model, and the current state of the environment S t Transition to the next state S t+1 ={G t ,SINR t ,B t-1 }, agent k obtains the next local observation And rewards: Where, μ1+μ2+μ3=1, η1+η2+η3=1, 0≤μ1,μ2,μ3,η1,η2,η3≤1, And transfer the state Stored in experience buffer pool D k In the middle, set S t+1 is the current state S t .

2. The energy harvesting cognitive Internet of Things resource allocation method based on deep reinforcement learning according to claim 1 is characterized in that: In step S2, at the beginning of each training round, a randomization scheme is used to update the positions of all nodes in the cognitive Internet of Things. The update of the channel gain of each link follows the Rayleigh channel fading model, and the initial state of the environment for this training round is set to: S0=S t | t=0 ={G t ,SINR t ,B t-1 } t=0 ={G0,SINR0,{B sT,max ,B J,max }} Where, The signal-to-interference-noise ratio set at PR, SR, and E ST and J battery capacity collection And there is B -1 ={B ST,max ,B J,max },in, is a set of subchannel numbers, is the subchannel gain, are the transmission powers of the secondary transmitter ST and the cooperative interferer J in the kth subchannel, B ST,max and B J,max The maximum battery capacities of ST and J respectively.

3. The energy harvesting cognitive Internet of Things resource allocation method based on deep reinforcement learning according to claim 1 is characterized in that: In step S4, the neural network parameters θ are updated as follows: k and From D k Randomly extract a set batch of samples and send them into the decision network to calculate the loss function: Where, is the target state-action value function, a′=argmax a′∈A Q(S t+1 ,a′;θ k ), Adam optimizer is used to minimize L(θ k ), thus realizing the parameter θ k Update; at the same time, every M consecutive time steps will θ k Copy to the target network weight parameters 4. The method for allocating energy harvesting cognitive IoT resources based on deep reinforcement learning according to claim 3 is characterized in that: The learning rate of the Adam optimizer is δ=0.001.

Citation Information

Patent Citations

  • Green cognitive radio power distribution method based on deep reinforcement learning

    CN114126021A

  • Cognitive radio network dynamic spectrum access method based on deep reinforcement learning

    CN115190489A