NOMA-CR network dynamic power distribution method and system based on deep Q learning

By building a two-layer network model and deep Q learning method, the problem of dynamic changes in the NOMA-CR network is solved, refined power adjustment and security protection are achieved, the security performance and energy efficiency of the network are improved, and communication quality and stability are ensured.

CN120343713APending Publication Date: 2025-07-18QUFU NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510463631.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to adapt to the dynamic changes brought about by rapid channel fading, node movement and multi-user collaboration in NOMA-CR networks, while ensuring the quality of communications of the main network, effectively reducing the risk of eavesdropping and taking into account the energy constraints of the relay nodes, and insufficient security protection.

Method used

A two-layer network model is built, a state space and discrete action space is designed, and a deep Q-learning method is used to perform dynamic power adjustment through a composite reward function, combining energy acquisition module and relay node energy management to achieve dynamic resource scheduling and security protection.

Benefits of technology

It realizes refined power adjustment in complex wireless environments, improves the security performance, energy efficiency and stability of the network, reduces the risk of eavesdropping, and improves communication quality and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343713A_ABST
    Figure CN120343713A_ABST
Patent Text Reader

Abstract

The invention discloses an NOMA-CR network dynamic power distribution method and system based on deep Q learning. The system comprises an energy acquisition module, a state acquisition module, a deep Q network module, an action execution module and a reward calculation module. The method comprises the following steps: firstly, designing a multi-dimensional state space containing a signal-to-noise ratio and an energy state, defining a discretization power regulation action space, and constructing a composite reward function integrating communication performance, a safety index and energy efficiency; then, an energy collection module is introduced, and cooperative transmission is activated only when node energy meets a preset threshold value; and finally, through dynamic interference power adjustment and main network interference tolerance threshold constraint, realizing joint optimization of the coexistence performance of the main and secondary networks and the secrecy throughput. According to the method, dual optimization of dynamic resource scheduling and security protection is realized, the communication quality, security performance, stability and energy efficiency of the network are improved, and the method has relatively high practical application value and technical popularization prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of wireless communication and intelligent resource allocation, and in particular to a dynamic power allocation method and system for a NOMA-CR network based on deep Q learning. Background Art

[0002] With the rapid development of technology, there are more and more portable intelligent terminal devices such as mobile phones and watches. At the same time, the proposal and development of the Internet of Things (IoT) have made it possible for everything to be interconnected. These intelligent devices and sensors have increased the demand for the number of devices that can be simultaneously accessed in a communication system. Under the premise of scarce spectrum resources, researching advanced spectrum access and power allocation technologies can effectively improve spectrum efficiency.

[0003] Power-domain non-orthogonal multiple access (NOMA) allows multiple users to multiplex in the power domain simultaneously, with the same frequency and the same code, which can significantly improve spectrum efficiency. On the other hand, cognitive radio (CR) technology allows secondary users to share the spectrum with primary users on the premise of interference suppression, which can also play a role in saving communication resources. Obviously, the two technologies of CR and NOMA can be combined through reasonable power allocation to achieve the effects of improving spectrum efficiency and accessing more users. For a NOMA-based CR network, it is necessary to comprehensively consider the interference between primary and secondary users, the respective quality of service (QoS) requirements of primary and secondary users, and the successive interference cancellation (SIC) constraints between users. Therefore, the power allocation problem in a NOMA-CR network is relatively complex. Traditional static power allocation and fixed interference management methods are difficult to adapt to the dynamic changes brought about by rapid channel fading, node movement, and multi-user cooperation, and there are also obvious deficiencies in security protection.

[0004] The invention patent CN 112770395 A discloses an optimal dynamic power allocation method, system, medium and terminal based on uplink NOMA, which can reasonably utilize the power allocated in each cluster in the uplink NOMA system, thereby reducing power waste and improving the throughput of the system; the invention patent CN 113691334 A discloses a cognitive radio dynamic power allocation method based on the cooperation of secondary user groups, which uses the Maddpg algorithm to solve the stability problem of the reinforcement learning algorithm in a dynamic and unstable environment. These existing methods cannot adapt to the dynamic changes brought about by fast channel fading, node movement and multi-user cooperation, and there are also obvious deficiencies in security protection. Especially in the NOMA-CR network, how to effectively reduce the eavesdropping risk while ensuring the communication quality of the primary network and taking into account the energy constraints of relay nodes has become a key problem to be solved urgently. Summary of the Invention

[0005] The purpose of the present invention is to provide a dynamic power allocation method and system for NOMA-CR networks based on deep Q learning, which have high security performance, strong anti-interference ability, high energy efficiency and high network stability.

[0006] The technical solution to achieve the purpose of the present invention is: a dynamic power allocation method for NOMA-CR networks based on deep Q learning, including the following steps:

[0007] Step 1: Construct a two-layer network model composed of a source node, a cooperative node and users, and describe each communication link using a Rayleigh fading and AWGN noise model;

[0008] Step 2: Design a state space, where the state vector includes the real-time signal-to-noise ratio information between the source node and the cooperative node and contains the dynamic changes of the energy state of the relay node;

[0009] Step 3: Construct a discretized action space, and discretize the power adjustment of the source node and the cooperative node into multiple combined actions;

[0010] Step 4: Design a composite reward function according to the communication performance reward, interference management reward and energy reward;

[0011] Step 5: Use a deep Q network for training, and through experience replay, target network update and ε-greedy strategy, realize the dynamic decision-making of the optimal power adjustment action in the state space;

[0012] Step 6: Adjust the transmission power of the source node and the cooperative node according to the selected action, and update the energy of the relay node according to the channel state and the source node power. Only when the energy reaches the preset threshold is the relay node allowed to participate in forwarding.

[0013] Furthermore, the double-layer network model consisting of source nodes, cooperative nodes, and users constructed in Step 1 is described by Rayleigh fading and AWGN noise models for each communication link, specifically as follows:

[0014] Step 1.1: Construct a double-layer network model, including a primary network and a secondary network. The primary network consists of a source node and a target user to ensure normal communication of the user. The secondary network consists of a source node, multiple relay nodes, a legitimate user, and an eavesdropper. All nodes use single antennas and operate in a half-duplex mode. Each relay node is randomly distributed between the source node and the user, forming a rectangular area with a set geometric range.

[0015] Step 1.2: All transmissions between nodes adopt an independent Rayleigh fading channel model, and the noise adopts an additive white Gaussian noise (AWGN) model with a mean of zero and a variance of σ 2 . Model the decode-and-forward (DF) and amplify-and-forward (AF) protocols respectively, analyze the signal-to-noise ratio (SINR) at each stage, and analyze the impact of the interference signal generated by the interfering relay on the eavesdropper and the interference suppression of the primary network communication link.

[0016] Furthermore, the state space described in Step 2 is a two-dimensional vector, representing the signal-to-noise ratio (SNR) of the source node and the cooperative node respectively, and the SNR of the source node and the cooperative node is discretized within a preset range.

[0017] Furthermore, the discrete action space constructed in Step 3 discretizes the power adjustment of the source node and the cooperative node into multiple combined actions, specifically as follows:

[0018] Construct a discrete action space, discretize the transmit powers of the source node and the cooperative node to form multiple action combinations. Each action combination corresponds to a set of power adjustment parameters, and the action index is mapped to the power selection of the source node and the cooperative node through integer division and remainder operations respectively.

[0019] Furthermore, the composite reward function designed in Step 4 based on communication performance rewards, interference management rewards, and energy rewards is specifically as follows:

[0020] Step 4.1: Design communication performance rewards, security rewards, and energy rewards, specifically as follows:

[0021] Communication performance rewards: Based on indicators such as the success rate of user signal decoding, whether the target SINR reaches a preset threshold, and the system's secure throughput (EST), a negative reward is given if communication fails.

[0022] Security rewards: Through a dynamic interference management strategy, a suppression reward is imposed on the eavesdropper's signal-to-noise ratio to ensure confidentiality performance.

[0023] Energy reward: The positive increment of the energy state of the relay node is multiplied by a preset coefficient and then added to the reward to encourage the system to consider the balance between energy harvesting and consumption during power allocation;

[0024] Step 4.2: Weightedly combine the communication performance reward, interference management reward, and energy reward to obtain a composite reward function, thereby guiding the deep Q-network to converge to the optimal dynamic decision-making strategy.

[0025] Furthermore, the training using the deep Q-network in Step 5 realizes the dynamic decision-making of the optimal power adjustment action in the state space through experience replay, target network update, and ε-greedy strategy, specifically as follows:

[0026] Step 5.1: Design a deep Q-network, adopting a multi-layer neural network structure. The input layer receives the state vectors SNR_S and SNR_R, followed by a hidden layer with 64 nodes activated by ReLU, a hidden layer with 128 nodes activated by ReLU, and an output layer with 676 nodes, corresponding to the Q-values of all actions;

[0027] Step 5.2: Train the deep Q-network:

[0028] (1) Initialize the network parameters and ε-greedy strategy parameters, and the exploration rate gradually decays from 1.0 to 0.1;

[0029] (2) In each time slot, select the corresponding action according to the current state, and after executing the action, the environment feedbacks the next state and the composite reward;

[0030] (3) Store the state transition samples in the experience replay pool, and randomly sample and update the network parameters using the mean squared error (MSE) loss function;

[0031] (4) Every fixed number of iterations, copy the evaluation network parameters to the target network to ensure the training stability;

[0032] (5) Repeat steps (2) to (4) until the system safety throughput EST tends to converge, and verify through simulation the superiority of the system performance under the DF protocol over the AF protocol.

[0033] Furthermore, in Step 6, the transmit powers of the source node and the cooperative node are adjusted according to the selected action, and the energy of the relay node is updated according to the channel state and the source node power. Only when the energy reaches the preset threshold is the relay node allowed to participate in forwarding, specifically as follows:

[0034] Step 6.1: Establish a set of candidate forwarding nodes for the relay node, and put the relay nodes that meet the conditions into the set of candidate forwarding nodes, specifically as follows:

[0035] The relay node must meet two conditions before participating in forwarding:

[0036] Signal decoding ability: All relays that can accurately decode the signals sent by the source node form the decoding set;

[0037] Energy threshold: Each relay node is equipped with a battery with infinite capacity. When the energy collected by a certain relay node reaches the preset threshold, the relay node is put into the set of candidate forwarding nodes;

[0038] Step 6.2: Using the Rayleigh channel model, calculate the energy harvesting amount E of the relay node according to the channel gain and the source node transmission power h , the formula is:

[0039] E h = η × P S × |h| 2 × T s

[0040] where η is the energy conversion efficiency, P S is the source node power, |h| 2 is the channel gain, T s is the time slot length;

[0041] When the energy collected by the relay node reaches or exceeds the set threshold, the relay node is allowed to participate in the decoding and forwarding of information, so as to realize energy management and extend the system operation life.

[0042] A dynamic power allocation system for a NOMA-CR network based on deep Q learning, which is used to implement the dynamic power allocation method for the NOMA-CR network based on deep Q learning. The system includes an energy harvesting module, a state acquisition module, a deep Q network module, an action execution module and a reward calculation module, where:

[0043] The energy harvesting module is used to collect and update the energy state of the relay node according to the channel state and the source node transmission power;

[0044] The state acquisition module is used to collect the signal-to-noise ratio information of the source node and the cooperative node in real time;

[0045] The deep Q network module is used to output the optimal action strategy according to the collected state information;

[0046] The action execution module is used to dynamically adjust the transmission power of the source node and the cooperative node according to the information;

[0047] The reward calculation module is used to calculate the composite reward and feedback it to the deep Q network to realize continuous parameter optimization.

[0048] Further, the energy harvesting module calculates the energy harvesting amount of the relay node according to the channel state and the source node transmission power, and when the energy harvested by the relay node reaches a preset threshold, allows the relay node to participate in signal decoding and forwarding.

[0049] Further, the deep Q-network module adopts an experience replay and target network update strategy, and uses an ε-greedy strategy to balance exploration and exploitation to achieve the optimal decision of network dynamic power allocation.

[0050] Compared with the prior art, the present invention has the following remarkable advantages: (1) discretize the transmission power adjustment into multiple combined actions, enabling the system to achieve refined power adjustment in the face of a complex wireless environment, thus breaking through the limitations of traditional fixed power allocation; (2) the composite reward function takes into account three dimensions of communication performance, security, and energy management, and can more comprehensively guide the deep Q-network to achieve global optimality. By jointly considering the user decoding success rate, the achievement of the target SINR, the dynamic interference suppression effect, and the relay energy increment, the dual optimization of dynamic resource scheduling and security protection is realized; (3) using the deep Q-network, through experience replay, target network update, and ε-greedy strategy, effectively cope with the complex dynamic changes in the high-dimensional state space, break through the limitations of traditional methods in the fast fading channel and multi-user cooperation environment, achieve dynamic adaptability of power allocation under different interference conditions, and improve the security performance and energy efficiency of the network; (4) adopt an energy harvesting module to calculate the energy harvested by the relay node according to the actual channel conditions and the source node transmission power, and only allow the relay to participate in forwarding when the energy reaches a preset threshold, fundamentally solving the problem of communication interruption caused by insufficient energy in traditional systems. Combined with the dynamic relay selection strategy, only select relay nodes that can ensure good decoding ability and sufficient energy, thereby further reducing system interference and improving the overall network stability; (5) achieve collaborative optimization in dynamic power allocation, security protection, and energy management, improve communication quality and security, and have the significant advantage of the rapid convergence of the system equivalent security throughput EST with iteration, with high practical application value and technical promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a schematic flowchart of the method for dynamic power allocation of the NOMA-CR network based on deep Q learning of the present invention.

[0052] Figure 2 is a schematic structural diagram of the double-layer network model constructed in the embodiment of the present invention.

[0053] Figure 3 is a relationship diagram between the EST_SUM of the system in the embodiment of the present invention and the number of iterations in the same scenario.

[0054] Figure 4 This is a graph showing the relationship between the SINR of the system decoding the signal m times and the number of iterations when using the DF and AF protocols in the embodiments of the present invention.

[0055] Figure 5 In the embodiments of the present invention, the relay node forwarding power is P R =-16dB, and this is a graph showing the influence relationship of the source node transmission power P S on EST_SUM.

[0056] Figure 6 In the embodiments of the present invention, when the fixed source node power is P S =-16dB, this is a graph showing the influence relationship of the relay node forwarding power P R on EST_SUM.

[0057] Figure 7 This is a graph showing the variation relationship between the equivalent total throughput EST_SUM of the system and the source node transmission power PS in the embodiments of the present invention.

[0058] Figure 8 This is a graph showing the variation relationship between the equivalent total throughput EST_SUM of the system and the relay node forwarding power PR in the embodiments of the present invention.

[0059] Figure 9 This is a graph showing the influence relationship of the codeword rate on the system performance in the embodiments of the present invention.

[0060] Figure 10 This is a graph showing the influence relationship of the secrecy rate on the system performance in the embodiments of the present invention. Detailed implementation manners

[0061] A NOMA-CR network dynamic power allocation system based on deep Q learning according to the present invention is characterized in that it includes an energy harvesting module, a state acquisition module, a deep Q network module, an action execution module, and a reward calculation module;

[0062] The energy harvesting module is used to collect and update the energy state of the relay node according to the channel state and the source node transmission power;

[0063] The state acquisition module is used to collect the signal-to-noise ratio information of the source node and the cooperative node in real time;

[0064] The deep Q network module is used to output the optimal action strategy according to the collected state information;

[0065] The action execution module is used to dynamically adjust the transmission powers of the source node and the cooperative node according to the information;

[0066] The reward calculation module is used to calculate the composite reward and feedback it to the deep Q network to realize continuous optimization of the parameters.

[0067] As a specific example, the energy harvesting module calculates the energy harvesting amount of the relay node according to the channel state and the source node transmission power, and allows it to participate in signal decoding and forwarding when the energy of the relay node reaches a preset threshold.

[0068] As a specific example, the deep Q-network module adopts an experience replay and target network update strategy, and uses an ε-greedy strategy to balance exploration and exploitation to achieve the optimal decision for network dynamic power allocation.

[0069] As Figure 1 shown, a method for dynamic power allocation in a NOMA-CR network based on deep Q-learning of the present invention includes the following steps:

[0070] Step 1: Construct a two-layer network model composed of a source node, a cooperative node, and users. Each communication link is described by a Rayleigh fading and AWGN noise model, specifically as follows:

[0071] Step 1.1: Construct a two-layer network model, including a primary network and a secondary network. The primary network is composed of a source node and a target user to ensure normal communication of the user; the secondary network is composed of a source node, multiple relay nodes, a legitimate user, and an eavesdropper; all nodes use single antennas and operate in a half-duplex mode; each relay node is randomly distributed between the source node and the user to form a rectangular area with a certain geometric range;

[0072] Step 1.2: The transmission between all nodes adopts an independent Rayleigh fading channel model, and the noise adopts an additive white Gaussian noise AWGN model with an average value of zero and a variance of σ 2 ; model the decode-and-forward DF and amplify-and-forward AF protocols respectively, analyze the signal-to-noise ratio SINR in each stage, and analyze the influence of the interference signal generated by the interfering relay on the eavesdropper and the interference suppression of the main network communication link.

[0073] Step 2: Design a state space, where the state vector includes the real-time signal-to-noise ratio information of the source node and the cooperative node, and contains the dynamic change of the energy state of the relay node;

[0074] As a specific example, the state space is a two-dimensional vector, which respectively represents the signal-to-noise ratio SNR of the source node and the cooperative node, and discretizes the signal-to-noise ratio SNR of the source node and the cooperative node within a preset range.

[0075] Step 3: Construct a discretized action space, and discretize the power adjustment of the source node and the cooperative node into 26×26 combined actions, specifically as follows:

[0076] Construct a discretized action space, discretize the transmission powers of the source node and the cooperative nodes to form 26×26 action combinations, each action combination corresponding to a set of power adjustment parameters, and the action index is mapped to the power selection of the source node and the cooperative nodes through integer division and remainder operations respectively.

[0077] Step 4: Design a composite reward function according to the communication performance reward, interference management reward, and energy reward, as follows:

[0078] Step 4.1: Design the communication performance reward, security reward, and energy reward, as follows:

[0079] Communication performance reward: Based on the user signal decoding success rate, whether the target SINR reaches the preset threshold, and the system security throughput EST, if the communication fails, a large negative reward is given;

[0080] Security reward: Through a dynamic interference management strategy, impose an inhibitory reward on the signal-to-noise ratio of the eavesdropper to ensure the secrecy performance of the system;

[0081] Energy reward: The positive increment of the energy state of the relay node will be multiplied by a preset coefficient and then added to the reward to encourage the system to fully consider the balance between energy harvesting and consumption during power allocation;

[0082] Step 4.2: Perform weighted combination of the communication performance reward, interference management reward, and energy reward to obtain a composite reward function, thereby guiding the deep Q-network to converge to the optimal dynamic decision-making strategy.

[0083] Step 5: Use the deep Q-network for training, and through experience replay, target network update, and ε-greedy strategy, realize the dynamic decision-making of the optimal power adjustment action in the state space, as follows:

[0084] Step 5.1: Design a deep Q-network, adopt a multi-layer neural network structure, the input layer receives the state vectors SNR_S and SNR_R, followed by a hidden layer with 64 nodes activated by ReLU, a hidden layer with 128 nodes activated by ReLU, and an output layer with 676 nodes, corresponding to the Q-values of all actions;

[0085] Step 5.2: Train the deep Q-network:

[0086] (1) Initialize the network parameters and ε-greedy strategy parameters, and the exploration rate gradually decays from 1.0 to 0.1;

[0087] (2) In each time slot, select the corresponding action according to the current state, and after executing the action, the environment feedbacks the next state and the composite reward;

[0088] (3) Store the state transition samples in the experience replay pool, and randomly sample and then update the network parameters using the mean square error (MSE) loss function.

[0089] (4) Every fixed iteration period, copy the evaluation network parameters to the target network to ensure training stability;

[0090] (5) Repeat the above process until the system EST converges, and verify through simulation that the system performance under the DF protocol is superior to that of the AF protocol.

[0091] Step 6: Adjust the transmission powers of the source node and the cooperative node according to the selected action, and update the energy of the relay node according to the channel state and the source node power. Only when the energy reaches the preset threshold is the relay node allowed to participate in forwarding, as follows:

[0092] Step 6.1: Establish a set of candidate forwarding nodes for the relay node, and put the relay nodes that meet the conditions into the set of candidate forwarding nodes, as follows:

[0093] The relay node must meet two conditions before participating in forwarding:

[0094] Signal decoding ability: All relays that can accurately decode the signals sent by the source node form a decoding set;

[0095] Energy threshold: Each relay node is equipped with a battery with infinite capacity. When the energy collected by a certain relay node reaches the preset threshold, put this relay node into the set of candidate forwarding nodes;

[0096] Step 6.2: Use the Rayleigh channel model to calculate the energy harvesting amount of the relay node according to the channel gain and the source node transmission power. The formula is:

[0097] E h =η×P S ×∣h∣ 2 ×T s

[0098] where η is the energy conversion efficiency, P S is the source node power, ∣h∣ 2 is the channel gain, and T s is the time slot length;

[0099] When the energy collected by the relay node reaches or exceeds the set threshold, allow this relay node to participate in the decoding and forwarding of information, so as to achieve energy management and extend the system operation life.

[0100] The following further elaborates on the present invention in conjunction with the accompanying drawings and specific embodiments.

[0101] Embodiment

[0102] A dynamic power allocation system for NOMA-CR network based on deep Q-learning provided in this embodiment includes an energy harvesting module, a state acquisition module, a deep Q-network module, an action execution module, and a reward calculation module;

[0103] As Figure 1 shown, a dynamic power allocation method for NOMA-CR network based on deep Q-learning includes the following steps:

[0104] Step 1: Construct a two-layer network model consisting of a source node, a cooperative node, and users. As Figure 2 shown, each communication link is described by a Rayleigh fading and AWGN noise model;

[0105] Step 2: Design the state space, where the state vector includes the real-time SNR information of the source node and the cooperative node and contains the dynamic changes of the energy state of the relay node;

[0106] Step 3: Construct a discretized action space, and discretize the power adjustment of the source node and the cooperative node into multiple combined actions;

[0107] Step 4: Design a composite reward function according to the communication performance reward, interference management reward, and energy reward;

[0108] Step 5: Use the deep Q-network for training, and through experience replay, target network update, and ε-greedy strategy, realize the dynamic decision-making of the optimal power adjustment action in the state space;

[0109] Step 6: Adjust the transmission power of the source node and the cooperative node according to the selected action, and update the energy of the relay node according to the channel state and the source node power. Only when the energy reaches the preset threshold is the relay node allowed to participate in forwarding.

[0110] Figure 3 is the relationship between the EST_SUM of the system and the number of iterations in the same scenario. This figure is drawn based on the average results of 1000 simulation runs. It can be seen that the EST_SUM of both protocols gradually increases and converges as the number of iterations increases, and reaches its maximum value at about 200 iterations. In addition, during the deep Q-learning process, the EST_SUM obtained by the DF protocol is always greater than the EST_SUM obtained by the AF protocol. Therefore, the DF protocol can obtain better security performance through deep Q-learning.

[0111] Figure 4 shows the variation relationship of the SINR of the system decoding the signal m times with the number of iterations when using the decode-and-forward DF and amplify-and-forward AF protocols. This Figure 4 is drawn based on the average results of 1000 simulation runs, and each simulation run contains 500 iterations. From Figure 4It can be seen that due to the adoption of the ε-greedy strategy to avoid falling into local optima, the SINR values of both protocols show certain fluctuations during the iterative process. Nevertheless, the overall trend shows a gradually increasing characteristic and converges after approximately 175 iterations, reaching their respective maximum values. In addition, from Figure 4 It can be seen that during the iterative process, the decoded SINR obtained by the DF protocol is always higher than that of the AF protocol. The reason for this phenomenon is that the amplification gain of the AF relay system is fixed and cannot be adaptively adjusted according to the signal strength, and the relay node will also amplify the noise while amplifying the signal, resulting in signal distortion. In contrast, the DF relay system significantly improves the signal reliability by decoding and re-encoding the signal. Therefore, the DF protocol has obvious advantages in improving the communication quality of edge users.

[0112] Figure 5 shows the influence on EST_SUM when the forwarding power of the fixed relay node is P R =-16 dB and the transmission power P S of the source node varies in the range from -40 dB to 30 dB. From Figure 5 it can be seen that as P S increases, EST_SUM shows a trend of first increasing and then decreasing. When PS is lower than -20 dB, due to the high probability of transmission interruption caused by insufficient power, the EST_SUM of the system approaches zero; in the medium power region -20 dB ≤ PS ≤ 0 dB, the increase of P S can improve the signal-to-noise ratio SINR of the destination node, thus significantly increasing the secrecy capacity. However, if P S is too high, for example, exceeding 0 dB, the SINR of the eavesdropping node will increase sharply, resulting in a decrease in the secrecy capacity and ultimately a reduction in EST_SUM. This non-monotonic relationship indicates that there is an optimal P S value, such as PS ≈ -10 dB, which can maximize the EST_SUM of the system. In addition, in the range from -20 dB to 0 dB, increasing the number of relays, such as from 4 to 20, can significantly reduce the outage probability through the diversity gain, thereby increasing EST_SUM.

[0113] Figure 6 shows the influence of the relay node forwarding power P S on EST_SUM when the power of the fixed source node is P R =-16 dB. From Figure 6 it can be seen that as PR increases, EST_SUM also shows a trend of first increasing and then decreasing. In the low P R region, that is, the P R <-20 dB region, the relay power is insufficient to support effective forwarding, resulting in transmission failure; when P RWhen it is between -20 dB and 0 dB, the increase in relay power enhances the reception quality of the destination node, but too high P R region, that is, when P R > 0 dB, it will cause the SINR of the eavesdropping link to rise rapidly, thus weakening the secrecy performance. Similar to Figure 4 that, increasing the number of relays can further improve EST_SUM within this range, but too many relays may introduce resource allocation complexity and potential interference problems. The experimental results show that when P R ≈ -5 dB, the system can achieve the optimal EST_SUM.

[0114] Figure 7 shows the relationship between the equivalent total throughput EST_SUM of the system and the transmit power PS of the source node, and analyzes the impact of different interference powers on the system performance. It can be observed from Figure 7 that the EST_SUM of the system first increases and then decreases as the transmit power PS of the source node increases. Specifically, when PS starts to increase from a lower value, the EST_SUM of the system gradually increases, indicating that increasing the transmit power can significantly improve the system performance within a certain range. However, when PS exceeds a certain critical value, about -26 dB, the EST_SUM begins to gradually decrease, indicating that too high transmit power may have a negative impact on the system performance.

[0115] It is worth noting that at different interference powers, the EST_SUM of the system reaches the maximum value of 0.14 when PS = -26 dB. This result shows that at this transmit power point, the change in interference power has little effect on the EST_SUM of the system. This phenomenon can be explained by the following mechanism: when the transmit power is high, the decoding signal-to-noise ratio SNR of the eavesdropping node E also increases significantly, and further increasing the interference power has limited effect on reducing the decoding ability of E. Therefore, when PS = -26 dB, the system reaches the optimal balance point of performance, and the change in interference power no longer has a significant impact on the EST_SUM of the system.

[0116] In addition, Figure 7 it can also be seen the non-linear relationship between the transmit power and the system performance. When PS is too low, the EST_SUM of the system is low, indicating that the system performance is not fully exerted; while when PS is too high, the EST_SUM decreases, indicating that too high transmit power may lead to a decrease in resource allocation efficiency, thus affecting the overall performance. This feature provides an important guidance for the actual system design, that is, when selecting the transmit power of the source node, too high or too low power settings should be avoided to maximize the equivalent total throughput of the system.

[0117] Figure 8It shows the relationship between the equivalent total throughput EST_SUM of the system and the relay node forwarding power PR, and analyzes the influence of different interference powers on the system performance. From Figure 8 it can be seen that the EST_SUM of the system also shows a trend of first rising and then falling as the relay node forwarding power PR increases. Specifically, when PR starts to increase from a lower value, the EST_SUM of the system gradually rises, indicating that increasing the forwarding power can significantly improve the system performance within a certain range. However, when PR exceeds a certain critical value, the EST_SUM begins to gradually decrease, indicating that too high a forwarding power may have a negative impact on the system performance.

[0118] Under different interference powers, the PR values at which the EST_SUM of the system reaches the maximum are different. Specifically: when the interference power is 20 dB, the EST_SUM of the system reaches the maximum value of 0.14 at PR = -15 dB. When the interference power is 40 dB, the EST_SUM of the system reaches the maximum value of 0.14 at PR = -7 dB.

[0119] This result indicates that there is a complex interaction between the relay node forwarding power and the interference power. As the interference power increases, the system requires a higher forwarding power to offset the influence of interference and thus maintain a higher EST_SUM. However, too high a forwarding power may lead to resource waste and performance degradation, so in the actual system design, it is necessary to dynamically adjust the forwarding power according to the interference power.

[0120] Figure 9 and Figure 10 shows the influence of the codeword rate and the secrecy rate on the system performance, and the EST that can be achieved after the deep Q-learning process under different parameter settings. In Figure 9 , the secrecy rates of both users are set to τ m,s = τ n,s = 0.4 bits / s / Hz, while the codeword rate ranges from 0.5 bits / s / Hz to 1.6 bits / s / Hz. In Figure 10 , the codeword rates of both users are fixed at τ m,t = τ n,t = 1.2 bits / s / Hz, and the secrecy rate ranges from 0 bits / s / Hz to 1.4 bits / s / Hz. It can be seen that the EST of the system increases with the increase of the codeword rate and the secrecy rate until it decreases after a certain point. At τ m,t = τ n,t = 1.3 bits / s / Hz and τ m,s = τ n,s = 0.5 bits / s / Hz, the EST reaches the maximum value. Therefore, appropriately setting the parameter values of the codeword rate and the secrecy rate can effectively improve the security performance of the system.

[0121] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A dynamic power allocation method for NOMA-CR networks based on deep Q-learning, characterized in that, It includes the following steps: Step 1: Construct a two-layer network model composed of a source node, cooperative nodes, and users. Each communication link is described by a Rayleigh fading and AWGN noise model; Step 2: Design a state space, where the state vector includes the real-time signal-to-noise ratio information of the source node and cooperative nodes, and contains the dynamic changes of the energy state of the relay node; Step 3: Construct a discretized action space, and discretize the power adjustment of the source node and cooperative nodes into multiple combined actions; Step 4: Design a composite reward function according to the communication performance reward, interference management reward, and energy reward; Step 5: Use a deep Q-network for training, and through experience replay, target network update, and ε-greedy strategy, realize the dynamic decision-making of the optimal power adjustment action in the state space; Step 6: Adjust the transmission power of the source node and cooperative nodes according to the selected action, and update the energy of the relay node according to the channel state and the source node power. Only when the energy reaches the preset threshold is the relay node allowed to participate in forwarding.

2. The dynamic power allocation method for the NOMA-CR network based on deep Q-learning according to claim 1, wherein For the construction of the two-layer network model composed of a source node, cooperative nodes, and users described in Step 1, each communication link is described by a Rayleigh fading and AWGN noise model, which is specifically as follows: Step 1.1: Construct a two-layer network model, including a primary network and a secondary network. The primary network is composed of a source node and a target user to ensure the normal communication of the user; the secondary network is composed of a source node, multiple relay nodes, legitimate users, and eavesdroppers; all nodes use single antennas and operate in a half-duplex mode; each relay node is randomly distributed between the source node and the user to form a rectangular area with a set geometric range; Step 1.2: The transmission between all nodes adopts an independent Rayleigh fading channel model, and the noise adopts an additive Gaussian white noise AWGN model with an average value of zero and a variance of σ2; model the decode-and-forward DF and amplify-and-forward AF protocols respectively, analyze the signal-to-noise ratio SINR in each stage, and analyze the impact of the interference signal generated by the interfering relay on the eavesdropper and the interference suppression of the primary network communication link.

3. The dynamic power allocation method for the NOMA-CR network based on deep Q-learning according to claim 1, characterized in that The state space described in Step 2 is a two-dimensional vector, which respectively represents the signal-to-noise ratio SNR of the source node and cooperative nodes, and the signal-to-noise ratio SNR of the source node and cooperative nodes is discretized within a preset range.

4. The dynamic power allocation method for NOMA-CR network based on deep Q learning according to claim 1, characterized in that For the construction of the discretized action space described in Step 3, the power adjustment of the source node and cooperative nodes is discretized into multiple combined actions, which is specifically as follows: Construct a discretized action space, discretize the transmission power of the source node and cooperative nodes to form multiple action combinations. Each action combination corresponds to a set of power adjustment parameters, and the action index is mapped to the power selection of the source node and cooperative nodes through integer division and remainder operations respectively.

5. The dynamic power allocation method for NOMA-CR network based on deep Q-learning according to claim 1, wherein For the design of the composite reward function according to the communication performance reward, interference management reward, and energy reward described in Step 4, which is specifically as follows: Step 4.1: Design the communication performance reward, security reward, and energy reward, which is specifically as follows: Communication performance reward: Based on the indicators of the user signal decoding success rate, whether the target SINR reaches the preset threshold, and the system security throughput EST, a negative reward is given if the communication fails; Security Reward: By means of a dynamic interference management strategy, a suppression reward is imposed on the signal-to-noise ratio of eavesdroppers to ensure the secrecy performance; Energy Reward: The positive increment of the energy state of the relay node is multiplied by a preset coefficient and then added to the reward to encourage the system to consider the balance between energy harvesting and consumption during power allocation; Step 4.2: The communication performance reward, the interference management reward, and the energy reward are combined by weighting to obtain a composite reward function, thereby guiding the deep Q-network to converge to the optimal dynamic decision-making strategy.

6. The dynamic power allocation method for the NOMA-CR network based on deep Q-learning according to claim 1, wherein The training using the deep Q-network described in Step 5 realizes the dynamic decision-making of the optimal power adjustment action in the state space through experience replay, target network update, and ε-greedy strategy, specifically as follows: Step 5.1: Design a deep Q-network, which adopts a multi-layer neural network structure. The input layer receives the state vectors SNR_S and SNR_R, followed by a hidden layer with 64 nodes activated by ReLU, a hidden layer with 128 nodes activated by ReLU, and an output layer with 676 nodes, corresponding to the Q-values of all actions; Step 5.2: Train the deep Q-network: (1) Initialize the network parameters and the ε-greedy strategy parameters, and the exploration rate gradually decays from 1.0 to 0.1; (2) In each time slot, select the corresponding action according to the current state, and after executing the action, the environment feedbacks the next state and the composite reward; (3) Store the state transition samples in the experience replay pool, and randomly sample and update the network parameters using the mean square error (MSE) loss function; (4) Every fixed iteration period, copy the evaluation network parameters to the target network to ensure the training stability; (5) Repeat steps (2) to (4) until the system security throughput EST converges, and verify the advantage that the system performance under the DF protocol is better than that of the AF protocol through simulation.

7. The dynamic power allocation method for NOMA-CR network based on deep Q-learning according to claim 1, characterized in that The adjustment of the transmission power of the source node and the cooperative node according to the selected action described in Step 6, and the update of the relay node energy according to the channel state and the source node power. Only when the energy reaches the preset threshold is the relay node allowed to participate in forwarding, specifically as follows: Step 6.1: Establish a set of candidate forwarding nodes for the relay node, and put the relay nodes that meet the conditions into the set of candidate forwarding nodes, specifically as follows: The relay node must meet two conditions before participating in forwarding: Signal decoding ability: All relays that can accurately decode the signals sent by the source node form a decoding set; Energy threshold: Each relay node is equipped with a battery with infinite capacity. When the energy collected by a certain relay node reaches the preset threshold, this relay node is put into the set of candidate forwarding nodes; Step 6.2: Using the Rayleigh channel model, calculate the energy harvesting amount E of the relay node according to the channel gain and the transmission power of the source node h , and the formula is: E h = η × P S × |h| 2 × T s Among them, η is the energy conversion efficiency, P S is the source node power, ∣h∣ 2 is the channel gain, T s is the time slot length; When the energy collected by the relay node reaches or exceeds the set threshold, the relay node is allowed to participate in the decoding and forwarding of information, thereby realizing energy management and extending the system operation life.

8. A dynamic power allocation system for NOMA-CR networks based on deep Q-learning, characterized in that, This system is used to implement the dynamic power allocation method for the NOMA-CR network based on deep Q learning described in any one of claims 1 to 7. The system includes an energy harvesting module, a state acquisition module, a deep Q-network module, an action execution module, and a reward calculation module, where: The energy harvesting module is used to collect and update the energy state of the relay node according to the channel state and the transmission power of the source node; The state acquisition module is used to collect the signal-to-noise ratio information of the source node and the cooperative node in real time; The deep Q-network module is used to output an optimal action policy according to the collected state information; The action execution module is used to dynamically adjust the transmission powers of the source node and the cooperative node according to the information; The reward calculation module is used to calculate the composite reward and feedback it to the deep Q-network to continuously optimize the parameters.

9. The NOMA-CR network dynamic power allocation system based on deep Q-learning according to claim 8, characterized in that, The energy harvesting module calculates the energy harvesting amount of the relay node according to the channel state and the transmission power of the source node, and when the energy harvested by the relay node reaches a preset threshold, allows the relay node to participate in signal decoding and forwarding.

10. The NOMA-CR network dynamic power allocation system based on deep Q-learning according to claim 8, characterized in that, The deep Q-network module adopts an experience replay and target network update strategy, and uses an ε-greedy strategy to balance exploration and exploitation to achieve optimal decision-making for network dynamic power allocation.

Citation Information

Patent Citations

  • Optimal dynamic power distribution method and system based on uplink NOMA, medium and terminal

    CN112770395A

  • Cognitive radio dynamic power distribution method based on secondary user group cooperation

    CN113691334A