A joint uplink and downlink channel access and energy collection method for wireless energy-carrying communication
By optimizing the channel access and energy collection of wireless energy-carrying communication networks through multi-layer machine learning, the problem of device energy evolution and channel gain information not being taken into account is solved, and a significant improvement in uplink and downlink rates is achieved. It is suitable for communication and energy collection of IoT devices.
Patent Information
- Application Number
- CN202411627490.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-11-14
AI Technical Summary
Existing wireless energy-carrying communication networks do not consider the energy evolution of devices and imperfect channel gain information in channel access optimization, resulting in limited uplink and downlink rates, and traditional methods fail to effectively optimize the power division ratio of the receiver.
A multi-layer machine learning method is used to construct a time-varying channel gain model and an energy evolution model. Combined with the Markov decision model and the Q-learning algorithm, channel access and energy collection are optimized. Power allocation and transmission strategies are adjusted through centralized and distributed Q-learning to achieve optimal access in dynamic channel states.
It improves the uplink and downlink communication rates, achieving an average sum rate of 44b/s/Hz, which is 6 times better than traditional methods. It is suitable for channel access and energy collection of IoT devices, and improves the communication and energy collection efficiency of devices.
Smart Images

Figure CN119521263B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wireless communications, and in particular to a combined uplink and downlink channel access and energy collection method for wireless energy-carrying communications. Background Art
[0002] Future IoT networks will consist of low-power devices that sense their environment and transmit data to a gateway. The gateway can then use the data from the devices to instruct them to perform sensing tasks or control actuators. Wireless power-carrying communication technology has emerged by enabling devices to simultaneously transfer RF energy using the existing spectrum used for data transmission. A growing number of IoT networks equipped with RF-powered devices operate over a shared medium. Unlike traditional networks where devices have wireless power, in RF-powered IoT networks, devices must first harvest RF energy before transmitting and / or receiving data. In these scenarios, channel access is a key technology that facilitates uplink and downlink transmissions on the same channel. Because collisions may occur when devices upload their data to the gateway, determining the uplink and downlink channel access parameters given an unknown number of competing devices is a key issue.
[0003] Furthermore, the device's lifespan and the amount of data collected are limited by available energy. Therefore, to ensure long-term device operation, it is necessary to manage the available energy for devices that rely on hybrid access points for wireless energy harvesting. Wireless energy-carrying communications support both time-switching and power-splitting modes. When using a power splitter, the receiver divides the power of the received signal between its energy harvester and the data decoder. Furthermore, there are trade-offs between energy harvesting and information decoding, which affect the uplink and downlink rates, respectively. Specifically, less energy harvested results in lower uplink transmission power, which undoubtedly compromises the uplink rate. Conversely, less power allocated to information decoding directly compromises the downlink rate. Therefore, a key issue in energy harvesting is optimizing the receiver's power split ratio.
[0004] Currently, traditional wireless energy-carrying communication network channel access optimization mainly considers the pre-allocation of time slots / subcarriers and only optimizes one time frame. It adopts an ideal linear RF energy conversion model, does not consider the energy evolution of the device and imperfect channel gain information, and lacks consideration for actual conditions. Summary of the Invention
[0005] The purpose of the present invention is to provide a joint uplink and downlink channel access and energy collection method for wireless energy-carrying communication, which learns the optimal channel access method and energy collection method under different states in each stage through multi-layer machine learning to obtain the best uplink and downlink communication and rate.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] A method for joint uplink and downlink channel access and energy collection for wireless energy-carrying communication, comprising the following steps:
[0008] Step S1: construct a time-varying channel gain model to obtain the channel gain between the user and the hybrid access point in each time frame;
[0009] Step S2: constructing an energy evolution model, collecting energy through a radio frequency energy harvester equipped with each user, and obtaining the user's energy level according to the energy evolution model;
[0010] Step S3: constructing a Markov decision model for the wireless energy-carrying communication downlink, where the downlink state is constructed as the channel state of all devices, the downlink action is constructed as the power allocation of the superimposed signal, and the reward is constructed as the downlink throughput;
[0011] Step S4: construct a wireless energy communication uplink Markov decision model, where the uplink state is constructed as the channel state of a single device and its current power, the uplink action is constructed as the transmission probability and transmission time slot selection of a single device, and the reward is constructed as the uplink throughput;
[0012] Step S5: Centralized Q learning is used in the downlink layer. The hybrid access point acts as an agent. The hybrid access point selects the action corresponding to the maximum Q value in the current state according to the channel status of all devices in the current time frame, determines the power allocation of the signal, and then calculates the downlink sum rate of the current time frame as a reward based on the transmission results.
[0013] Step S6: Distributed Q learning is used in the uplink layer. Each device acts as an agent. Each device selects the action corresponding to the maximum Q value in the current state according to its own channel state and real-time power in the current time frame, determines its transmission probability and time slot selection, and then calculates the sum rate of the uplink in the current time frame as a reward based on the transmission result.
[0014] In step S7, the energy harvesting of step S2 is optimized using stateless Q learning. The system acts as an agent and selects the action with the highest probability in each cycle, controls the uplink frame size and downlink power split ratio, and then collects the rewards of the uplink and downlink in the current cycle as the stateless reward R κ , The system updates the Q value and probability mass function according to the reward value until the system converges, thereby optimizing energy harvesting.
[0015] Furthermore, in step S1, time is divided into frames and indexed by t, and the channel gain is obtained according to the following formula:
[0016]
[0017] Where, represents the channel gain between user i and the hybrid access point at time frame t, d i represents the distance between user i and the hybrid access point, n represents the path loss exponent, and λ represents an exponential random variable with mean 1.
[0018] Furthermore, in step S2, the energy level evolves according to the following formula:
[0019]
[0020] Where B max Indicates the battery capacity of each user, represents the energy collected by user i in time frame t, represents the energy consumed by user i in uplink transmission in time frame t, where
[0021]
[0022] In the formula, η represents the energy conversion efficiency, θ represents the power division ratio, represents the received power at the user, τ d Indicates the length of the downlink period in each frame.
[0023] Please note that users can only transmit if they have enough energy to transmit at least one packet, and each transmission uses their entire current energy. Assuming the smallest packet is 21 bytes and transmitting 1 bit requires 18nJ of energy, transmitting one packet requires at least 21×8×18=3024nJ.
[0024] Furthermore, in step S3, the downlink Markov decision model specifically includes: the downlink state is constructed as the channel state of all devices, that is, All possible states constitute the downlink state set, namely The downlink action is constructed as the power allocation of the superimposed signal, i.e. All possible actions constitute the downlink action set, i.e. Rewards are structured as downlink throughput.
[0025] Furthermore, in step S4, the uplink Markov decision model specifically includes: the uplink state is constructed as the channel state of a single device and its current power All possible states of each device constitute the uplink state set of the device Uplink actions are constructed as single device transmission probability and transmission time slot selection All possible actions of each device constitute the uplink action set of the device Rewards are structured as uplink throughput.
[0026] Furthermore, in step S5, during the system warm-up phase, the hybrid access point randomly selects a downlink action according to its state, and updates the state-action corresponding Q value according to the following Bellman equation:
[0027] Q(s t ,a t )=(1-α)Q(s t ,a t )+α(r(s t ,a t )+γmaxQ(s t+1 ,a t+1 ))
[0028] Where, Q(s t ,a t ) is state s t Next take action a t The corresponding Q value, r(s t ,a t ) is state s t Next take action a t The reward obtained, maxQ(s t+1 ,a t+1 ) is the maximum Q value corresponding to the next time slot state predicted by the system, α is the learning parameter, γ is the discount factor, γ∈[0,1].
[0029] Furthermore, in step S5, the power allocation should satisfy the requirement that the SIC of the device is successfully decoded and Where P is the total transmit power of the hybrid access point, is the transmission power allocated to user i in time frame t,
[0030] The sum rate of the downlink is obtained according to the following formula:
[0031]
[0032] Where W is the bandwidth, θ is the power split ratio, i is the current user, N is the total number of users, is the channel gain of user i in time frame t, is the signal transmission power allocated to user i in time frame t, and user j is the user other than user i, is the transmission power of the signal allocated to user j in time frame t, and n0 is the noise power.
[0033] Furthermore, in step S6, during the system warm-up phase, each device first randomly selects an upward action according to its state, and updates the corresponding Q value according to the Bellman equation.
[0034] Furthermore, in step S6, the sum rate of the uplink is obtained according to the following formula:
[0035]
[0036] Where W is the bandwidth, i is the current user, N is the total number of users, is the channel gain of user i in time frame t, is the signal transmission power of user i in time frame t, user j is the user other than user i, is the channel gain of user j in time frame t, is the signal transmission power of user j in time frame t, and n0 is the noise power.
[0037] Furthermore, in step S7, the system updates the Q value according to the reward value using the following formula:
[0038] Q(a)←Q(a)+λ(r(a)-Q(a))
[0039] Where Q(a) represents the system Q value corresponding to action a, λ represents the learning rate, λ∈[0,1], and r(a) represents the stateless reward function value corresponding to action a;
[0040] The probability mass function is:
[0041]
[0042] Where a i represents the current action, a represents all possible actions, Pr(a i ) indicates selecting action a i The probability of Q(a i ) indicates action a i The corresponding system Q value, T represents the continuous time frame.
[0043] By adopting the above technical solution, the present invention can achieve the following beneficial effects:
[0044] (1) The present invention divides the channel access method into three levels: downlink channel access based on non-orthogonal multiple access (NOMA), uplink channel access based on dynamic frame time slot Aloha with successive interference cancellation (SIC), and downlink wireless energy collection. Then, the channel access and energy collection parameters in each scenario are determined through multi-layer learning, fully considering the energy evolution of wireless energy-carrying communication users, the changes in user energy collection and data transmission with time-varying channel gain information, and the trade-offs of different uplink and downlink link transmission ratios. Uplink users and downlink hybrid access points can learn the optimal channel access method under different states in each stage through multi-layer machine learning. When the channel model is unknown, the channel state is accurately judged through a continuous learning process, thereby selecting the optimal channel access method to obtain the best uplink and downlink communication and rate.
[0045] (2) The present invention allows hybrid access points and devices to learn the optimal transmission strategy through a multi-layer Q learning (Multi-Q) solution. The downlink hybrid access point uses downlink NOMA to allocate different powers to the data of all users and superimpose them into a composite signal for transmission to all users. Each user collects energy from the received signal according to power segmentation and uses the SIC decoder to decode the information. Uplink users use frame time slot Aloha for uplink transmission. The hybrid access point does not allocate fixed time slots to the device and is equipped with a SIC decoder for information decoding. For each system state, the hybrid access point learns to optimize its downlink transmission power allocation, the device learns the time slot selection and transmission probability of each frame in the uplink, and the system learns the frame size and the optimal power segmentation ratio corresponding to different frame sizes. Multi-Q does not assume non-causal information and state transition probability, but takes into account the nonlinear energy conversion model, the energy evolution of the device and the imperfect causal channel gain from a practical perspective.
[0046] (3) Simulation results show that Multi-Q achieves an average sum rate of 44b / s / Hz, which is 6 times that of Aloha, 2.3 times that of TDMA, and 30% more than polling. This is because Multi-Q, a learning-based approach, can flexibly schedule the system to respond to different network conditions. Therefore, compared with Aloha, TDMA, etc., Multi-Q is able to obtain the best transmission strategy. The present invention can be applied to a variety of IoT deployments, such as smart agriculture, smart transportation, smart cities, etc., to better provide charging solutions and communications for agricultural sensors, parking sensors, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1is a flowchart of the wireless energy-carrying communication system in an embodiment of the present invention;
[0048] Figure 2 It is a time frame structure diagram in the present invention;
[0049] Figure 3 It is a flow chart of the multi-layer learning algorithm in the present invention;
[0050] Figure 4 is a flowchart of the centralized Q learning method at the downlink layer in the present invention;
[0051] Figure 5 This is a flow chart of the uplink layer distributed Q learning in the present invention;
[0052] Figure 6 It is the stateless Q learning flow chart of the present invention;
[0053] Figure 7 is the rate convergence curve of the uplink and downlink in the present invention;
[0054] Figure 8 is the sum rate convergence curve of different learning parameters in the present invention;
[0055] Figure 9 The effects of different HAP transmission powers on the average sum rate per frame, the average uplink rate per frame, and the average downlink rate per frame in the present invention are respectively;
[0056] Figure 10 The influence of different user positions on the average sum rate per frame, the average uplink rate per frame, and the average downlink rate per frame in the present invention;
[0057] Figure 11 This is the influence of different power split ratios on the average sum rate per frame, the average uplink rate per frame, and the downlink rate per frame in the present invention. DETAILED DESCRIPTION
[0058] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0059] like Figure 1 As shown, the wireless energy-carrying communication network in this embodiment includes a hybrid access point (HAP) that provides services to N energy-harvesting users / devices. Each device is denoted as i, where i∈[1,2,…,N]. The HAP uses NOMA in the downlink, and the device uses frame time slot Aloha for uplink transmission. The HAP is equipped with a SIC decoder for uplink information decoding. Time is divided into frames and indexed by t, with reference to Figure 2At the beginning of each time frame t, HAP sends pilot symbols for channel estimation. Afterwards, each frame is divided into a downlink period and an uplink period, each with a length of τ d and τ u During the downlink period, the device uses power splitting to divide the received power into two parts, namely energy collection and information decoding. After that, there are M time slots in the uplink, and the device selects a time slot from the M time slots to transmit its data packet.
[0060] The present invention provides a combined uplink and downlink channel access and energy collection method for wireless energy-carrying communication, comprising the following steps:
[0061] Step S1: construct a time-varying channel gain model to obtain the channel gain between the user and the hybrid access point in each time frame; the channel gain is obtained according to the following formula:
[0062]
[0063] Where, represents the channel gain between user i and the hybrid access point at time frame t, d i represents the distance between user i and the hybrid access point, n represents the path loss exponent, and λ represents an exponential random variable with mean 1.
[0064] Step S2: constructing an energy evolution model, collecting energy through a radio frequency energy harvester equipped with each user, and obtaining the energy level of the user according to the energy evolution model.
[0065] Each user is equipped with a RF energy harvester, such as the P2110B RF energy harvester. Represents the received power at the user, which is calculated as In this formula, is the channel gain of user i in time frame t, and P is the transmitting power. According to power division, Part of the power is transferred to the energy harvester. The RF energy conversion process is nonlinear and is a function of the received power. Therefore, this application adopts a nonlinear energy harvesting model based on reality. The energy conversion efficiency is expressed as η, which ranges from [0,1]. The energy conversion efficiency is calculated as in,
[0066]
[0067] Ω=1 / (1+e ab )
[0068]
[0069] Where M represents the maximum energy harvesting power, and the values of a and b depend on the circuit of the RF energy harvester.
[0070] make Expressed as the energy collected by user i in time frame t, the collected energy is calculated as:
[0071]
[0072] In the formula, η represents the energy conversion efficiency, θ represents the power division ratio, represents the received power at the user, τ d Indicates the length of the downlink period in each frame.
[0073] Energy levels evolve according to the following formula:
[0074]
[0075] Where B max Indicates the battery capacity of each user, represents the energy collected by user i in time frame t, represents the energy consumed by user i in uplink transmission in time frame t.
[0076] Step S3: construct a Markov decision model for the wireless energy-carrying communication downlink. The downlink state is constructed as the channel state of all devices, i.e. All possible states constitute the downlink state set, namely The downlink action is constructed as the power allocation of the superimposed signal, i.e. All possible actions constitute the downlink action set, i.e. Rewards are structured as downlink throughput.
[0077] Step S4: construct a wireless energy communication uplink Markov decision model, where the uplink state is constructed as the channel state of a single device and its current power. All possible states of each device constitute the uplink state set of the device Uplink actions are constructed as single device transmission probability and transmission time slot selection All possible actions of each device constitute the uplink action set of the device Rewards are structured as uplink throughput.
[0078] Step S5: Centralized Q learning is used in the downlink layer. The hybrid access point acts as an agent. The hybrid access point selects the action corresponding to the maximum Q value in the current state according to the channel status of all devices in the current time frame, determines the power allocation of the signal, and then calculates the downlink sum rate of the current time frame as a reward based on the transmission results.
[0079] First, initialize the centralized Q-learning parameters α, γ∈[0,1] and the centralized Q table. During the system warm-up phase, the hybrid access point randomly selects a downlink action based on its state and updates the Q value corresponding to the state-action according to the following Bellman equation:
[0080] Q(s t ,a t )=(1-α)Q(s t ,a t )+α(r(s t ,a t )+γmaxQ(s t+1 ,a t+1 ))
[0081] Where, Q(s t ,a t ) is state s t Next take action a t The corresponding Q value, r(s t ,a t ) is state s t Next take action a t The reward obtained (i.e., downlink throughput in time frame t), maxQ(s t+1 ,a t+1 ) is the maximum Q value corresponding to the next time slot state predicted by the system, α is the learning parameter, γ is the discount factor, γ∈[0,1].
[0082] After the warm-up period, at the beginning of each time slot, the hybrid access point will Select the action corresponding to the maximum Q value with probability (1-ε), and decay the ε value, that is, Determine the power distribution of its signal And ensure that the power allocation meets the device's SIC successful decoding and In this formula, P is the total transmit power of the hybrid access point, is the transmission power allocated to user i in time frame t, and then the downlink sum rate is calculated as the reward based on its transmission results, that is, the downlink throughput.
[0083] The sum rate of the downlink is obtained according to the following formula:
[0084]
[0085] Where W is the bandwidth, θ is the power split ratio, i is the current user, N is the total number of users, is the channel gain of user i in time frame t, is the signal transmission power allocated to user i in time frame t, and user j is the user other than user i, is the transmission power of the signal allocated to user j in time frame t, and n0 is the noise power.
[0086] Furthermore, the device continues to update the Q value according to the Bellman formula based on the current selected action and the downlink throughput value. This step is repeated for each subsequent time slot until the channel access decision process is completed.
[0087] Step S6: Distributed Q learning is used in the uplink layer. Each device acts as an agent. Each device selects the action corresponding to the maximum Q value in the current state according to its own channel state and real-time power in the current time frame, determines its transmission probability and time slot selection, and then calculates the sum rate of the uplink in the current time frame as a reward based on the transmission result.
[0088] First, each device initializes its independent distributed Q learning parameters α, γ∈[0,1] and distributed Q table. During the system warm-up phase, each device first randomly selects an upward action based on its state and updates the corresponding Q value according to the Bellman formula:
[0089] Q(s t ,a t )=(1-α)Q(s t ,a t )+α(r(s t ,a t )+γmax Q(s t+1 ,a t+1 ))
[0090] Where, Q(s t ,a t ) is state s t Next take action a t The corresponding Q value, r(s t ,a t ) is state s t Next take action a t The reward obtained (i.e., downlink throughput in time frame t), maxQ(s t+1 ,a t+1 ) predicts the maximum Q value corresponding to the next time slot state for each device, α is the learning parameter, γ is the discount factor, γ∈[0,1].
[0091] After the warm-up, at the beginning of each time slot, each device will be charged according to its channel status and real-time power. Select the action corresponding to the maximum Q value with probability (1-ε), and decay the ε value, that is, Determine its transmission probability and time slot selection Then the sum rate of the uplink is calculated as a reward based on its transmission results, that is, the uplink throughput.
[0092] The sum rate of the uplink is obtained according to the following formula:
[0093]
[0094] Where W is the bandwidth, i is the current user, N is the total number of users, is the channel gain of user i in time frame t, is the signal transmission power of user i in time frame t, user j is the user other than user i, is the channel gain of user j in time frame t, is the signal transmission power of user j in time frame t, and n0 is the noise power.
[0095] Furthermore, the device continues to update the Q value according to the Bellman formula based on the current selected action and the uplink throughput value. This step is repeated for each subsequent time slot until the channel access decision process is completed.
[0096] In step S7, the energy harvesting of step S2 is optimized using stateless Q learning. The system acts as an agent and selects the action with the highest probability in each cycle, controls the uplink frame size and downlink power split ratio, and then collects the rewards of the uplink and downlink in the current cycle as the stateless reward R κ , The system updates the Q value and probability mass function according to the reward value until the system converges, thereby optimizing energy harvesting.
[0097] First, the system initializes the learning rate λ∈[0,1], probability mass function (PMF) and stateless Q table of stateless Q learning. The probability mass function calculates the probability of taking each action a i The probability is expressed as:
[0098]
[0099] Where a i represents the current action, a represents all possible actions, Pr(a i ) indicates selecting action a i The probability of Q(a i ) indicates action a i The corresponding system Q value, T represents the continuous time frame.
[0100] In the warm-up phase, the system randomly selects an action. After the warm-up, at the beginning of each cycle κ, the system will select the action a with the highest probability with probability (1-ε) κ = argmax Pr(a), and decays the value of ε. κ= [M, θ] controls the uplink frame size M and the downlink power split ratio θ. Afterwards, the system collects the uplink and downlink rewards during period κ to obtain the cumulative sum rate of the uplink and downlink during this period as the stateless reward of period κ:
[0101] Then, the system updates the Q value based on the reward value using the following formula:
[0102] Q(a)←Q(a)+λ(r(a)-Q(a))
[0103] Where Q(a) represents the system Q value corresponding to action a, λ represents the learning rate, λ∈[0,1], and r(a) represents the stateless reward function value corresponding to action a;
[0104] The probability mass function is:
[0105]
[0106] Where a i represents the current action, a represents all possible actions, Pr(a i ) indicates selecting action a i The probability of Q(a i ) indicates action a i The corresponding system Q value, T represents the continuous time frame.
[0107] The present invention adopts multi-layer Q learning, namely Multi-Q. The downlink layer learns the downlink superposition signal power distribution, the uplink layer learns the uplink device time slot selection and transmission probability, and the stateless layer learns the system time frame size and power division ratio. Figure 3 This figure shows the Multi-Q framework. Traditional Q-learning is used for the downstream and upstream layers, while stateless Q-learning is used for the stateless layers. All layers use ε-greedy for action selection. Therefore, after the warmup phase, each agent randomly selects an action with probability ε, after which the value of ε is decayed to ensure convergence.
[0108] Specifically, at the beginning of each cycle κ, the system uses stateless Q learning to learn the probability mass function PMF to calculate the probability of taking each action. The action a with the maximum probability is selected. κ= [M,θ] to determine the power split ratio θ for energy collection in the current cycle and the uplink frame size M. Among them, θ represents the proportion of received power dedicated to information decoding. The remaining (1-θ) part of the received power will be sent to the energy harvester. During the warm-up period, the system randomly determines the uplink frame size and downlink power split ratio. For each frame size and power split ratio, the Q table will update the Q value based on the sum rate. After several cycles, the system will select the frame size and power split ratio with the highest Q value. Frame sizes and power allocation ratios that achieve high system sum rates will obtain higher Q values. Each time an action is selected, its Q value will be updated based on its past rewards, current rewards, and possible future rewards.
[0109] During each downlink time frame of this cycle, HAP acts as an agent and adopts downlink layer centralized Q learning. HAP Select the action corresponding to the maximum Q value, that is Determine the power distribution of its signal All signals are then superimposed according to this power allocation, and the resulting composite signal is transmitted to all users. The Q table updates the Q values corresponding to different power allocations under different channel conditions based on the downlink throughput achieved by the selected action. Power allocations that achieve high throughput will have high Q values. Each time a power allocation is selected, its Q value is updated based on its past throughput, current throughput, and predicted future throughput. After several cycles, the HAP will primarily select the power allocation with the highest Q value, aiming for high downlink throughput. Therefore, as the learning process converges, the optimal power allocation for each given channel state will achieve the highest Q value.
[0110] During each uplink time frame, users use frame slot Aloha for uplink transmission, and the hybrid access point is equipped with a SIC decoder to decode the information. Each device is an intelligent agent that learns its uplink behavior independently. Each device learns its uplink behavior based on its channel status and real-time power consumption. Select the action corresponding to the maximum Q value: Determine its transmission probability and time slot selection Then each user uses the transmission probability ρ that he has learned independently i Select the learned time slot δ i Uplink transmissions use all available energy. This means that, in each frame, each device learns to select a specific uplink transmission slot and transmission probability. Similar to the downlink layer, the Q-table updates the Q value of each action-state pair until convergence. Therefore, for a given channel state, the uplink transmission slot and transmission probability with the highest transmission rate will achieve the highest Q value.
[0111] In order to further illustrate the effectiveness and superiority of the present invention, a simulation experiment is conducted on the suboptimal decision fusion rule adopted by the present invention, and the suboptimal decision fusion rule is compared with the existing optimal decision fusion rule.
[0112] First, we studied the convergence, with the user placed at distances of 1m, 5m, and 9m from the HAP, and ran the simulator for 200 iterations, with each iteration containing 150 frames. Figure 7 The uplink and downlink rates are plotted in . The warm-up period is 15,000 frames. Figure 7 , we can see that the uplink and downlink rates converged after 140 iterations. Specifically, the downlink rate converged to around 34b / s / Hz, and the uplink rate converged to around 13.66b / s / Hz.
[0113] Secondly, we study the uplink and downlink layer learning rates, including the uplink learning rate α u , downlink learning rate α d , no state learning rate λ, discount factor γ and warmup period. Figure 8 It can be seen that each learning parameter combination converges to a different sum rate. Due to the short warm-up period, the system will experience more randomness during convergence. When the stateless layer uses a suitable learning rate λ, the system converges faster, from Figure 8 It can be seen that when the learning rate λ = 0.1, the system converges fastest and more stably. Therefore, considering all factors, this solution prefers λ = 0.1, α u =0.1,α d =0.1, γ=0.9, warmup=10000.
[0114] Next, we will compare Multi-Q with polling, TDMA, and Aloha. Polling protocols are used for both uplink and downlink transmissions; that is, the HAP transmits downlink data to each device in turn, and devices transmit uplink data to the HAP in turn. TDMA and slotted Aloha are used only for uplink transmissions. Both protocols consider downlink NOMA and uniform power distribution. During the uplink, TDMA allocates a dedicated time slot to each device. With slotted Aloha, devices with sufficient energy randomly compete for uplink time slots, and the HAP is equipped with a SIC decoder for information decoding. We will measure and compare the performance of these protocols in terms of average system sum rate, average downlink transmission rate, and average uplink transmission rate.
[0115] In addition, we studied different HAP transmit powers, device locations, and power allocation ratios. Each simulation had 30,000 time frames, and the results of the last 300 frames after convergence were collected and plotted as the average of 10 simulation runs. The TDMA frame size was 3, and the path loss exponent was n = 2.7. In terms of computational complexity, Multi-Q involves three layers of Q-learning. Analyzing the computational complexity of each server at time t, for both the downlink and uplink layers, each server needs to determine the Q-value for each state-action pair; for the stateless layer, the server needs to determine the Q-value for each action. For each layer, the Q-value is updated solely based on the server's reward. Furthermore, Multi-Q is adaptable to larger networks. However, the computational complexity of the downlink layer may increase with scale because it requires global information on the server.
[0116] For different HAP transmission powers, Figure 9 (a) shows the average sum rate of uplink and downlink. In all methods, the total rate of uplink and downlink increases with the increase of HAP transmission power. Specifically, both uplink and downlink rates increase with the increase of HAP transmission power, as shown in Figure 9 (b) shows the increase in uplink rate because the user collects more energy at a higher HAP transmit power. Therefore, the user is able to transmit at a higher transmit power. Similarly, higher HAP transmit power results in higher downlink rates. Multi-Q performance is best when the HAP transmit power is changed from 1W to 5W.
[0117] from Figure 9 (a) It can be clearly seen that Multi-Q consistently achieves the highest sum rate when the HAP transmit power ranges from 1 W to 5 W. Furthermore, Multi-Q achieves an average rate of 44.3 b / s / Hz, which is six times that of Aloha, 2.3 times that of TDMA, and 30% higher than polling. Figure 9(b) shows that Multi-Q achieves an average downlink transmission rate of 31.9 b / s / Hz, the highest among all methods and 6.6 b / s / Hz higher than polling. In contrast, TDMA and Slotted Aloha achieve a zero downlink rate. This is because both TDMA and Slotted Aloha use a uniform power allocation for each user during the downlink period, which can lead to decoding failures. Furthermore, Multi-Q outperforms polling in both the uplink and downlink. Specifically, Multi-Q achieves an average uplink and downlink rate of 12.4 b / s / Hz and 31.9 b / s / Hz, respectively, which are 6.6 b / s / Hz and 4.1 b / s / Hz higher than polling. This is because Multi-Q users utilize the entire downlink period to receive data, while polling users only receive data when the HAP polls for data. On the downlink, polling can encounter idle slots when users lack sufficient energy. However, Multi-Q can dynamically adjust frame size and transmission probability based on user battery level and channel conditions to avoid idle slots. Overall, Multi-Q has a significant advantage in sum rate, which is 6 times that of the frame-slot Aloha method. It also shows a significant advantage in downlink rate, which is 26% higher than polling and 100% higher than TDMA and frame-slot Aloha.
[0118] For different user locations, we consider five groups of user locations. The distances (in meters) from each user to the HAP are as follows: [5, 5, 5], [4, 5, 6], [3, 5, 7], [2, 5, 8], and [1, 5, 9]. The total distance from each group of devices to the HAP is 15 meters. The users are moved to different locations to obtain different channel gain conditions. The SIC decoding is improved with the channel gain difference. Specifically, the HAP transmission power is 3 watts, and the frame size of Aloha and TDMA is 3. The path loss exponent is n = 2.7. Figure 10As shown in (a), the average uplink and downlink sum rates increase when the user distances vary more. This is because when users are placed at different distances from the HAP, their channel gains experience significant differences. Consequently, the energy collected by the users varies significantly. This also means that one user located near the HAP will transmit at high power, while another user located further away will transmit at low power. This difference in transmit power helps increase the number of successful SIC decodings. Overall, Multi-Q outperforms all other methods. When user locations vary significantly, Multi-Q achieves an average sum rate of 32 b / s / Hz across locations. Meanwhile, the average sum rates for Aloha, polling, and TDMA are 2.5 b / s / Hz, 29 b / s / Hz, and 4 b / s / Hz, respectively. When users are 1, 5, and 9 meters away from the HAP, Multi-Q achieves a sum rate of 44 b / s / Hz, six times that of Aloha, three times that of TDMA, and 30% higher than that of polling. The reason why Multi-Q performs better is that the downlink rate of Aloha and TDMA is zero, e.g. Figure 10 (b) Because both Aloha and TDMA use uniform power allocation in the downlink, they always experience glitches during downlink transmission. Multi-Q outperforms polling because it simultaneously learns the frame size and transmission probability to avoid idle time slots. Overall, Multi-Q consistently achieves the highest sum rate across different user locations, averaging 11 times that of Aloha, 7 times that of TDMA, and 10% higher than polling. Furthermore, Multi-Q demonstrates a significant advantage in downlink rate, exceeding TDMA and Aloha by 100% and polling by 5%.
[0119] For different initial power split ratios, we set the users at distances of 1 m, 5 m, and 9 m from the HAP. The HAP transmission power is 3 W, the frame size of TDMA and Aloha is 3, and the path loss exponent is n = 2.7. The power split ratio is gradually changed from 0 to 1 with a step size of 0.1. Initially, the power split ratio is zero, which means that all the received power is used for energy harvesting. After that, the power split ratio is increased to 0.1. Therefore, 10% of the power is directed for data reception and the remaining 90% is used for energy harvesting. The power split ratio is then increased in steps of 0.1 until it reaches 1.0. Reference Figure 11 (a), Multi-Q achieves a sum rate of approximately 44b / s / Hz; it is able to converge to the optimal power split ratio starting from any initial ratio. The polled sum rate continues to increase within the power split ratio range up to 0.9. This is because when the power split ratio is increased, more power is allocated to downlink data reception and less power is used for energy harvesting. Figure 11(b) As can be seen, polling achieves an average downlink rate of 25.3 b / s / Hz and an uplink rate of 5.7 b / s / Hz. It can be seen that the downlink rate of polling is approximately five times the uplink rate. Therefore, even though the uplink rate of polling is reduced, the sum rate of polling increases because the increase in downlink rate is greater than the decrease in uplink rate. When the power split ratio is increased from 0.9 to 1.0, the sum rate of polling decreases because all downlink power is allocated to the energy harvester. When the power split ratio is increased from 0 to 1.0, the sum rate of Aloha and TDMA decreases because the users use a uniform power allocation for the downlink, resulting in no difference between each user's transmissions and no ability for each user to decode the received packets. Furthermore, a higher power split ratio results in less energy harvesting and uplink transmit power, resulting in a significantly lower uplink rate.
[0120] We also studied the performance of Aloha with different frame sizes. Figure 11 As shown in (a), Aloha achieves the highest sum rate when the frame size is 1. Although larger frames mean fewer collisions, a frame size of 1 outperforms larger frames. This is because we consider SIC, which allows decoding of multiple concurrent transmissions in the same time slot. In addition, a smaller frame size extends the transmission period, so users can transmit for a longer time. A frame size of 1 means that all transmissions occur in one time frame, but only transmissions whose SINR of the transmitting user power meets the SIC decoding conditions will be successful. In the Aloha method, users do not learn transmission probabilities to avoid SIC decoding failures, and all users can only transmit concurrently in the same time slot. Multi-Q improves the success probability of SIC decoding at the receiver by learning the optimal time frame size, uplink transmission probability, and time slot selection. Therefore, Multi-Q performs the best among all methods.
[0121] Multi-Q achieves an average sum rate of 44 b / s / Hz, 9.8 times higher than TDMA, 6.5 times higher than Aloha when the frame size is 1, and 1.4 times higher than polling. This is because Multi-Q simultaneously learns the power split ratio, frame size, uplink transmission probability, uplink timeslot selection, and downlink power allocation for all users. Specifically, Multi-Q learns the optimal frame size and timeslot selection to avoid uplink decoding failures and idle timeslots. Learning the downlink power allocation for each user enhances SIC decoding success. It also learns the power split ratio to balance uplink and downlink rates. Overall, for different preset split ratios, Multi-Q shows a significant advantage over all other methods in terms of sum rate. Furthermore, Multi-Q consistently achieves the highest downlink and uplink rates of 31.5 b / s / Hz and 12.5 b / s / Hz, respectively.
[0122] The present invention fully considers the actual wireless energy-carrying communication network, adopts nonlinear radio frequency energy conversion and energy evolution of the device, and considers the causal channel gain model. Therefore, this method has good applicability in real wireless energy-carrying communication networks.
[0123] Finally, it should be noted that the parts of the present invention that are not described in detail are all prior art. Those skilled in the art will understand that the above description is only a preferred embodiment of the invention and is not intended to limit the invention. Although the invention has been described in detail with reference to the above examples, those skilled in the art can still modify the technical solutions described in the above examples or replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, etc. made within the spirit and principles of the invention should be included in the scope of protection of the invention.
Claims
1. A method for joint uplink and downlink channel access and energy collection for wireless energy-carrying communication, characterized in that: The following steps are involved: Step S1: construct a time-varying channel gain model to obtain the channel gain between the user and the hybrid access point in each time frame; Step S2: constructing an energy evolution model, collecting energy through a radio frequency energy harvester equipped with each user, and obtaining the user's energy level according to the energy evolution model; Step S3: constructing a Markov decision model for the wireless energy-carrying communication downlink, where the downlink state is constructed as the channel state of all devices, the downlink action is constructed as the power allocation of the superimposed signal, and the reward is constructed as the downlink throughput; Step S4: construct a wireless energy communication uplink Markov decision model, where the uplink state is constructed as the channel state of a single device and its current power, the uplink action is constructed as the transmission probability and transmission time slot selection of a single device, and the reward is constructed as the uplink throughput; Step S5: Centralized Q learning is used in the downlink layer. The hybrid access point acts as an agent. The hybrid access point selects the action corresponding to the maximum Q value in the current state according to the channel status of all devices in the current time frame, determines the power allocation of the signal, and then calculates the downlink sum rate of the current time frame as a reward based on the transmission results. Step S6: Distributed Q learning is used in the uplink layer. Each device acts as an agent. Each device selects the action corresponding to the maximum Q value in the current state according to its own channel state and real-time power in the current time frame, determines its transmission probability and time slot selection, and then calculates the sum rate of the uplink in the current time frame as a reward based on the transmission result. In step S7, the energy harvesting of step S2 is optimized using stateless Q learning. The system acts as an agent and selects the action with the highest probability in each cycle, controls the uplink frame size and downlink power split ratio, and then collects the rewards of the uplink and downlink in the current cycle as the stateless reward R κ , The system updates the Q value and probability mass function according to the reward value until the system converges, thereby optimizing energy harvesting.
2. The method for joint uplink and downlink channel access and energy collection for wireless energy-carrying communication according to claim 1, characterized in that: In step S1, time is divided into frames and indexed by t, and the channel gain is obtained according to the following formula: Where, represents the channel gain between user i and the hybrid access point at time frame t, d i represents the distance between user i and the hybrid access point, n represents the path loss exponent, and λ represents an exponential random variable with mean 1.
3. The method for joint uplink and downlink channel access and energy collection for wireless energy-carrying communication according to claim 2, characterized in that: In step S2, the energy level evolves according to the following formula: Where B max Indicates the battery capacity of each user, represents the energy collected by user i in time frame t, represents the energy consumed by user i in uplink transmission in time frame t, in, In the formula, η represents the energy conversion efficiency, θ represents the power division ratio, represents the received power at the user, τ d Indicates the length of the downlink period in each frame.
4. The method for joint uplink and downlink channel access and energy collection for wireless energy-carrying communication according to claim 1, characterized in that: In step S3, the downlink Markov decision model specifically includes: the downlink state is constructed as the channel state of all devices, that is, All possible states constitute the downlink state set, namely The downlink action is constructed as the power allocation of the superimposed signal, i.e. All possible actions constitute the downlink action set, i.e. Rewards are structured as downlink throughput.
5. The method for joint uplink and downlink channel access and energy collection for wireless energy-carrying communication according to claim 4, characterized in that: In step S4, the uplink Markov decision model specifically includes: the uplink state is constructed as the channel state of a single device and its current power All possible states of each device constitute the uplink state set of the device Uplink actions are constructed as single device transmission probability and transmission time slot selection All possible actions of each device constitute the uplink action set of the device Rewards are structured as uplink throughput.
6. The method for joint uplink and downlink channel access and energy collection for wireless energy-carrying communication according to claim 5, characterized in that: In step S5, during the system warm-up phase, the hybrid access point randomly selects a downlink action based on its state and updates the Q value corresponding to the state-action according to the following Bellman equation: Q(s t ,a t )=(1-α)Q(s t ,a t )+α(r(s t ,a t )+γmaxQ(s t+1 ,a t+1 )) Where, Q(s t ,a t ) is state s t Next take action a t The corresponding Q value, r(s t ,a t ) is state s t Next take action a t The reward obtained, maxQ(s t+1 ,a t+1 ) is the maximum Q value corresponding to the next time slot state predicted by the system, α is the learning parameter, γ is the discount factor, γ∈[0,1].
7. The method for joint uplink and downlink channel access and energy collection for wireless energy-carrying communication according to claim 6, characterized in that: In step S5, the power allocation should satisfy the requirement that the SIC of the device is successfully decoded and Where P is the total transmit power of the hybrid access point, is the transmission power allocated to user i in time frame t, The sum rate of the downlink is obtained according to the following formula: Where W is the bandwidth, θ is the power split ratio, i is the current user, N is the total number of users, is the channel gain of user i in time frame t, is the signal transmission power allocated to user i in time frame t, and user j is the user other than user i, is the transmission power of the signal allocated to user j in time frame t, and n0 is the noise power.
8. The method for joint uplink and downlink channel access and energy collection for wireless energy-carrying communication according to claim 7, characterized in that: In step S6, during the system warm-up phase, each device first randomly selects an upward action according to its state, and updates the corresponding Q value according to the Bellman equation.
9. The method for joint uplink and downlink channel access and energy collection for wireless energy-carrying communication according to claim 8, characterized in that: In step S6, the sum rate of the uplink is obtained according to the following formula: Where W is the bandwidth, i is the current user, N is the total number of users, is the channel gain of user i in time frame t, is the signal transmission power of user i in time frame t, user j is the user other than user i, is the channel gain of user j in time frame t, is the signal transmission power of user j in time frame t, and n0 is the noise power.
10. The method for joint uplink and downlink channel access and energy collection for wireless energy-carrying communication according to claim 1, characterized in that: In step S7, the system updates the Q value according to the reward value using the following formula: Q(a)←Q(a)+λ(r(a)-Q(a)) Where Q(a) represents the system Q value corresponding to action a, λ represents the learning rate, λ∈[0,1], and r(a) represents the stateless reward function value corresponding to action a; The probability mass function is: Where a i represents the current action, a represents all possible actions, Pr(a i ) indicates selecting action a i The probability of Q(a i ) indicates action a i The corresponding system Q value, T represents the continuous time frame.