A dynamic resource allocation method for a dual-port rectenna-oriented swipt system

CN122846402APending Publication Date: 2026-09-29SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611127686.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0006]本发明提供一种面向双端口整流天线的SWIPT系统动态资源分配方法,以解决现有SWIPT资源分配方法存在的非线性优化困难、在线策略部署不便、联合优化复杂度高等问题

Benefits of technology

1. 规避非凸优化难题:通过将功率分流比ρ解析求解,将其从优化变量中移除,外层仅需优化离散时间参数,从根本上避免了非线性能量收集模型带来的非凸性,降低了求解复杂度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122846402A_ABST
    Figure CN122846402A_ABST
Patent Text Reader

Abstract

The application discloses a dynamic resource allocation method for a SWIPT system facing a two-port rectification antenna. Through a double-layer heterogeneous framework of "outer-layer discrete-time reinforcement learning optimization + inner-layer analytical power splitting", offline optimal configuration of charging time and communication time is realized, and a power splitting ratio is dynamically adjusted in real time according to a channel state, so as to meet fixed rate requirements and long-term energy neutral constraints.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless power-carrying communication technology, specifically to a dynamic resource allocation method for a SWIPT system with a dual-port rectified antenna. Background Technology

[0002] With the large-scale deployment of the Internet of Things (IoT), wearable devices, and environmental monitoring sensors, the issue of long-term maintenance-free power supply for low-power nodes is becoming increasingly prominent. Traditional battery power supply suffers from drawbacks such as limited capacity, high replacement costs, and difficult maintenance. In unmanned environments or implantable medical scenarios, frequent battery replacements are neither practical nor economical. Simultaneous Wireless Information and Power Transfer (SWIPT) technology, which transmits information and energy simultaneously through the same radio frequency signal, provides a feasible path to extend node lifespan.

[0003] The existing SWIPT resource allocation method has the following main technical defects: 1) Difficulties in nonlinear optimization: The conversion efficiency of actual rectifier circuits varies nonlinearly with input power (approximately piecewise linear in the low-power region). Traditional convex optimization and alternating optimization methods face inherent nonconvexity when dealing with nonlinear energy harvesting models, making the solution complex and heavily reliant on accurate channel statistics. Most studies use linear models to simplify the process, leading to discrepancies between the theoretical optimal solution and actual hardware performance.

[0004] 2) Disconnect between online policy and engineering deployment: In recent years, Deep Reinforcement Learning (DRL) has been introduced into the field of SWIPT resource allocation. However, the vast majority of these works output online policy networks, requiring nodes to perform neural network inference in real time based on channel conditions in each time slot. This is difficult for IoT nodes with extremely limited energy and computing power (such as sensors equipped with low-power MCUs) to handle. In actual deployment, nodes require a set of long-term stable fixed configuration parameters during the firmware burning stage, rather than continuous online decision-making.

[0005] In summary, a SWIPT resource allocation method needs to be proposed that can avoid the difficulties of nonlinear and nonconvex optimization, output fixed parameters, and accurately satisfy the dual constraints of communication rate and energy neutrality. Summary of the Invention

[0006] This invention provides a dynamic resource allocation method for a SWIPT system with a dual-port rectified antenna, in order to solve the problems of nonlinear optimization difficulties, inconvenient online strategy deployment, and high complexity of joint optimization in existing SWIPT resource allocation methods.

[0007] This invention provides a dynamic resource allocation method for a SWIPT system with a dual-port rectified antenna, the method comprising: A dual-port SWIPT system model is established, comprising a transmitter and a receiver. The transmitter uses a single antenna to transmit radio frequency (RF) signals, and the receiver uses a dual-port rectified antenna to receive RF signals. Each port is connected to a single-pole double-throw (SPDT) RF switch. Each switch has two selectable paths, path A and path B. Path A is used to send 100% of the received power to the energy harvesting branch, and path B is used to distribute the received power to the energy harvesting branch and the information decoding branch according to the power split ratio ρ. During the charging phase, the receiver's MCU controls the two switches to switch to path A, and during the communication phase, the receiver's MCU controls the two switches to switch to path B. Construct a DDQN network, define the state space, action space, and reward function, and train the DDQN network offline. Obtain the optimal charging time through the offline trained DDQN network. and optimal communication time The parameters are then stored in the receiver's MCU. Based on the obtained optimal charging time and optimal communication time Parameters: At the start of the charging phase, the receiving MCU controls two switches to switch to path A, and the charging phase continues. After a certain time, the communication phase begins. First, based on the target communication rate and the current instantaneous received power, the power split ratio ρ is calculated analytically using Shannon's formula. Then, the receiving MCU controls two switches to switch to path B and distributes power according to the power split ratio ρ. The communication phase continues. The event will end after a set time.

[0008] Furthermore, a DDQN network is constructed, defining the state space, action space, and reward function. The DDQN network is then trained offline to obtain the optimal charging time. and optimal communication time The parameters are then embedded into the MCU, specifically including: state space Defined as a four-dimensional vector:

[0009] Where: t is time. This represents the instantaneous received power, reflecting the current channel quality. This represents the current energy stored in the supercapacitor, reflecting the remaining energy level. This represents the length of the data queue, reflecting the degree of data backlog. This is a virtual rate queue used to constrain long-term average rates; each component is divided by its corresponding normalized benchmark. , , , This unifies the magnitude of all dimensions to .

[0010] Furthermore, a DDQN network is constructed, defining the state space, action space, and reward function. The DDQN network is then trained offline to obtain the optimal charging time. and optimal communication time The parameters are then embedded into the MCU, specifically including: Action space: Under energy neutrality constraint The theoretical minimum ratio of charging time to communication time is as follows. ,in For the power consumption of the receiving end MCU, For rectification efficiency, For instantaneous received power, This represents the expectation of random channel fading. For charging time, For communication time; Construct a discrete set of candidate actions while satisfying the energy neutrality constraint. :

[0011] That is, when t0=0.1, t1∈{1500,1700,1900}; when t0=0.2, t1∈{2900,3100,3300}, for a total of 6 discrete actions.

[0012] Furthermore, a DDQN network is constructed, defining the state space, action space, and reward function. The DDQN network is then trained offline to obtain the optimal charging time. and optimal communication time The parameters are then embedded into the MCU, specifically including: reward function for:

[0013] in For net energy reward, For efficiency-oriented items, This is a queue penalty item.

[0014] Furthermore, Net Energy Reward Linear piecewise design is adopted:

[0015] Net energy The energy collected during the charging phase is The energy consumed during the communication phase is ; Efficiency Guide This is used to guide agents to avoid extreme duty cycles; Queue penalty item This is used to prevent data queues from becoming too long.

[0016] Furthermore, a DDQN network is constructed, defining the state space, action space, and reward function. The DDQN network is then trained offline to obtain the optimal charging time. and optimal communication time The parameters are then embedded into the MCU, specifically including: After training is complete, save the main network parameters. , initial state Input the DDQN network, calculate the Q-value for each action, and select the action with the largest Q-value:

[0017] Extracting the optimal action Corresponding charging time and communication time And it is embedded into the firmware of the receiving MCU.

[0018] Furthermore, the DDQN network is a fully connected neural network, including an input layer, two hidden layers, and an output layer. The input layer includes 4 nodes, corresponding to a four-dimensional state vector; each of the two hidden layers has 64 nodes, with ReLU as the activation function; and the output layer includes 6 nodes, corresponding to the Q-values ​​of 6 discrete actions.

[0019] Furthermore, based on the target communication rate and the current instantaneous received power, an analytical method based on Shannon's formula is used to solve for the power split ratio ρ, specifically including: By Shannon's formula Back-engineering to meet the target communication rate Minimum signal-to-noise ratio required:

[0020] The minimum received power required for the information decoding branch is:

[0021] in This represents the equivalent noise power spectral density. Let the power allocated to the energy harvesting branch be... The power obtained by the information decoding branch is To ensure that the communication rate meets the standard, it must meet the following requirements. To maximize energy harvesting, the minimum permissible power of the information decoding branch should be taken, i.e.:

[0022] The MCU obtains the current instantaneous received power through RSSI measurement. .

[0023] This invention provides a dynamic resource allocation method for a SWIPT system with a dual-port rectified antenna, which has the following advantages: 1. Avoiding the problem of non-convex optimization: By analytically solving the power split ratio ρ, it is removed from the optimization variables. The outer layer only needs to optimize the discrete time parameters, which fundamentally avoids the non-convexity caused by the nonlinear energy harvesting model and reduces the solution complexity.

[0024] 2. Output parameters can be fixed and adapted for engineering deployment: DDQN offline training is used, and only the optimal time parameters are extracted after training convergence. , This is then embedded into the MCU firmware. During actual operation, the node does not require any online neural network inference; it only needs to execute a timer state machine and simple analytical formulas, resulting in extremely low power consumption and computing power overhead, which meets the engineering constraints of IoT nodes.

[0025] 3. Precisely meets fixed rate requirements: The inner-layer analytical power splitting method uses the closed-form solution of Shannon's formula to precisely lock the communication rate to the target value, rather than the "maximize rate" or "variable rate" mode, which is suitable for scenarios that require stable service quality.

[0026] 4. Energy-neutral self-powered: By pre-screening the action space through energy-neutral constraints, the selected time parameters are ensured to meet the long-term energy balance, and the system can operate permanently by relying on the collected radio frequency energy without external batteries. Attached Figure Description

[0027] Figure 1 This is a flowchart of the outer discrete-time optimization process in a dynamic resource allocation method for a dual-port rectified antenna SWIPT system, provided in one embodiment of the present invention. Figure 2 This is a convergence analysis curve of DDQN training in a dynamic resource allocation method for a dual-port rectified antenna SWIPT system provided in one embodiment of the present invention; Figure 3Long-period simulation results of a dynamic resource allocation method for a SWIPT system oriented towards a dual-port rectified antenna, provided in an embodiment of the present invention; Figure 4 The power shunt ratio ρ distribution histogram is provided in a dynamic resource allocation method for a SWIPT system with a dual-port rectified antenna, as an embodiment of the present invention. Detailed Implementation

[0028] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings. Similar elements in different embodiments are referred to by associated similar element reference numerals. In the following embodiments, many details are described to facilitate a better understanding of the invention. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other elements, materials, or methods. In some cases, certain operations related to the present invention are not shown or described in the specification. This is to avoid obscuring the core parts of the invention with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; they can fully understand the related operations based on the description in the specification and general technical knowledge in the art.

[0029] Furthermore, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can be rearranged or adjusted in a manner obvious to those skilled in the art. Therefore, the various orders in the specification and drawings are only for the clear description of a particular embodiment and do not imply a necessary order, unless otherwise stated that a particular order must be followed.

[0030] The first embodiment of this invention provides a dynamic resource allocation method for a SWIPT system oriented towards a dual-port rectified antenna. The following is in conjunction with... Figure 1 and Figure 2 A detailed explanation is provided. The main steps are as follows: Step S1: Establish a two-port SWIPT system model The system includes a transmitter and a receiver. The transmitter operates in the 2.45 GHz ISM band and transmits radio frequency signals using a single antenna. The receiver uses a dual-port rectified antenna, with each port connected to an ultra-low power single-pole double-throw (SPDT) RF switch. Each switch provides two selectable paths: Path A (Bypass Straight-through): The antenna port is directly connected to the energy harvesting (EH) stage circuit, and 100% of the received power enters the EH bus; Path B (Power Divider Path): The radio frequency signal passes through an adjustable power divider, and the received power is distributed to the EH branch (ρ times) and the information decoding (ID) branch (1-ρ times) according to the power division ratio ρ (0≤ρ≤1).

[0031] The EH outputs of the two ports are electrically connected in parallel to the same DC bus, and the ID outputs are connected in parallel to the same baseband signal bus.

[0032] The system dynamic channel model comprehensively considers free-space path loss (Friis formula), log-normal shadowing fading (standard deviation 6dB), and Rayleigh fading (non-line-of-sight multipath environment). The instantaneous received power P_r is expressed as: P_r = P_t · G_t · G_r · PL_FS · _shadow · ξ_Rayleigh. Where P_t is the transmit power, G_t and G_r are the antenna gain, and PL_FS is the free space path loss. _shadow is the shadow fading factor, and ξ_Rayleigh is the Rayleigh fading power factor.

[0033] The rectification efficiency η(P_r) adopts a piecewise linear model: when the input power corresponding to P_r is below -26dBm, η≈0.43; it linearly increases to 0.50 in the range of -26dBm to -20dBm; it linearly increases to 0.61 in the range of -20dBm to -10dBm; and it saturates at 0.61 when it is above -10dBm.

[0034] Step S2: Construct a two-tier resource allocation framework This invention proposes a two-layer heterogeneous framework that combines outer-layer discrete-time optimization with inner-layer analytical power splitting.

[0035] S2.1 Outer Layer: Construction of Discretized Candidate Action Sets for Charging Time and Communication Time Charging time and communication time These are key parameters that need to be embedded into the MCU firmware. First, the feasible regions for both parameters are determined based on the energy neutrality constraint. Energy neutrality requires that the energy collected by the system during long-term operation is not less than the energy consumed.

[0036] in This is the power consumption of the receiver MCU (typical value 19.47mW). For rectification efficiency, This represents the expected fading of the random channel. From this, we obtain the theoretical minimum ratio of charging time to communication time:

[0037] To reduce the action space dimension of reinforcement learning while satisfying the constraints of equation (2), this invention discretizes the continuous feasible region into a finite set of candidate actions. (Communication time) Two typical values ​​are chosen: 0.1s is suitable for bursty small data packets, and 0.2s is suitable for slightly larger data packets; for each... Charging time Three representative values ​​are selected, corresponding to the three scenarios of "just satisfying energy neutrality," "moderate engineering margin," and "sufficient engineering margin," respectively. The specific candidate action set is as follows:

[0038] Equation (3) gives a total of Each discrete action, and all combinations of actions automatically satisfy the energy neutrality constraint (2).

[0039] S2.2 Outer Layer: Design of DDQN State Space and Reward Function The state space of a DDQN agent is defined as a four-dimensional vector:

[0040] The physical meanings of each component are as follows: The instantaneous received power (W) reflects the current channel quality; The current energy stored in the supercapacitor (J) reflects the remaining energy level. The length of the data queue (in bits) reflects the degree of data backlog. This is a virtual rate queue (bit) used to constrain the long-term average rate. Each component is divided by the normalized baseline. , J、 , This unifies the magnitude of all dimensions to The location is nearby, which is beneficial for neural network training.

[0041] The reward function is designed with the core objective of guiding the agent to choose the action with the maximum net energy, supplemented by very light auxiliary terms to avoid extreme parameter selection:

[0042] Net Energy Reward This is the primary reward. Let the energy collected during the charging phase be... The energy consumed during the communication phase is Net energy The net energy reward uses a linear, segmented design:

[0043] When net energy is positive, a higher slope (100) is used to encourage the agent to prioritize actions with larger energy surplus; when net energy is negative, a lower slope (20) is used to give a moderate penalty without causing drastic fluctuations in training.

[0044] Efficiency Guide Used to guide agents to avoid extreme duty cycles, it has a small weight and does not affect the primary objective of maximizing net energy. Queue penalty term. This is used to prevent excessive data queue backlog. If the supercapacitor runs out of energy after the communication phase ends ( ), immediately terminate the current episode and give The punishment guides the intelligent agent to avoid inactions.

[0045] S2.3 Outer Layer: DDQN Network Structure and Training Process DDQN network is a fully connected neural network with the following structure: The input layer has 4 nodes corresponding to a four-dimensional state vector; the two hidden layers each have 64 nodes, with ReLU activation function; the output layer has 6 nodes corresponding to the Q-values ​​of 6 discrete actions.

[0046] DDQN updates network parameters by minimizing the following loss function:

[0047] in Here, E represents the expected value, which in this formula refers to the average loss calculated from multiple samples randomly drawn from the experience replay pool. It is an experience tuple representing a complete record of an agent's interaction with the environment; s: the current state of the environment; a: the action the agent decides to perform in the current state s; r: the immediate reward value from the environment after performing action a; The state of the environment at the next moment after action a is performed; In the current main network, the predicted Q-value is calculated for state s and action a; Target value The computation is performed using a dual-network decoupling method with DDQN:

[0048] In the formula Discount factor ( ), Main network parameters (used to select actions). These are the target network parameters (used to evaluate action value). It is the main network (with parameter θ) for the next state and actions The Q-value is estimated. r is the immediate reward (i.e., single-step reward) returned by the environment immediately after the agent performs action a. In the Bellman equation, it represents the direct gain of the current step, while γ... Q( The first part represents the estimate of future earnings, and the second part, when added together, constitutes the total target value y. The state at the next moment. In the next state, the main network (parameters) The optimal action is selected. DDQN decouples action selection from action evaluation, effectively mitigating the Q-value overestimation problem of standard DQN.

[0049] The hyperparameter settings are as follows: optimizer Adam, learning rate Target network soft update coefficient ( Experience replay pool capacity: 200,000; Mini-batch size: 128; Exploration strategy adopted: - Greedy, initial Multiply by 0.998 per round to a minimum of 0.05; maximum training rounds: 1500, with a maximum of 400 steps per round.

[0050] The training process is as follows: At the start of each round, the environment is reset to obtain the initial state. Each step is based on... - Greedy strategy from Select one action to execute, and the environment simulator will simulate charging. Seconds and communication The full cycle of seconds is calculated according to formula (5), and the reward is returned to the next state. Experience is then transferred. Store the samples in the experience replay pool. When the number of samples in the pool exceeds the size of the mini-batch, randomly sample the mini-batch, calculate the target value according to equation (8), calculate the loss according to equation (7), update the main network parameters using gradient descent, and then softly update the target network. Update the state after each step, and terminate the round early if the energy is exhausted. Decrease the exploration rate after each round. After training reaches 1500 rounds, save the main network parameters. .

[0051] S2.4 Outer Layer: Extraction of Optimal Time Parameters After training is complete, the initial state ( Take the statistical average value -25dBm. J, , Input the Q-network and calculate the Q-value for each action:

[0052] Extracting the optimal action Corresponding charging time and communication time It is embedded into the firmware of the receiving MCU.

[0053] S2.5 Inner Layer: Analytical Power Split Ratio Calculation Based on Shannon's Formula After the outer layer training converges, and The timing allocation is already embedded in the firmware and is macroscopically determined. However, during the communication phase, the instantaneous received power... As channel fading fluctuates in real time, the power shunt ratio needs to be dynamically adjusted to ensure that the rate of each frame accurately meets the target. This invention proposes an analytical calculation method based on Shannon's formula to achieve real-time closed-form solution of the power shunt ratio.

[0054] By Shannon's formula Back-engineering to meet the target rate Minimum signal-to-noise ratio required:

[0055] Further, the minimum received power required for the ID branch is obtained:

[0056] in Let be the equivalent noise power spectral density (including receiver noise figure). Assume the power ratio allocated to the EH branch by the adjustable power divider is . Then the power obtained by the ID branch is To ensure that the communication rate meets the standard, the following must be met: To maximize energy harvesting, the minimum permissible ID branch power should be selected, i.e.:

[0057] when At that time, the ID branch obtained The power is just enough to meet the 1Mbps speed requirement, and the remaining... All power is used for energy harvesting, achieving an optimal allocation where "communication just meets the requirements, and the remaining power is fully harvested." When At times, the system cannot guarantee the communication rate and prioritizes all power allocation to energy harvesting. Under the typical parameters of this invention (average received power approximately -25 dBm), dBm), in most cases The result is close to 1, which is consistent with the theoretical expectation of equation (12).

[0058] Step S3: System Workflow The receiving MCU (such as CC2650) uses the curing time parameter extracted in step S2.4. and Periodic operation. The correspondence between the two hardware paths and the two stages is as follows:

[0059] At the start of the charging phase, the MCU controls the two SPDT switches to switch to path A. In this path, the antenna port is directly connected to the EH downstream circuitry, bypassing any power divider, and 100% of the received power enters the EH bus. The MCU then enters deep sleep mode to reduce power consumption, keeping only the RTC timer running, waiting... Wake up in seconds.

[0060] RTC time reached A wake-up interrupt is generated at the specified second, waking up the MCU and initiating the communication phase. The MCU first obtains the current instantaneous received power through RSSI measurement. Calculate the power shunt ratio according to formula (12). Through DAC or SPI interface The value is written to the adjustable power divider. Then, the MCU controls the two SPDT switches to switch to path B, and the RF signal passes through the adjustable power divider according to... distribute: After rectification in the EH branch, the rectified current is stored in a supercapacitor. Demodulate the data in the ID branch. Communication phase continues. End in seconds.

[0061] After the communication phase ends, the system automatically returns to the charging phase and repeats the above "charging" process. seconds → communication The node operates in a periodic mode of "seconds" to achieve long-term self-powered operation. During the entire operation, the node does not need to perform any neural network inference, but only needs to execute timer interrupts and simple floating-point operations of equation (12). The power consumption and computing power overhead are extremely low, which meets the engineering constraints of IoT nodes.

[0062] Step S4: Long-term performance verification The simulation was run for over 500 cycles in a simulation environment, and the average communication rate, total collected energy, total consumed energy, and changes in supercapacitor energy storage were statistically analyzed. The method of this invention can achieve an average communication rate that accurately reaches the target value, with total collected energy exceeding total consumed energy, resulting in a positive net energy surplus and satisfying long-term energy neutrality.

[0063] Application example: The method of the present invention will be clearly described below with reference to specific application examples. The simulations and experiments in this embodiment were both completed in the MATLAB environment.

[0064] S1. Hardware configuration and parameter settings for dual-port SWIPT system In this embodiment, the transmitter operates in the 2.45GHz ISM band, uses a single antenna, and sets the transmit power to 32dBm (linear value approximately 1.58W). The receiver employs a dual-port rectified antenna, optimized through HFSS full-wave electromagnetic simulation, with the port return loss S... 11 S 22 All are better than -16dB, port isolation S 12 Better than -19dB. The antenna gain is 4.95dBi in pure energy harvesting mode (ρ=1). The receiver control center uses the TI CC2650 ultra-low-power wireless MCU, with a typical receiver power consumption of P_RX=19.47mW and power consumption below 1µW in deep sleep mode. The supercapacitor has a maximum energy storage of E_max=100J and an initial energy storage of E_init=15J.

[0065] S2, Dynamic Channel and Energy Harvesting Model Parameter Settings In this embodiment, the channel model comprehensively considers large-scale path loss, shadowing fading, and Rayleigh fading. The free-space path loss is calculated using the Friis formula: PL_FS = (λ / (4πd))², where λ = 0.1224m (corresponding to 2.45GHz), and the communication distance d = 10m, resulting in PL_FS ≈ 9.43 × 10⁻⁶. -7 The shadow fading is modeled using a log-normal distribution with a standard deviation σ_s = 6 dB, i.e., S ~ N(0, σ_s²). _shadow=10^(S / 10). Rayleigh fading power factor ξ_Rayleigh=-ln(U), U~Uniform(0,1), this generation method guarantees a mean of 1. Instantaneous received power P_r = P_t·G_t·G_r·PL_FS· Substituting _shadow·ξ_Rayleigh with P_t=1.58W, G_t=1, and G_r=3.1251, the average received power in free space is approximately -23.4dBm, and the mean power after superimposed fading is approximately -25dBm (linear value 3.16×10). -6 W).

[0066] The rectification efficiency η(P_r) adopts a piecewise linear model: when the input power corresponding to P_r is below -26dBm, η=0.43; it linearly increases to 0.50 in the range of -26dBm to -20dBm; it linearly increases to 0.61 in the range of -20dBm to -10dBm; and it saturates at 0.61 above -10dBm. Experiments show that the error between this model and the measured data of the CMOS bridge rectifier is less than 5%.

[0067] S3 and DDQN offline training implementation steps This embodiment sets up an environment simulator in MATLAB and performs DDQN training according to the following steps.

[0068] S3.1 State Space, Action Space, and Reward Function The state space is a four-dimensional vector: s(t) = [P_r(t) / P_norm, E_cap(t) / E_max, Q(t) / Q_norm, Z(t) / Z_norm] Where P_r(t) is the instantaneous received power, E_cap(t) is the current energy stored in the supercapacitor, Q(t) is the data queue length, and Z(t) is the virtual rate queue. All components are normalized to a range of [0,1].

[0069] Action space: under energy neutral constraints Under the condition [η(P_r)P_r]·t1≥ P_RX·t0, the theoretical minimum ratio t1 / t0≥14320. In this embodiment, t0∈{0.1s,0.2s} is selected, corresponding to t1: when t0=0.1s, t1∈{1500,1700,1900}s; when t0=0.2s, t1∈{2900,3100,3300}s, for a total of 6 discrete actions.

[0070] Reward function: R_total = R_energy + R_efficiency + R_queue. The net energy reward R_energy uses a linear piecewise design: if the net energy E_net = E_harv,charge - P_RX·t0≥0, then R_energy=100·E_net (pruned to [0,2]); otherwise, R_energy=20·E_net (pruned to [-2,0]). The efficiency guideline R_efficiency = -0.05·t1,norm + 0.1·t0,norm is used to avoid extreme duty cycles. The queue penalty R_queue= -0.001·(Q / 2e6).

[0071] S3.2 DDQN Network Structure and Hyperparameters The Q network is a fully connected neural network: 4 nodes in the input layer, 64 nodes in each of the two hidden layers (using ReLU activation function), and 6 nodes in the output layer (corresponding to 6 actions). The network structure can be represented as 4→64→64→6.

[0072] Hyperparameter settings: Optimizer Adam, learning rate 1×10 -6 Discount factor γ = 0.99; target network soft update coefficient τ = 0.001; experience replay pool capacity 200,000; mini-batch size 128. The exploration strategy uses an ε-greedy approach: initial ε = 1.0, multiplied by 0.998 each round, with a minimum value of 0.05. The maximum number of training rounds is 1500, with a maximum of 400 steps per round.

[0073] S3.3 Training Process and Convergence Criterion like Figure 2 As shown, the reward curves were recorded during the training process. Figure 2 a) Initial state Q0 ( Figure 2 (b) Bellman error ( Figure 2 c) and exploration rate ( Figure 2 (d) After approximately 300 rounds of training, the reward entered the positive interval. The average reward for the last 100 rounds was 135.79 ± 9.96, with a coefficient of variation of 7.33%. A two-sample t-test showed p = 0.7508, indicating no significant upward trend in reward. The initial Q0 standard deviation was 0.0673, indicating stable value estimation. The average Bellman error for the last 1000 steps was 2.30 × 10⁻⁶. - ¹±2.40×10 - ², p=0.084, the error has stabilized. The final exploration rate ε=0.05, the exploration phase ends. After training, the initial state (P_r average value -25dBm, E_cap=15J, Q=0) is input into DDQN, the action with the largest Q value is selected, and the optimal time parameters t1=3100s, t0=0.2s, duty cycle 0.0065% are obtained. These parameters are fixed in the constant table of the CC2650 firmware.

[0074] S4. Calculation and Real-time Control of Power Split Ratio During the communication phase, the MCU measures the instantaneous received power P_r (linear value) via RSSI. Based on the target rate R_target = 1 Mbps, bandwidth B = 1 MHz, and equivalent noise power spectral density N0 (thermal noise -174 dBm / Hz + noise figure 5 dB), we get N0 ≈ 1.26 × 10⁻⁶. - ¹ 7 The minimum signal-to-noise ratio (SNR) is calculated as SNR_min = 2^(R_target / B) - 1 = 1 (W / Hz). Therefore, the minimum received power required for the ID branch is P_comm,min = SNR_min·N0B ≈ -109dBm. The power shunt ratio is calculated as follows: if P_r > P_comm,min, then ρ = 1 - P_comm,min / P_r; otherwise, ρ = 1. The MCU writes the ρ value to the adjustable power divider via a DAC or SPI interface.

[0075] S5, Long-cycle Simulation Verification The system was run for 500 cycles using optimal time parameters t1=3100s and t0=0.2s. Each cycle consisted of a charging phase (t1 seconds) and a communication phase (t0 seconds), with the instantaneous received power P_r generated independently for each phase. Key simulation results are shown in Table 1. Figure 3 and Figure 4 As shown. Figure 3 The two values ​​are the instantaneous rates ( ). Figure 3 a) Cumulative average rate ( Figure 3 (b) Energy balance comparison ( Figure 3 c) Changes in supercapacitor energy storage ( Figure 3 (d).

[0076] Table 1 Long-period simulation parameters

[0077] Figure 3 (a) and 3(b) show that the instantaneous rate of all communication cycles is stable at around 1 Mbps, and the cumulative average rate converges rapidly to 1.000 Mbps. Figure 3 (c) The bar chart shows that the total energy collected over 500 cycles is 2.804 J, the total energy consumed is 1.947 J, and the net surplus is 0.857 J, thus satisfying the long-term energy neutrality condition. Figure 3 (d) Demonstrates that the supercapacitor energy storage continuously increases from 15J to approximately 42J without falling below the initial value, verifying that the system will not be interrupted due to energy depletion. Figure 4 The distribution histogram of the power split ratio ρ is given. All samples are above 0.9999999, with an average value of 0.99999994, which is consistent with the theoretical prediction.

[0078] The above examples illustrate the present invention only to aid in understanding it and are not intended to limit the scope of the invention. Those skilled in the art can make various simple deductions, modifications, or substitutions based on the principles of this invention.

Claims

1. A dynamic resource allocation method for a SWIPT system with a dual-port rectified antenna, characterized in that, The method includes: A dual-port SWIPT system model is established, comprising a transmitter and a receiver. The transmitter uses a single antenna to transmit radio frequency (RF) signals, and the receiver uses a dual-port rectified antenna to receive RF signals. Each port is connected to a single-pole double-throw (SPDT) RF switch. Each switch has two selectable paths, path A and path B. Path A is used to send 100% of the received power to the energy harvesting branch, and path B is used to distribute the received power to the energy harvesting branch and the information decoding branch according to the power split ratio ρ. During the charging phase, the receiver's MCU controls the two switches to switch to path A, and during the communication phase, the receiver's MCU controls the two switches to switch to path B. Construct a DDQN network, define the state space, action space, and reward function, and train the DDQN network offline. Obtain the optimal charging time through the offline trained DDQN network. and optimal communication time The parameters are then stored in the receiver's MCU. Based on the obtained optimal charging time and optimal communication time Parameters: At the start of the charging phase, the receiving MCU controls two switches to switch to path A, and the charging phase continues. After a certain time, the communication phase begins. First, based on the target communication rate and the current instantaneous received power, the power split ratio ρ is calculated analytically using Shannon's formula. Then, the receiving MCU controls two switches to switch to path B and distributes power according to the power split ratio ρ. The communication phase continues. The event will end after a set time.

2. The dynamic resource allocation method for a SWIPT system oriented towards a dual-port rectified antenna as described in claim 1, characterized in that, Construct a DDQN network, define the state space, action space, and reward function, and train the DDQN network offline. Obtain the optimal charging time through the offline trained DDQN network. and optimal communication time The parameters are then embedded into the MCU, specifically including: state space Defined as a four-dimensional vector: Where: t is time. This represents the instantaneous received power, reflecting the current channel quality. This represents the current energy stored in the supercapacitor, reflecting the remaining energy level. This represents the length of the data queue, reflecting the degree of data backlog. This is a virtual rate queue used to constrain long-term average rates; each component is divided by its corresponding normalized benchmark. , , , This unifies the magnitude of all dimensions to .

3. The dynamic resource allocation method for a SWIPT system oriented towards a dual-port rectified antenna as described in claim 2, characterized in that, Construct a DDQN network, define the state space, action space, and reward function, and train the DDQN network offline. Obtain the optimal charging time through the offline trained DDQN network. and optimal communication time The parameters are then embedded into the MCU, specifically including: Action space: Under energy neutrality constraint The theoretical minimum ratio of charging time to communication time is as follows. ,in For the power consumption of the receiving end MCU, For rectification efficiency, For instantaneous received power, This represents the expectation of random channel fading. For charging time, For communication time; Construct a discrete set of candidate actions while satisfying the energy neutrality constraint. : That is, when t0=0.1, t1∈{1500,1700,1900}; when t0=0.2, t1∈{2900,3100,3300}, for a total of 6 discrete actions.

4. The dynamic resource allocation method for a SWIPT system oriented towards a dual-port rectified antenna as described in claim 3, characterized in that, Construct a DDQN network, define the state space, action space, and reward function, and train the DDQN network offline. Obtain the optimal charging time through the offline trained DDQN network. and optimal communication time The parameters are then embedded into the MCU, specifically including: reward function for: in For net energy reward, For efficiency-oriented items, This is a queue penalty item.

5. A dynamic resource allocation method for a SWIPT system with a dual-port rectified antenna as described in claim 4, characterized in that, Net Energy Reward Linear piecewise design is adopted: Net energy The energy collected during the charging phase is The energy consumed during the communication phase is Clipping refers to limiting the calculated raw net energy within a predefined range to prevent training oscillations caused by extreme values. Efficiency Guide This is used to guide agents to avoid extreme duty cycles; , These are the normalized values ​​of charging time and communication time within their respective ranges; Queue penalty item This is used to prevent data queues from becoming too long.

6. The dynamic resource allocation method for a SWIPT system oriented towards a dual-port rectified antenna as described in claim 4, characterized in that, Construct a DDQN network, define the state space, action space, and reward function, and train the DDQN network offline. Obtain the optimal charging time through the offline trained DDQN network. and optimal communication time The parameters are then embedded into the MCU, specifically including: After training is complete, save the main network parameters. , initial state Input the DDQN network, calculate the Q-value for each action, and select the action with the largest Q-value: Extracting the optimal action Corresponding charging time and communication time And it is embedded into the firmware of the receiving MCU.

7. A dynamic resource allocation method for a SWIPT system oriented towards a dual-port rectified antenna as described in claim 1, characterized in that, The DDQN network is a fully connected neural network, consisting of an input layer, two hidden layers, and an output layer. The input layer has four nodes, corresponding to a four-dimensional state vector; each of the two hidden layers has 64 nodes, with ReLU as the activation function; and the output layer has six nodes, corresponding to the Q-values ​​of six discrete actions.

8. A dynamic resource allocation method for a SWIPT system with a dual-port rectified antenna as described in claim 1, characterized in that, Based on the target communication rate and the current instantaneous received power, the power shunt ratio ρ is solved using an analytical method based on Shannon's formula, specifically including: By Shannon's formula Back-engineering to meet the target communication rate Minimum signal-to-noise ratio required : in, It refers to communication bandwidth; The minimum received power required for the information decoding branch is : in This represents the equivalent noise power spectral density. Let the power allocated to the energy harvesting branch be... The power obtained by the information decoding branch is To ensure that the communication rate meets the standard, it must meet the following requirements. To maximize energy harvesting, the minimum permissible power of the information decoding branch should be taken, i.e.: The MCU obtains the current instantaneous received power through RSSI measurement. .