Reinforcement Learning for Wireless Powered Communication Network Time Allocation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional wireless powered communication networks (WPCNs) face reduced total throughput due to inability to dynamically adjust communication times of base stations and communication nodes, as specific models and parameters vary with time, making existing optimization methods ineffective.
Innovation Solution
A communication time allocation method using reinforcement learning, where a base station obtains weight vectors and models eigenvectors for communication nodes to determine optimal time allocations that maximize estimated throughput while satisfying constraints, and updates weight vectors based on actual throughput, allowing for dynamic adjustments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If convex optimization algorithms or Lagrange multiplier methods are used for throughput optimization, then optimization can be achieved, but the method requires known model forms and cannot adapt to time-varying parameters
Solution Approach 1:
The patent transforms the static optimization problem into a dynamic reinforcement learning framework where the base station continuously learns optimal communication time allocations through interaction with the time-varying wireless environment. The policy network dynamically adjusts communication times based on current channel conditions and node states, enabling adaptation to time-varying parameters without requiring known model forms.
Solution Approach 2:
The patent implements a feedback mechanism where the base station observes actual throughput outcomes from communication operations and uses this feedback to update its policy through reinforcement learning. The reward signal based on total throughput provides continuous feedback that guides the learning process, allowing the system to adapt to changing conditions over time without requiring explicit model knowledge.
2Productivity
If communication times are fixed without dynamic adjustment, then system complexity is reduced, but total throughput of the WPCN is significantly reduced
Solution Approach 1:
The patent enables the base station to autonomously determine optimal communication time allocations through reinforcement learning without requiring complex centralized coordination or manual configuration. The base station learns and adapts communication schedules independently based on observed environmental feedback, achieving high throughput while maintaining relatively simple system architecture compared to traditional optimization approaches.
3Adaptability or versatility
If reinforcement learning is used for communication time allocation, then adaptability to time-varying parameters is improved, but computational complexity increases
Solution Approach 1:
The patent segments the reinforcement learning problem into manageable components: state representation (channel conditions, node states), action space (communication time allocations for each node), and reward function (total throughput). This segmentation allows the base station to process information in discrete, computationally tractable steps while maintaining high adaptability to time-varying wireless conditions.
Data Source
AI summary
The disclosure provides a communication time allocation method using reinforcement learning for a wireless powered communication network and a base station. The method includes: determining a communication time allocation corresponding to the t-th time block according to an objective function associated with the total estimated throughput of the communication nodes; requesting each communication node to perform specific communication behaviors according to the corresponding communication time interval in the t-th time block; obtaining the actual throughput of each communication node in the t-th time block; generating the weight vector of each communication node in the (t+1)-th time block according to the actual throughput, the weight vector, and the estimated throughput of each communication node in the t-th time block.


