Reinforcement Learning for Wireless Powered Communication Network Time Allocation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional wireless powered communication networks (WPCNs) face reduced total throughput due to inability to dynamically adjust communication times of base stations and communication nodes, as specific models and parameters vary with time, making existing optimization methods ineffective.

Innovation Solution

A communication time allocation method using reinforcement learning, where a base station obtains weight vectors and models eigenvectors for communication nodes to determine optimal time allocations that maximize estimated throughput while satisfying constraints, and updates weight vectors based on actual throughput, allowing for dynamic adjustments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If convex optimization algorithms or Lagrange multiplier methods are used for throughput optimization, then optimization can be achieved, but the method requires known model forms and cannot adapt to time-varying parameters

Engineering Contradiction:
Improvethroughput optimizationVSAvoidadaptability to time-varying parameters
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent transforms the static optimization problem into a dynamic reinforcement learning framework where the base station continuously learns optimal communication time allocations through interaction with the time-varying wireless environment. The policy network dynamically adjusts communication times based on current channel conditions and node states, enabling adaptation to time-varying parameters without requiring known model forms.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements a feedback mechanism where the base station observes actual throughput outcomes from communication operations and uses this feedback to update its policy through reinforcement learning. The reward signal based on total throughput provides continuous feedback that guides the learning process, allowing the system to adapt to changing conditions over time without requiring explicit model knowledge.

Inventive Principle:
Principle #23Feedback

2Productivity

If communication times are fixed without dynamic adjustment, then system complexity is reduced, but total throughput of the WPCN is significantly reduced

Engineering Contradiction:
Improvetotal throughputVSAvoidcomplexity of time allocation mechanism
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent enables the base station to autonomously determine optimal communication time allocations through reinforcement learning without requiring complex centralized coordination or manual configuration. The base station learns and adapts communication schedules independently based on observed environmental feedback, achieving high throughput while maintaining relatively simple system architecture compared to traditional optimization approaches.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If reinforcement learning is used for communication time allocation, then adaptability to time-varying parameters is improved, but computational complexity increases

Engineering Contradiction:
Improveadaptability to time-varying parametersVSAvoidcomputational complexity of base station
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the reinforcement learning problem into manageable components: state representation (channel conditions, node states), action space (communication time allocations for each node), and reward function (total throughput). This segmentation allows the base station to process information in discrete, computationally tractable steps while maintaining high adaptability to time-varying wireless conditions.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11323167B2Communication time allocation method using reinforcement learning for wireless powered communication network and base station
Publication Date: 2022.05.03 NATIONAL TSING HUA UNIVERSITY
  • US11323167B2 patent drawing
  • US11323167B2 patent drawing
  • US11323167B2 patent drawing

AI summary

The disclosure provides a communication time allocation method using reinforcement learning for a wireless powered communication network and a base station. The method includes: determining a communication time allocation corresponding to the t-th time block according to an objective function associated with the total estimated throughput of the communication nodes; requesting each communication node to perform specific communication behaviors according to the corresponding communication time interval in the t-th time block; obtaining the actual throughput of each communication node in the t-th time block; generating the weight vector of each communication node in the (t+1)-th time block according to the actual throughput, the weight vector, and the estimated throughput of each communication node in the t-th time block.