Reinforcement Learning Precoder Selection Policy for MIMO Transmitters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current precoder selection methods for multi-antenna transmitters in wireless communication networks are complex and impractical for frequency-selective wideband systems, often relying on approximate methods that fail to provide accurate results even in favorable channel conditions.

Innovation Solution

The implementation of reinforcement learning to adapt an action value function for precoder selection, using reward information to optimize precoder choice without requiring detailed knowledge of the underlying system or channel model, enabling near-optimal policy learning for challenging MIMO environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If reinforcement learning is applied to adapt an action value function for precoder selection, then precoder selection accuracy and data transmission efficiency are improved, but computational complexity and training time increase

Engineering Contradiction:
Improveprecoder selection accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-training the reinforcement learning agent offline to develop an optimized precoder selection policy before actual data transmission. The action value function is adapted during offline training using simulated channel conditions, allowing the system to make accurate precoder selections during online operation without performing complex real-time learning computations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an action value function as an intermediary between the channel state observations and precoder selections. This function Q(s,a) serves as a learned mediator that maps channel states to optimal precoder actions, simplifying the decision-making process during online operation while maintaining high selection accuracy through offline learning.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If reinforcement learning is used to learn precoder selection policy, then adaptability to dynamic channel conditions is improved, but training time and computational resources increase

Engineering Contradiction:
Improveadaptability to channel conditionsVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs the computationally intensive reinforcement learning training in advance during an offline phase, separating the learning process from real-time operation. This allows the system to adapt to various channel conditions through extensive simulation before deployment, while online operation requires only simple policy execution without additional training time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses simulated channel conditions and virtual training environments to create copies of real-world scenarios for offline learning. The reinforcement learning agent is trained on synthesized channel states and precoder outcomes, allowing it to learn adaptive policies without requiring actual real-time channel variations during training.

Inventive Principle:
Principle #26Copying

3Device complexity

If approximate methods are used for precoder selection, then computational complexity is reduced, but selection accuracy deteriorates even in favorable channel conditions

Engineering Contradiction:
Improvecomputational complexityVSAvoidprecoder selection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent pre-computes an optimized precoder selection policy through offline reinforcement learning, storing the learned action value function for rapid online queries. This eliminates the need for complex real-time calculations while achieving superior accuracy compared to approximate methods, as the optimal policy is determined in advance through exhaustive exploration of the action space.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses itself to learn the optimal precoder selection policy through self-play reinforcement learning during offline training. The agent learns from its own experiences and rewards without requiring external guidance or complex real-time computations, achieving high accuracy through autonomous offline learning.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11968005B2Provision of precoder selection policy for a multi-antenna transmitter
Publication Date: 2024.04.23 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US11968005B2 patent drawing
  • US11968005B2 patent drawing
  • US11968005B2 patent drawing

AI summary

Method and device(s) for providing precoder selection policy for a multi-antenna transmitter arranged to transmit data over a communication channel of a wireless communication network. Machine learning in the form of reinforcement learning is applied involving adaptation of an action value function configured to compute an action value based on information indicative of a precoder of the multi-antenna transmitter and of a state relating to at least the communication channel. The adaptation being further based on reward information provided by a reward function, indicative of how successfully data is transmitted over the communication channel, and the precoder selection policy is provided based on the adapted action value function resulting from the reinforcement learning.