Multi-Agent Actor-Critic Learning for Adaptive Beamforming Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Next-generation wireless communication networks face challenges in managing complex multi-agent systems due to non-stationarity and partial observability, leading to sub-optimal decision-making in multi-cell multi-user MIMO environments.

Innovation Solution

A multi-agent deep reinforcement learning-based framework is employed, utilizing actor and critic neural networks to learn optimal precoding policies through centralized learning with decentralized execution, allowing agents to adapt their policies based on feedback from the environment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional modeling and management methods are used for wireless communication networks, then the system structure is simple and easy to manage, but the network cannot meet the challenging requirements of next-generation services (higher data rates, lower latency, higher energy efficiency)

Engineering Contradiction:
Improvenetwork performanceVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical/mathematical modeling and manual management methods with machine learning-based intelligent systems. Actor-critic neural networks are deployed at network nodes to autonomously learn and optimize network operations, substituting conventional control mechanisms with data-driven intelligent agents that can handle the complexity of next-generation wireless networks

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces dynamic adaptability through machine learning models that continuously learn from network operations and adjust their behavior in real-time. The actor-critic architecture enables network nodes to dynamically optimize their actions based on changing network conditions, service requirements, and environmental factors, transforming static network management into a dynamic adaptive system

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If actor neural networks are trained with incomplete local information only, then the training process is simple and fast, but the decision-making quality becomes sub-optimal due to partial observability

Engineering Contradiction:
Improvedecision-making qualityVSAvoidinformation processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges local observations from individual network nodes with global network state information through the critic network. The critic receives inputs from multiple actor networks and environmental observations, integrating distributed information sources to form a comprehensive view of the network state, thereby enabling more accurate decision-making without requiring each node to process all information independently

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The critic network serves as an intermediary that bridges the gap between local actor networks and the global network environment. It processes information from multiple sources, evaluates the quality of actions taken by actors, and provides guidance feedback to improve their decision-making, effectively mediating between partial local observations and optimal global decisions

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If centralized learning is implemented across multiple actor neural networks, then the overall system performance improves, but the training complexity and computational requirements increase significantly

Engineering Contradiction:
Improvesystem performanceVSAvoidtraining complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the centralized learning problem into distributed actor-critic pairs deployed at individual network nodes. Each actor network is responsible for local decision-making while the critic network provides centralized guidance. This segmentation allows parallel training of multiple actors independently while maintaining coordinated optimization through shared critic feedback, reducing overall training complexity compared to fully centralized approaches

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The critic network is designed as a universal component that serves multiple actor networks simultaneously. A single critic can evaluate and provide feedback to multiple actors, making the learning process more efficient. The critic's policy gradient estimator can be shared across the system, reducing redundant computations and enabling scalable training of large multi-agent systems

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12355524B2Multi-agent policy machine learning
Publication Date: 2025.07.08 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US12355524B2 patent drawing
  • US12355524B2 patent drawing
  • US12355524B2 patent drawing

AI summary

There is disclosed a method of operating a beam-forming wireless communication system, the system has a plurality of radio nodes, an actor neural network being associated to each radio node, wherein further to each actor neural network, there is associated a critic network. The method includes training each actor neural network, for controlling at least one associated radio node, based on learning feedback provided by its associated critic network, the learning feedback being based on operation information provided be the actor neural network for the critic network. The disclosure also pertains to related devices and methods.