Quantum Multi-Agent Meta Reinforcement Learning for Non-Stationary Environments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multi-agent reinforcement learning faces challenges with abnormal rewards and training convergence issues due to non-stationarity and credit assignment problems in environments with multiple agents.

Innovation Solution

A quantum multi-agent meta reinforcement learning apparatus that applies a learnable axis to a quantum circuit, using a state encoding unit to convert observation values into quantum states, and a quantum circuit unit that updates parameters through angle learning and noise addition, with a measurement unit for local axis learning and continuous parameter initialization using an axis memory.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multi-agent reinforcement learning is performed in a fully centralized method, then high rewards can be obtained by interacting with other agents, but abnormal rewards are invited and training convergence is hindered

Engineering Contradiction:
Improvetraining convergenceVSAvoidreward quality
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the learning process into two distinct phases: meta-learning phase where a quantum circuit learns from multiple single-hop offloading environments, and execution phase where the learned quantum circuit is applied to specific multi-agent environments. This segmentation allows the system to separate the learning of generalizable features from environment-specific interactions, thereby avoiding abnormal rewards while maintaining training convergence.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary learning in single-hop offloading environments before applying the learned quantum circuit to multi-agent environments. The state encoding unit and quantum circuit unit are trained in advance on simplified environments, enabling the system to pre-acquire useful patterns and avoid convergence issues when facing complex multi-agent scenarios.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If multi-agent reinforcement learning considers non-stationarity characteristic and credit-assignment between agents, then learning can be progressed, but the problem complexity increases

Engineering Contradiction:
Improvelearning capabilityVSAvoidproblem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a quantum circuit as an intermediary between observation values and policy decisions. The state encoding unit converts observations into quantum states, which are then processed by the quantum circuit unit. This intermediary representation simplifies the handling of non-stationarity and credit-assignment problems by transforming complex multi-agent interactions into quantum state transformations.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter representation from classical multi-agent state spaces to quantum state parameters. By encoding observations into quantum states with parameters like amplitudes and phases, the system can represent and learn from complex environmental dynamics more efficiently, reducing the apparent complexity of non-stationarity and credit-assignment issues.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If a quantum circuit is applied to different environments including multiple agents, then the system can adapt to changing conditions, but more parameters are required for learning

Engineering Contradiction:
Improveenvironment adaptabilityVSAvoidparameter quantity
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent designs a universal quantum circuit that can function across different single-hop offloading environments and multi-agent environments. The state encoding unit and quantum circuit unit are trained on diverse environments during the meta-learning phase, enabling the same quantum circuit to generalize to various settings without requiring environment-specific parameters, thus achieving multi-functionality with a fixed parameter set.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240104390A1Apparatus and method for quantum multi-agent meta reinforcement learning
Publication Date: 2024.03.28 KOREA UNIV RES & BUSINESS FOUND
  • US20240104390A1 patent drawing
  • US20240104390A1 patent drawing
  • US20240104390A1 patent drawing

AI summary

The present invention relates to a quantum multi-agent meta reinforcement learning apparatus, which receives at least one observation value from different single-hop offloading environments, and the apparatus includes: a state encoding unit for calculating an angle along each axis by encoding the at least one observation value, and converting the angle along each axis into a quantum state; a quantum circuit unit for learning the angle along each axis, and overlapping the learned base layer using a controlled X (CX) gate; and a measurement unit for learning the overlapped base layer and measuring an axis parameter. Through the apparatus, the non-stationarity characteristic and credit-assignment problem of the conventional multi-agent reinforcement learning can be solved.