Multi-Agent Reinforcement Learning via Kalman Consensus Filters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems of autonomous distributed agents using reinforcement learning face inefficiencies due to errors in communication and detection, which hinder their ability to perform predetermined tasks effectively.

Innovation Solution

A system and method that incorporates a first, second, and third autonomous agent with detectors and communication components, utilizing reinforcement learning and Kalman consensus filters to account for errors in communication and detection, allowing agents to perform tasks by sharing local state information and maintaining a global average through distributed consensus.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional reinforcement learning is used for multi-agent systems, then agents can perform predetermined tasks, but communication errors and detection errors reduce system efficiency and reliability

Engineering Contradiction:
Improvetask execution reliabilityVSAvoidcommunication and detection information accuracy
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent implements feedback mechanisms where agents continuously exchange state information and update their beliefs about other agents' states based on received information and observations. This feedback loop allows agents to compensate for communication and detection errors by iteratively refining their understanding of the system state, thereby improving task execution reliability despite information loss.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces belief states as intermediary representations that mediate between raw sensor data and control decisions. These belief states act as intermediaries that filter and interpret noisy communication and detection signals, allowing agents to make robust decisions even when direct observations are inaccurate or incomplete.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If agents share local state information to achieve global consensus, then task coordination improves, but communication channel errors propagate through the network

Engineering Contradiction:
Improvetask coordination efficiencyVSAvoiderror propagation in communication network
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent applies beforehand cushioning by having agents maintain and update belief states about other agents' conditions before making critical decisions. These pre-computed belief states act as cushions that absorb the impact of communication errors, allowing agents to anticipate potential information inaccuracies and adjust their behavior accordingly, thereby reducing error propagation while maintaining coordination efficiency.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentUS11321635B2Method for performing multi-agent reinforcement learning in the presence of unreliable communications via distributed consensus
Publication Date: 2022.05.03 THE UNITED STATES OF AMERICA AS REPRESENTED BY THE SECRETARY OF THE NAVY
  • US11321635B2 patent drawing
  • US11321635B2 patent drawing
  • US11321635B2 patent drawing

AI summary

A system is provided for performing a predetermined function within a total area of operation, wherein the system includes a plurality of autonomous agents. Each autonomous agent is able to detect respective local parameters. Each autonomous agent uses a Kalman filter component to establish an environment state based a plurality of state measurements over time. The output of the Kalman filter component within a respective agent is applied to reinforcement learning by an actor-critic task controller, within the respective agent, to determine a subsequent action to be performed by the respective agent in accordance with a reward function. Each agent includes a Kalman consensus filter that addresses errors of the plurality of state measurements over time.