Human Collaborative Agent Device for Multi-Agent Learning Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Human-in-the-loop multi-agent reinforcement learning methods fail to correct behaviors and incorporate interpretability, hindering the learning of desired behaviors.

Innovation Solution

A human collaborative agent device processes environmental information, presents inferred behaviors and their reasons to users, and acquires user corrections, enabling behavior correction and interpretability in multi-agent learning systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If human-in-the-loop multi-agent reinforcement learning is implemented without behavior correction mechanisms, then the system maintains autonomous operation, but the ability to learn desired behaviors is hindered

Engineering Contradiction:
Improvelearning effectivenessVSAvoidbehavior correction capability
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system implements a feedback mechanism where users can provide corrections to agent behaviors. The behavior correction unit receives user inputs about desired behaviors and uses this feedback to adjust and improve agent performance, enabling the system to learn from human guidance while maintaining autonomous operation

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

A behavior correction unit acts as an intermediary between users and the multi-agent system. This intermediary component translates user feedback into actionable corrections, bridging the gap between human intent and autonomous agent behavior without requiring direct user control

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If behavior correction mechanisms are added to multi-agent reinforcement learning, then learning of desired behaviors is enabled, but system complexity increases

Engineering Contradiction:
Improvebehavior correction capabilityVSAvoidsystem structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The behavior correction unit is designed to handle multiple types of corrections and work with different agent types uniformly. It can process various forms of user feedback and apply corrections across diverse multi-agent scenarios, reducing the need for separate correction mechanisms for each specific case

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If interpretability information is presented to users for behavior correction, then behavior correction accuracy is improved, but information processing load increases

Engineering Contradiction:
Improvebehavior correction accuracyVSAvoidinformation processing burden
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system segments interpretability information into distinct components such as behavior reasons, environmental information, and correction suggestions. This segmentation allows users to process information in manageable chunks and enables selective presentation of relevant information based on the specific correction context

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240403595A1Human collaborative agent device, system, multi-agent learning method, and, storage medium
Publication Date: 2024.12.05 MITSUBISHI ELECTRIC CORP
  • US20240403595A1 patent drawing
  • US20240403595A1 patent drawing
  • US20240403595A1 patent drawing

AI summary

A multi-agent learning method includes: by an agent, presenting acquired information to a human; and, by the human, correcting a behavior of the agent.