Robot Action Policy Updating with Weighted Multi-Agent Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for updating a robot's policy for action control are limited in precision and adaptability, as they often rely on single-agent learning data, which may not account for diverse environmental situations, leading to suboptimal performance in untrained scenarios.

Innovation Solution

A method involving heterogeneous agents to generate weighted learning datasets, creating a weighted learning database that adjusts the policy to increase reward values for robot actions, and further updates the policy using direct learning data from real-world interactions, ensuring the robot can adapt to various situations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If single-agent learning data is used for policy updates, then the learning process is simple, but the precision and adaptability of the policy are limited

Engineering Contradiction:
Improvepolicy update precisionVSAvoidlearning system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines learning data from multiple heterogeneous agents into a unified weighted learning database. Different agents (e.g., simulation agents, real robot agents, human operators) contribute their respective learning data, which is then integrated with appropriate weighting to update the policy. This merging approach improves policy precision by incorporating diverse perspectives while managing complexity through systematic integration protocols.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The learning database is designed to universally accept learning data from various sources and agent types. The system can process simulation data, real-world robot data, and human demonstration data within a single framework, making the learning system multi-functional and adaptable to different data sources without requiring separate processing pipelines for each agent type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If single-agent learning data is used for policy updates, then the system is easy to implement, but the adaptability to diverse environmental situations is poor

Engineering Contradiction:
Improvepolicy adaptabilityVSAvoidlearning system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges learning data from heterogeneous agents operating in different environments and under different conditions. This combination exposes the policy to a wider variety of situations, improving its adaptability. The weighted integration mechanism manages the complexity of combining multiple data sources by assigning appropriate importance weights to each agent's contributions based on their reliability and relevance.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If weighted learning database from multiple agents is used, then the policy precision and adaptability improve, but the system complexity increases

Engineering Contradiction:
Improvepolicy performance reliabilityVSAvoidlearning system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies local quality by assigning different weight values to learning data from different agents based on their specific characteristics, reliability, and relevance to the current task. Rather than treating all agent data uniformly, the system selectively emphasizes data from agents that are more reliable or relevant in specific contexts, improving policy reliability while managing complexity through targeted differentiation.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts the weight parameters associated with each agent's learning data. These weight parameters can be modified based on the current task requirements, environmental conditions, and the demonstrated reliability of each agent. This parameter adjustment mechanism allows the system to adapt to different scenarios while maintaining a manageable level of complexity through systematic parameter control.

Inventive Principle:
Principle #35Parameter changes

4Loss of information

If multiple heterogeneous agents generate learning data, then the learning comprehensiveness improves, but the data integration complexity increases

Engineering Contradiction:
Improvelearning information completenessVSAvoiddata integration complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent implements a unified weighted learning database that merges learning data from multiple heterogeneous agents. This consolidation ensures that information from all agents is captured and integrated, preventing information loss. The weighted integration approach manages the complexity of combining diverse data sources by providing a systematic method for prioritizing and combining their contributions.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11631028B2Method of updating policy for controlling action of robot and electronic device performing the method
Publication Date: 2023.04.18 SAMSUNG ELECTRONICS CO LTD
  • US11631028B2 patent drawing
  • US11631028B2 patent drawing
  • US11631028B2 patent drawing

AI summary

A tendency of an action of a robot may vary based on learning data used for training. The learning data may be generated by an agent performing an identical or similar task to a task of the robot. An apparatus and method for updating a policy for controlling an action of a robot may update the policy of the robot using a plurality of learning data sets generated by a plurality of heterogeneous agents, such that the robot may appropriately act even in an unpredicted environment.