Reward Function Sharing for Multi-Target Reinforcement Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In system optimization based on reinforcement learning, managing multiple systems with different objects is inefficient due to lack of cooperation among individual learning devices, preventing effective reinforcement learning.

Innovation Solution

A computer system that manages reward function information for multiple control targets, allowing for comparison and updating of rewards to determine optimal actions, enabling efficient sharing of rewards across systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If individual learning devices are used for each control target, then each system can be optimized independently, but the learning devices do not cooperate with one another, preventing efficient reinforcement learning

Engineering Contradiction:
Improveindependent optimization capabilityVSAvoidlearning efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent merges individual learning devices into a unified multi-agent reinforcement learning system where multiple agents share a common learning framework. The agents cooperate by exchanging information about their respective control targets and jointly optimizing their policies, thereby achieving both independent optimization capability and improved learning efficiency through collaborative learning.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal learning framework that can handle multiple different control targets simultaneously. The multi-agent system uses a common architecture and learning algorithms that are adaptable to various control targets, allowing the system to maintain versatility for different optimization tasks while achieving efficient cooperative learning through shared experience and knowledge transfer.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If a master appliance manages multiple electric appliances, then system optimization is achieved, but management is not possible when each of a plurality of systems having different objects is optimized

Engineering Contradiction:
Improvesystem optimization efficiencyVSAvoidmanagement capability for different objects
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent segments the centralized master appliance management into distributed multi-agent systems. Each agent is responsible for a specific control target and can independently learn and optimize its own policy while maintaining communication with other agents. This segmentation allows the system to manage multiple different objects effectively, as each agent is specialized for its specific target while contributing to overall system optimization through cooperative learning.

Inventive Principle:
Principle #1Segmentation

3Speed

If rewards are shared across multiple control targets, then optimization speed is improved, but the complexity of managing reward functions increases

Engineering Contradiction:
Improveoptimization speedVSAvoidreward function management complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary mechanism in the form of a centralized coordination module or communication protocol that manages reward sharing among multiple agents. This intermediary handles the complexity of reward function management by providing a standardized interface for agents to exchange reward information, reconcile conflicting objectives, and coordinate their learning processes, thereby enabling fast optimization without requiring each agent to independently manage complex reward interactions.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11372379B2Computer system and control method
Publication Date: 2022.06.28 HITACHI LTD
  • US11372379B2 patent drawing
  • US11372379B2 patent drawing
  • US11372379B2 patent drawing

AI summary

A computer system includes a processor and a memory connected to the processor, and manages pieces of reward function information for defining rewards for states and actions of the control targets for each of the control targets. The pieces of reward function information includes first reward function information for defining the reward of a first control target and second reward function information for defining the reward of a second control target. When updating the first reward function information, the processor compares the rewards of the first reward function information and the second reward function information with each other, specifies a reward, which is reflected in the first reward function information from rewards set in the second reward function information, updates the first reward function information on the basis of the specified reward, and decides an optimal action of the first control target by using the first reward function information.