Reward Function Sharing for Multi-Target Reinforcement Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In system optimization based on reinforcement learning, managing multiple systems with different objects is inefficient due to lack of cooperation among individual learning devices, preventing effective reinforcement learning.
Innovation Solution
A computer system that manages reward function information for multiple control targets, allowing for comparison and updating of rewards to determine optimal actions, enabling efficient sharing of rewards across systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If individual learning devices are used for each control target, then each system can be optimized independently, but the learning devices do not cooperate with one another, preventing efficient reinforcement learning
Solution Approach 1:
The patent merges individual learning devices into a unified multi-agent reinforcement learning system where multiple agents share a common learning framework. The agents cooperate by exchanging information about their respective control targets and jointly optimizing their policies, thereby achieving both independent optimization capability and improved learning efficiency through collaborative learning.
Solution Approach 2:
The patent creates a universal learning framework that can handle multiple different control targets simultaneously. The multi-agent system uses a common architecture and learning algorithms that are adaptable to various control targets, allowing the system to maintain versatility for different optimization tasks while achieving efficient cooperative learning through shared experience and knowledge transfer.
2Productivity
If a master appliance manages multiple electric appliances, then system optimization is achieved, but management is not possible when each of a plurality of systems having different objects is optimized
Solution Approach 1:
The patent segments the centralized master appliance management into distributed multi-agent systems. Each agent is responsible for a specific control target and can independently learn and optimize its own policy while maintaining communication with other agents. This segmentation allows the system to manage multiple different objects effectively, as each agent is specialized for its specific target while contributing to overall system optimization through cooperative learning.
3Speed
If rewards are shared across multiple control targets, then optimization speed is improved, but the complexity of managing reward functions increases
Solution Approach 1:
The patent introduces an intermediary mechanism in the form of a centralized coordination module or communication protocol that manages reward sharing among multiple agents. This intermediary handles the complexity of reward function management by providing a standardized interface for agents to exchange reward information, reconcile conflicting objectives, and coordinate their learning processes, thereby enabling fast optimization without requiring each agent to independently manage complex reward interactions.
Data Source
AI summary
A computer system includes a processor and a memory connected to the processor, and manages pieces of reward function information for defining rewards for states and actions of the control targets for each of the control targets. The pieces of reward function information includes first reward function information for defining the reward of a first control target and second reward function information for defining the reward of a second control target. When updating the first reward function information, the processor compares the rewards of the first reward function information and the second reward function information with each other, specifies a reward, which is reflected in the first reward function information from rewards set in the second reward function information, updates the first reward function information on the basis of the specified reward, and decides an optimal action of the first control target by using the first reward function information.


