Reinforcement Learning Reward Weight Adjustment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing reinforcement learning technologies face limitations in supporting the use of machine learning models trained with reinforcement learning, as they often fail to ensure the intended behavior of robots, requiring large amounts of training data and struggling to balance reward weights effectively.

Innovation Solution

An information processing apparatus that acquires and adjusts machine learning models using reinforcement learning, combining reinforcement learning with supervised learning to adjust reward weights based on a small amount of training data, allowing the model to output intended actions by optimizing reward ratios through alternating phases of reinforcement learning and reward adjustment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If reinforcement learning is used to train machine learning models, then the model can learn from rewards and optimize behavior, but the model fails to output intended actions and requires large amounts of training data

Engineering Contradiction:
Improvelearning efficiencyVSAvoidtraining data amount
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent combines reinforcement learning with supervised learning by introducing a reward adjustment unit that receives training data and adjusts reward weights. This hybrid approach allows the model to leverage both reinforcement learning's reward optimization and supervised learning's ability to learn from labeled examples, thereby achieving intended actions with smaller training data amounts.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements feedback mechanisms where the reward adjustment unit receives training data feedback and modifies reward weights accordingly. This feedback loop enables the system to learn from training data and adjust the reinforcement learning process, improving the model's ability to output intended actions while reducing the quantity of training data needed.

Inventive Principle:
Principle #23Feedback

2Reliability

If reinforcement learning optimizes reward weights, then the model can achieve desired behaviors, but the process is complex and difficult to control

Engineering Contradiction:
Improvebehavior accuracyVSAvoidtraining process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the training process into distinct functional units: a reinforcement learning unit that handles reward optimization and a separate reward adjustment unit that processes training data and adjusts reward weights. This segmentation simplifies the overall process by dividing complex tasks into manageable components, making the system more controllable while maintaining behavior accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The reward adjustment unit acts as an intermediary between the reinforcement learning unit and the machine learning model. It receives training data, processes it to determine optimal reward weights, and feeds these weights back to the reinforcement learning unit. This intermediary role simplifies the training process by providing a dedicated layer for reward optimization, reducing overall system complexity while improving behavior accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If the model uses reinforcement learning to learn actions, then it can optimize based on rewards, but it cannot ensure the intended behavior is achieved

Engineering Contradiction:
Improvebehavior flexibilityVSAvoidintended behavior assurance
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent uses feedback from training data to adjust reward weights, ensuring the model learns intended behaviors. The reward adjustment unit processes training data and modifies reward weights to align with desired outcomes, providing a feedback mechanism that guarantees the model achieves intended behaviors while maintaining the flexibility to adapt to different scenarios.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the parameter of reward weights dynamically based on training data. By adjusting these weights through the reward adjustment unit, the system can adapt to different training scenarios and ensure intended behaviors are achieved. This parameter change approach maintains behavior flexibility while improving reliability through data-driven weight optimization.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240086714A1Information processing apparatus and information processing method
Publication Date: 2024.03.14 SONY GROUP CORP
  • US20240086714A1 patent drawing
  • US20240086714A1 patent drawing
  • US20240086714A1 patent drawing

AI summary

An information processing apparatus (100) includes: an acquisition unit (153) that acquires a machine learning model trained with reinforcement learning such that, when first state information indicating a first state has been input, the model will output first action information indicating a first action corresponding to the first state, based on a plurality of rewards weighted by a weight of each of the rewards; a reception unit (151) that receives training data being a set of second state information indicating a second state and second action information indicating a second action corresponding to the second state; and a display unit (156) that displays information regarding the weight of each of the rewards estimated by training the machine learning model in which the weight of each of the rewards is defined as a part of a connection coefficient of the machine learning model such that, when the second state information included in the training data and a value based on the weight of each of the rewards have been input, the model will output the second action information included in the training data.