Reward Control Unit Automates Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Reinforcement learning devices face challenges in configuring rewards for actions in business environments, as they are typically unilateral and require manual adjustment, leading to inefficiencies and high computational costs, limiting the ability to optimize models effectively.

Innovation Solution

A data-based reinforcement learning approach that defines a reward as the difference in overall variation caused by actions, using a reward control unit to calculate and standardize the individual variation rate of metrics such as rate of return, limit exhaustion rate, and loss rate, providing a reward between 0 and 1, thereby automating the reward configuration process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If arbitrary reward points are assigned and manually adjusted while watching learning results, then the reinforcement learning model can be optimized, but massive time and computing resources are consumed for trial and error

Engineering Contradiction:
Improvemodel optimization effectivenessVSAvoidtime and computing resources for trial and error
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system automatically determines reward values by analyzing business data and calculating variation rates, eliminating the need for manual reward configuration. The reward control unit self-adjusts reward parameters based on learned patterns from historical data, allowing the system to serve itself without continuous human intervention and massive trial-and-error iterations.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements automated feedback loops where learning results are continuously analyzed to adjust reward values. The reward control unit receives feedback from performance metrics and automatically modifies reward parameters, creating a closed-loop system that reduces manual adjustment cycles and accelerates model optimization without consuming excessive computational resources.

Inventive Principle:
Principle #23Feedback

2Productivity

If reward points are unilaterally determined and assigned to actions, then the reinforcement learning process can proceed, but users need to repeat and experiment reward configurations conforming to business objectives

Engineering Contradiction:
Improvereinforcement learning process efficiencyVSAvoidreward configuration complexity
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The reward control unit automatically determines appropriate reward values by analyzing business data and calculating variation rates, eliminating the need for users to manually configure and experiment with different reward settings. The system self-adjusts reward parameters to align with business objectives without requiring user intervention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system dynamically changes reward parameters based on learned patterns from business data. Instead of using fixed unilateral reward assignments, the reward values are continuously adjusted according to variation rates calculated from historical performance data, allowing the system to adapt to changing business conditions automatically.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If reinforcement learning proceeds on the basis of rewards determined unilaterally in connection with metric accomplishment, then the learning process can be simplified, but only one action pattern can be taken to accomplish the metric

Engineering Contradiction:
Improvelearning process complexityVSAvoidaction pattern diversity
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The system transitions from static unilateral reward determination to dynamic reward adjustment based on learned patterns. The reward control unit continuously modifies reward values according to variation rates derived from business data, enabling the system to adapt to different situations and discover multiple effective action patterns rather than being constrained to a single approach.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes reward parameters dynamically based on learned insights from business data. By calculating variation rates and adjusting reward values accordingly, the system enables diverse action patterns to be explored and rewarded, increasing adaptability while maintaining manageable complexity through data-driven parameter adjustment.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20220230097A1Device and method for data-based reinforcement learning
Publication Date: 2022.07.21 AGILESODA INC
  • US20220230097A1 patent drawing
  • US20220230097A1 patent drawing
  • US20220230097A1 patent drawing

AI summary

Disclosed is a device for data-based reinforcement learning. The disclosure allows an agent to learn a reinforcement learning model so as to maximize a reward for an action selectable according to a current state in a random environment, wherein a difference between a total variation rate and an individual variation rate for each action is provided as a reward for the agent.