State-Based Reward Control Learning for Multi-State Targets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing control devices using reinforcement learning struggle to provide appropriate reward values when the state of the control target is divided into multiple states, leading to inadequate learning of control details.
Innovation Solution
A control device with a state data acquisition unit, state category identification unit, and reward generation unit that calculates reward values based on selected formulas specific to each state category, enabling more accurate learning of control details.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a constant reward value (+1 or -1) determined by a single rule is used in reinforcement learning, then the control device can implement simple reward evaluation, but it cannot provide appropriate reward values when the state of the control target is divided into multiple states, leading to inadequate learning of control details
Solution Approach 1:
The patent segments the reward evaluation process by dividing states into multiple state categories and assigning different reward calculation formulas to each category. This segmentation allows the system to provide appropriate reward values for different states while maintaining manageable complexity through modular formula selection.
Solution Approach 2:
The patent changes the parameters of reward evaluation by selecting different calculation formulas based on state categories. Instead of using a fixed constant reward value, the system dynamically adjusts the reward calculation parameters according to the current state category, enabling accurate reward assignment across multiple states.
2Measurement precision
If different reward calculation formulas are selected for each state category, then appropriate reward values can be given for multiple states, but the device complexity increases due to the need for formula selection and category identification
Solution Approach 1:
The patent applies preliminary action by pre-defining multiple reward calculation formulas and state categories before the reinforcement learning process begins. This preparation allows the system to quickly select appropriate formulas during operation without complex real-time decision-making, balancing accuracy with manageable complexity.
Solution Approach 2:
The patent introduces an intermediary mechanism (the state category identification and formula selection unit) that bridges the gap between raw state data and reward calculation. This intermediary simplifies the overall structure by organizing the complexity into a clear workflow: state identification → category classification → formula selection → reward calculation.
Data Source
AI summary
A control device capable of more appropriately learning a control detail of a control target in accordance with a state of the control target is provided. A control device according to the present disclosure includes a state data acquisition unit to acquire state data indicating a state of a control target, a state category identification unit to identify a state category to which a state indicated by the state data belongs among a plurality of state categories indicating classifications of states of the control target on the basis of the state data, a reward generation unit to calculate a reward value of a control detail for the control target on the basis of the state category and the state data, and a control learning unit to learn the control detail on the basis of the state data and the reward value.


