Operation Rule Learning with Time-Step Smoothed Evaluation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The learning of operation rules for controlled objects becomes difficult when conditions relating to operation are set, as existing methods face challenges in reducing the complexity and variability of evaluation functions across time steps.
Innovation Solution
An operation rule determination device and method that alter the evaluation function to reduce differences between time steps, using a second evaluation function derived from a first evaluation function that reflects operation conditions, allowing for staged learning and adjustment of penalty thresholds to stabilize and ease the learning process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conditions relating to operation are set in the learning of operation rules, then the reliability of operation control is improved, but the difficulty of learning increases
Solution Approach 1:
The learning process is divided into two distinct stages: first learning with the first evaluation function that does not reflect operation conditions, then learning with the second evaluation function that does reflect operation conditions. This segmentation allows the system to progressively incorporate constraints without overwhelming the learning process, thereby improving reliability while managing learning difficulty.
Solution Approach 2:
The system performs preliminary learning using the first evaluation function before introducing the constraints of the second evaluation function. This preliminary action establishes a baseline operation rule that can later be refined with operational conditions, making the overall learning process more manageable while achieving reliable controlled operation.
2Reliability
If the first evaluation function reflecting operation conditions is used directly for learning, then the operation constraints are satisfied, but the learning process becomes unstable due to large differences in evaluation function between time steps
Solution Approach 1:
The evaluation function is segmented into two versions: the first evaluation function for initial learning without operational constraints, and the second evaluation function for refined learning with constraints. This segmentation stabilizes the learning process by avoiding large fluctuations that would occur if the constraint-heavy second evaluation function were used from the beginning.
Solution Approach 2:
The system performs preliminary learning with the simpler first evaluation function before transitioning to the more complex second evaluation function. This preliminary action with a stable, constraint-free evaluation function establishes a foundation that prevents instability when operational constraints are later introduced.
Data Source
AI summary
An operation rule determination device includes: an evaluation function setting unit that sets a second evaluation function that has been altered from a first evaluation function in which a condition relating to operation of a controlled object is reflected, such that a difference in an evaluation function between time steps of evaluation relating to the operation of the controlled object is reduced; and a learning unit that performs learning on an operation rule of the controlled object using the second evaluation function, and performs learning on the operation rule of the controlled object using a learning result and the first evaluation function.


