Operation Rule Learning with Time-Step Smoothed Evaluation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The learning of operation rules for controlled objects becomes difficult when conditions relating to operation are set, as existing methods face challenges in reducing the complexity and variability of evaluation functions across time steps.

Innovation Solution

An operation rule determination device and method that alter the evaluation function to reduce differences between time steps, using a second evaluation function derived from a first evaluation function that reflects operation conditions, allowing for staged learning and adjustment of penalty thresholds to stabilize and ease the learning process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conditions relating to operation are set in the learning of operation rules, then the reliability of operation control is improved, but the difficulty of learning increases

Engineering Contradiction:
Improvereliability of operation controlVSAvoiddifficulty of learning
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The learning process is divided into two distinct stages: first learning with the first evaluation function that does not reflect operation conditions, then learning with the second evaluation function that does reflect operation conditions. This segmentation allows the system to progressively incorporate constraints without overwhelming the learning process, thereby improving reliability while managing learning difficulty.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary learning using the first evaluation function before introducing the constraints of the second evaluation function. This preliminary action establishes a baseline operation rule that can later be refined with operational conditions, making the overall learning process more manageable while achieving reliable controlled operation.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If the first evaluation function reflecting operation conditions is used directly for learning, then the operation constraints are satisfied, but the learning process becomes unstable due to large differences in evaluation function between time steps

Engineering Contradiction:
Improvesatisfaction of operation constraintsVSAvoidstability of learning process
Core Design Contradiction:
ReliabilityVSStability of the object's composition

Solution Approach 1:

The evaluation function is segmented into two versions: the first evaluation function for initial learning without operational constraints, and the second evaluation function for refined learning with constraints. This segmentation stabilizes the learning process by avoiding large fluctuations that would occur if the constraint-heavy second evaluation function were used from the beginning.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary learning with the simpler first evaluation function before transitioning to the more complex second evaluation function. This preliminary action with a stable, constraint-free evaluation function establishes a foundation that prevents instability when operational constraints are later introduced.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240345547A1Operation rule determination device, operation rule determination method, and recording medium
Publication Date: 2024.10.17 NEC CORP
  • US20240345547A1 patent drawing
  • US20240345547A1 patent drawing
  • US20240345547A1 patent drawing

AI summary

An operation rule determination device includes: an evaluation function setting unit that sets a second evaluation function that has been altered from a first evaluation function in which a condition relating to operation of a controlled object is reflected, such that a difference in an evaluation function between time steps of evaluation relating to the operation of the controlled object is reduced; and a learning unit that performs learning on an operation rule of the controlled object using the second evaluation function, and performs learning on the operation rule of the controlled object using a learning result and the first evaluation function.