Risk-Aware Operation Rule Learning Using Cumulative Reward Frequency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Reinforcement learning systems lack a method to determine operations considering risk, which is essential for safe decision-making in dynamic environments.

Innovation Solution

An operation rule determination device and method that includes an environment execution unit and a risk-considered history generation unit, which calculates a cumulative degree of operations and adjusts the degree of desirability based on risk information, allowing for risk-aware decision-making.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If reinforcement learning is used to determine operations, then decision-making speed and adaptability are improved, but risk consideration capability deteriorates

Engineering Contradiction:
Improvedecision-making speedVSAvoidrisk consideration capability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the reward signal into two distinct components: cumulative reward (long-term outcome) and frequency of cumulative reward (short-term risk indicator). This segmentation allows the reinforcement learning system to process risk information separately from standard reward information, enabling risk-aware decision-making without sacrificing learning speed. The frequency component specifically tracks how often certain cumulative reward thresholds are exceeded, providing a risk metric that can be independently optimized.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces frequency of cumulative reward as an intermediary variable that bridges the gap between standard reinforcement learning and risk consideration. This intermediary transforms the raw cumulative reward data into a risk indicator that can be integrated into the learning process. By using this intermediary, the system can maintain its standard reinforcement learning framework while adding risk awareness through the frequency metric, which modulates the learning updates based on risk levels.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If risk assessment is incorporated into the decision-making process, then safety is improved, but system complexity increases

Engineering Contradiction:
ImprovesafetyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent makes the frequency of cumulative reward calculation serve multiple functions: it acts as a risk indicator, a learning signal modifier, and a decision-making constraint. By designing this single mechanism to perform multiple roles, the patent avoids adding separate complex risk assessment modules. The frequency calculation integrates seamlessly with the existing reinforcement learning framework, demonstrating multi-functionality that reduces overall system complexity while maintaining safety improvements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes the parameter representation by introducing frequency as an additional dimension to the cumulative reward parameter. Instead of creating a completely new risk assessment architecture, the system modifies the existing reward parameter space by adding frequency information. This parameter change allows the system to incorporate risk assessment using the same computational infrastructure, minimizing the increase in system complexity while enhancing safety capabilities.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12093001B2Operation rule determination device, method, and recording medium using frequency of a cumulative reward calculated for series of operations
Publication Date: 2024.09.17 NEC CORP
  • US12093001B2 patent drawing
  • US12093001B2 patent drawing
  • US12093001B2 patent drawing

AI summary

An operation rule determination device includes an environment execution unit that obtains a state of a control target after each operation and the degree associated with the state for a series of operations on the control target, by using degree information in which the state and the degree of desirability of the state are associated with each other, and a risk-considered history generation unit that calculates a cumulative degree obtained by accumulating the obtained degree for the series of operations, and, when the cumulative degree satisfies a condition, reduces the degree associated with the state after the series of operations in the degree information.