Risk-Parameterized Reinforcement Learning for Adaptive Device Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current reinforcement learning methods for autonomous devices, such as autonomous driving robots, lack the ability to adaptively adjust their risk measures based on environmental characteristics, leading to suboptimal decision-making in uncertain situations.

Innovation Solution

A method that utilizes a risk-measure parameter associated with device control, allowing the learning model to selectively set this parameter according to environmental characteristics, thereby determining actions that are either risk-averse or risk-seeking, using a quantile regression method to predict actions and rewards.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If reinforcement learning methods use fixed risk measures, then the model structure remains simple, but the device cannot adaptively adjust to different environmental conditions

Engineering Contradiction:
Improveadaptive controlVSAvoidmodel complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by introducing a risk-measure parameter that can be selectively set based on environmental characteristics. This parameter modifies the decision-making behavior of the learning model without changing the fundamental model structure, enabling adaptive control across different risk environments while maintaining model simplicity.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The learning model is designed with universality by incorporating a risk-measure parameter that allows it to function in multiple risk scenarios (risk-averse, risk-neutral, risk-seeking) using the same model structure. This multi-functionality enables the device to adapt to various environmental conditions without requiring separate models for each scenario.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If the device operates in uncertain environments with fixed risk measures, then decision-making is consistent, but safety and performance deteriorate

Engineering Contradiction:
ImprovesafetyVSAvoidcontrol mechanism complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent enhances reliability by dynamically adjusting the risk-measure parameter based on environmental uncertainty. In high-uncertainty environments, the parameter shifts toward risk-averse settings, improving safety. In low-uncertainty environments, it allows risk-seeking behavior for better performance, all while using the same control mechanism.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The control mechanism incorporates dynamics by allowing the risk-measure parameter to change according to environmental conditions. This dynamic adjustment enables the device to respond appropriately to varying levels of uncertainty, improving both safety and performance without requiring complex separate control systems.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If reinforcement learning models are retrained for different risk measures, then adaptability improves, but training time and computational resources increase

Engineering Contradiction:
Improverisk adaptationVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent eliminates the need for retraining by using parameter changes instead. The risk-measure parameter can be selectively set to different values (risk-averse, risk-neutral, risk-seeking) based on environmental characteristics, allowing the same trained model to adapt to different risk scenarios instantly without additional training time or computational resources.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The learning model achieves universality by being trained once to handle multiple risk measures through the risk-measure parameter. This single model can be deployed across various risk environments, providing risk adaptation capability without requiring separate training processes for each scenario, thus saving time and computational resources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20220198225A1Method and system for determining action of device for given state using model trained based on risk-measure parameter
Publication Date: 2022.06.23 NAVER CORP
  • US20220198225A1 patent drawing
  • US20220198225A1 patent drawing
  • US20220198225A1 patent drawing

AI summary

A method of determining an action of a device for a given situation, implemented by a computer system, includes for a learning model that learns a distribution of rewards according to the action of the device for the situation using a risk-measure parameter associated with control of the device, selectively setting a value of the risk-measure parameter in accordance with an environment in which the device is controlled; and determining the action of the device for the given situation when controlling the device in the environment, based on the set value of the risk-measure parameter.