Learning Device Using Neural ODE Gradient Model for Control Policy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current reinforcement learning methods for control targets are inefficient, requiring extensive time to learn optimal control policies and reward functions, especially when lacking prior learning results.

Innovation Solution

A learning device and method that utilizes a neural ordinary differential equation as a gradient model to learn the relationship between a control target's state, control actions, and temporal changes, enabling faster learning by leveraging previous task results for initial policy and reward function determination.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If reinforcement learning is performed without prior learning results, then control over the control target can be learned, but the learning time becomes excessively long

Engineering Contradiction:
Improvelearning timeVSAvoidlearning efficiency
Core Design Contradiction:
Loss of timeVSProductivity

Solution Approach 1:

The patent applies preliminary action by learning a gradient model beforehand that captures the temporal changes of the control target's state. This pre-learned model serves as a foundation for subsequent control learning, eliminating the need to start from scratch and significantly reducing the time required to learn optimal control policies.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The gradient model acts as an intermediary between the control target's state and the control policy. By introducing this intermediate representation that encodes temporal dynamics, the system can leverage it to accelerate the reinforcement learning process and achieve faster convergence to optimal control strategies.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If extensive reinforcement learning is performed to achieve accurate control policies, then control accuracy improves, but the time required for learning increases significantly

Engineering Contradiction:
Improvecontrol accuracyVSAvoidlearning time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary learning of the gradient model to capture temporal changes in the control target's state. This pre-computed model provides accurate representations of system dynamics that can be directly utilized during control policy learning, achieving high control accuracy without requiring extensive learning time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the approach by introducing a gradient model parameterization that represents temporal changes. This parameter change enables the system to encode complex temporal dynamics in a compact form, allowing for accurate control policies to be learned more efficiently by leveraging the pre-learned gradient model parameters.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240330702A1Learning device, learning method, and storage medium
Publication Date: 2024.10.03 NEC CORP
  • US20240330702A1 patent drawing
  • US20240330702A1 patent drawing
  • US20240330702A1 patent drawing

AI summary

A learning device includes at least one memory configured to store instructions; and at least one processor configured to execute the instructions to: perform reinforcement learning of control over a control target; use data used in the reinforcement learning to learn a model that shows the relationship between a state relating to the control target, control over the control target, and a temporal change in the state relating to the control target; and use the model and the result of the reinforcement learning to learn control over the control target.