Policy Training Device for Constrained Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Constrained reinforcement learning faces inefficiencies in training agents for various constraint conditions, requiring separate agents for each condition, which increases training time and storage costs.

Innovation Solution

A policy training device that uses constrained reinforcement learning to train a single agent by changing a parameter related to the constraint condition for every predetermined number of training operations, allowing the agent to adapt to multiple constraint conditions without needing separate agents for each condition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If separate agents are trained for each constraint condition in constrained reinforcement learning, then the agent can satisfy specific constraint conditions, but the training time and storage costs increase

Engineering Contradiction:
Improveconstraint condition satisfactionVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies universality by training a single reinforcement learning agent to handle multiple constraint conditions simultaneously. Instead of creating separate specialized agents for each constraint condition, the system uses one universal agent that learns to satisfy various constraint conditions through parameter changes during training, reducing both training time and storage requirements while maintaining the ability to meet specific constraints

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If separate agents are trained for each constraint condition in constrained reinforcement learning, then the agent can satisfy specific constraint conditions, but the storage requirements increase

Engineering Contradiction:
Improveconstraint condition satisfactionVSAvoidstorage requirements
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent reduces storage requirements by using a single universal agent model instead of storing multiple separate agent models for different constraint conditions. The universal agent achieves constraint satisfaction through dynamic parameter adjustment during training and execution, eliminating the need to store redundant agent copies for each constraint scenario

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If a single agent is trained to handle various constraint conditions, then training efficiency improves, but the agent must adapt to multiple conditions which increases complexity

Engineering Contradiction:
Improvetraining efficiencyVSAvoidagent adaptation complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent manages agent complexity through parameter changes by modifying constraint parameters during the training process rather than creating structurally different agents for each condition. This allows the single agent to adapt to various constraint conditions by adjusting training parameters, maintaining training efficiency while controlling the complexity through systematic parameter variation rather than architectural complexity

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250103957A1Policy training device, policy training method, and communication system
Publication Date: 2025.03.27 1FINITY INC
  • US20250103957A1 patent drawing
  • US20250103957A1 patent drawing
  • US20250103957A1 patent drawing

AI summary

A policy training device that trains, through first reinforcement learning, a first agent configured to output a first action of a control object according to an input of a first state of the control object, includes a memory, and processor circuitry coupled to the memory and configured to change a first parameter regarding a constraint condition in the first reinforcement learning for every predetermined number of times of a training operation in the first reinforcement learning, and train the first agent by using the first parameter as at least a part of the first state and by ensuring that the constraint condition is satisfied.