Multimodal Reinforcement Learning With Critical-Period Guidance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Real-time reinforcement learning in large-scale real-world environments is costly and limited, and existing methods do not effectively leverage critical periods for enhanced learning in artificial intelligence agents.

Innovation Solution

A method for reinforcement learning using a multimodal artificial intelligence agent that divides frames in captured images into sections and applies varying guidance types, including moderate-level guidance during critical periods, to enhance learning efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If reinforcement learning is performed in large-scale real-world environments, then learning capability is improved, but cost and time consumption increase significantly

Engineering Contradiction:
Improvelearning capabilityVSAvoidtime consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The training process is segmented into two distinct phases: pre-training in a virtual environment and fine-tuning in a real-world environment. This segmentation allows the agent to learn basic capabilities efficiently in simulation before adapting to real-world complexities, significantly reducing the time and cost required for pure real-world training while maintaining learning effectiveness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The agent performs preliminary learning actions in a virtual environment before deploying to real-world settings. By pre-training on simulated data and experiences, the agent acquires foundational knowledge and skills that transfer to real-world applications, reducing the exploratory time and resource consumption needed during actual deployment.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If guidance level is increased during critical periods, then learning effect is enhanced, but system complexity increases

Engineering Contradiction:
Improvelearning effectVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The guidance level is made dynamic rather than static, automatically adjusting based on the agent's performance and the identification of critical periods. The system transitions between different guidance levels (from high guidance during critical periods to lower guidance during stable phases), optimizing learning efficiency without requiring complex manual intervention or system reconfiguration.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements feedback mechanisms that monitor the agent's learning progress and performance metrics to identify critical periods. Based on this feedback, the guidance level is automatically adjusted - increasing guidance during critical periods when the agent shows rapid learning potential, and reducing guidance during stable phases, thereby enhancing learning effect without proportionally increasing system complexity.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12505659B2Computing apparatus and method for performing reinforcement learning using multimodal artificial intelligence agent
Publication Date: 2025.12.23 SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
  • US12505659B2 patent drawing
  • US12505659B2 patent drawing
  • US12505659B2 patent drawing

AI summary

Disclosed herein are a computing apparatus and method for performing reinforcement learning using a multimodal artificial intelligence agent. The method for performing reinforcement learning using a multimodal artificial intelligence agent includes: dividing frames, included in images acquired by capturing a virtual environment, into a plurality of sections; and performing reinforcement learning by applying any one of a plurality of guidance types to each of the plurality of sections and then allowing a multimodal artificial intelligence agent to interact with the virtual environment through the images. The plurality of guidance types is classified into three or more types according to their guidance level. Performing the reinforcement learning is performing reinforcement learning by applying a moderate-level guidance type to the sections of predetermined critical periods and also applying any one of the plurality of guidance types to the other sections.