Reinforcement Learning Agent Pre-training via Knowledge Store Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Reinforcement learning software agents often fail to make optimal decisions in complex environments and are limited in performance due to their inability to learn from past decisions and adapt to dynamic situations.

Innovation Solution

A reinforcement learning model is trained using data from a knowledge store, including semi-structured information, and tested in a sandbox environment with limited connectivity to an external environment, allowing the software agent to autonomously perform actions and improve its decision-making processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If reinforcement learning is used to enable autonomous decision-making, then the software agent can operate independently, but the agent fails to make optimal decisions in complex situations

Engineering Contradiction:
Improveautonomous decision-makingVSAvoiddecision optimality
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-training the reinforcement learning agent in a simulated environment before deployment to the real environment. The agent learns optimal policies through extensive training in the simulator, which prepares it to handle complex situations effectively when deployed, thus resolving the contradiction between autonomous operation and decision optimality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses a domain-specific simulator that creates a virtual copy of the real environment. This simulator replicates the complex environment's behavior, allowing the agent to train and learn from virtual experiences that mirror real-world scenarios, thereby improving decision optimality while maintaining autonomous operation capabilities.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If traditional reinforcement learning paradigms are used, then the agent can learn from interactions, but the agent performs poorly in complex situations with high fan out of actions

Engineering Contradiction:
Improvelearning capabilityVSAvoidperformance in complex situations
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system performs preliminary training in a simulator before real-world deployment. During this pre-training phase, the agent learns from extensive simulated interactions and develops policies that handle complex situations effectively, thereby improving productivity and performance in real complex environments while maintaining adaptability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The domain-specific simulator acts as an intermediary between the training process and real-world deployment. It provides a controlled environment where the agent can learn from simulated interactions and develop robust policies, serving as a bridge that translates learning capabilities into effective real-world performance in complex situations.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If the agent operates in a real environment directly, then the agent can process real requests, but the agent cannot learn from past decisions efficiently

Engineering Contradiction:
Improvereal environment operationVSAvoidtraining time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent applies preliminary action by conducting extensive training in a simulated environment before deploying the agent to handle real requests. This pre-training phase allows the agent to learn from vast amounts of simulated past decisions and interactions, significantly reducing the time needed to learn effective policies when deployed to real environments.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The simulator creates virtual copies of real environment scenarios, allowing the agent to train on replicated past decisions and situations without consuming real time or resources. This copying approach enables efficient learning from historical data and simulated experiences while maintaining the ability to operate in real environments.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11797820B2Data augmented training of reinforcement learning software agent
Publication Date: 2023.10.24 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11797820B2 patent drawing
  • US11797820B2 patent drawing
  • US11797820B2 patent drawing

AI summary

Techniques are provided for reinforcement learning software agents enhanced by external data. A reinforcement learning model supporting the software agent may be trained based on information obtained from one or more knowledge stores, such as online forums. The trained reinforcement learning model may be tested in an environment with limited connectivity to an external environment to meet performance criteria. The reinforcement learning software agent may be deployed with the tested and trained reinforcement learning model within an environment to autonomously perform actions to process requests.