Reward Function for Neural Network Exploration Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current exploration strategies in reinforcement learning lack incentive for agents to explore novel situations, leading to limited diversity in behaviors and inadequate training of neural networks in dynamic environments.

Innovation Solution

Implementing a reward-based strategy that differentiates between novel and previously explored agent states, using a neural network to process simulations and assign rewards for unexplored states, thereby promoting exploration and expanding the agent's knowledge base.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional reinforcement learning is used without reward-based exploration strategies, then the training process is simpler, but the agent fails to explore novel situations leading to limited behavioral diversity

Engineering Contradiction:
Improvebehavioral diversityVSAvoidtraining system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a feedback mechanism where the agent receives rewards based on whether it transitions to novel or previously explored states. This feedback loop encourages exploration by providing positive reinforcement for discovering new states, thereby increasing behavioral diversity without requiring complex manual intervention

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system enables self-service exploration by automatically identifying novel states and assigning rewards without external intervention. The neural network autonomously determines which states are novel and provides appropriate rewards, allowing the agent to self-drive the exploration process

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If exploration strategies are implemented to maximize diversity, then behavioral diversity improves, but the incentive mechanism for exploring novel situations becomes insufficient

Engineering Contradiction:
Improveexploration incentiveVSAvoidnovel state identification accuracy
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent replaces manual or mechanical methods of identifying novel states with a neural network-based system. The neural network automatically processes state information and determines novelty, providing accurate identification of new states while enabling the reward mechanism to effectively incentivize exploration

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of information

If the neural network is trained with novel unexplored agent states, then the knowledge base expands, but the training data requirements increase

Engineering Contradiction:
Improveknowledge base coverageVSAvoidtraining data volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The system performs preliminary identification of novel states during the exploration process and prepares training data incrementally. By identifying and storing novel state information as it encounters them during exploration, the system builds the knowledge base progressively without requiring large batches of pre-collected training data

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11699062B2System and method for implementing reward based strategies for promoting exploration
Publication Date: 2023.07.11 HONDA MOTOR CO LTD
  • US11699062B2 patent drawing
  • US11699062B2 patent drawing
  • US11699062B2 patent drawing

AI summary

A system and method for implementing reward based strategies for promoting exploration that include receiving data associated with an agent environment of an ego agent and a target agent and receiving data associated with a dynamic operation of the ego agent and the target agent within the agent environment. The system and method also include implementing a reward function that is associated with exploration of at least one agent state within the agent environment. The system and method further include training a neural network with a novel unexplored agent state.