Reinforcement Learning Action Verification via Domain Knowledge Conflict Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Reinforcement learning models are susceptible to bias and contamination, leading to flawed decision-making, and existing verification methods are time-consuming and lack transparency, especially in complex domains like communication networks and robotics, where flawed actions can have negative outcomes.

Innovation Solution

A computer-implemented method that classifies inputs to a reinforcement learning model as supportive or resistant to proposed actions and compares these classifications to domain knowledge to determine conflicts, allowing for the initiation of actions that do not contradict domain knowledge, thereby distinguishing between flawed logic and novel insights.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a verification step is employed to verify actions proposed by the reinforcement learning model, then the reliability of action execution is improved, but the time consumption and operational complexity increase

Engineering Contradiction:
Improvereliability of action executionVSAvoidtime consumption for verification
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by classifying inputs as supportive or resistant to proposed actions before execution. This pre-classification enables rapid verification by comparing input classifications against domain knowledge, allowing most actions to be verified automatically without time-consuming manual review while maintaining reliability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary verification mechanism that acts as a mediator between the reinforcement learning model and action execution. This intermediary automatically compares input classifications with domain knowledge to determine whether proposed actions conflict with established knowledge, reducing the need for direct human operator involvement while ensuring reliable execution.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If the reinforcement learning model operates in a black-box manner to maximize performance, then the productivity is improved, but the transparency and verifiability of decisions deteriorate

Engineering Contradiction:
Improveperformance of reinforcement learning modelVSAvoidtransparency of decision-making
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent applies segmentation by breaking down the black-box decision-making process into interpretable components. Specifically, it segments the input data into classified categories (supportive or resistant to proposed actions), making the model's reasoning process transparent and verifiable against domain knowledge while maintaining the model's high-performance capabilities.

Inventive Principle:
Principle #1Segmentation

3Reliability

If manual operator verification is performed for each proposed action, then the reliability is improved, but the productivity and speed of action execution deteriorate

Engineering Contradiction:
Improvereliability of action executionVSAvoidspeed of action execution
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies partial action by implementing a tiered verification approach. Instead of requiring full manual verification for all actions, it automatically verifies actions where input classifications align with domain knowledge and only flags for manual review those cases where conflicts are detected. This partial automation maintains reliability for routine actions while preserving speed, and engages human operators only when necessary.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20240249199A1Verifying an action proposed by a reinforcement learning model
Publication Date: 2024.07.25 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US20240249199A1 patent drawing
  • US20240249199A1 patent drawing
  • US20240249199A1 patent drawing

AI summary

The present disclosure provides a computer-implemented method for determining whether to perform an action proposed by a model. The model is developed using a reinforcement learning process. The method comprises classifying at least one of a plurality of inputs to the model as being supportive or resistant to an action proposed by the model. The method further comprises comparing the classification of the at least one of the plurality of inputs to domain knowledge to determine whether or not the proposed action conflicts with the domain knowledge, and, in response to determining that the proposed action does not conflict with the domain knowledge, initiating the proposed action. In this context, the domain knowledge is indicative of a relationship between the proposed action and the at least one of the plurality of inputs.