Abstract Reinforcement Learning for IT Process Automation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The IT infrastructure space is complex and requires highly skilled engineers to set up and troubleshoot, hindering transversal innovation and increasing costs, while existing reinforcement learning mechanisms are cumbersome and expensive due to domain-specific coding needs.

Innovation Solution

An abstract reinforcement learning model that generates infrastructure process automation candidates based on user intent, allowing non-expert users to define states, transitions, and rewards without programming, using a graphical interface and abstract domain models like SBVR to decouple vendor-specific syntax from processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If reinforcement learning is applied to IT infrastructure automation, then productivity is improved, but device complexity increases due to domain-specific coding requirements

Engineering Contradiction:
Improveautomation of infrastructure operationsVSAvoidcomplexity of reinforcement learning implementation
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary layer (abstraction model with states, transitions, and rewards) that mediates between the complex reinforcement learning algorithms and the domain-specific IT infrastructure operations. This intermediary translates high-level operational intents into machine learning problems without requiring experts to code domain-specific algorithms, thus improving productivity while managing complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a universal reinforcement learning framework that can be applied across multiple IT infrastructure domains (networking, security, data centers) through a common abstraction model. This universal approach eliminates the need for separate domain-specific coding implementations, reducing device complexity while maintaining high productivity across diverse infrastructure operations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If expert engineers are used to set up and troubleshoot IT infrastructure, then reliability is improved, but loss of substance increases due to high costs

Engineering Contradiction:
Improveinfrastructure operation reliabilityVSAvoidcost of ownership
Core Design Contradiction:
ReliabilityVSLoss of substance

Solution Approach 1:

The patent enables self-service automation where the reinforcement learning system autonomously performs infrastructure setup, configuration, and troubleshooting without requiring expert engineers. The system learns optimal operations through the abstraction model and executes them automatically, maintaining reliability while eliminating the ongoing cost of expert personnel.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses preliminary action by pre-defining the abstraction model structure (states, transitions, rewards) that captures domain expertise beforehand. This preliminary formulation allows the system to automatically handle complex operations without requiring experts to be present during execution, reducing operational costs while maintaining reliability through encoded best practices.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If domain-specific coding is required for reinforcement learning, then manufacturing precision is improved, but ease of manufacture deteriorates

Engineering Contradiction:
Improveprecision of automation processesVSAvoidease of creating automation processes
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent segments the complex task of creating domain-specific reinforcement learning code into manageable components: defining states, transitions, and rewards through the abstraction model. This segmentation allows non-experts to configure precise automation processes by selecting and combining predefined elements rather than writing complex domain-specific code from scratch.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent enables parameter changes by allowing users to modify the abstraction model parameters (state definitions, transition conditions, reward structures) to achieve manufacturing precision without changing the underlying code. This parameter-based configuration maintains precise automation while dramatically improving ease of manufacture for non-programming users.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20260044808A1Systems and Methods for Autogeneration of Information Technology Infrastructure Process Automation and Abstraction of the Universal Application of Reinforcement Learning to Information Technology Infrastructure Components and Interfaces
Publication Date: 2026.02.12 UBIQUBE (IRELAND) LTD
  • US20260044808A1 patent drawing
  • US20260044808A1 patent drawing
  • US20260044808A1 patent drawing

AI summary

Information defining a plurality of states, a plurality of transitions, an initial state, and a final state is received from a user. The user may also provide additional information including pre-conditions and post-conditions for one or more transitions. Context information including one or more context variables and context variable values is generated based on the information provided by the user. A first plurality of possible paths between the initial state and the final state is automatically identified, wherein each path traverses at least one state and at least one transition. A second plurality of paths is identified from among the plurality of paths, based on the context information and the pre-conditions defined by the user. A Q-value is determined for each path in the second plurality of paths, using the rewards. A path having a highest Q-value is selected and presented to the user as a BPM. An acceptance or rejection of the proposed BPM is received from the user. Reward values associated with transitions in the selected path are updated, if the user accepts the proposed BPM.