Abstract Reinforcement Learning for IT Process Automation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The IT infrastructure space is complex and requires highly skilled engineers to set up and troubleshoot, hindering transversal innovation and increasing costs, while existing reinforcement learning mechanisms are cumbersome and expensive due to domain-specific coding needs.
Innovation Solution
An abstract reinforcement learning model that generates infrastructure process automation candidates based on user intent, allowing non-expert users to define states, transitions, and rewards without programming, using a graphical interface and abstract domain models like SBVR to decouple vendor-specific syntax from processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If reinforcement learning is applied to IT infrastructure automation, then productivity is improved, but device complexity increases due to domain-specific coding requirements
Solution Approach 1:
The patent introduces an intermediary layer (abstraction model with states, transitions, and rewards) that mediates between the complex reinforcement learning algorithms and the domain-specific IT infrastructure operations. This intermediary translates high-level operational intents into machine learning problems without requiring experts to code domain-specific algorithms, thus improving productivity while managing complexity.
Solution Approach 2:
The patent creates a universal reinforcement learning framework that can be applied across multiple IT infrastructure domains (networking, security, data centers) through a common abstraction model. This universal approach eliminates the need for separate domain-specific coding implementations, reducing device complexity while maintaining high productivity across diverse infrastructure operations.
2Reliability
If expert engineers are used to set up and troubleshoot IT infrastructure, then reliability is improved, but loss of substance increases due to high costs
Solution Approach 1:
The patent enables self-service automation where the reinforcement learning system autonomously performs infrastructure setup, configuration, and troubleshooting without requiring expert engineers. The system learns optimal operations through the abstraction model and executes them automatically, maintaining reliability while eliminating the ongoing cost of expert personnel.
Solution Approach 2:
The patent uses preliminary action by pre-defining the abstraction model structure (states, transitions, rewards) that captures domain expertise beforehand. This preliminary formulation allows the system to automatically handle complex operations without requiring experts to be present during execution, reducing operational costs while maintaining reliability through encoded best practices.
3Manufacturing precision
If domain-specific coding is required for reinforcement learning, then manufacturing precision is improved, but ease of manufacture deteriorates
Solution Approach 1:
The patent segments the complex task of creating domain-specific reinforcement learning code into manageable components: defining states, transitions, and rewards through the abstraction model. This segmentation allows non-experts to configure precise automation processes by selecting and combining predefined elements rather than writing complex domain-specific code from scratch.
Solution Approach 2:
The patent enables parameter changes by allowing users to modify the abstraction model parameters (state definitions, transition conditions, reward structures) to achieve manufacturing precision without changing the underlying code. This parameter-based configuration maintains precise automation while dramatically improving ease of manufacture for non-programming users.
Data Source
AI summary
Information defining a plurality of states, a plurality of transitions, an initial state, and a final state is received from a user. The user may also provide additional information including pre-conditions and post-conditions for one or more transitions. Context information including one or more context variables and context variable values is generated based on the information provided by the user. A first plurality of possible paths between the initial state and the final state is automatically identified, wherein each path traverses at least one state and at least one transition. A second plurality of paths is identified from among the plurality of paths, based on the context information and the pre-conditions defined by the user. A Q-value is determined for each path in the second plurality of paths, using the rewards. A path having a highest Q-value is selected and presented to the user as a BPM. An acceptance or rejection of the proposed BPM is received from the user. Reward values associated with transitions in the selected path are updated, if the user accepts the proposed BPM.


