Self-Healing Automation Module for Production Incident Resolution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional software application development and production management tools require excessive manual effort, leading to error-prone and time-consuming processes for resolving production incidents, especially in complex environments with many dependencies, resulting in prolonged resolution times and increased developer involvement.

Innovation Solution

Implementing an automation module that triggers self-healing processes, automatically identifies and mitigates job failures, and redirects incidents to self-healing queues, minimizing human intervention and enabling real-time changes in job flow execution, thereby reducing manual intervention and improving overall production support efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual processes are used for production incident resolution, then domain knowledge can be applied through run books, but the process becomes extremely time consuming and error prone

Engineering Contradiction:
Improveresolution accuracyVSAvoidresolution time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-configuring automation templates with mitigation steps for known error codes before incidents occur. When an incident happens, the pre-configured templates are automatically applied, eliminating the need for manual run book execution and reducing resolution time while maintaining accuracy through predefined validated procedures.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system enables self-service by allowing automation to autonomously identify errors, select appropriate templates, and execute mitigation steps without human intervention. The automation module continuously monitors job failures, automatically correlates them with error codes, and applies corrective actions, freeing human operators from repetitive manual tasks while ensuring consistent application of domain knowledge.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If developers are involved in production support to identify root causes, then accurate diagnosis can be achieved, but significant delay is added to resolving production issues

Engineering Contradiction:
Improveroot cause identification accuracyVSAvoidincident resolution time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system replaces the mechanical process of manual root cause analysis by developers with an automated diagnostic engine. The automation module uses pre-configured error codes and correlation rules to automatically identify root causes by analyzing job failure patterns and comparing them against known error signatures, achieving both accuracy and speed without requiring developer intervention for routine incidents.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces an intermediary automation layer between job failures and developer involvement. This intermediary automatically correlates failures with error codes, determines root causes, and executes mitigations, only escalating to developers when automation cannot resolve the issue. This intermediary process filters out routine incidents that would otherwise require developer time.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of manufacture

If production support manually executes recovery instructions from run books, then issues can be addressed following established procedures, but the process remains fragmented and time consuming

Engineering Contradiction:
Improverecovery process standardizationVSAvoidincident resolution throughput
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The system achieves universality by creating a unified automation platform that handles multiple types of job failures across different data flows and systems through a single set of configurable templates. Instead of separate run books for each scenario, the automation module uses error codes and correlation rules to universally apply appropriate mitigation strategies, standardizing the recovery process while increasing throughput through automated execution.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system merges fragmented recovery procedures into a unified automated workflow. Multiple run books and recovery instructions are consolidated into configurable automation templates that are automatically selected and executed based on error code correlation. This merging eliminates the need for manual navigation through separate procedures, standardizing the process and dramatically increasing resolution throughput through automated orchestration.

Inventive Principle:
Principle #5Merging (Combining)

4Reliability

If ad-hoc requests are sent to trigger data flow jobs in specific order, then data quality issues can be addressed, but significant delay is added to resolving production issues

Engineering Contradiction:
Improvedata quality correctionVSAvoidcorrection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-configuring dependency relationships and execution sequences in the automation templates. When data quality issues are detected, the automation module automatically triggers the required data flow jobs in the correct order based on pre-defined dependencies, eliminating the need for manual ad-hoc requests and significantly reducing correction time while ensuring data quality reliability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11693727B2Systems and methods to identify production incidents and provide automated preventive and corrective measures
Publication Date: 2023.07.04 JPMORGAN CHASE BANK NA
  • US11693727B2 patent drawing
  • US11693727B2 patent drawing
  • US11693727B2 patent drawing

AI summary

Various methods, apparatuses/systems, and media for identifying production incidents and implementing automated preventive and corrective measures are disclosed. A processor automatically triggers, in response to a generated incident of a job/process/host failure, a self-healing service. The processor identifies an application to which the event generated belongs to by accessing a database that stores the application and host details; fetches functional identification (ID) of the application from the database, identifies the type of job failure or service degradation; automatically executes, by utilizing predefined micro services, the steps required for mitigation; records, in response to executing, outcome of the mitigation in the database along with output at each stage of execution; and evaluates the outcome of the mitigation by executing health checks using micro services to determine whether the failed job or process or host is healthy; and closes the incident based on healthy determination.