Automated Adaptive System for Production Environment Failure Remediation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying and remediating issues in production environments are inefficient and time-consuming, often requiring manual analysis and intervention, which can lead to bottlenecks and poor consumer experience.

Innovation Solution

An automated and adaptive system that performs frequent health checks on hypervisors, analyzes failure cases, and automatically remediates issues by matching patterns to stored or predefined rules, using a database to identify and execute remediation plans without manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual analysis and debugging methods are used to identify and fix problems in production systems, then issues can be resolved with human judgment and adaptability, but the process becomes time-consuming and creates bottlenecks

Engineering Contradiction:
Improveproblem resolution accuracyVSAvoidtime to identify and fix problems
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by continuously collecting system data, analyzing failure patterns, and preparing remediation plans before actual failures occur. Health checks are performed proactively, and when failures are detected, the system quickly matches them against stored patterns to immediately execute pre-prepared remediation actions, eliminating the time delay associated with manual analysis and debugging

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system enables self-service by automating the entire problem identification and resolution process. The automated health check system detects failures, analyzes patterns, selects appropriate remediation plans, and executes fixes without human intervention. This self-service capability allows the system to resolve issues autonomously, significantly reducing the time loss associated with manual problem-solving while maintaining reliable resolution through structured pattern matching

Inventive Principle:
Principle #25Self-service

2Speed

If frequent health checks are performed to quickly identify issues, then problem detection speed increases, but system resource consumption and complexity increase

Engineering Contradiction:
Improveissue detection speedVSAvoidsystem complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The system applies segmentation by dividing the health check process into distinct, manageable components: data collection modules that gather specific system metrics, pattern analysis modules that compare failures against stored templates, and remediation execution modules that apply predefined fixes. This segmentation allows frequent health checks to be performed efficiently without overwhelming system complexity, as each component handles a specific aspect of the monitoring and remediation process independently

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system utilizes parameter changes by dynamically adjusting health check frequencies and thresholds based on system conditions and learned failure patterns. Rather than performing uniform frequent checks on all systems, the automated system adapts monitoring parameters to prioritize critical components and adjust check intensity based on historical data, thereby maintaining fast issue detection while managing system complexity through intelligent parameter optimization

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11126494B2Automated, adaptive, and auto-remediating system for production environment
Publication Date: 2021.09.21 PAYPAL INC
  • US11126494B2 patent drawing
  • US11126494B2 patent drawing
  • US11126494B2 patent drawing

AI summary

A first computer system identifies a failure case from a collected information, the collected information corresponds to one or more functions of a second system. The first computer system analyzes the failure case to determine a failure pattern. The first computer system determines whether the determined failure pattern corresponds to a stored failure pattern of a plurality of stored failure patterns in a database. In response to determining that the determined failure pattern corresponds to the stored failure pattern, the first computer system determines a remediation plan corresponding to the stored failure pattern, and utilizing the remediation plan to automatically remediate the failure case.