Snapback Framework for AI Model Failure Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing application lifecycle management technologies lack pre-built or reusable support for automated failure recovery and continuous improvement of AI models post-deployment, requiring custom-built solutions for each application.

Innovation Solution

A configurable snapback framework that implements runtime detection and restoration of failure conditions in deployed AI models, using failure detection components and associated model restoration pipelines, which plugs into existing lifecycle pipelines and delivers new model artifacts for improved versions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If custom-built solutions are implemented for each application to support automated failure recovery and continuous improvement of AI models, then the reliability and functionality of individual applications are improved, but the device complexity and development time increase significantly

Engineering Contradiction:
Improvefailure recovery capabilityVSAvoidcustom-built solution complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a universal snapback framework that provides failure detection and restoration capabilities across multiple AI model applications. The framework includes reusable components such as failure condition modules, snapback modules, and model restoration pipelines that can be deployed once and applied to numerous different AI models and applications, eliminating the need to build custom solutions for each application while maintaining reliability and failure recovery capabilities

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The framework segments the failure recovery functionality into independent, modular components including failure condition detection modules, snapback modules, and model restoration pipelines. These segmented components can be independently configured, deployed, and reused across different applications, reducing overall system complexity while providing comprehensive failure recovery coverage

Inventive Principle:
Principle #1Segmentation

2Productivity

If manual processes are used for model restoration and failure recovery, then the ease of operation and understanding are maintained, but the productivity and time required for model recovery decrease

Engineering Contradiction:
Improvemodel restoration speedVSAvoidmanual operation simplicity
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The framework performs preliminary actions by pre-configuring failure condition modules and snapback modules before failures occur. The system proactively monitors for failure conditions, maintains restoration pipelines ready to execute, and automatically triggers restoration operations when failures are detected, eliminating the need for manual intervention and significantly accelerating model recovery time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The snapback framework implements self-service by automatically detecting failure conditions, triggering appropriate snapback modules, and executing model restoration operations without human intervention. The system monitors its own operational status, identifies failures, and performs recovery actions autonomously, thereby increasing productivity while maintaining operational simplicity through automation

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11294759B2Detection of failure conditions and restoration of deployed models in a computing environment
Publication Date: 2022.04.05 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11294759B2 patent drawing
  • US11294759B2 patent drawing
  • US11294759B2 patent drawing

AI summary

A computer-implemented method includes obtaining data associated with execution of a model deployed in a computing environment. At least a portion of the obtained data is analyzed to detect one or more failure conditions associated with the model. One or more restoration operations are executed to generate one or more restoration results to address one or more detected failure conditions. At least a portion of the one or more restoration results is sent to the computing environment in which the model is deployed.