Resiliency Services for Distributed Cloud Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing distributed computer architectures face challenges in providing reliable failover systems due to the complexity and inter-dependencies of cloud-based services, leading to high costs and potential human errors in manually creating fault models, which can result in critical failures.

Innovation Solution

A system and method utilizing a service handling engine to analyze metric values of dependency services, enabling resiliency services through circuit breakers and fallback techniques based on faulty metrics, and automatically disabling them once the services recover, thus providing proactive and efficient failure management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual fault models are created using circuit breaker and fallback techniques, then resiliency services can be provided, but the cost and time increase prohibitively when functionality, architecture, and dependencies change often

Engineering Contradiction:
Improveresiliency service availabilityVSAvoidfault model creation and maintenance time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables self-service by automatically generating and updating fault models through continuous monitoring of service metrics and dependencies. The fault detection engine autonomously identifies faltering services and updates circuit breaker configurations without manual intervention, allowing the system to adapt to architecture changes automatically

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary action by proactively monitoring service health metrics and detecting potential failures before they impact customers. Circuit breakers are configured in advance based on detected anomalies, and fallback services are pre-identified and ready to activate immediately when failures occur

Inventive Principle:
Principle #10Preliminary action

2Reliability

If manual fault models are created, then some resiliency can be achieved, but human error and lack of knowledge about dependencies result in critical failures being missed

Engineering Contradiction:
Improvefailure detection accuracyVSAvoidsystem monitoring and detection complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system implements feedback by continuously monitoring service metrics, detecting anomalies, and using this information to dynamically update circuit breaker configurations. The fault detection engine receives ongoing data from service endpoints and automatically adjusts resiliency parameters based on observed service behavior and dependency health

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The fault detection engine acts as an intermediary between service endpoints and circuit breakers. It collects metrics from multiple services, analyzes dependency health, and translates this information into appropriate circuit breaker actions, eliminating the need for manual fault model creation while accurately capturing complex service relationships

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated monitoring and resiliency management is implemented, then operational efficiency improves, but system complexity increases

Engineering Contradiction:
Improveoperational efficiencyVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system achieves universality by designing a fault detection engine that can monitor multiple service types and dependencies through a unified interface. The same core mechanisms handle diverse service metrics, dependency relationships, and failure scenarios, reducing operational complexity while maintaining high productivity across the distributed system

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10592359B2System and a method for providing on-demand resiliency services
Publication Date: 2020.03.17 COGNIZANT TECH SOLUTIONS INDIA PVT LTD
  • US10592359B2 patent drawing
  • US10592359B2 patent drawing
  • US10592359B2 patent drawing

AI summary

A system and method for handling one or more dependency services hosted by one or more dependency servers for an upstream service hosted by an administrative server in a distributed computer architecture is provided. The present invention provides for identifying any abnormality in the behavior of the dependency services on the basis of metric values associated with service-parameters of said dependency services. Further, the resiliency services are enabled in order to handle one or more faltering dependency services based on the faulty metric values associated with the service-parameters. Yet further, the one or more faltering dependency services are continuously monitored, and one or more resiliency services are withdrawn once the fault in said dependency services is resolved. Yet further, the present invention provides a conversational bot interface for managing the administrative server and associated dependency services.