Distributed Error Impact Quantification Using Callee Incidents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing distributed computing environments lack effective methods to accurately quantify and measure the impact of errors or incidents on availability, performance, and correctness, making it difficult to prioritize and mitigate issues effectively.

Innovation Solution

A method and system for detecting errors in a distributed computing environment, associating them with callee incidents, and quantifying impact using service level metrics, allowing for improved measurement and prioritization of errors based on their severity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If error detection and quantification mechanisms are implemented in distributed computing environments, then measurement precision of error impact is improved, but device complexity increases

Engineering Contradiction:
Improveerror impact measurementVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces intermediary components including error detection agents deployed across distributed systems, central error quantification servers, and intermediary data structures that bridge local error detection with global impact assessment. These intermediaries handle the complex tasks of error tracking, correlation, and quantification, allowing the overall system to achieve precise error measurement without requiring every component to be complex

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The error quantification system is segmented into multiple independent components: error detection modules at individual system levels, error collection and transmission mechanisms, central quantification processing units, and reporting systems. This segmentation allows each component to perform a specific function with simple logic, while the aggregation of these components achieves comprehensive error impact measurement across the distributed environment

Inventive Principle:
Principle #1Segmentation

2Reliability

If comprehensive error tracking across distributed systems is implemented, then reliability of error impact assessment is improved, but loss of time for error processing increases

Engineering Contradiction:
Improveerror impact assessmentVSAvoiderror processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements preliminary error detection and logging at the source systems before errors propagate through the distributed environment. Error agents continuously monitor and capture error events as they occur, storing them in local buffers with minimal processing. This preliminary action ensures errors are captured reliably without requiring time-consuming analysis at each subsequent stage

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system establishes feedback loops where error quantification results are continuously updated and fed back to stakeholders. The error impact assessment is an iterative process that refines measurements over time, providing progressively more accurate information without requiring complete re-analysis of all error data. This feedback mechanism maintains reliability while managing processing time through incremental updates

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12625755B2Errors in distributed computing environments
Publication Date: 2026.05.12 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12625755B2 patent drawing
  • US12625755B2 patent drawing
  • US12625755B2 patent drawing

AI summary

Embodiments of the present invention provide concepts for quantifying impact of one or more errors in a distributed computing environment. A processor may detect, at a caller entity of the distributed computing environment, an error resulting from a request from the caller entity to a callee entity of the distributed computing environment. The processor may associate the detected error with a callee incident, the callee incident describing an abnormal operating condition of the callee entity. The processor may quantify an impact of the error based on callee incident associated with the detected error and a service level metric of the distributed computing environment.