Transaction Mirroring Fault Tolerance via Timeout Validation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data storage systems face challenges in maintaining fault tolerance during multiple node failures, particularly in transactions using the two-phase commit protocol, which can lead to indeterminate transactions and system downtime.

Innovation Solution

Implementing a transaction management system with a state monitoring and update component that marks secondary participant nodes as invalid if they fail to respond within a threshold time, allowing the primary participant node to continue transactions and maintain fault tolerance even with multiple node failures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the system waits for confirmation from all participant nodes before committing a transaction, then transaction consistency is improved, but system availability and productivity deteriorate due to node failures causing transaction delays

Engineering Contradiction:
Improvetransaction consistencyVSAvoidsystem availability
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary actions by sending the commit command to all participant nodes in advance, including secondary nodes. The primary node then waits for confirmations within a timeout period. If confirmations are not received within the threshold time, the transaction is committed anyway based on the preliminary preparations already made at other nodes, thus maintaining consistency while improving availability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements beforehand cushioning by establishing timeout mechanisms and predefined commit protocols that allow the transaction to proceed even if secondary nodes fail to respond. This cushioning mechanism ensures that a single node failure does not block the entire transaction, thereby maintaining system availability while preserving transaction consistency through the predefined commit logic.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

2Productivity

If the system marks failed nodes as invalid immediately, then system availability is improved by allowing transactions to continue, but reliability may worsen due to potential premature failure declarations

Engineering Contradiction:
Improvesystem availabilityVSAvoidnode failure detection accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system applies preliminary action by implementing a timeout threshold mechanism before marking a node as invalid. The node is only marked as invalid if it fails to respond within the predetermined threshold time. This preliminary waiting period ensures that transient network issues do not cause premature failure declarations, maintaining reliability while enabling timely failure detection for actual node failures.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If the system implements timeout-based failure detection, then ease of operation is improved by automating failure handling, but device complexity increases due to additional monitoring and timeout management mechanisms

Engineering Contradiction:
Improveautomatic failure handlingVSAvoidmonitoring mechanism complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system implements self-service by enabling the primary participant node to automatically detect node failures through timeout mechanisms and autonomously mark failed nodes as invalid without requiring external intervention. The monitoring component continuously tracks response times and automatically triggers failure declaration when thresholds are exceeded, providing automated failure handling that simplifies operation despite the added monitoring complexity.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11669516B2Fault tolerance for transaction mirroring
Publication Date: 2023.06.06 EMC IP HLDG CO LLC
  • US11669516B2 patent drawing
  • US11669516B2 patent drawing
  • US11669516B2 patent drawing

AI summary

Systems and methods facilitating fault tolerance for transaction mirroring are described herein. A method as described herein can include receiving a commit command for a data transaction from an initiator node of the system, wherein the data transaction is associated with a first failure domain, and wherein the commit command is directed to a primary participant node and a secondary participant node of the system; determining whether a response to the commit command has been received at the primary participant node from the secondary participant node in response to the receiving; and, in response to determining that the response to the commit command was not received at the primary participant node, indicating that the secondary participant node is invalid in a data store associated with a second failure domain that is distinct from the first failure domain.