Fault-Tolerant Data Distribution Service for HPC Clusters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing high performance and grid computing systems for hydrocarbon exploration and production lack proper quality of service (QoS) control, leading to inefficiencies and data accuracy issues due to limitations in communication libraries and interconnects, resulting in wasted processing time and resource duplication upon node failures.

Innovation Solution

A data processing system with a designated master publisher node and subscriber processor nodes, utilizing a quality of service profile with ownership strength and deadline intervals, implements the Data Distribution Service (DDS) standard to ensure fault-tolerant and predictable data processing, allowing for continuous operation even in case of node failures by designating a new master publisher based on ownership strength.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If standard communication libraries (MPI, PVM) are used for data distribution, then device complexity is reduced, but quality of service control and predictability deteriorate

Engineering Contradiction:
Improvecommunication library complexityVSAvoidquality of service control
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces a Quality of Service Manager as an intermediary component that mediates between the data distribution service and communication infrastructure. This manager enables applications to specify service quality requirements and translates them into actionable control parameters, resolving the contradiction by adding a dedicated mediation layer that provides QoS control without requiring fundamental changes to existing communication libraries.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the operational parameters of the data distribution service by introducing quality of service profiles with adjustable parameters such as ownership strength, deadline intervals, and data accuracy failure rates. These parameter changes enable fine-grained control over communication behavior, allowing the system to adapt to different quality requirements while maintaining compatibility with standard communication infrastructure.

Inventive Principle:
Principle #35Parameter changes

2Speed

If high-speed interconnects (InfiniBand, Myrinet, Quadrics, Gigabit Ethernet) are used, then processing speed is improved, but communication latency predictability and bandwidth reliability deteriorate

Engineering Contradiction:
Improvedata processing speedVSAvoidcommunication latency predictability
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent implements feedback mechanisms through quality of service monitoring that continuously tracks communication performance metrics including latency and bandwidth utilization. The system uses this feedback information to dynamically adjust data distribution strategies and prioritize critical communications, thereby achieving predictable performance even on high-speed interconnects with variable conditions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces dynamic quality of service adjustment capabilities that allow the system to adapt communication parameters in real-time based on current network conditions and application requirements. This dynamic behavior enables the system to maintain predictable latency and bandwidth utilization by adjusting data transmission characteristics on-the-fly, compensating for the inherent variability of high-speed interconnects.

Inventive Principle:
Principle #15Dynamics

3Reliability

If node failure occurs in existing HPC systems, then system reliability is improved through fault tolerance, but processing time increases due to resubmitting data sets

Engineering Contradiction:
Improvefault toleranceVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements preliminary action through quality of service profiles that pre-establish ownership strength assignments and data distribution strategies before node failures occur. The system prepares backup data paths and maintains quality of service state information that can be quickly activated upon failure detection, enabling rapid failover without requiring complete resubmission of data sets and minimizing processing time loss.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent ensures continuity of useful action by implementing quality of service monitoring that tracks processing state and maintains data accuracy failure rate information. When node failures occur, the system can continue processing by redistributing only the affected data portions while preserving the state of successfully processed data, thereby maintaining continuous productive operation rather than requiring complete restart.

Inventive Principle:
Principle #20Continuity of useful action

4Reliability

If data accuracy failure rate is increased for fault tolerance, then system reliability is improved, but data processing efficiency deteriorates

Engineering Contradiction:
Improvefault toleranceVSAvoiddata processing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies local quality by implementing quality of service profiles that assign different ownership strength values to different data elements and processing tasks. This allows the system to concentrate fault tolerance resources on critical data portions with high ownership strength while using more efficient processing strategies for less critical data, thereby maintaining high data processing efficiency while achieving the required level of system reliability through localized quality differentiation.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9674033B2High performance and grid computing with liveliness and deadlines fault tolerant data distributor quality of service
Publication Date: 2017.06.06 SAUDI ARABIAN OIL CO
  • US9674033B2 patent drawing
  • US9674033B2 patent drawing
  • US9674033B2 patent drawing

AI summary

High performance computing (HPC) and grid computing processing for seismic and reservoir simulation are performed without impacting or losing processing time in case of failures. A Data Distribution Service (DDS) standard is implemented in High Performance Computing (HPC) and grid computing platforms, to avoid the shortcomings of current Message Passing Interface (MPI) communication between computing modules, and provide quality of service (QoS) for such applications. QoS properties of the processing can be controlled. Multiple data publishers or master nodes of a cluster have access to the same data source. Each of these publishers has an ownership strength quality of service, and the publisher with the highest ownership strength number is the designated publisher of the data to subscriber processor nodes of the cluster. If the designated data publisher prematurely terminates or crashes for some reason, then the publisher node with the next highest ownership strength measure is designated as data publisher and continues publishing data to subscribers. The QoS properties include ability to switch to a different designated publisher when the input data is being processed in real time, and a specified deadline time has passed during which data has not been published.