Virtualized Grid Cluster for Fault Tolerant Batch Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current batch processing solutions in virtualized environments face challenges with scalability, fault tolerance, and resource management, leading to delays and potential data integrity issues, especially in heterogeneous and distributed environments.

Innovation Solution

A virtualized grid cluster system using grid computing and virtualization technologies, with a centralized storage repository, grid manager, and policy engine, that provisions resources on-demand, manages job queues, and ensures fault-tolerant execution by deploying virtual machines and scheduling jobs based on priority and resource availability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If dedicated clustering technologies are used for batch processing, then high availability and performance are achieved, but scalability beyond local spatial environment is limited

Engineering Contradiction:
Improvehigh availabilityVSAvoidscalability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system segments batch processing jobs into independent units that can be distributed across a heterogeneous grid of computing resources. Each job is broken down into executable tasks that can be assigned to different nodes in the grid, enabling scalability beyond traditional homogeneous clusters while maintaining fault tolerance through redundant task assignment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The grid manager implements a universal scheduling framework that can handle diverse job types, resource configurations, and failure scenarios through a single system. The policy-based scheduling mechanism provides multi-functional capability to manage both homogeneous and heterogeneous resources, ensuring high availability while enabling broad scalability across different computing environments.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If homogeneous computing clusters are used, then management simplicity is maintained, but ability to meet enterprise demands for heterogeneous distributed environment is insufficient

Engineering Contradiction:
Improvemanagement simplicityVSAvoidheterogeneous environment support
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The grid manager acts as an intermediary layer between the heterogeneous grid resources and the batch processing jobs. It abstracts the complexity of managing diverse computing nodes through a unified interface, allowing simple policy-based scheduling while supporting heterogeneous environments. The message bus serves as another intermediary for standardized communication between grid components.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If batch processing is executed during business hours, then real-time business intelligence is provided, but critical customer front-end applications are impacted due to high resource cost

Engineering Contradiction:
Improvereal-time processing speedVSAvoidfront-end application performance
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system dynamically adjusts resource allocation based on real-time conditions. The policy engine monitors system state and dynamically schedules batch jobs when resources are available, enabling flexible execution that can adapt to varying workloads. This dynamic scheduling allows batch processing to utilize otherwise idle resources without impacting front-end applications during peak business hours.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The scheduling policy changes system parameters such as job execution timing, resource allocation, and priority levels based on real-time conditions. By adjusting these parameters dynamically, the system can execute batch processing during periods when front-end applications have lower resource demands, thus maintaining both productivity and reliability.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If traditional grid computing is used, then fault tolerance is achieved, but resource provisioning efficiency and scalability are reduced

Engineering Contradiction:
Improvefault toleranceVSAvoidresource provisioning efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The grid system implements self-service mechanisms where the grid manager automatically monitors resource availability, provisions computing resources on-demand, and schedules jobs without manual intervention. The system self-adjusts to failures by automatically reassigning tasks to healthy nodes, maintaining fault tolerance while improving resource provisioning efficiency through automated decision-making.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes operational parameters such as resource provisioning timing, job scheduling priorities, and failure response strategies based on real-time system state. These parameter adjustments enable the system to maintain fault tolerance while optimizing resource utilization and improving overall productivity in the grid environment.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9483314B2Systems and methods for fault tolerant batch processing in a virtual environment
Publication Date: 2016.11.01 INFOSYS LTD
  • US9483314B2 patent drawing
  • US9483314B2 patent drawing
  • US9483314B2 patent drawing

AI summary

A system for fault tolerant batch processing in a virtual environment is configured to perform batch job execution, the system includes computing devices configured as a virtualized grid cluster by means of a virtualization platform, the cluster includes a centralized storage repository, a grid manager deployed on an instantiated virtual machine and a message bus whereby data and messages are exchanged between the grid manager and one or more grid nodes. The grid manager is configured to manage one or more incoming job requests, queue one or more of the received job requests in a job execution queue and monitor one or more virtual grid nodes.