RAIS Distributed Parity Mapping for Scalable Fault Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional RAID architectures rely on a centralized controller, creating a single point of failure and scalability constraints, leading to performance bottlenecks and increased system downtime.
Innovation Solution
Implement a redundant array of independent servers (RAIS) that distributes parity computation, redundancy management, and data recovery across multiple servers, eliminating reliance on a centralized controller and enabling parallelized operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a centralized RAID controller is used to manage data distribution and parity computations, then data reliability is improved through redundancy mechanisms, but system scalability is limited and performance bottlenecks occur due to the single point of failure
Solution Approach 1:
The patent segments the centralized RAID controller functionality into distributed components across multiple servers. Each server maintains local redundancy management while participating in distributed parity computations, eliminating the single point of failure and enabling system scalability without compromising data reliability
2Reliability
If a centralized RAID controller manages all parity computations and data distribution, then fault tolerance is achieved, but performance bottlenecks occur due to centralized processing limitations
Solution Approach 1:
Parity computation tasks are segmented and distributed across multiple servers rather than centralized. Each server performs local parity computations independently, enabling parallel processing that eliminates performance bottlenecks while maintaining fault tolerance through distributed redundancy
Solution Approach 2:
The system dynamically distributes computational workloads across available servers based on current system state and resource availability. This dynamic allocation enables the system to scale performance with added resources while maintaining consistent fault tolerance levels
3Reliability
If redundancy management is centralized in a single controller, then data integrity is maintained through coordinated error correction, but system downtime increases during controller failures or maintenance
Solution Approach 1:
Redundancy management is segmented into distributed functions across multiple servers, eliminating the single point of failure. Each server can independently manage its local redundancy and participate in system-wide error correction, ensuring data integrity continues during maintenance or controller failures
Solution Approach 2:
Each server performs self-service redundancy management and can independently handle error correction for its local data. This self-sufficiency ensures that system operations continue without interruption during controller maintenance or failures, eliminating forced downtime
Data Source
AI summary
Various examples, systems, controllers, and methods are disclosed relating to distributed redundant storage across systems in a RAIS environment. Some systems can include a local array of drives configured to store data blocks and at least one configuration data structure including at least one mapping function for at least one logical block address. Some systems can include processing circuitry configured to perform a plurality of first operations to maintain distributed redundancy across a plurality of the systems and perform a plurality of second operations on the data blocks mapped to the local array of drives or remote data blocks on at least one remote array of drives of at least one remote system. The system and the remote system can share a distributed data mapping of the data blocks and share a distributed parity mapping of parity blocks and remote parity blocks.


