Scalable Data Storage Network Topology with Distributed RAID
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage systems face challenges in achieving high capacity, performance, and data availability while being scalable and fault-tolerant, particularly in handling large volumes of digital data and increasing client loads without complex hardware and software replacements.
Innovation Solution
The Omneon Content Library (OCL) system employs a scalable architecture with automatic load balancing, high-speed network switching, data caching, and replication to provide non-stop operation, redundancy, and easy expansion, distributing file slices across multiple drives and servers to prevent data loss and improve read performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If a distributed architecture is used to reduce cost, then storage capacity is improved, but system reliability deteriorates due to increased failure chances
Solution Approach 1:
The system segments storage into distributed units across multiple servers and uses RAID to segment disk drives into redundant groups. This allows the system to achieve high storage capacity through distribution while maintaining reliability through localized redundancy - if one drive fails, the RAID group can reconstruct the data without affecting the entire system.
Solution Approach 2:
The system implements beforehand cushioning by pre-configuring RAID redundancy and replication mechanisms before failures occur. The redundant array of inexpensive disks (RAID) is set up in advance to cushion against disk failures, and data replication ensures that if a server or drive fails, data availability is maintained through prior prepared redundancy.
2Quantity of substance
If storage capacity is increased to handle large volumes of data, then data availability is improved, but system complexity increases
Solution Approach 1:
The system manages large storage capacity by segmenting data across multiple servers and storage drives. This segmentation allows the system to scale capacity while distributing the management complexity across multiple independent units, each with its own RAID configuration and failure isolation mechanisms.
Solution Approach 2:
The system employs universal RAID groups that can serve multiple servers and storage needs. The RAID configuration provides multi-functional capability - it enables data redundancy, performance improvement through parallel access, and failure isolation, thereby managing complexity while scaling capacity.
3Reliability
If redundancy is implemented to maintain data availability during failures, then system reliability is improved, but hardware cost increases
Solution Approach 1:
The system implements redundancy through copying data across multiple RAID groups and servers. Instead of using expensive single-point-of-failure elimination techniques, the system creates copies of data and critical components across distributed units, achieving reliability through replication rather than through expensive redundant hardware at every level.
Solution Approach 2:
The system uses beforehand cushioning by implementing RAID redundancy and data replication in advance. These pre-configured redundancy mechanisms protect against failures without requiring expensive real-time backup systems or redundant hardware at every component level, thereby maintaining reliability while controlling hardware cost.
4Speed
If high-speed network links are used to improve data access performance, then data transfer speed is improved, but system scalability deteriorates
Solution Approach 1:
The system segments data access across multiple servers and storage drives, allowing high-speed network links to serve multiple clients simultaneously. This segmentation enables the system to scale by adding more servers and storage units without requiring proportional increases in network bandwidth between all components, as each unit can operate at high speed independently.
Solution Approach 2:
The system transitions from point-to-point high-speed connections to a distributed architecture where high-speed network links connect servers to a shared storage infrastructure. This dimensional change allows scalability - more servers can be added to the network without requiring dedicated high-speed links between all pairs of components, as they share the high-speed infrastructure.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A data storage system has a number of server groups, where each group has data storage servers. A file is stored in the system by being spread across two or more of the servers. The servers are communicatively coupled to internal packet switches. An external packet switch is communicatively coupled to the internal packet switches. Client access to each of the servers is through one of the internal packet switches and the external packet switch. Other embodiments are also described and claimed.