IO Server Redundant Data Management via Remote NIC
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current information processing systems, particularly IO servers, face reliability challenges due to a single-point failure in their controllers, which can lead to data loss and system instability, especially as system scale increases for improved performance through parallel processing.
Innovation Solution
The system implements a redundant data management mechanism where each IO server can store and retrieve redundant data directly through a remote NIC command execution mechanism, allowing for data segmentation and distribution across multiple servers without CPU intervention, enabling data recovery through exclusive OR operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If redundant data is retained in multiple information processing apparatuses to ensure reliability, then system reliability is improved, but memory usage quantity and overhead increase
Solution Approach 1:
The patent segments data into multiple pieces and distributes them across different information processing apparatuses. Instead of storing complete redundant copies of all data in each apparatus, the data is divided into segments (e.g., through hashing or encoding techniques), and each apparatus stores only its assigned segments. This segmentation approach maintains system reliability through distribution while significantly reducing the memory burden on individual apparatuses.
Solution Approach 2:
The patent transitions from a single-dimension storage model (storing complete data copies) to a multi-dimensional distributed storage model. Data is organized across multiple dimensions including spatial distribution (different apparatuses), segment identification (hash values or indices), and redundancy levels. This dimensional transformation enables efficient retrieval and reconstruction of original data from distributed segments while optimizing memory utilization across the system.
2Reliability
If redundant data is retained in multiple information processing apparatuses to ensure reliability, then system reliability is improved, but overhead due to data transfer processes increases
Solution Approach 1:
The patent performs preliminary actions by pre-segmenting data and pre-distributing it to appropriate information processing apparatuses before failures occur. Hash values or segment identifiers are calculated in advance, and data segments are proactively placed in optimal locations. This preliminary organization eliminates the need for complex real-time computation and data movement during failure recovery, significantly reducing overhead and transfer time when reliability events occur.
3Productivity
If system scale is increased for parallel processing to improve performance, then processing capability is improved, but reliability becomes more critical due to increased complexity
Solution Approach 1:
The patent implements self-service mechanisms where each information processing apparatus autonomously manages its stored data segments. Apparatuses independently track which segments they hold, respond to recovery requests without requiring centralized coordination, and can even initiate recovery operations themselves. This decentralized self-service approach scales efficiently with system size while maintaining reliability, as each node operates independently yet contributes to overall system resilience.
Data Source
AI summary
An apparatuses includes a processor, a storage unit, and a communication unit to access the storage unit without intermediary of the processor and to access a second apparatus of the plurality of information processing apparatuses via a communication unit of the second apparatus. The communication unit of a first apparatus of the plurality of information processing apparatuses executes at least one of a process of storing redundant data which is generated by making redundant data stored in the storage unit of the first apparatus in the storage unit of the second apparatus via the communication unit of the second apparatus, and a process of acquiring redundant data which is generated by making redundant data stored in the storage unit of the second apparatus via the communication unit of the second apparatus, and storing the acquired data in the storage unit of the first apparatus.


