NAND Storage Offloading ECC and Flash Translation for Big Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional distributed storage systems for big data analysis are inefficient due to high costs, power consumption, and latency, requiring extensive infrastructure and complex networking, which becomes increasingly problematic as big data analysis grows.
Innovation Solution
A distributed storage system utilizing simplified NAND cards that offload error correction encoding and flash translation layer functionality, eliminating the need for DRAM interfaces and processors, and using parallel decompression engines to efficiently handle read-intensive data access requests.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional distributed storage systems are used to store replicas of original data source, then storage capacity and data availability are improved, but deployment cost, power consumption, and access latency increase significantly
Solution Approach 1:
The patent extracts and removes unnecessary components from traditional distributed storage systems. Specifically, it eliminates the need for separate storage servers, complex networking infrastructure, and associated control software by integrating storage functionality directly into the computing nodes through simple local storage devices, thereby reducing system complexity while maintaining data availability
Solution Approach 2:
The patent creates simplified copies of storage functionality at each computing node rather than using complex centralized storage systems. Each node maintains local storage copies of data, eliminating the need for complex distributed storage infrastructure while ensuring data availability through local access
2Quantity of substance
If conventional distributed storage systems with multiple servers are deployed, then storage capacity and redundancy are improved, but power consumption and operational costs increase
Solution Approach 1:
The patent merges storage and computing functions into a single integrated system at each node. By combining what were previously separate components (computing servers and storage servers) into unified computing nodes with local storage, the system reduces the total number of components and their associated power consumption while maintaining equivalent storage capacity
Solution Approach 2:
Each computing node serves its own storage needs through local storage devices, eliminating the need for power-consuming data transfer over networks and reducing reliance on external storage infrastructure. The system performs storage operations locally without requiring energy-intensive centralized storage resources
3Ease of operation
If data is transferred through Ethernet over long distances for big data analysis, then data accessibility is improved, but transfer overhead and latency increase considerably
Solution Approach 1:
The patent segments the monolithic distributed storage system into distributed local storage units at each computing node. This segmentation allows data to be accessed locally without requiring long-distance network transfers, thereby reducing access latency while maintaining data accessibility through distributed architecture
Solution Approach 2:
The system performs preliminary local storage of data at each computing node before analysis is needed. By having data readily available in local storage rather than requiring remote retrieval, the system eliminates transfer overhead and reduces access latency while maintaining ease of data access
4Productivity
If intermediate results are stored in memory-style medium to accelerate data analysis, then processing speed is improved, but data transfer overhead through Ethernet increases
Solution Approach 1:
The patent extracts the need for intermediate result storage in separate memory-style media by enabling direct local storage and processing at computing nodes. This eliminates the requirement for additional data transfer to and from external storage systems, reducing transfer overhead while maintaining processing speed through local access
Data Source
AI summary
One embodiment facilitates data access in a storage device. During operation, the system obtains, by the storage device, a file from an original physical media separate from the storage device, wherein the file comprises compressed data which has been previously encoded based on an error correction code (ECC). The system stores, on a physical media of the storage device, the obtained file as a read-only replica. In response to receiving a request to read the file, the system decodes, by the storage device based on the ECC, the replica to obtain ECC-decoded data, wherein the ECC-decoded data is subsequently decompressed by a computing device associated with the storage device and returned as the requested file.


