Selective Data Range Retrieval in Cloud Backup Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cloud-based storage systems require users to download entire files, which can be time-consuming and inefficient, especially when only a subset of data is needed, such as a specific email or dataset, leading to unnecessary resource usage.
Innovation Solution
Implementing a client-side application that allows users to specify and retrieve specific data ranges, such as byte ranges, from a cloud datacenter, enabling efficient retrieval and assembly of requested data without needing to download the entire file, using tools like FUSE to facilitate this process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the entire file is downloaded from the cloud datacenter to the local user machine, then the user can access the complete data, but the retrieval time and resource usage become unacceptably long
Solution Approach 1:
The patent divides the file into multiple data ranges (e.g., byte ranges) that can be independently identified and retrieved. Instead of downloading the entire file, the system segments the data retrieval process into discrete range requests, allowing selective recovery of only the necessary portions of the file to the user machine.
2Reliability
If the entire file is downloaded from the cloud datacenter to the local user machine, then the user can access the complete data, but the resource usage becomes excessive
Solution Approach 1:
The system extracts only the specific data ranges needed by the user from the complete file stored in the cloud datacenter. By identifying and retrieving only the relevant byte ranges rather than the entire file, the patent reduces network bandwidth consumption, storage requirements, and processing resources at the user machine while still providing the necessary data for recovery operations.
3Productivity
If selective data range retrieval is implemented, then retrieval time and resource usage are reduced, but the system complexity increases
Solution Approach 1:
The patent introduces a file system tool as an intermediary component that manages the complexity of selective data range retrieval. This tool acts as a mediator between the user application and the cloud datacenter, handling the identification of required data ranges, formulation of range requests, and assembly of retrieved data segments. By localizing this complexity in a dedicated file system tool, the overall system architecture remains manageable while enabling efficient selective recovery operations.
Data Source
AI summary
In one example, a method includes receiving, at a datacenter, a request from a client, where the request identifies a data range required by an application residing at the client, and the data range embraces less than all the contents of a file, backed up at the datacenter, with which the data range is associated. The example method further includes accessing the data in the data range, and transmitting data in the data range to the client, where the data transmitted to the client from the datacenter comprises respective portions of multiple incremental backups stored at the datacenter.


