Virtual Data Files for Faster Unstructured Data Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data storage and backup solutions incur high costs and extended recovery downtimes due to the inefficient handling of inactive unstructured data, which delays access to active data during disasters or ransomware attacks, especially in cloud-based systems.
Innovation Solution
Implementing Virtual Data Files (VDFs) that serve as pointers to unstructured data, allowing rapid restoration by sending VDFs first, followed by the actual data upon request, thereby reducing storage needs and accelerating data recovery.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional backup solutions store all unstructured data (active and inactive) using high-cost storage, then data protection is ensured, but storage costs increase and recovery time extends
Solution Approach 1:
The patent segments unstructured data into active and inactive portions, applying different storage strategies to each. Active data is stored in high-cost storage with fast access, while inactive data is stored in low-cost storage with slower access. This segmentation resolves the contradiction by enabling fast recovery of critical active data while reducing overall storage costs through economical storage of inactive data.
Solution Approach 2:
The patent extracts inactive data from the primary high-cost storage system and stores it separately in low-cost storage. This extraction allows the high-cost storage to be dedicated to active data that requires fast recovery, thereby reducing recovery time for critical data while lowering overall storage costs by moving inactive data to economical storage.
2Reliability
If all unstructured data is restored from cloud-based backup solutions, then complete data recovery is achieved, but recovery time extends to days due to large data volumes
Solution Approach 1:
The patent performs preliminary actions by pre-separating and pre-storing active and inactive data in different storage locations before a disaster occurs. When recovery is needed, active data can be immediately restored from low-cost storage without waiting for the entire backup to be transferred, achieving complete data recovery while dramatically reducing recovery time.
3Reliability
If high-cost storage and backup solutions are used for all data, then data protection is ensured, but storage costs increase
Solution Approach 1:
The patent applies local quality by assigning different storage qualities to different data portions based on their characteristics. Active data receives high-quality fast storage, while inactive data is stored in low-quality economical storage. This resolves the contradiction by ensuring data protection for active data while reducing overall storage costs through differentiated storage quality.
Solution Approach 2:
The patent changes the storage parameter (cost/quality) based on the data state (active/inactive). By dynamically adjusting storage parameters according to data activity levels, the system ensures adequate protection for active data while minimizing costs for inactive data, resolving the contradiction between data protection and storage cost.
4Ease of operation
If inactive data is stored alongside active data in the same storage system, then simplified management is achieved, but storage efficiency decreases
Solution Approach 1:
The patent segments the storage system into separate storage locations for active and inactive data. This segmentation improves storage efficiency by allowing each storage location to be optimized for its specific data type, while maintaining ease of operation through automated management that handles the segmentation transparently.
Data Source
AI summary
A data management system includes: a transceiver; a memory; and a processor communicatively coupled to the transceiver and the memory and configured to: receive, via the transceiver, a copy data request for unstructured data; access, via the transceiver in response to the copy data request, a plurality of backed-up files of unstructured data stored in a first data storage device; send, in response to the copy data request, a plurality of Virtual Data Files (VDFs) to a second data storage device, the processor being configured to respond to receipt of information from each of the plurality of VDFs to retrieve a respective backed-up file of unstructured data of the plurality of backed-up files of unstructured data stored in the first data storage device.


