Live Browse Cache for Backup Copy Indexing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data storage systems experience significant performance degradation during live browse and file indexing operations when using cloud storage and tape media, leading to slow data retrieval and access, making these features impractical for many applications.
Innovation Solution
Implementing a live browse cache that pre-fetched and stores key data blocks, allowing for faster retrieval during subsequent operations, and using a cacheable flag system to identify and prioritize data blocks for caching, especially for block-level backup copies of virtual machines and file systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is stored on cloud storage or tape media for backup, then storage capacity and cost-effectiveness are improved, but data retrieval speed and access performance deteriorate
Solution Approach 1:
The system performs preliminary actions by pre-fetching and caching data blocks from cloud storage or tape media into local storage before they are actually needed for live browse or file indexing operations. This advance preparation eliminates the need for slow on-demand retrieval during critical operations, directly resolving the speed bottleneck while maintaining the use of cost-effective cloud/tape storage for the master backup copies.
Solution Approach 2:
The patent introduces an intermediary component (cache storage system) between the cloud storage/tape media and the live browse/file indexing operations. This intermediary pre-loads data into a faster access medium, acting as a buffer that decouples the slow storage medium from the fast access requirement, thereby resolving the contradiction between using slow but cost-effective storage and maintaining fast access performance.
2Ease of operation
If live browse and file indexing operations are performed on backup copies stored on slow media, then data accessibility is improved, but operation speed and productivity deteriorate
Solution Approach 1:
The system performs preliminary actions by pre-fetching data blocks that are likely to be accessed during live browse or file indexing operations and storing them in a cache. This advance preparation ensures that when users initiate these operations, the data is already available in fast storage, maintaining ease of operation while dramatically improving operation speed and productivity.
Solution Approach 2:
The system implements self-service by automatically detecting which data blocks will be needed for upcoming operations and pre-loading them into cache without requiring manual intervention. The system serves itself by anticipating data access patterns and proactively preparing the necessary data, thereby improving both accessibility and speed without additional human effort.
3Speed
If data blocks are cached for faster retrieval, then access speed is improved, but storage complexity and memory usage increase
Solution Approach 1:
The system employs feedback mechanisms to monitor cache performance, hit rates, and data access patterns. Based on this feedback, the cache management system dynamically adjusts which data blocks to retain or evict, optimizing cache utilization. This feedback-driven approach automates cache management decisions, reducing complexity while maintaining high access speeds through intelligent, adaptive cache policies.
Solution Approach 2:
The patent applies parameter changes by dynamically adjusting cache size, eviction policies, and pre-fetch thresholds based on system conditions, workload characteristics, and performance metrics. These parameter adjustments allow the system to optimize between access speed and storage complexity, adapting to different operational scenarios without requiring complex manual configuration or management.
Data Source
AI summary
An illustrative approach accelerates file indexing operations for block-level backup copies in a data storage management system. A cache storage area is maintained for locally storing and serving key data blocks, thus relying less on retrieving data on demand from the backup copy. File indexing operations are used for populating the cache storage area for speedier retrieval during subsequent live browsing of the same backup copy, and vice versa. The key data blocks cached while file indexing and/or live browsing an earlier backup copy help to pre-fetch corresponding data blocks of later backup copies, thus producing a beneficial learning cycle. The approach is especially beneficial for cloud and tape backup media, and is available for a variety of data sources and backup copies, including block-level backup copies of virtual machines (VMs) and block-level backup copies of file systems, including UNIX-based and Windows-based operating systems and corresponding file systems.


