Filtered Reference Copy for Secondary Storage Data Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Accessing and managing backup data across multiple storage devices is time-consuming and resource-intensive due to slower network connections and storage media, especially when dealing with large volumes of data and complex retention policies.
Innovation Solution
A data storage system that creates a filtered, global reference copy of secondary data by using media agents to identify and filter data based on user-defined criteria, allowing for efficient retention policies and access, while maintaining actual or reference copies of files in their native format.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If backup data is stored across multiple secondary storage devices, then data protection and retention policies are improved, but access time and resource consumption increase due to slower network connections and storage media
Solution Approach 1:
The patent segments backup data across multiple secondary storage devices while creating a filtered reference copy that organizes and indexes this distributed data. The reference copy acts as a centralized access point that virtualizes the segmented storage, allowing users to access backup data without directly interacting with the distributed storage infrastructure, thus maintaining data protection while reducing access time.
Solution Approach 2:
The patent introduces a filtered reference copy as an intermediary layer between users and the distributed secondary storage devices. This reference copy contains filtered and indexed information about the backup data, serving as a mediator that enables efficient access to distributed data without requiring direct access to multiple slow storage devices, thereby resolving the contradiction between data protection and access speed.
2Reliability
If backup data is stored across multiple secondary storage devices, then data protection is improved, but resource consumption increases due to slower network connections and storage media
Solution Approach 1:
The patent segments the workload of managing backup data across multiple storage devices by creating a filtered reference copy that consolidates access operations. Instead of consuming resources to access each distributed storage device directly, the system consumes minimal resources to query the reference copy, significantly reducing network and processing resource consumption while maintaining data protection across multiple devices.
Solution Approach 2:
The filtered reference copy serves as an intermediary that reduces resource consumption by eliminating the need to directly access multiple slow storage devices for routine operations. The reference copy contains pre-filtered and indexed information that can be queried efficiently, reducing network traffic and processing resources while the underlying distributed storage continues to provide data protection.
3Productivity
If a filtered reference copy is created from secondary storage data, then data access efficiency is improved, but system complexity increases due to additional filtering and indexing mechanisms
Solution Approach 1:
The patent creates a filtered reference copy that is a simplified representation of the actual backup data. This copy contains only the essential filtered and indexed information needed for efficient access, rather than duplicating the entire complex backup infrastructure. The reference copy can be stored in memory or on fast storage, providing efficient access while the complex distributed storage system remains in the background, thus improving productivity with minimal additional complexity.
4Ease of operation
If media agents are used to identify and filter data based on user-defined criteria, then ease of operation is improved, but device complexity increases due to distributed filtering workload
Solution Approach 1:
The patent implements media agents that autonomously perform data identification and filtering based on user-defined criteria stored in the reference copy. The media agents self-organize the filtering workload across the distributed system, automatically querying the reference copy and retrieving relevant data without requiring complex centralized coordination. This self-service approach improves ease of operation while distributing complexity across multiple agents rather than concentrating it in a single complex component.
Data Source
Figure 1A
Figure 1B
Figure 1C
AI summary
The method creates a user browsable filtered representation of secondary copy data in a networked data storage system. The method includes creating a reference copy based on files that meet a filtering criteria and storing the reference copy in a reference copy data store. The reference copy includes a copy of each of the files in a native format associated with a corresponding application that generated the file. The reference copy data store includes an index, the index comprising information regarding the files in the reference copy. The method comprises creating a user browsable filtered representation to be provided via a user interface, the user browsable filtered representation including a listing of the files in the reference copy.