VM Virtual Disk Block-Level Backup Indexing for Granular Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional approaches to VM backups do not offer the desired granularity and are hampered by slow backup and recovery procedures, especially in cloud and non-cloud data centers, lacking efficient methods for granular searching and browsing of virtual machine data.

Innovation Solution

A streamlined approach that analyzes block-level backup copies of VM virtual disks to create coarse and fine indexes, enabling granular searching and browsing without restoring the VM data files, using a virtual machine content indexer to extract filenames, metadata, and content without staging.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional VM backup approaches are used, then backup simplicity is maintained, but granular searching and browsing capability is lost

Engineering Contradiction:
Improvegranular searching and browsing capabilityVSAvoidindexing system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the backup data into block-level units and creates separate indexes (coarse index for filenames/metadata, fine index for content) that can be independently queried. This segmentation enables granular searching without requiring full restore operations, resolving the contradiction between search capability and system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces indexing structures as intermediary layers between the backup data and user queries. The coarse and fine indexes act as mediators that translate user search requests into efficient data retrieval operations, enabling granular searching without direct interaction with the complex backup storage system.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If full restore operations are performed for browsing, then complete data access is achieved, but time and resources are consumed

Engineering Contradiction:
Improvesearch and browse speedVSAvoidrestore time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary indexing of backup data block-level, creating coarse and fine indexes before actual search operations. This preliminary action enables fast search and browse operations without requiring time-consuming full restore operations, directly resolving the contradiction between productivity and time loss.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts only the necessary information (filenames, metadata, content snippets) from the backup data and stores it in separate indexes. This extraction allows users to search and browse using only the extracted index data rather than requiring full restore of the entire backup, significantly reducing time and resource consumption.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If block-level indexing is performed, then granular access is enabled, but backup performance may be impacted

Engineering Contradiction:
Improvedata granularityVSAvoidbackup speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent performs indexing as a preliminary action during the backup process itself, rather than as a separate post-processing step. By indexing block-level data as it is being backed up, the system enables granular access while minimizing impact on backup performance, resolving the contradiction between measurement precision and productivity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements continuous indexing during the backup operation, maintaining the useful action of data processing throughout the backup process. This continuous approach ensures that granular indexing is performed without interrupting or significantly delaying the backup operation, preserving both data granularity and backup speed.

Inventive Principle:
Principle #20Continuity of useful action

4Ease of operation

If indexing is performed online, then data accessibility is improved, but source VM performance is impacted

Engineering Contradiction:
Improvedata accessibilityVSAvoidVM performance
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent performs indexing as a preliminary action during the backup process, before the backup data needs to be accessed. This timing allows indexing to be completed in advance without impacting source VM performance during normal operations, while still enabling improved data accessibility when needed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the indexing operation from the source VM operations by performing indexing on backup copies rather than on the live VM data. This segmentation allows indexing to occur independently without impacting source VM performance, while still providing improved data accessibility through the indexed backup data.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12367107B2Content indexing of files in virtual disk block-level backup copies
Publication Date: 2025.07.22 COMMVAULT SYSTEMS INC
  • US12367107B2 patent drawing
  • US12367107B2 patent drawing
  • US12367107B2 patent drawing

AI summary

A streamlined approach analyzes block-level backups of VM virtual disks and creates both coarse and fine indexes of backed up VM data files in the block-level backups. The indexes (collectively the “content index”) enable granular searching by filename, by file attributes (metadata), and/or by file contents, and further enable granular live browsing of backed up VM files. Thus, by using the illustrative data storage management system, ordinary block-level backups of virtual disks are “opened to view” through indexing. Any block-level copies can be indexed according to the illustrative embodiments, including file system block-level copies. The indexing occurs offline in an illustrative data storage management system, after VM virtual disks are backed up into block-level backup copies, and therefore the indexing does not cut into the source VM's performance. The disclosed approach is widely applicable to VMs executing in cloud computing environments and/or in non-cloud data centers. The illustrative content indexing is accomplished without restoring the VM data files being indexed to a staging location.