Virtual Disk Block-Level Indexing for Granular VM File Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional approaches to virtual machine (VM) backups lack granularity and are hampered by slow backup and recovery procedures, failing to provide efficient data management and indexing of VM-generated data in cloud and non-cloud environments.

Innovation Solution

A data storage management system that analyzes block-level backup copies of VM virtual disks to create coarse and fine indexes, enabling granular searching and browsing of backed up VM files without restoring them to a staging location, using a virtual machine content indexer to extract filenames, metadata, and content for indexing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional VM backup approaches are used, then backup simplicity is maintained, but data granularity and search capability are insufficient

Engineering Contradiction:
Improvedata granularityVSAvoidbackup system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments backup data into block-level units and creates multiple index types (coarse index for filenames/metadata, fine index for content) to achieve granular search capability without requiring complete file restoration. This segmentation allows precise data location and retrieval while maintaining system manageability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces indexing structures as intermediary components between backup data and user queries. The coarse and fine indexes act as mediators that enable granular searching and browsing without directly exposing the complex block-level backup structure, thus improving data accessibility while hiding system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If full file restoration is performed for indexing, then content search capability is improved, but storage resources and time are consumed

Engineering Contradiction:
Improvecontent search capabilityVSAvoidstorage resources
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the necessary indexing information (filenames, metadata, and selective content samples) from backup files without restoring complete files. The fine index extracts content signatures or samples to enable search capability while storing minimal data, thus achieving content search without consuming full storage resources.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial indexing by creating coarse and fine indexes that contain selective information about backup files rather than complete file contents. This partial action provides sufficient search capability while significantly reducing storage requirements compared to full restoration approaches.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If indexing is performed online during VM operation, then search capability is improved, but source VM performance is impacted

Engineering Contradiction:
Improvesearch capabilityVSAvoidsource VM performance
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent performs indexing as a preliminary offline action after backup completion but before search operations. The coarse and fine indexes are built in advance using backup data, enabling fast search capability without impacting source VM performance during operation. This preliminary indexing separates the indexing workload from production environments.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent maintains continuous backup operations while performing indexing in between backup cycles, ensuring that source VM operations continue uninterrupted. The indexing process utilizes backup windows or offline periods, maintaining the continuity of useful VM operations while still providing ongoing search capability through updated indexes.

Inventive Principle:
Principle #20Continuity of useful action

4Speed

If block-level backup copying is used, then backup speed is improved, but data accessibility and browsing capability are reduced

Engineering Contradiction:
Improvebackup speedVSAvoiddata accessibility
Core Design Contradiction:
SpeedVSEase of operation

Solution Approach 1:

The patent performs preliminary indexing of block-level backup data to create searchable structures before data retrieval operations. By pre-building coarse and fine indexes during backup windows or offline periods, the system maintains fast block-level backup speeds while enabling easy data accessibility through rapid index-based searching without requiring full file restoration.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250307079A1Content indexing of files in virtual disk block-level backup copies
Publication Date: 2025.10.02 COMMVAULT SYSTEMS INC
  • US20250307079A1 patent drawing
  • US20250307079A1 patent drawing
  • US20250307079A1 patent drawing

AI summary

A streamlined approach analyzes block-level backups of VM virtual disks and creates both coarse and fine indexes of backed up VM data files in the block-level backups. The indexes (collectively the “content index”) enable granular searching by filename, by file attributes (metadata), and/or by file contents, and further enable granular live browsing of backed up VM files. Thus, by using the illustrative data storage management system, ordinary block-level backups of virtual disks are “opened to view” through indexing. Any block-level copies can be indexed according to the illustrative embodiments, including file system block-level copies. The indexing occurs offline in an illustrative data storage management system, after VM virtual disks are backed up into block-level backup copies, and therefore the indexing does not cut into the source VM's performance. The disclosed approach is widely applicable to VMs executing in cloud computing environments and/or in non-cloud data centers. The illustrative content indexing is accomplished without restoring the VM data files being indexed to a staging location.