Offline Content Indexing for Historical Data Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional content indexing systems burden computer systems, fail to account for offline or archived data, and are insufficient for locating historical or deleted content, leading to interruptions and manual searches during legal discovery requests.

Innovation Solution

An offline content indexing system that creates an index from secondary copies of data, such as backups, snapshots, or change journals, allowing for the association of metadata and tags, enabling search across multiple copies without impacting the primary system and providing access to historical or deleted content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional content indexing is performed on primary systems, then content search capability is improved, but system availability and performance deteriorate due to resource consumption

Engineering Contradiction:
Improvecontent search capabilityVSAvoidsystem availability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent creates a copy of the primary data to a secondary location, then performs indexing operations on this copy rather than the primary system. This allows comprehensive content indexing to be performed without impacting the availability or performance of the primary production systems, resolving the contradiction between search capability and system availability

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent segments the indexing operation from the primary system by performing it on a separate copy at a secondary location. This separation allows the indexing process to run independently without consuming primary system resources, maintaining system availability while still enabling content search functionality

Inventive Principle:
Principle #1Segmentation

2Reliability

If content indexing is deferred to off hours, then system performance is maintained, but indexing completeness deteriorates due to backup window interruptions

Engineering Contradiction:
Improvesystem performanceVSAvoidindexing completeness
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

By working on a copy of the data rather than the primary system, the patent eliminates the need to coordinate indexing operations with backup schedules. The indexing process can run continuously on the copy without being interrupted by backup windows, achieving complete indexing while maintaining primary system performance

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs data copying to a secondary location in advance, creating a ready-to-index copy. This preliminary action separates the copying operation from the indexing operation, allowing indexing to proceed without being blocked by backup activities and ensuring indexing completeness

Inventive Principle:
Principle #10Preliminary action

3Reliability

If manual searches are performed for historical content, then compliance requirements are met, but time consumption increases significantly

Engineering Contradiction:
Improvecompliance capabilityVSAvoidsearch time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary indexing of all content including historical and archived data, creating a searchable index in advance. When compliance searches are needed, the pre-created index enables rapid retrieval of historical content without manual searching, reducing search time while maintaining compliance capability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

By pre-indexing historical and archived content that would otherwise be inaccessible, the patent prepares search capabilities in advance. This eliminates the need for time-consuming manual retrieval and searching when compliance requests arise, significantly reducing response time while ensuring compliance requirements are met

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7882077B2Method and system for offline indexing of content and classifying stored data
Publication Date: 2011.02.01 COMMVAULT SYSTEMS INC
  • US7882077B2 patent drawing
  • US7882077B2 patent drawing
  • US7882077B2 patent drawing

AI summary

A method and system for creating an index of content without interfering with the source of the content includes an offline content indexing system that creates an index of content from an offline copy of data. The system may associate additional properties or tags with data that are not part of traditional indexing of content, such as the time the content was last available or user attributes associated with the content. Users can search the created index to locate content that is no longer available or based on the associate attributes.