Backup Data Indexing for eDiscovery Compliance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in quickly and effectively accessing and producing responsive data for eDiscovery from large backup data sets while ensuring compliance with production rules, particularly in managing privileged, confidential, and third-party data.

Innovation Solution

A method and system that process backup data sets to extract metadata, identify and index eDiscovery-relevant data items, and manage their retention and access according to eDiscovery policies, ensuring chain of custody and confidentiality, using a computing system with electronic storage and processors to facilitate eDiscovery data storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If backup data sets are stored and retained for long periods to ensure data availability for eDiscovery, then data completeness and reliability are improved, but storage costs and data management complexity increase

Engineering Contradiction:
Improvedata availability for eDiscoveryVSAvoiddata management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments backup data into two distinct categories: eDiscovery data and non-eDiscovery data. This segmentation is achieved through metadata tagging that identifies data items relevant to eDiscovery requests. By separating these data types, the system can apply different retention and management policies to each segment, reducing overall management complexity while maintaining data availability for eDiscovery purposes.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by pre-tagging and indexing backup data with metadata that identifies eDiscovery relevance before actual eDiscovery requests are made. This advance preparation includes categorizing data by potential eDiscovery criteria (e.g., date ranges, custodians, document types), so that when eDiscovery requests occur, the system can quickly retrieve relevant data without having to process entire backup sets, thereby reducing management complexity during active eDiscovery operations.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If entire backup data sets are processed and reviewed to identify eDiscovery responsive data, then completeness of data identification is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvecompleteness of data identificationVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary processing of backup data by extracting and storing metadata that describes each data item's characteristics (e.g., creation date, author, document type, custodian). This metadata is indexed and made searchable before eDiscovery requests are received. When eDiscovery requests arrive, the system queries this pre-processed metadata to quickly identify responsive data, avoiding the need to scan and analyze entire backup data sets at the time of requests, thus reducing processing time while maintaining identification completeness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces metadata as an intermediary layer between the raw backup data and the eDiscovery search process. Instead of directly searching through large volumes of backup data, the system searches through the compressed, structured metadata that represents the backup data. This intermediary layer acts as an efficient index that enables rapid identification of eDiscovery responsive data without requiring complete processing of the underlying backup sets.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If all backup data is retained and accessed for eDiscovery requests, then data completeness is improved, but access control and confidentiality protection become more difficult

Engineering Contradiction:
Improvedata completenessVSAvoidaccess control
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system segments backup data into eDiscovery responsive data and non-responsive data based on metadata analysis. This segmentation enables differentiated access control where eDiscovery responsive data can be accessed and produced according to eDiscovery rules, while non-responsive data (including privileged and confidential information) remains protected and inaccessible. The metadata tags that identify eDiscovery relevance serve as the basis for this segmentation and subsequent access control policies.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses metadata as an intermediary mechanism to enforce access control without requiring direct examination of confidential backup data. The metadata layer contains information about data characteristics and eDiscovery relevance, allowing the system to make access decisions based on this intermediate information rather than exposing or processing the actual confidential data. This intermediary layer protects confidentiality while enabling appropriate access to responsive data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10339099B2Techniques for leveraging a backup data set for eDiscovery data storage
Publication Date: 2019.07.02 COBALT IRON
  • US10339099B2 patent drawing
  • US10339099B2 patent drawing
  • US10339099B2 patent drawing

AI summary

Techniques for facilitating electronic discovery (eDiscovery) data storage in a backup environment are disclosed. In one particular embodiment, the technique(s) may be realized as a method of operating a computing system to facilitate electronic discovery (eDiscovery) data storage in a backup environment. The method may comprise storing, using electronic storage, a backup data set associated with an organization, processing, using at least one computer processor, the backup data set to extract metadata associated with data items in the backup data set, processing the metadata to identify a subset of the data items that are associated with eDiscovery, and generating an index of the metadata that identifies the subset of the data items in the electronic storage that are associated with the eDiscovery.