Content-Based Dataset Management for Legal Hold Compliance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data management systems struggle to efficiently handle the lifecycle management of large-scale datasets across multiple locations and environments, particularly in managing access and compliance for sensitive data, as they rely on location-based controls which are insufficient for the complexity and volume of data generated from diverse sources.
Innovation Solution
Implementing a dataset management system that uses metadata to create logical datasets spanning multiple storage devices and environments, allowing for content-based data protection policies that automatically track and manage data regardless of location, and applying Role-Based Access Control (RBAC) and Access Control List (ACL) rules based on dataset metadata for secure access.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If location-based access controls are used to manage data security, then access control implementation is simple, but the system cannot effectively manage data across multiple locations and environments
Solution Approach 1:
The patent introduces a dataset as an intermediary layer between the data storage system and the access control system. This dataset contains metadata about multiple data locations and acts as a mediator that translates content-based access rules into location-specific access controls, enabling unified management without increasing system complexity
Solution Approach 2:
The patent transitions from managing data access based on physical location (single dimension) to managing access based on data content and metadata (adding a new dimension). This allows the system to control access to data regardless of its physical location, enabling versatile data management across multiple environments
2Quantity of substance
If hierarchical directory structures are used to organize data access, then implementation is straightforward, but the system becomes insufficient for managing huge amounts of files from diverse sources
Solution Approach 1:
The patent segments the data management system into distinct components: a data catalog that stores metadata about data locations, and datasets that group data by content rather than physical location. This segmentation allows the system to handle huge volumes of files from diverse sources while maintaining ease of operation through logical organization
3Productivity
If manual data lifecycle management is used, then policy creation is simple, but the system cannot scale to handle increasing data volumes
Solution Approach 1:
The patent implements self-service automation where the system automatically discovers data locations through data catalogs, identifies data that matches dataset criteria, and applies lifecycle policies without manual intervention. This automation enables the system to scale to handle increasing data volumes while maintaining simple policy creation through content-based rules
Data Source
AI summary
Enforcing a legal hold procedure in a system by scanning multiple data sources to identify data objects for processing as a unitary group with respect to common access and control processes of the legal hold to preserve the data for a defined period of time and protected against modification and unauthorized access. The metadata is stored in a static dataset that defines a single data access unit for the referenced data. A user query regarding a referenced data object is processed, and accesses the data through the dataset as a single unit based on data content rather than data location in a file directory of the system. The data may be sensitive data and the legal hold procedure may be implemented as court rules in accordance with Federal Rules of Civil Procedure.


