Data Catalog for Content-Based Data Protection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data management systems struggle to scale with the exponential growth of large and dynamic datasets, leading to challenges in managing data lifecycles efficiently.

Innovation Solution

Implementing a dataset lifecycle management system that uses metadata to create logical datasets across multiple storage devices and environments, allowing for content-based data protection policies that can scale with increasing data volumes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional manual data management approaches are used, then data protection policies can be customized for specific data, but the system cannot scale to handle increasing data volumes

Engineering Contradiction:
Improvedata-specific protection policy customizationVSAvoiddata management scalability
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system enables self-service data management by allowing data to be automatically categorized and protected based on its intrinsic properties. Metadata automatically identifies data characteristics (such as type, sensitivity, access patterns) and applies appropriate protection policies without human intervention, enabling the system to scale while maintaining customized protection for each data element

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the fundamental parameter of data management from manual policy assignment to automated policy application based on data parameters. By analyzing metadata parameters (data type, sensitivity level, access frequency, ownership), the system dynamically determines and applies appropriate protection policies, enabling scalability while preserving data-specific customization

Inventive Principle:
Principle #35Parameter changes

2Productivity

If data management responsibilities are moved to data creators, then management capacity increases, but complexity of managing lifecycle rules across distributed creators increases

Engineering Contradiction:
Improvedata management capacityVSAvoidlifecycle rule management complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system implements a universal data management platform that handles lifecycle rules for all data creators through a single interface. This multi-functional system can manage diverse data types and creators while applying consistent governance frameworks, reducing the complexity burden on individual data creators while maintaining high management capacity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces an intermediary layer between data creators and protection policies. This intermediary automatically translates data creator needs into appropriate protection rules by analyzing metadata and applying predefined policy templates, enabling distributed data creation while simplifying lifecycle rule management through automated mediation

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If lifecycle rules are made data-specific, then data protection accuracy improves, but the time and resources required to create and maintain rules increase

Engineering Contradiction:
Improvedata protection accuracyVSAvoidrule creation and maintenance time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-defining protection policy templates based on common data characteristics and compliance requirements. When new data is created, the system matches it against these pre-configured templates using metadata analysis, automatically applying appropriate protection rules without requiring time-consuming manual rule creation while maintaining high protection accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms that learn from applied protection policies and their outcomes. By analyzing metadata patterns and protection effectiveness across the data portfolio, the system automatically refines and updates protection rules, reducing the time required to create and maintain accurate data-specific protection policies over time

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12321240B2Data catalog for dataset lifecycle management system for content-based data protection
Publication Date: 2025.06.03 DELL PROD LP
  • US12321240B2 patent drawing
  • US12321240B2 patent drawing
  • US12321240B2 patent drawing

AI summary

Providing content based data protection for data stored in a large-scale data storage system by scanning data stored in one or more databases for discovery of metadata, and extracting the discovered metadata, for storage in a data catalog, the data catalog having a scanning function performing the scanning step, and comprising a database storing the metadata in one or more tables. A protection policy is defined to commonly protect content data referenced by metadata in the data catalog, and applied to the referenced content data to perform a data protection operation the content data. Datasets stored in the catalog are generated by running queries on the catalog, where a query comprises metadata selectors as tags applied to the catalog, where the tags define at least one of a file type, name, location, creation time, or file characteristic.