Content-Based Dataset Classification for Scalable Data Sensitivity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data management systems struggle to efficiently classify large-scale enterprise data based on its sensitivity, as they rely on physical or directory location rather than content, leading to time-intensive and inefficient processing of sensitive data.

Innovation Solution

A data sensitivity classification method that classifies data assets based on their content using content-based datasets, allowing simultaneous classification of multiple data objects, and applies protection policies based on data type rather than location.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If data sensitivity classification is performed on a per data asset basis or per location basis by investigating filesystems to their fullest depths, then data sensitivity can be accurately identified, but the process becomes very time and process intensive

Engineering Contradiction:
Improvedata sensitivity classification accuracyVSAvoidclassification time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts the classification task from individual data assets and elevates it to the dataset level. By classifying datasets as sensitive or non-sensitive at a higher abstraction level, the system avoids time-intensive investigation of every individual data asset while still achieving accurate sensitivity identification through content-based dataset classification

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary classification at the dataset level before individual data assets need to be processed. By pre-classifying datasets based on their content, the system eliminates the need for repeated time-consuming sensitivity investigations of individual assets within those datasets, significantly reducing overall classification time

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If data classification is performed on individual data objects one at a time, then each object can be accurately classified, but the process cannot efficiently handle large volumes of data

Engineering Contradiction:
Improveclassification accuracyVSAvoidclassification throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent merges multiple individual data object classifications into a single dataset-level classification operation. By combining the classification of numerous data objects within a dataset and applying the classification result to the entire dataset, the system achieves both accurate classification and high throughput, efficiently handling large volumes of data without processing each object individually

Inventive Principle:
Principle #5Merging (Combining)

3Manufacturing precision

If lifecycle rules are made very data specific requiring knowledge of data location and creator, then data management can be precise, but the operating model cannot scale to handle increasing data volumes

Engineering Contradiction:
Improvedata management precisionVSAvoidoperating model scalability
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces a new dimension of classification by moving from location-based and creator-based rules to content-based dataset classification. This dimensional shift allows the system to maintain precise data management through content understanding while achieving scalability by operating at the dataset level rather than requiring detailed knowledge of individual data locations and creators

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12423457B2Data sensitivity classification using content-based datasets
Publication Date: 2025.09.23 DELL PROD LP
  • US12423457B2 patent drawing
  • US12423457B2 patent drawing
  • US12423457B2 patent drawing

AI summary

Providing content-based sensitivity classification to content data in a data processing system, by defining sensitivity designations for data objects stored in the system, and creating datasets by grouping metadata for data objects that are intended to classified with a same sensitivity designation, wherein each dataset spans multiple storage devices of different storage types, and wherein each dataset defines a single data sensitivity unit for the data objects referenced by a respective dataset. A sensitivity classifier component tags each dataset with a sensitivity tag to specify a protection or control operation on the data objects referenced by the respective dataset, and a processing component then processes data objects for each tagged dataset in accordance with the specified protection or control operation.