Data Preparation Engine for Secure Compliant Distributed Data Curation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data retrieval techniques lack visibility into retrieved data, especially sensitive information, leading to inefficiencies in data discovery, compliance issues, and security risks, and relying on synthetic datasets compromises the effectiveness and reliability of downstream applications.

Innovation Solution

A data preparation engine that performs real-time data discovery, sanitizes sensitive information, and ensures compliance by integrating role-based access control and continuous monitoring across distributed storage systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data is retrieved from distributed storage systems using current techniques, then data availability is improved, but data security and compliance are worsened due to lack of visibility into sensitive information

Engineering Contradiction:
Improvedata availabilityVSAvoiddata security
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary classification and identification of sensitive data during the data ingestion phase, before the data is used in downstream applications. Data classification labels are assigned and stored with the data, enabling security controls to be applied proactively rather than reactively, thus maintaining both availability and security

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary data classification and labeling layer between the distributed storage systems and downstream applications. This intermediary component analyzes data, assigns sensitivity labels, and enables security policies to be enforced without blocking legitimate data access, thus resolving the contradiction between availability and security

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If data is scattered across multiple distributed storage systems, then data storage capacity is improved, but data discovery efficiency is worsened

Engineering Contradiction:
Improvedata storage capacityVSAvoiddata discovery efficiency
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent implements a universal data classification framework that operates across multiple distributed storage systems simultaneously. The same classification algorithms, sensitivity labels, and discovery mechanisms work consistently across different storage platforms, enabling efficient data discovery without requiring system-specific implementations

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system employs feedback mechanisms where data classification results and usage patterns are continuously analyzed to improve discovery efficiency. The system learns from query patterns and data access behaviors to optimize data location and retrieval, reducing the time required to discover relevant data across distributed systems

Inventive Principle:
Principle #23Feedback

3Productivity

If sensitive data is not properly identified and sanitized, then data processing speed is improved, but compliance with regulatory policies is worsened

Engineering Contradiction:
Improvedata processing speedVSAvoidcompliance risk
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary identification and classification of sensitive data during the ingestion phase, assigning sensitivity labels that enable downstream applications to process data appropriately. This preliminary action allows compliant data processing to proceed at full speed without requiring retroactive security checks, thus maintaining both processing speed and compliance

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies different processing treatments to different portions of data based on their sensitivity classification. Sensitive data receives appropriate sanitization and access controls, while non-sensitive data is processed freely, thus maintaining overall processing speed while ensuring compliance for sensitive portions

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20260080099A1Data Preparation Engine(s) For Curating Secure And Compliant Data Collections From Distributed Sources
Publication Date: 2026.03.19 NETAPP INC
  • US20260080099A1 patent drawing
  • US20260080099A1 patent drawing
  • US20260080099A1 patent drawing

AI summary

Various embodiments of the present technology generally relate to systems and methods for providing a data preparation engine for curating secure and compliant data collections from distributed storage systems. In an aspect, a data preparation engine receives a query from a client device and determines files from one or more distributed sources based on the query. The data preparation engine determines sensitive data within the files and anonymizes the sensitive data while preserving context and integrity of the underlying information. The data preparation engine generates a data collection including the files with anonymized sensitive data. The data collection may then be deployed to downstream applications or workflows, such as used to generate curated data sets for training of artificial intelligence applications. Once deployed, the data preparation engine may continuously monitor the distributed sources for changes to data within the files and automatically update data collections in real-time.