ML-Based Digital Content Classification for Sensitive Data Routing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional data storage systems face challenges in identifying and managing sensitive information, especially in decentralized and complex storage architectures, leading to difficulties in data security and compliance due to misclassified digital items and unstructured data.

Innovation Solution

A machine learning-based method for classifying digital content, which involves accessing a digital content corpus, computing in-migration content classifications, and migrating items to designated storage locations based on sensitivity inferences, using a system comprising a data handling service, access and discovery subsystem, feature identification and classification subsystem, and content route handling subsystem.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Difficulty of detecting and measuring

If traditional on-premises data storage and nonintegrated storage architectures are used, then data storage flexibility is maintained, but identifying and managing sensitive information becomes difficult

Engineering Contradiction:
Improvedifficulty in identifying sensitive informationVSAvoiddecentralized storage architecture complexity
Core Design Contradiction:
Difficulty of detecting and measuringVSDevice complexity

Solution Approach 1:

The patent introduces a centralized data governance platform as an intermediary layer between decentralized storage systems and users. This platform provides unified access points, centralized classification services, and coordinated security management, thereby reducing the difficulty of detecting and managing sensitive information across complex decentralized architectures without requiring changes to the underlying storage systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If centralized cloud-based storage architectures are implemented, then data security and compliance are improved, but migration and classification of digital content becomes complex

Engineering Contradiction:
Improvedata security and complianceVSAvoidmigration and classification process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements preliminary classification and risk assessment actions before data migration to the cloud. The system analyzes digital content characteristics, identifies sensitive information, determines appropriate storage locations and security measures in advance, and prepares migration configurations. This staged approach simplifies the overall migration process while maintaining high security standards.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes key parameters such as data classification labels, security levels, and storage location assignments during migration. By transforming data attributes and reconfiguring metadata parameters, the system adapts decentralized storage data to centralized cloud architectures, reducing migration complexity while enhancing security and compliance.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If machine learning-based classification is applied to digital content, then classification accuracy is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveclassification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the classification process into multiple stages: preliminary metadata filtering, intermediate content analysis, and final verification. By dividing the workload and applying different classification algorithms at different stages, the system achieves high accuracy for sensitive content identification while reducing overall processing time through parallel processing and prioritization.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12259847B2Systems and methods for intelligent digital item discovery and machine learning-informed handling of digital items and digital item governance
Publication Date: 2025.03.25 DRYVIQ INC
  • US12259847B2 patent drawing
  • US12259847B2 patent drawing
  • US12259847B2 patent drawing

AI summary

Systems and methods of computing classifications for and migrating digital content that includes accessing a digital content corpus within a source data storage system; in response to accessing the digital content corpus, for each distinct item of digital content of the plurality of distinct items of digital content: computing, via one or more digital content machine learning classification models, a content classification inference; identifying automated digital content handling tasks of a plurality of distinct digital content handling tasks based on the content classification inference; executing the automated content handling tasks identified for each distinct item of digital content, wherein executing the automated content handling tasks includes: designating a storage location within a target data storage system based on the in-migration content classification inference; and migrating a respective item of digital content from the source data storage system to the designated storage location within the target data storage system.