Automated Data Classification via Machine Learning Pattern Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data management systems require intense manual effort for cataloging and maintaining authorized data classification rules and element tagging, leading to errors and potential sanctions and fines, especially in legal holds and eDiscovery processes.

Innovation Solution

Implementing machine learning systems that analyze data streams to determine patterns, rule sets, and entities, tagging relevant data objects, and using database graphs to enforce legal holds and archival requirements, thereby automating the data management process and reducing manual errors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual cataloging and tagging of data classification rules is performed, then data management compliance is maintained, but intense manual effort and human errors occur leading to potential sanctions and fines

Engineering Contradiction:
Improvedata management complianceVSAvoidmanual effort
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system enables self-service automated data classification and tagging through machine learning models that automatically analyze data streams, identify patterns, and apply appropriate classification rules without requiring manual intervention, thereby maintaining compliance while eliminating intensive manual effort

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process of cataloging and tagging data with an automated computational system using machine learning algorithms that process data streams, detect patterns, and automatically apply classification rules, substituting human manual operations with automated mechanical systems

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated machine learning systems are implemented to analyze data streams and tag data objects, then manual effort is reduced and errors are minimized, but system complexity increases

Engineering Contradiction:
Improvedata management efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex data management task into distinct modular components: data stream reception, pattern detection, rule set determination, entity identification, and automated tagging. Each component is implemented as a separate functional module that can be independently developed, maintained, and optimized, reducing overall system complexity while maintaining high productivity

Inventive Principle:
Principle #1Segmentation

3Reliability

If database graphs are used to track data relationships and enforce legal holds, then data integrity and compliance are enhanced, but query and search complexity increases

Engineering Contradiction:
Improvedata integrityVSAvoidquery complexity
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

The system creates simplified copies or projections of the complex database graph structure that can be queried efficiently. Instead of querying the entire graph structure directly, the system uses indexed representations and pre-computed relationship paths that maintain data integrity while enabling fast searches and queries through simplified access interfaces

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10423590B2System and method of managing data in a distributed computing environment
Publication Date: 2019.09.24 BANK OF AMERICA CORP
  • US10423590B2 patent drawing
  • US10423590B2 patent drawing
  • US10423590B2 patent drawing

AI summary

In one or more embodiments, one or more systems, processes, and/or methods may receive a first data stream and determine a pattern from the first data stream. At least one rule set based at least on the pattern may be determined. A second data stream, different from the first data stream may be received and entities may be determined, where each of the entities may be associated with respective data of the second data stream that satisfies the at least one rule set. At least one data object of the second data stream may be tagged, in response to determining the entities. In one or more embodiments, tagging the at least one data object may associate the at least one data object with at least one of the entities.