Predictive Data Structuring Using ML Classification Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data warehouse environments require manual intervention for structuring new data, which is time-consuming and lacks automation in predicting technical design component structuring, especially when dealing with complex data formats and security policies across multiple platforms.

Innovation Solution

A system that uses metadata and regulatory standards to predictively structure electronic data by forming feature column groups, identifying sensitive data, and outputting the structure in various formats, employing classification models like K-means and Naïve Bayes algorithms to match existing data architectures and comply with regulatory requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual intervention is used for data structuring, then data architecture consistency is maintained, but time consumption and labor effort increase significantly

Engineering Contradiction:
Improvedata architecture consistencyVSAvoidtime consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system enables automated self-service data structuring by training a machine learning model on existing data tables and metadata. The model automatically predicts and generates structured data outputs for new inputs without requiring manual architect intervention, thus maintaining consistency while reducing time consumption.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary actions by pre-training the machine learning model using historical data tables, metadata, and regulatory standards. This pre-processing creates a ready-to-use prediction model that can quickly structure new data without requiring real-time manual analysis, thereby reducing time loss while maintaining architectural consistency.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If automated prediction is implemented, then time efficiency improves, but system complexity increases due to machine learning model requirements

Engineering Contradiction:
Improvetime efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The machine learning model serves multiple functions: it predicts data column structures, identifies sensitive data, determines data types, and ensures regulatory compliance. This multi-functionality consolidates what would otherwise require multiple separate systems into a single automated prediction engine, improving productivity while managing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces metadata as an intermediary layer between raw data and the prediction model. By analyzing metadata patterns from existing data tables, the model learns structural relationships without directly processing complex raw data, thereby improving prediction efficiency while reducing the computational complexity of the core model.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If sensitive data identification is automated, then security policy compliance is ensured, but false positive rates may increase

Engineering Contradiction:
Improvesecurity policy complianceVSAvoididentification accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The system incorporates feedback mechanisms where the prediction model continuously learns from regulatory standards and security policies. By training on labeled data that includes sensitive data patterns and compliance requirements, the model refines its identification accuracy over time, ensuring security compliance while reducing false positives through iterative improvement.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system dynamically adjusts identification parameters based on the specific data context and regulatory requirements. By changing sensitivity thresholds and prediction parameters according to the data type and regulatory domain, the system maintains high compliance reliability while minimizing false positives through adaptive parameter tuning.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11892989B2System and method for predictive structuring of electronic data
Publication Date: 2024.02.06 BANK OF AMERICA CORP
  • US11892989B2 patent drawing
  • US11892989B2 patent drawing
  • US11892989B2 patent drawing

AI summary

Embodiments of the invention are directed to a system, method, or computer program product for an approach to predictive structuring of electronic data. The system uses data mining and data clustering techniques using classification models to organize feature column groups comprising feature columns. The system identifies and flags feature column groups and/or feature columns based on regulatory data standards provided by regulatory bodies. Thereafter, data objects are imported into the system and prediction algorithms are implemented to characterize the feature columns containing the data objects.