Extensible Data Objects for Machine Learning Explainability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning systems, particularly in financial and banking institutions, face challenges due to their 'black box' nature, making it difficult to interpret and explain their decision-making processes, and the complexity of datasets used, which can lead to hesitancy in relying on their accuracy and compliance with regulatory requirements.

Innovation Solution

The development of extensible data objects that represent unstructured data elements, augmented with structured insight features through machine learning-based analyses, allowing for the validation, visualization, and auditability of data inputs into machine learning models, facilitating feature engineering and improving data preprocessing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If machine learning models are used for decision-making, then productivity and accuracy are improved, but explainability and interpretability deteriorate due to black box nature

Engineering Contradiction:
Improvedecision-making efficiencyVSAvoiddecision-making transparency
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent introduces an intermediary layer between the black box machine learning model and the user. This intermediary takes the model's output and transforms it into human-explainable formats, such as natural language explanations, visualizations, and structured data representations. The intermediary acts as a bridge that preserves the model's accuracy while making its decision-making process transparent and interpretable.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the machine learning system into distinct components: the black box model, the intermediary explanation layer, and the data input/output interfaces. By segmenting the system, the complex model logic is isolated from the user interface, allowing the model to maintain its computational power while the segmented explanation layer provides transparency without compromising the model's productivity.

Inventive Principle:
Principle #1Segmentation

2Productivity

If complex datasets are used to improve model accuracy, then productivity is improved, but device complexity and difficulty of detection increase

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata structure complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent extracts and separates critical metadata and structured information from the complex unstructured datasets. By taking out these essential elements and representing them as simplified data objects with clear schemas, the system maintains the richness needed for high accuracy while presenting a simplified, easier-to-process interface that reduces perceived and actual system complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the representation parameters of complex data by transforming them into standardized data object formats with defined schemas. This parameter transformation allows the system to handle complex datasets internally while presenting simplified structures externally, thereby maintaining accuracy without proportionally increasing complexity in the user-facing interface.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If unstructured data is processed directly, then data volume is maintained, but manufacturing precision and measurement precision deteriorate

Engineering Contradiction:
Improvedata volumeVSAvoiddata structure standardization
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent applies preliminary action by pre-processing and structuring unstructured data into standardized data objects before they are used in machine learning models. This preliminary structuring step maintains the comprehensive data volume while improving precision through consistent formatting, typing, and schema enforcement, thereby preparing the data for more accurate and reliable processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250094441A1Extensible data objects for use in machine learning models
Publication Date: 2025.03.20 DEEPSEE AI INC
  • US20250094441A1 patent drawing
  • US20250094441A1 patent drawing
  • US20250094441A1 patent drawing

AI summary

Systems and methods are described herein for creating a data object for each of a plurality of imported unstructured data files. Each data object may expressly include one of the unstructured data files. Preprocessing subsystems and/or machine learning algorithms and subsystems process the data to generate or otherwise identify structured insight features. The system updates each data object to expressly include the structured insight features.