Extensible Data Objects for Machine Learning Explainability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning systems, particularly in financial and banking institutions, face challenges due to their 'black box' nature, making it difficult to interpret and explain their decision-making processes, and the complexity of datasets used, which can lead to hesitancy in relying on their accuracy and compliance with regulatory requirements.
Innovation Solution
The development of extensible data objects that represent unstructured data elements, augmented with structured insight features through machine learning-based analyses, allowing for the validation, visualization, and auditability of data inputs into machine learning models, facilitating feature engineering and improving data preprocessing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If machine learning models are used for decision-making, then productivity and accuracy are improved, but explainability and interpretability deteriorate due to black box nature
Solution Approach 1:
The patent introduces an intermediary layer between the black box machine learning model and the user. This intermediary takes the model's output and transforms it into human-explainable formats, such as natural language explanations, visualizations, and structured data representations. The intermediary acts as a bridge that preserves the model's accuracy while making its decision-making process transparent and interpretable.
Solution Approach 2:
The patent segments the machine learning system into distinct components: the black box model, the intermediary explanation layer, and the data input/output interfaces. By segmenting the system, the complex model logic is isolated from the user interface, allowing the model to maintain its computational power while the segmented explanation layer provides transparency without compromising the model's productivity.
2Productivity
If complex datasets are used to improve model accuracy, then productivity is improved, but device complexity and difficulty of detection increase
Solution Approach 1:
The patent extracts and separates critical metadata and structured information from the complex unstructured datasets. By taking out these essential elements and representing them as simplified data objects with clear schemas, the system maintains the richness needed for high accuracy while presenting a simplified, easier-to-process interface that reduces perceived and actual system complexity.
Solution Approach 2:
The patent changes the representation parameters of complex data by transforming them into standardized data object formats with defined schemas. This parameter transformation allows the system to handle complex datasets internally while presenting simplified structures externally, thereby maintaining accuracy without proportionally increasing complexity in the user-facing interface.
3Quantity of substance
If unstructured data is processed directly, then data volume is maintained, but manufacturing precision and measurement precision deteriorate
Solution Approach 1:
The patent applies preliminary action by pre-processing and structuring unstructured data into standardized data objects before they are used in machine learning models. This preliminary structuring step maintains the comprehensive data volume while improving precision through consistent formatting, typing, and schema enforcement, thereby preparing the data for more accurate and reliable processing.
Data Source
AI summary
Systems and methods are described herein for creating a data object for each of a plurality of imported unstructured data files. Each data object may expressly include one of the unstructured data files. Preprocessing subsystems and/or machine learning algorithms and subsystems process the data to generate or otherwise identify structured insight features. The system updates each data object to expressly include the structured insight features.


