Machine-Learning Data Stores for Unstructured Data Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data systems struggle to effectively utilize unstructured and obscured data, such as unstructured text, binary formats, and image data, which are not readily accessible for analysis, leading to suboptimal performance in search and machine-learning applications.
Innovation Solution
Implementing a data store management system that uses machine learning techniques to enhance data sets by identifying and adding hidden information, storing it as extensions in a common data format like FHIR, and providing access through enhanced data stores.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is stored in unstructured formats (text, binary, images), then data volume and storage flexibility are improved, but data accessibility and system performance deteriorate
Solution Approach 1:
The patent segments data into structured and unstructured portions, applying different processing methods to each. Structured data is stored in standardized formats for easy access, while unstructured data is enhanced through machine learning to extract meaningful attributes that are then integrated into the structured portion, enabling selective optimization of accessibility without sacrificing storage flexibility.
Solution Approach 2:
The patent introduces machine learning models as intermediary components that process unstructured data and generate enhanced structured representations. These models act as mediators between the unstructured data source and the structured data storage system, extracting features and attributes that bridge the gap between data volume preservation and accessibility improvement.
2Adaptability or versatility
If data is stored in unstructured formats, then storage flexibility is improved, but search and analysis capabilities deteriorate
Solution Approach 1:
The patent applies preliminary action by pre-processing unstructured data through machine learning enhancement before storage. Attributes, features, and metadata are extracted and embedded into the data structure during the ingestion phase, so that when search or analysis operations are performed later, the data is already optimized for these operations without requiring real-time processing.
Solution Approach 2:
The patent changes parameters of unstructured data by applying machine learning transformations that convert raw data into enhanced representations with extracted features, derived attributes, and structured metadata. This parameter transformation maintains the original data's versatility for storage while fundamentally improving its searchability and analytical utility.
3Ease of operation
If machine learning enhancement is applied to all data, then data accessibility is improved, but processing time and computational resources increase
Solution Approach 1:
The patent applies partial action by selectively enhancing only the portions of data that would benefit most from machine learning processing. Rather than uniformly processing all data, the system identifies and prioritizes unstructured data portions that contain hidden valuable information or require enhancement for specific use cases, applying computational resources efficiently to high-value targets.
4Loss of information
If hidden information is extracted and added as extensions, then information completeness is improved, but data structure complexity increases
Solution Approach 1:
The patent implements nesting by embedding extracted attributes and enhanced information within hierarchical data structures. Hidden information is organized in nested layers where core data elements contain embedded attributes, which themselves may contain further nested metadata. This nested organization preserves information completeness while managing complexity through hierarchical abstraction.
Data Source
AI summary
Data sets may be enhanced to create data stores. Request to create data stores may be received. As part of performing the request to create a data store, items stored in an extensible data format may be identified for machine learning enhancement. Machine learning models may be applied to generate additional data from data in the items. The additional data may be added to extend the items and store the extended items in a new data store.


