Machine-Learning Data Stores for Unstructured Data Access

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data systems struggle to effectively utilize unstructured and obscured data, such as unstructured text, binary formats, and image data, which are not readily accessible for analysis, leading to suboptimal performance in search and machine-learning applications.

Innovation Solution

Implementing a data store management system that uses machine learning techniques to enhance data sets by identifying and adding hidden information, storing it as extensions in a common data format like FHIR, and providing access through enhanced data stores.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is stored in unstructured formats (text, binary, images), then data volume and storage flexibility are improved, but data accessibility and system performance deteriorate

Engineering Contradiction:
Improvedata volumeVSAvoiddata accessibility
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent segments data into structured and unstructured portions, applying different processing methods to each. Structured data is stored in standardized formats for easy access, while unstructured data is enhanced through machine learning to extract meaningful attributes that are then integrated into the structured portion, enabling selective optimization of accessibility without sacrificing storage flexibility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces machine learning models as intermediary components that process unstructured data and generate enhanced structured representations. These models act as mediators between the unstructured data source and the structured data storage system, extracting features and attributes that bridge the gap between data volume preservation and accessibility improvement.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If data is stored in unstructured formats, then storage flexibility is improved, but search and analysis capabilities deteriorate

Engineering Contradiction:
Improvestorage flexibilityVSAvoidsearch capability
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent applies preliminary action by pre-processing unstructured data through machine learning enhancement before storage. Attributes, features, and metadata are extracted and embedded into the data structure during the ingestion phase, so that when search or analysis operations are performed later, the data is already optimized for these operations without requiring real-time processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes parameters of unstructured data by applying machine learning transformations that convert raw data into enhanced representations with extracted features, derived attributes, and structured metadata. This parameter transformation maintains the original data's versatility for storage while fundamentally improving its searchability and analytical utility.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If machine learning enhancement is applied to all data, then data accessibility is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvedata accessibilityVSAvoidprocessing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent applies partial action by selectively enhancing only the portions of data that would benefit most from machine learning processing. Rather than uniformly processing all data, the system identifies and prioritizes unstructured data portions that contain hidden valuable information or require enhancement for specific use cases, applying computational resources efficiently to high-value targets.

Inventive Principle:
Principle #16Partial or excessive action

4Loss of information

If hidden information is extracted and added as extensions, then information completeness is improved, but data structure complexity increases

Engineering Contradiction:
Improveinformation completenessVSAvoiddata structure
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent implements nesting by embedding extracted attributes and enhanced information within hierarchical data structures. Hidden information is organized in nested layers where core data elements contain embedded attributes, which themselves may contain further nested metadata. This nested organization preserves information completeness while managing complexity through hierarchical abstraction.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS12353379B1Enhancing data sets to create data stores
Publication Date: 2025.07.08 AMAZON TECH INC
  • US12353379B1 patent drawing
  • US12353379B1 patent drawing
  • US12353379B1 patent drawing

AI summary

Data sets may be enhanced to create data stores. Request to create data stores may be received. As part of performing the request to create a data store, items stored in an extensible data format may be identified for machine learning enhancement. Machine learning models may be applied to generate additional data from data in the items. The additional data may be added to extend the items and store the extended items in a new data store.