Unstructured Data Entity Embeddings for SQL Querying
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for extracting value from unstructured data, such as images, video, audio, and text, require deep learning expertise and expensive infrastructure, making it inaccessible for most businesses to leverage Natural Data effectively.
Innovation Solution
A system that enables average engineers and SQL users to process unstructured data by forming entities with relational attributes, using machine learning embedding models to compute numeric vectors, and applying SQL queries to these entities, embeddings, and enrichment models, allowing for easy extraction of value from unstructured data without the need for specialized ML/data talent or infrastructure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If deep learning models and infrastructure are used to extract value from unstructured data, then data extraction capability is improved, but device complexity and cost increase
Solution Approach 1:
The patent introduces an intermediary layer between unstructured data and SQL queries that automatically generates embedding models and enrichment models. This mediator handles the complex deep learning operations, allowing users to interact with unstructured data through simple SQL queries without needing to understand or manage the underlying ML infrastructure.
Solution Approach 2:
The system performs self-service by automatically training embedding models on unstructured data and generating enrichment models that predict properties of entities. The infrastructure autonomously manages model training, optimization, and maintenance, eliminating the need for specialized ML talent to continuously manage the system.
2Productivity
If specialized ML/data talent is hired to process unstructured data, then data processing capability is improved, but operational cost increases
Solution Approach 1:
The system is designed to be self-sufficient, automatically training embedding models and generating enrichment models without requiring specialized ML engineers. The infrastructure manages its own optimization and maintenance, converting operational costs primarily to computational resources rather than expensive specialized talent.
3Loss of energy
If unstructured data is processed using traditional methods, then infrastructure cost is reduced, but data value extraction is limited
Solution Approach 1:
The patent creates an intermediary layer that bridges inexpensive traditional infrastructure with powerful deep learning capabilities. This layer automatically generates embedding models and enrichment models, enabling sophisticated data extraction using standard SQL queries on conventional databases, thus achieving high value extraction without expensive specialized infrastructure.
4Measurement precision
If deep learning models are optimized for performance, then data extraction accuracy is improved, but device complexity increases
Solution Approach 1:
The system automatically optimizes embedding models and enrichment models through self-service mechanisms. The infrastructure handles model training, hyperparameter optimization, and performance tuning autonomously, achieving high data extraction accuracy while keeping the user-facing interface simple and the operational complexity low.
Data Source
AI summary
A non-transitory computer readable storage medium has instructions executed by a processor to receive from a network connection different sources of unstructured data. An entity is formed by combining one or more sources of the unstructured data, where the entity has relational data attributes. A representation for the entity is created, where the representation includes embeddings that are numeric vectors computed using machine learning embedding models, including trunk models, where a trunk model is a machine learning model trained on data in a self-supervised manner. An enrichment model is created to predict a property of the entity. A query is processed to produce a query result, where the query is applied to one or more of the entity, the embeddings, the machine learning embedding models, and the enrichment model.


