Vector Index Querying for Unstructured Data Repositories

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models face challenges in efficiently searching and accurately identifying relevant information in large, potentially unstructured data repositories, especially when unsupervised learning is applied, due to the difficulty in understanding relationships between data.

Innovation Solution

A system that generates serialized sequences of cell values from data tables, converts them into contextualized embeddings using a natural language model, and stores these embeddings in vector indices to determine unionability scores between tables, allowing for effective query-based retrieval of relevant data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional machine learning models are used to search large data repositories, then the system can handle basic queries, but the accuracy in identifying relevant information deteriorates when data is unstructured or relationships are complex

Engineering Contradiction:
Improveaccuracy in identifying relevant informationVSAvoidability to handle unstructured data and relationships
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent transforms tabular data into serialized sequences and converts them into contextualized embeddings, changing the parameter representation from discrete tabular format to continuous vector space. This enables the model to capture syntactic and semantic relationships while maintaining high accuracy in identifying relevant information across structured and unstructured data.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces contextualized embeddings as an intermediary representation between the raw tabular data and the query matching process. These embeddings serve as a mediator that captures relationships between data values, enabling accurate identification of relevant information without requiring the model to directly process complex unstructured relationships.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If the data repository size increases, then more information is available for queries, but the efficiency of searching and identifying relevant information deteriorates

Engineering Contradiction:
Improvedata repository sizeVSAvoidsearch efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent replaces traditional mechanical search methods (scanning and comparing tables) with a vector-based similarity search system. By converting tables into contextualized embeddings and storing them in vector indices, the system enables efficient approximate nearest neighbor search that scales to large data repositories while maintaining search efficiency through vector space operations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the search parameter from exact table matching to similarity-based vector search. This parameter change enables the system to efficiently handle large data repositories by finding semantically similar tables through vector distance calculations rather than exhaustive comparison, thereby maintaining productivity as data size increases.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If unsupervised learning is applied to understand relationships in data, then the system can handle unlabeled data, but the precision in capturing relationships deteriorates compared to supervised approaches

Engineering Contradiction:
Improveability to handle unlabeled dataVSAvoidprecision in capturing relationships
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by pre-training the natural language model on large amounts of unlabeled tabular data before fine-tuning for specific query tasks. This pre-training enables the model to learn general relationships and patterns from unlabeled data, and subsequent fine-tuning refines this precision for specific relationship types, combining the benefits of unsupervised learning with supervised precision.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12189590B1Systems and methods for querying large data repositories
Publication Date: 2025.01.07 RECRUIT
  • US12189590B1 patent drawing
  • US12189590B1 patent drawing
  • US12189590B1 patent drawing

AI summary

Disclosed embodiments relate to systems, methods, and computer readable storage media for performing dataset discovery. Some embodiments may include accessing a data repository having a plurality of tables having cell values arranged in one or more columns and one or more rows, generating serialized sequences of the cell values that correspond to particular columns of the plurality of tables, inputting the serialized sequences into a natural language model, converting, using the natural language model, the serialized sequences into contextualized embeddings associated with the plurality of tables, storing the contextualized embeddings associated with the plurality of tables in one or more vector indices, receiving a query table, or generating an output of one or more candidate tables from the plurality of tables that are unionable with the received query table.