Natural Language Query System with Interactive Ambiguity Resolution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing complexity and cost of managing and storing large volumes of data across various data storage systems pose challenges for organizations, as different data storage technologies offer varying performance benefits, leading to data dispersion across multiple locations and types of storage systems, making it difficult for users to interact with and analyze the data effectively.

Innovation Solution

Implementing a natural language query processing system that allows users to submit queries without knowing the underlying data storage system interfaces, using interactive assistance features like query assistance, auto-completion, and ambiguity handling to improve query performance and accuracy, and generating intermediate representations to execute queries across multiple data sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If data is stored across multiple different data storage systems to optimize performance and analysis benefits, then data storage performance is improved, but device complexity increases

Engineering Contradiction:
Improvedata storage performanceVSAvoiddata storage system complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent divides data storage across multiple specialized storage systems (object stores, database systems, data streams) where each handles specific types of data. This segmentation allows each storage system to be optimized for its specific function while the virtual data warehouse provides a unified interface, resolving the contradiction between performance optimization and system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The virtual data warehouse acts as an intermediary layer between users and the multiple underlying storage systems. It provides a unified interface that abstracts the complexity of multiple storage technologies, allowing users to query data without needing to understand the underlying distributed storage architecture, thus maintaining simplicity while enabling performance optimization.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If data is distributed across multiple storage locations and types, then data analysis capability is improved, but ease of operation deteriorates

Engineering Contradiction:
Improvedata analysis capabilityVSAvoiduser interaction difficulty
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The virtual data warehouse provides a universal interface that works with multiple different storage systems and data types through a single standardized query mechanism. Users can analyze diverse data (objects, relational data, streams) using the same natural language or SQL interface, eliminating the need to learn different interaction methods for each storage type while maintaining access to all data types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The virtual data warehouse serves as a mediator that translates user queries into appropriate operations on the underlying diverse storage systems. It handles the complexity of data distribution and retrieval automatically, allowing users to operate with simple, unified commands while the system manages the complexity of accessing data across multiple locations and formats.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If specialized data storage technologies are used for different data types, then manufacturing precision is improved, but device complexity increases

Engineering Contradiction:
Improvedata storage precisionVSAvoidstorage technology complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system segments data storage by type, using specialized storage technologies (object stores for unstructured data, relational databases for structured data, data streams for real-time data) where each is optimized for its specific data type. This segmentation enables high precision for each data type while the virtual data warehouse abstracts the complexity of managing multiple specialized systems.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12007988B2Interactive assistance for executing natural language queries to data sets
Publication Date: 2024.06.11 AMAZON TECH INC
  • US12007988B2 patent drawing
  • US12007988B2 patent drawing
  • US12007988B2 patent drawing

AI summary

Interactive assistances for executing natural language queries to data sets may be performed. A natural language query may be received. Candidate entity linkages may be determined between an entity recognized in the natural language query and columns in data sets. The candidate linkages may be ranked according to confidence scores which may be evaluated to detect ambiguity for an entity linkage. Candidate entity linkages may be provided to a user via an interface to select an entity linkage to use as part of completing the natural language query.