Learned Data Ontology Using Word Embeddings for Natural Language Queries
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Database systems face inaccuracies and inefficiencies in processing natural language queries due to the lack of structured or contextual data, leading to non-contextual query results, as they often rely on generic datasets that do not capture unique relationships within user-specific data.
Innovation Solution
The system processes entity-specific datasets to generate a master dataset, which is then used to create vector sets through word embeddings, capturing implicit relationships and contexts, allowing for accurate natural language query processing without manual mapping, and triggering relevant actions based on query similarities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If generic datasets are used for processing natural language queries, then the system can handle queries without manual configuration, but the query results become inaccurate and non-contextual
Solution Approach 1:
The system automatically processes entity-specific datasets to generate master datasets and vector sets without requiring manual configuration. The word embedding function autonomously captures implicit relationships and contexts from the data, enabling the system to serve itself by adapting to unique dataset characteristics while maintaining ease of operation.
Solution Approach 2:
The system transforms generic datasets into entity-specific master datasets by processing and reorganizing data parameters. By generating vector sets through word embeddings, the system changes the representation parameters of the data to capture contextual relationships, thereby improving query result accuracy while maintaining automated operation.
2Measurement precision
If manual mapping of natural language queries to database queries is performed, then query results can be contextualized, but the system complexity and maintenance requirements increase
Solution Approach 1:
The patent replaces the mechanical process of manual query mapping with an automated word embedding function. This function uses neural network-based vector representations to automatically capture semantic relationships and contextual meanings from entity-specific datasets, eliminating the need for manual configuration while maintaining high query result accuracy.
Solution Approach 2:
The word embedding function acts as an intermediary between the natural language query and the database. It processes the query and the master dataset to generate vector sets that capture implicit relationships, serving as a bridge that enables accurate contextualized results without requiring direct manual mapping between query languages and database schemas.
3Measurement precision
If entity-specific datasets are processed to generate master datasets with word embeddings, then query accuracy improves, but the processing time and computational resources increase
Solution Approach 1:
The system performs preliminary processing of entity-specific datasets to generate master datasets and vector sets before actual query execution. By pre-processing the data and creating embedded representations in advance, the system reduces the computational burden during query processing, thereby improving accuracy while minimizing real-time processing time.
Solution Approach 2:
The system processes only the necessary portions of the datasets required for the specific query rather than the entire dataset. By generating vector sets from relevant entity-specific data and using selective retrieval based on vector similarity, the system achieves high accuracy while reducing overall processing time and computational resource consumption.
Data Source
AI summary
Techniques described herein may support a learned ontology or meaning for user, organization, or customer specific data. According to the techniques described herein, a set of datasets corresponding to an entity may be processed to generate a master dataset including rows that include at least a field name and a value corresponding to the field. The master dataset is processed to generate a corpus of text strings that is input into a word embedding function which generates a set of vectors based on the corpus. Because the configuration of the text string positions values by field names and field values, implicit relationships and contexts are identified within the data using the word embedding function.


