Learned Data Ontology Using Word Embeddings for Natural Language Queries

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Database systems face inaccuracies and inefficiencies in processing natural language queries due to the lack of structured or contextual data, leading to non-contextual query results, as they often rely on generic datasets that do not capture unique relationships within user-specific data.

Innovation Solution

The system processes entity-specific datasets to generate a master dataset, which is then used to create vector sets through word embeddings, capturing implicit relationships and contexts, allowing for accurate natural language query processing without manual mapping, and triggering relevant actions based on query similarities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If generic datasets are used for processing natural language queries, then the system can handle queries without manual configuration, but the query results become inaccurate and non-contextual

Engineering Contradiction:
ImproveQuery processing without manual configurationVSAvoidQuery result accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system automatically processes entity-specific datasets to generate master datasets and vector sets without requiring manual configuration. The word embedding function autonomously captures implicit relationships and contexts from the data, enabling the system to serve itself by adapting to unique dataset characteristics while maintaining ease of operation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system transforms generic datasets into entity-specific master datasets by processing and reorganizing data parameters. By generating vector sets through word embeddings, the system changes the representation parameters of the data to capture contextual relationships, thereby improving query result accuracy while maintaining automated operation.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If manual mapping of natural language queries to database queries is performed, then query results can be contextualized, but the system complexity and maintenance requirements increase

Engineering Contradiction:
ImproveQuery result contextual accuracyVSAvoidSystem configuration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical process of manual query mapping with an automated word embedding function. This function uses neural network-based vector representations to automatically capture semantic relationships and contextual meanings from entity-specific datasets, eliminating the need for manual configuration while maintaining high query result accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The word embedding function acts as an intermediary between the natural language query and the database. It processes the query and the master dataset to generate vector sets that capture implicit relationships, serving as a bridge that enables accurate contextualized results without requiring direct manual mapping between query languages and database schemas.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If entity-specific datasets are processed to generate master datasets with word embeddings, then query accuracy improves, but the processing time and computational resources increase

Engineering Contradiction:
ImproveQuery result accuracyVSAvoidData processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary processing of entity-specific datasets to generate master datasets and vector sets before actual query execution. By pre-processing the data and creating embedded representations in advance, the system reduces the computational burden during query processing, thereby improving accuracy while minimizing real-time processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system processes only the necessary portions of the datasets required for the specific query rather than the entire dataset. By generating vector sets from relevant entity-specific data and using selective retrieval based on vector similarity, the system achieves high accuracy while reducing overall processing time and computational resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11675764B2Learned data ontology using word embeddings from multiple datasets
Publication Date: 2023.06.13 SALESFORCE INC
  • US11675764B2 patent drawing
  • US11675764B2 patent drawing
  • US11675764B2 patent drawing

AI summary

Techniques described herein may support a learned ontology or meaning for user, organization, or customer specific data. According to the techniques described herein, a set of datasets corresponding to an entity may be processed to generate a master dataset including rows that include at least a field name and a value corresponding to the field. The master dataset is processed to generate a corpus of text strings that is input into a word embedding function which generates a set of vectors based on the corpus. Because the configuration of the text string positions values by field names and field values, implicit relationships and contexts are identified within the data using the word embedding function.