Word Embedding Synonym Discovery for Data Curation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional database and data visualization systems struggle to discover and annotate data with synonyms, leading to inefficiencies in data preparation and curation, especially with the rapid growth of large and complex databases.

Innovation Solution

The implementation of a method that uses word embeddings and a word similarity model to semantically annotate data fields and values with automatically discovered synonyms, enabling effective data preparation and visualization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional database and data visualization systems are used, then data can be stored and visualized, but the systems are too slow and difficult to maintain due to the rapid growth of database sizes

Engineering Contradiction:
Improvedata curation efficiencyVSAvoiddata preparation time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary semantic annotation of data fields and values by automatically discovering synonyms using word embeddings and word similarity models. This advance preparation creates semantic indexes that enable faster querying and data manipulation operations, resolving the contradiction between data curation efficiency and data preparation time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system employs automated synonym discovery mechanisms that self-annotate data without requiring manual intervention. The word embedding models and similarity algorithms automatically process data fields and values, generating semantic annotations that improve data usability while reducing the time and resources needed for manual data preparation and curation.

Inventive Principle:
Principle #25Self-service

2Ease of operation

If manual data annotation with synonyms is performed, then data can be better queried and understood, but the process is too slow and difficult to maintain

Engineering Contradiction:
Improvedata query capabilityVSAvoidannotation maintenance complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system replaces manual mechanical annotation processes with automated computational methods. Word embedding models and word similarity algorithms automatically discover synonyms and generate semantic annotations for data fields and values, eliminating the need for manual annotation while improving query capability and reducing maintenance complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the approach to synonym discovery by using vector space representations (word embeddings) and similarity threshold parameters. By adjusting similarity thresholds and using different embedding models, the system can automatically adapt to varying data characteristics without requiring manual re-annotation, thereby improving ease of operation while managing complexity.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If semantic annotations are not automatically discovered, then data sources can be used as-is, but data preparation and curation become inefficient

Engineering Contradiction:
Improvedata visualization efficiencyVSAvoiddata preparation ease
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

The system performs preliminary semantic annotation by automatically discovering synonyms using word embeddings before data visualization operations. This advance preparation creates semantic indexes that improve data visualization efficiency by enabling faster data retrieval, filtering, and manipulation, while the automation maintains ease of data preparation.

Inventive Principle:
Principle #10Preliminary action

4Quantity of substance

If databases grow larger and more complex, then more data can be analyzed, but the systems become too slow and difficult to maintain

Engineering Contradiction:
Improvedata volumeVSAvoiddata processing speed
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The system performs preliminary semantic annotation and synonym discovery for all data fields and values in advance, creating optimized semantic indexes. This preliminary processing enables faster querying and data manipulation operations on large datasets, resolving the contradiction between handling increased data volume and maintaining processing speed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The automated synonym discovery system continuously processes new data as it is added to the database, self-annotating fields and values without requiring manual intervention. This self-service approach maintains data processing speed even as database size grows, as the system automatically adapts to new data without manual re-annotation efforts.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250103814A1Automatic Synonyms Using Word Embedding and Word Similarity Models
Publication Date: 2025.03.27 TABLEAU SOFTWARE INC
  • US20250103814A1 patent drawing
  • US20250103814A1 patent drawing
  • US20250103814A1 patent drawing

AI summary

A computer system receives a user input to specify a natural language command. The computer system, in response to receiving the user input, generates a semantic interpretation for the natural language command using a trained word model, based on semantic annotations for a published data source. The trained word model is trained to identify similar words based on a plurality of processed representations corresponding to a set of words of a natural language. The computer system queries the published data source based on the semantic interpretation, thereby retrieving a dataset, and generates and displays a data visualization based on the retrieved dataset.