Word Embedding Synonym Discovery for Data Curation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional database and data visualization systems struggle to discover and annotate data with synonyms, leading to inefficiencies in data preparation and curation, especially with the rapid growth of large and complex databases.
Innovation Solution
The implementation of a method that uses word embeddings and a word similarity model to semantically annotate data fields and values with automatically discovered synonyms, enabling effective data preparation and visualization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional database and data visualization systems are used, then data can be stored and visualized, but the systems are too slow and difficult to maintain due to the rapid growth of database sizes
Solution Approach 1:
The system performs preliminary semantic annotation of data fields and values by automatically discovering synonyms using word embeddings and word similarity models. This advance preparation creates semantic indexes that enable faster querying and data manipulation operations, resolving the contradiction between data curation efficiency and data preparation time.
Solution Approach 2:
The system employs automated synonym discovery mechanisms that self-annotate data without requiring manual intervention. The word embedding models and similarity algorithms automatically process data fields and values, generating semantic annotations that improve data usability while reducing the time and resources needed for manual data preparation and curation.
2Ease of operation
If manual data annotation with synonyms is performed, then data can be better queried and understood, but the process is too slow and difficult to maintain
Solution Approach 1:
The system replaces manual mechanical annotation processes with automated computational methods. Word embedding models and word similarity algorithms automatically discover synonyms and generate semantic annotations for data fields and values, eliminating the need for manual annotation while improving query capability and reducing maintenance complexity.
Solution Approach 2:
The system changes the approach to synonym discovery by using vector space representations (word embeddings) and similarity threshold parameters. By adjusting similarity thresholds and using different embedding models, the system can automatically adapt to varying data characteristics without requiring manual re-annotation, thereby improving ease of operation while managing complexity.
3Productivity
If semantic annotations are not automatically discovered, then data sources can be used as-is, but data preparation and curation become inefficient
Solution Approach 1:
The system performs preliminary semantic annotation by automatically discovering synonyms using word embeddings before data visualization operations. This advance preparation creates semantic indexes that improve data visualization efficiency by enabling faster data retrieval, filtering, and manipulation, while the automation maintains ease of data preparation.
4Quantity of substance
If databases grow larger and more complex, then more data can be analyzed, but the systems become too slow and difficult to maintain
Solution Approach 1:
The system performs preliminary semantic annotation and synonym discovery for all data fields and values in advance, creating optimized semantic indexes. This preliminary processing enables faster querying and data manipulation operations on large datasets, resolving the contradiction between handling increased data volume and maintaining processing speed.
Solution Approach 2:
The automated synonym discovery system continuously processes new data as it is added to the database, self-annotating fields and values without requiring manual intervention. This self-service approach maintains data processing speed even as database size grows, as the system automatically adapts to new data without manual re-annotation efforts.
Data Source
AI summary
A computer system receives a user input to specify a natural language command. The computer system, in response to receiving the user input, generates a semantic interpretation for the natural language command using a trained word model, based on semantic annotations for a published data source. The trained word model is trained to identify similar words based on a plurality of processed representations corresponding to a set of words of a natural language. The computer system queries the published data source based on the semantic interpretation, thereby retrieving a dataset, and generates and displays a data visualization based on the retrieved dataset.


