Dhatu Vector Word Embedding for Semantic Disambiguation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Natural language processing (NLP) systems face challenges in accurately capturing the complex and context-dependent meanings of words due to the complexity of human communication, leading to difficulties in generating accurate word embeddings that reflect semantic and syntactic features.

Innovation Solution

The method employs a language-morphology-based lexical semantic extraction using Sanskrit morphological rules and Dhatus to represent English words as sparse and low-dimensional Dhatu vectors, leveraging the morphological semantics of Sanskrit to disambiguate meanings and improve interpretability in NLP tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If neural networks are trained on text corpora to predict word meanings, then word embeddings can be generated, but the complexity of human communication makes it difficult to accurately capture all semantic and syntactic features

Engineering Contradiction:
Improveaccuracy of word embeddingVSAvoidcomplexity of capturing word meaning
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the word meaning into distinct semantic features by mapping words to multiple Dhatus (semantic atoms). Each Dhatu represents a specific semantic component, allowing the system to break down complex word meanings into manageable, identifiable parts. This segmentation enables more precise capture of semantic features while reducing the complexity of processing overall meaning.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms word representations from traditional dense vectors to sparse high-dimensional vectors based on Dhatu mappings. This dimensional transformation allows the system to represent word meanings in a space where semantic relationships are more explicitly captured through sparsity patterns, improving measurement precision while managing complexity through structured representation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If neural networks are trained to predict word meanings in context, then word embeddings can be learned, but most words have multiple context-dependent meanings that are difficult to capture

Engineering Contradiction:
Improvecontext-dependent meaning captureVSAvoidaccuracy of semantic representation
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments word meanings into discrete Dhatus, where each Dhatu corresponds to a specific semantic component. This allows the system to represent multiple context-dependent meanings as separate, identifiable components rather than attempting to capture them as a monolithic meaning, thereby improving adaptability while maintaining precision through structured decomposition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent enables dynamic selection of relevant Dhatus based on context. By mapping words to multiple possible Dhatus and selecting the appropriate ones based on contextual information, the system can adapt its representation to match the specific meaning intended in different contexts, improving both versatility and precision simultaneously.

Inventive Principle:
Principle #15Dynamics

3Ease of manufacture

If traditional neural network embeddings are used, then word representations can be generated, but interpretability of the vector representations is poor

Engineering Contradiction:
Improvegeneration of word embeddingsVSAvoidinterpretability of embeddings
Core Design Contradiction:
Ease of manufactureVSEase of operation

Solution Approach 1:

The patent introduces Dhatus as intermediary semantic atoms that mediate between raw word representations and meaningful interpretations. These Dhatus serve as interpretable intermediaries that bridge the gap between computational processing and human understanding, allowing the system to maintain ease of embedding generation while significantly improving interpretability through meaningful semantic decomposition.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240220719A1Methods and Systems for Graphically Organizing the Meanings of Words
Publication Date: 2024.07.04 ZOHO OFFICE SUITE
  • US20240220719A1 patent drawing
  • US20240220719A1 patent drawing
  • US20240220719A1 patent drawing

AI summary

Described are methods and systems for graphically organizing words and their meanings. The words, in an input language such as English, are added as nodes to a graph. Synonyms of each input word taken from a second language, such as Sanskrit, are likewise added as nodes to the graph and connected to the corresponding input-word nodes via synonym edges. Root elements of the synonyms, semantic meanings, are also added to the graph as meaning nodes. The resulting graph depicts linguistic and semantic interrelations between input words, their second-language synonyms, and the meanings derived therefrom.