Semi-supervised Word Embeddings for Industry-Specific Semantic Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In industries lacking extensive documentation, creating custom word embeddings for specific domains is challenging due to the absence of relevant corpora, leading to inefficient data mining and semantic analysis, as existing search engines return irrelevant results from large databases.
Innovation Solution
A method to generate a custom corpus by creating a domain graph, gathering seed data, identifying domain-related data, querying it, and creating word embeddings using a semi-supervised enhancement model, which narrows the dataset to industry-specific data, allowing for accurate semantic analysis and data mining.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a large corpus of general documents is used for training word embeddings, then the general language understanding is improved, but the industry-specific semantic accuracy deteriorates due to lack of relevant domain data
Solution Approach 1:
The patent segments the general document corpus into industry-specific subsets by creating a domain graph that categorizes documents according to industry taxonomies. This segmentation allows the system to focus on relevant domain data while maintaining the ability to handle multiple industries through the graph structure.
Solution Approach 2:
The patent performs preliminary action by pre-processing and curating industry-specific document corpora before training word embeddings. The system creates domain graphs and gathers seed data in advance, preparing industry-tailored training datasets that improve semantic accuracy for specific domains before the actual embedding training occurs.
2Measurement precision
If industry-specific corpora are created from scratch, then the semantic analysis accuracy for specific domains is improved, but the time and resources required for data collection and processing increase
Solution Approach 1:
The patent applies preliminary action by pre-curating and organizing industry-specific document corpora before training. The system creates domain graphs, identifies seed documents, and prepares training datasets in advance, which reduces the time required for actual embedding training and deployment in specific industries.
Solution Approach 2:
The patent uses an intermediary approach by introducing domain graphs as a mediating structure between general document corpora and industry-specific word embeddings. The domain graphs serve as intermediaries that organize and filter general documents into industry-relevant subsets, reducing the time needed to create industry-specific corpora from scratch.
3Quantity of substance
If general search engines query large databases, then the coverage of search results is improved, but the relevance of results to specific industries deteriorates due to irrelevant data
Solution Approach 1:
The patent segments the large database into industry-specific subsets using domain graphs that organize documents according to industry taxonomies. This segmentation maintains comprehensive coverage within each industry domain while filtering out irrelevant data from other domains, thus improving result relevance without sacrificing coverage.
Solution Approach 2:
The patent applies local quality by tailoring the search index and word embeddings to specific industry domains. The system creates domain-specific representations that optimize search results for each industry's unique terminology and context, ensuring high relevance for industry-specific queries while maintaining the ability to handle multiple industries.
Data Source
AI summary
A method, computer system, and a computer program product for generating a custom corpus is provided. The present invention may include generating a domain graph. The present invention may also include gathering seed data based on the generated domain graph. The present invention may then include identifying domain related data based on the gathered seed data. The present invention may further include querying the domain related data. The present invention may also include creating word embeddings for the domain related data. The present invention may then include evaluating the domain related data.


