Semi-supervised Word Embeddings for Industry-Specific Semantic Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In industries lacking extensive documentation, creating custom word embeddings for specific domains is challenging due to the absence of relevant corpora, leading to inefficient data mining and semantic analysis, as existing search engines return irrelevant results from large databases.

Innovation Solution

A method to generate a custom corpus by creating a domain graph, gathering seed data, identifying domain-related data, querying it, and creating word embeddings using a semi-supervised enhancement model, which narrows the dataset to industry-specific data, allowing for accurate semantic analysis and data mining.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a large corpus of general documents is used for training word embeddings, then the general language understanding is improved, but the industry-specific semantic accuracy deteriorates due to lack of relevant domain data

Engineering Contradiction:
Improvesemantic analysis accuracyVSAvoidindustry-specific adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the general document corpus into industry-specific subsets by creating a domain graph that categorizes documents according to industry taxonomies. This segmentation allows the system to focus on relevant domain data while maintaining the ability to handle multiple industries through the graph structure.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary action by pre-processing and curating industry-specific document corpora before training word embeddings. The system creates domain graphs and gathers seed data in advance, preparing industry-tailored training datasets that improve semantic accuracy for specific domains before the actual embedding training occurs.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If industry-specific corpora are created from scratch, then the semantic analysis accuracy for specific domains is improved, but the time and resources required for data collection and processing increase

Engineering Contradiction:
Improveindustry-specific semantic accuracyVSAvoidcorpus creation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-curating and organizing industry-specific document corpora before training. The system creates domain graphs, identifies seed documents, and prepares training datasets in advance, which reduces the time required for actual embedding training and deployment in specific industries.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses an intermediary approach by introducing domain graphs as a mediating structure between general document corpora and industry-specific word embeddings. The domain graphs serve as intermediaries that organize and filter general documents into industry-relevant subsets, reducing the time needed to create industry-specific corpora from scratch.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If general search engines query large databases, then the coverage of search results is improved, but the relevance of results to specific industries deteriorates due to irrelevant data

Engineering Contradiction:
Improvesearch result coverageVSAvoidindustry-specific relevance
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent segments the large database into industry-specific subsets using domain graphs that organize documents according to industry taxonomies. This segmentation maintains comprehensive coverage within each industry domain while filtering out irrelevant data from other domains, thus improving result relevance without sacrificing coverage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by tailoring the search index and word embeddings to specific industry domains. The system creates domain-specific representations that optimize search results for each industry's unique terminology and context, ensuring high relevance for industry-specific queries while maintaining the ability to handle multiple industries.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11651159B2Semi-supervised system to mine document corpus on industry specific taxonomies
Publication Date: 2023.05.16 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11651159B2 patent drawing
  • US11651159B2 patent drawing
  • US11651159B2 patent drawing

AI summary

A method, computer system, and a computer program product for generating a custom corpus is provided. The present invention may include generating a domain graph. The present invention may also include gathering seed data based on the generated domain graph. The present invention may then include identifying domain related data based on the gathered seed data. The present invention may further include querying the domain related data. The present invention may also include creating word embeddings for the domain related data. The present invention may then include evaluating the domain related data.