Semi-Supervised Entity Chain Generation via Topic Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current entity generation techniques are either supervised or unsupervised, relying on annotated textual corpora or co-occurrences, which makes them less reliable and inconsistent due to the need for well-defined rules and randomness in untrained data analysis.

Innovation Solution

A semi-supervised method utilizing a topic modeling framework that combines structured and unstructured data, eliminating the need for manual annotations and co-occurrences, and integrating similarity metrics to produce portable and consistent results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If supervised learning is used for entity generation, then reliability is improved through annotated training data, but device complexity increases due to well-defined rules and manual annotations

Engineering Contradiction:
Improveentity generation reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces a semi-supervised learning framework that acts as an intermediary between supervised and unsupervised approaches. This framework uses a small amount of annotated data to guide the extraction of entity chains from large volumes of unstructured data, reducing the need for extensive manual annotations while maintaining reliability through the semi-supervised approach that combines rules-based methods with machine learning.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If unsupervised learning is used for entity generation, then ease of operation is improved by eliminating manual annotations, but reliability deteriorates due to randomness in untrained data analysis

Engineering Contradiction:
Improveoperation simplicityVSAvoidentity generation consistency
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-processing unstructured data through feature extraction and transformation before applying the semi-supervised learning model. This preliminary structuring of data allows the system to operate with minimal manual intervention while maintaining consistency through the organized input data that guides the entity chain extraction process.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If traditional entity generation methods are used, then manufacturing precision is maintained through established processes, but productivity decreases due to inability to handle diverse data sources

Engineering Contradiction:
Improveentity extraction accuracyVSAvoiddata processing throughput
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent implements universality by creating a semi-supervised learning framework that can handle multiple types of data sources (structured and unstructured) and various entity types simultaneously. The system extracts entity chains from diverse sources including text documents, databases, and other data formats, maintaining accuracy while significantly increasing productivity through this multi-functional approach.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11080615B2Generating chains of entity mentions
Publication Date: 2021.08.03 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11080615B2 patent drawing
  • US11080615B2 patent drawing
  • US11080615B2 patent drawing

AI summary

Aspects of the present invention disclose a method for analyzing data from a plurality of data sources. The method includes extracting features of data received from a first source and from a second source by analyzing the data received from the first source of data and from the second source. The method includes processors determining a topic modeling framework, wherein the topic modeling framework detects a semantic structure of the features of the data received from the first data source and the second source. The method includes processors applying the topic modeling framework to the data received from the first source of data the second source of data. The method includes generating a final entity output, wherein the final entity output includes a cluster of entity mentions that the applied topic modeling framework extracts from the first source of data and the second source of data are combined.