Automated Relationship Type Candidate Extraction from Annotated Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Defining and identifying relationship types between entity types in machine learning datasets is challenging due to the domain-specific nature of these definitions, requiring manual or automated annotation and the need for extensive definitions, which increases with the number of entity types.

Innovation Solution

A computer-implemented method that analyzes documents annotated with entity types, counts co-occurring pairs, and identifies candidates for relationship types and their labels by parsing sentences to determine relationships between entity types, allowing for the storage and output of candidates with or without labels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation is used to define relationship types, then precision of relationship definitions is improved, but productivity and time consumption deteriorate

Engineering Contradiction:
Improveprecision of relationship definitionsVSAvoidproductivity of relationship type definition
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs preliminary automated extraction of candidate relationship types and labels from training data before final definition is needed. This preliminary action identifies potential relationships and prepares candidate definitions, which are then refined or selected, significantly reducing the time and effort required for manual relationship type definition while maintaining precision.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If the number of entity types increases, then versatility of the machine learning system is improved, but the number of relationship type definitions required increases, making the system more complex

Engineering Contradiction:
Improveversatility of machine learning systemVSAvoidcomplexity of relationship type definitions
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements a universal relationship extraction framework that can handle any number of entity types without requiring proportional increases in manually defined relationship types. The automated extraction process discovers relationships between any entity type combinations present in the data, making the system scalable and adaptable to domains with varying numbers of entity types without increasing complexity proportionally.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system enables self-service relationship type discovery by automatically extracting candidate relationship types and labels from training data. This self-service mechanism eliminates the need for manual definition of relationship types for each new entity type combination, allowing the system to adapt to increased versatility without proportionally increasing complexity.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If extensive manual definition of relationship types is performed, then measurement precision of relationships is improved, but loss of time and productivity worsen

Engineering Contradiction:
Improveprecision of relationship identificationVSAvoidtime for relationship type definition
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system copies relationship patterns and labels from existing training data to identify new relationship types. By extracting and replicating relationship structures from annotated examples, the system maintains measurement precision through pattern consistency while dramatically reducing the time required to define relationship types compared to creating each definition from scratch.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11163806B2Obtaining candidates for a relationship type and its label
Publication Date: 2021.11.02 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11163806B2 patent drawing
  • US11163806B2 patent drawing
  • US11163806B2 patent drawing

AI summary

The present invention may be a method, a computer system, and/or a computer program product. An embodiment of the present invention provides a computer-implemented method for obtaining one or more candidates for a relationship type and its label. The method comprises the following steps: analyzing a document annotated with entity types, the analysis comprising counting the number of pairs of co-occurring entity types in each sentence in the document, and judging whether there exists, in the document, a candidate for a label of a relationship type which shows relationship between or among the co-occurring entity types and, if the judgment is positive, storing a candidate for the relationship type and a candidate for its label; and outputting a result of the analysis. The method may further comprise, if the judgment is negative, storing a candidate for the relationship type without a candidate for its label.