Web-scale Entity Relationship Extraction via Bootstrapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current tools are inadequate for automatically detecting and extracting entity relationships from the vast amount of data on the Web, as they rely on pre-specified relations and human-tagged examples, making it difficult to identify new relationships and create a web-scale relationship graph efficiently.

Innovation Solution

The use of discriminative and probabilistic models, such as Markov Logic Networks, for iterative relationship extraction, starting with initial seeds and learning new models to identify and extract new relationship tuples, which can be clustered for open information extraction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual detection methods are used to identify entity relationships, then relationship extraction accuracy can be maintained, but the time consumption becomes too great to handle web-scale data effectively

Engineering Contradiction:
Improverelationship extraction accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system uses automatically discovered patterns from extracted relationships to improve its own extraction capability iteratively. The bootstrapping process allows the system to self-enhance by learning from its own outputs, eliminating the need for continuous manual intervention while maintaining improving accuracy over time

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical detection processes with automated computational models including discriminative models, probabilistic models, and Markov Logic Networks. These automated systems process web-scale data at machine speed while maintaining relationship extraction accuracy through learned patterns rather than human judgment

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of manufacture

If existing extraction tools are used that rely on pre-specified relations and human-tagged examples, then extraction can be performed with available tools, but the system cannot identify new relationships or handle web-scale data effectively

Engineering Contradiction:
Improveextraction capabilityVSAvoidability to identify new relationships
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary extraction with initial seeds and patterns, then uses these results to generate new patterns and expand its capability. This preliminary action creates a foundation that enables subsequent identification of new relationships without requiring complete pre-specification

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The extraction system is designed to be dynamic and adaptive, continuously evolving its patterns and models through iterative bootstrapping. The system transitions from static pre-specified relations to dynamic pattern discovery, allowing it to adapt to new relationship types and handle web-scale data effectively

Inventive Principle:
Principle #15Dynamics

3Quantity of substance

If comprehensive relationship extraction is performed across the entire web, then complete relationship graphs can be created, but the computational resources and time required become prohibitive

Engineering Contradiction:
Improveamount of relationship data extractedVSAvoidextraction efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The system uses a small set of initial seed relationships and patterns to bootstrap the extraction process, rather than attempting to process all web data from scratch. This partial action approach allows the system to achieve web-scale extraction by iteratively expanding from a manageable initial subset

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The iterative bootstrapping process maintains continuous useful action by constantly refining and expanding patterns based on previously extracted relationships. Each iteration builds upon the previous one, creating a continuous improvement cycle that efficiently processes large volumes of data without redundant computation

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS8918348B2Web-scale entity relationship extraction
Publication Date: 2014.12.23 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8918348B2 patent drawing
  • US8918348B2 patent drawing
  • US8918348B2 patent drawing

AI summary

Techniques for displaying a relationship graph are described herein. In one example, a search term may be used to obtain a plurality of documents from a network, such as the Internet. A plurality of entities, and relationships between at least some of those entities, may be extracted from the documents. In an example user interface, representations of a plurality of entities may be displayed, such as by shapes (e.g., circles) labeled to identify people or organizations. Edges (e.g., lines) may be used to connect different representations of entities and to thereby indicate a relationship between the connected entities. In a particular example, input from movement of a cursor over an edge may result in display of a description of a relationship between the connected entities. In a further particular example, size of each entity may be related to a number of connections each has with others.