Web-scale Entity Relationship Extraction via Bootstrapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current tools are inadequate for automatically detecting and extracting entity relationships from the vast amount of data on the Web, as they rely on pre-specified relations and human-tagged examples, making it difficult to identify new relationships and create a web-scale relationship graph efficiently.
Innovation Solution
The use of discriminative and probabilistic models, such as Markov Logic Networks, for iterative relationship extraction, starting with initial seeds and learning new models to identify and extract new relationship tuples, which can be clustered for open information extraction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual detection methods are used to identify entity relationships, then relationship extraction accuracy can be maintained, but the time consumption becomes too great to handle web-scale data effectively
Solution Approach 1:
The system uses automatically discovered patterns from extracted relationships to improve its own extraction capability iteratively. The bootstrapping process allows the system to self-enhance by learning from its own outputs, eliminating the need for continuous manual intervention while maintaining improving accuracy over time
Solution Approach 2:
The patent replaces manual mechanical detection processes with automated computational models including discriminative models, probabilistic models, and Markov Logic Networks. These automated systems process web-scale data at machine speed while maintaining relationship extraction accuracy through learned patterns rather than human judgment
2Ease of manufacture
If existing extraction tools are used that rely on pre-specified relations and human-tagged examples, then extraction can be performed with available tools, but the system cannot identify new relationships or handle web-scale data effectively
Solution Approach 1:
The system performs preliminary extraction with initial seeds and patterns, then uses these results to generate new patterns and expand its capability. This preliminary action creates a foundation that enables subsequent identification of new relationships without requiring complete pre-specification
Solution Approach 2:
The extraction system is designed to be dynamic and adaptive, continuously evolving its patterns and models through iterative bootstrapping. The system transitions from static pre-specified relations to dynamic pattern discovery, allowing it to adapt to new relationship types and handle web-scale data effectively
3Quantity of substance
If comprehensive relationship extraction is performed across the entire web, then complete relationship graphs can be created, but the computational resources and time required become prohibitive
Solution Approach 1:
The system uses a small set of initial seed relationships and patterns to bootstrap the extraction process, rather than attempting to process all web data from scratch. This partial action approach allows the system to achieve web-scale extraction by iteratively expanding from a manageable initial subset
Solution Approach 2:
The iterative bootstrapping process maintains continuous useful action by constantly refining and expanding patterns based on previously extracted relationships. Each iteration builds upon the previous one, creating a continuous improvement cycle that efficiently processes large volumes of data without redundant computation
Data Source
AI summary
Techniques for displaying a relationship graph are described herein. In one example, a search term may be used to obtain a plurality of documents from a network, such as the Internet. A plurality of entities, and relationships between at least some of those entities, may be extracted from the documents. In an example user interface, representations of a plurality of entities may be displayed, such as by shapes (e.g., circles) labeled to identify people or organizations. Edges (e.g., lines) may be used to connect different representations of entities and to thereby indicate a relationship between the connected entities. In a particular example, input from movement of a cursor over an edge may result in display of a description of a relationship between the connected entities. In a further particular example, size of each entity may be related to a number of connections each has with others.


