Automated Fact Extraction Using Title and Contextual Patterns

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The vast volume of web-based documents makes it impractical for human editors to efficiently identify and extract objects and facts, necessitating an automated solution for mass fact extraction.

Innovation Solution

A system and method that select a source object and document, identify title and contextual patterns, and associate objects with identified facts by applying these patterns to a set of documents, creating or merging objects within a fact repository.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human editors manually identify and extract objects and facts from documents, then extraction accuracy and quality are improved, but productivity and scalability are worsened due to the vast volume of documents

Engineering Contradiction:
Improvefact extraction accuracyVSAvoidfact extraction throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces the manual mechanical process of human editors extracting facts with an automated computational system. The system uses pattern recognition algorithms to automatically identify objects and facts from documents, substituting human cognitive processing with computer-based pattern matching and extraction mechanisms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service fact extraction by automatically learning and applying patterns from source documents without requiring human intervention for each extraction task. The pattern recognition system autonomously processes documents, identifies objects and facts, and populates the fact repository without manual editing.

Inventive Principle:
Principle #25Self-service

2Productivity

If automated pattern recognition is used to extract facts from documents, then productivity and scalability are improved, but measurement precision and extraction accuracy are worsened

Engineering Contradiction:
Improvefact extraction throughputVSAvoidfact extraction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary pattern learning by analyzing source documents to identify title patterns and contextual patterns before extracting facts from target documents. This preliminary action of pattern recognition and classification prepares the system to accurately extract facts from new documents by matching them against learned patterns.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system incorporates feedback mechanisms where extracted facts are validated and used to refine pattern recognition. The fact repository stores extracted information that can be referenced to improve future extractions, creating a feedback loop that enhances accuracy over time as the system processes more documents.

Inventive Principle:
Principle #23Feedback

3Productivity

If the system processes a large set of documents to improve coverage, then productivity is improved, but device complexity and processing requirements increase

Engineering Contradiction:
Improvevolume of facts extractedVSAvoidpattern matching system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the document processing task into distinct phases: pattern learning phase (analyzing source documents to identify title and contextual patterns) and fact extraction phase (applying patterns to target documents). This segmentation allows the system to handle large volumes of documents by reusing learned patterns rather than reanalyzing all documents from scratch.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The pattern recognition system serves multiple functions: it identifies title patterns from document headings, extracts contextual patterns from document content, classifies documents into categories, and guides fact extraction. This multi-functionality reduces overall system complexity by consolidating multiple processing tasks into a unified pattern-based framework.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8812435B1Learning objects and facts from documents
Publication Date: 2014.08.19 GOOGLE LLC
  • US8812435B1 patent drawing
  • US8812435B1 patent drawing
  • US8812435B1 patent drawing

AI summary

A system, method, and computer program product for learning objects and facts from documents. A source object and a source document are selected and a title pattern and a contextual pattern are identified based on the source object and the source document. A set of documents matching the title pattern and the contextual pattern are selected. For each document in the selected set, a name and one or more facts are identified by applying the title pattern and the contextual pattern to the document. Objects are identified or created based on the identified names and associated with the identified facts.