Information Extraction Model for Document Metadata Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Authors often fail to completely and effectively define metadata in electronic documents, which hinders document processing applications like search and retrieval.

Innovation Solution

An information extraction model is trained on format and linguistic features within labeled documents to identify and extract relevant information, such as titles and authors, improving metadata accuracy and completeness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If authors manually define metadata in documents, then metadata can be customized to specific needs, but authors seldom define document metadata completely and effectively

Engineering Contradiction:
Improvemetadata customizationVSAvoidmetadata completeness
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system automatically extracts metadata from documents without requiring author intervention. The information extraction model processes documents and generates metadata fields (title, author, date, etc.) autonomously, making the system self-serving rather than relying on manual author definition.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the approach from manual metadata definition to automated extraction based on document format features. By detecting formatting patterns, font characteristics, and structural elements, the system transforms undefined metadata fields into extracted information, effectively completing metadata definition without author input.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If automated information extraction is implemented, then metadata completeness improves, but the system complexity increases

Engineering Contradiction:
Improvemetadata accuracyVSAvoidextraction system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces manual metadata definition processes with automated computational extraction. Instead of relying on mechanical manual input, the system uses information extraction models that process document formatting features, font characteristics, and structural patterns to automatically generate accurate metadata.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces an information extraction model as an intermediary between the raw document and the required metadata. This intermediary component processes document format features and linguistic characteristics, transforming them into structured metadata fields, thereby managing system complexity through a dedicated processing layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS7469251B2Extraction of information from documents
Publication Date: 2008.12.23 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7469251B2 patent drawing
  • US7469251B2 patent drawing
  • US7469251B2 patent drawing

AI summary

An information extraction model is trained on format features identified within labeled training documents. Information from a document is extracted by assigning labels to units based on format features of the units within the document. A begin label and end label are identified and the information is extracted between the begin label and the end label. The extracted information can be used in various document processing tasks such as ranking.