Content Extraction Document for Automated Data Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Accessing and identifying content within structured electronic documents in an automated and efficient manner is challenging, especially when the source is not available or impractical to access.

Innovation Solution

A method and system utilizing a content extraction document (CED) that defines a common expression and data structure definition to extract and structure content elements from structured electronic documents, employing XPath and WSDL to identify and extract content elements, allowing for automated processing and presentation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated content identification methods are implemented, then productivity is improved, but device complexity increases

Engineering Contradiction:
Improvecontent identification efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces a content extraction document (CED) as an intermediary component that mediates between the structured electronic document and the extraction system. The CED contains pre-defined XPath expressions and data structure definitions that enable automated content identification without requiring complex processing logic in the extraction system itself, thus improving productivity while managing system complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies preliminary action by pre-defining content extraction rules, XPath expressions, and data structure mappings in the CED before the actual extraction process. This preparation work is done in advance, allowing the extraction system to simply follow pre-established instructions, thereby improving efficiency without increasing operational complexity

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If content is accessed from the original source, then measurement precision is improved, but loss of time increases

Engineering Contradiction:
Improvecontent accuracyVSAvoidaccess time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates a virtual copy of the content extraction configuration in the CED, which mirrors the structure and content requirements of the original document. This copy enables automated extraction without needing to physically access or retrieve the original source document, thus maintaining content accuracy while eliminating access time delays

Inventive Principle:
Principle #26Copying

Solution Approach 2:

By pre-defining all extraction paths, expressions, and data structure mappings in the CED before extraction is needed, the system prepares everything in advance. This preliminary configuration allows immediate content extraction when needed, eliminating the time loss associated with source access while ensuring accurate content retrieval

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8661335B2Methods and systems for identifying content elements
Publication Date: 2014.02.25 MALIKIE INNOVATIONS LTD
  • US8661335B2 patent drawing
  • US8661335B2 patent drawing
  • US8661335B2 patent drawing

AI summary

A method of identifying content of interest in a structured electronic document by an electronic device having a processor, an input device, and a display device, includes rendering a structured electronic document to the display device; receiving through the input device at least two separate indications of content elements within the rendered structured electronic document; and identifying with the processor a common characteristic of the indicated content elements, and identifying any further content element within the rendered structured electronic document sharing the common characteristic with the indicated content elements.