Content Extraction Document for Automated Data Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Accessing and identifying content within structured electronic documents in an automated and efficient manner is challenging, especially when the source is not available or impractical to access.
Innovation Solution
A method and system utilizing a content extraction document (CED) that defines a common expression and data structure definition to extract and structure content elements from structured electronic documents, employing XPath and WSDL to identify and extract content elements, allowing for automated processing and presentation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated content identification methods are implemented, then productivity is improved, but device complexity increases
Solution Approach 1:
The patent introduces a content extraction document (CED) as an intermediary component that mediates between the structured electronic document and the extraction system. The CED contains pre-defined XPath expressions and data structure definitions that enable automated content identification without requiring complex processing logic in the extraction system itself, thus improving productivity while managing system complexity
Solution Approach 2:
The patent applies preliminary action by pre-defining content extraction rules, XPath expressions, and data structure mappings in the CED before the actual extraction process. This preparation work is done in advance, allowing the extraction system to simply follow pre-established instructions, thereby improving efficiency without increasing operational complexity
2Measurement precision
If content is accessed from the original source, then measurement precision is improved, but loss of time increases
Solution Approach 1:
The patent creates a virtual copy of the content extraction configuration in the CED, which mirrors the structure and content requirements of the original document. This copy enables automated extraction without needing to physically access or retrieve the original source document, thus maintaining content accuracy while eliminating access time delays
Solution Approach 2:
By pre-defining all extraction paths, expressions, and data structure mappings in the CED before extraction is needed, the system prepares everything in advance. This preliminary configuration allows immediate content extraction when needed, eliminating the time loss associated with source access while ensuring accurate content retrieval
Data Source
AI summary
A method of identifying content of interest in a structured electronic document by an electronic device having a processor, an input device, and a display device, includes rendering a structured electronic document to the display device; receiving through the input device at least two separate indications of content elements within the rendered structured electronic document; and identifying with the processor a common characteristic of the indicated content elements, and identifying any further content element within the rendered structured electronic document sharing the common characteristic with the indicated content elements.


