Context Analysis Module for Automated Document Rule Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional information extraction methods require both engineering and domain knowledge, making it difficult for domain experts to extract information from documents without programming skills, especially in complex development environments.

Innovation Solution

A system and method that uses a context analysis module to determine discriminative sequences from annotated documents, generating proposed rules or feature sets that do not require programming knowledge, allowing domain experts to extract information using rule-based or feature-based methods.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional manual extractor creation is used, then extraction accuracy can be achieved, but the complexity of the development environment requires engineering knowledge that most domain experts do not possess

Engineering Contradiction:
Improveease of extractor creationVSAvoiddevelopment environment complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces an automated rule generation system that acts as an intermediary between domain experts and the complex extraction development environment. This system automatically generates extraction rules from annotated documents, eliminating the need for domain experts to directly navigate complex development environments while preserving extraction accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables domain experts to perform extractor creation themselves by providing tools that automatically generate extraction rules from their annotated documents. This self-service approach eliminates the need for external extractor engineers, allowing domain experts to independently create extractors using their domain knowledge without requiring programming or engineering expertise.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If automated rule generation is implemented, then ease of use for domain experts improves, but the system complexity increases

Engineering Contradiction:
Improveusability for domain expertsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the extraction system into distinct functional modules: document annotation interface, rule generation engine, and extraction execution component. This segmentation allows the complex automated rule generation functionality to be isolated and managed separately, improving usability for domain experts while containing system complexity in specific modular components.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The automated rule generation system is designed to handle multiple document types and extraction scenarios through a universal interface. By creating a multi-functional system that can process various document formats and generate different types of extraction rules, the patent improves adaptability for domain experts while managing system complexity through standardized processing pipelines.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Manufacturing precision

If manual extractor engineering is required, then extraction precision can be maintained, but productivity decreases due to the need for specialized engineering knowledge

Engineering Contradiction:
Improveextraction precisionVSAvoidextractor development speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system performs preliminary analysis of annotated documents to automatically generate extraction rules before the actual extraction process. This preliminary action captures the precision requirements from domain expert annotations and translates them into executable rules, maintaining extraction precision while significantly accelerating the extractor development process by eliminating manual rule engineering steps.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a feedback mechanism where the system learns from domain expert annotations and automatically refines extraction rules. This feedback loop maintains high extraction precision by incorporating domain expert corrections and preferences while improving productivity through automated iterative rule generation, eliminating the need for repeated manual engineering cycles.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10102193B2Information extraction and annotation systems and methods for documents
Publication Date: 2018.10.16 OPEN TEXT CORPORATION
  • US10102193B2 patent drawing
  • US10102193B2 patent drawing
  • US10102193B2 patent drawing

AI summary

Information extraction and annotation systems and methods for use in annotating and determining annotation instances are provided herein. Exemplary methods include receiving annotated documents, the annotated documents comprising annotated fields, analyzing the annotated documents to determine contextual information for each of the annotated fields, determining discriminative sequences using the contextual information, generating a proposed rule or a feature set using the discriminative sequences and annotated fields, and providing the proposed rule or the feature set to a document annotator.