Conditional Random Field Passage Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for identifying relevant passages in electronic documents are limited by their dependence on pre-defined tokens or features, and struggle with accuracy and efficiency, especially when dealing with poorly recognized text or foreign languages.

Innovation Solution

A method using conditional random fields to label sentences as relevant or background information, based on a stream of features including text, tokens, word vectors, and layout, which allows for the identification of passages even when the concept is expressed differently or in poor quality text.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If simple word search or pre-defined token methods are used, then the system is easy to implement, but the accuracy and versatility of passage identification deteriorates

Engineering Contradiction:
Improveease of implementationVSAvoidaccuracy of passage identification
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent transforms the search approach from simple keyword matching to a probabilistic framework using conditional random fields. By changing the parameters from binary match/no-match to probability-based labeling with multiple features (text content, tokens, word vectors, layout, typography), the system achieves higher accuracy while maintaining implementation feasibility through algorithmic automation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent combines multiple feature types (text, tokens, word vectors, layout information, typography) into a composite feature stream that feeds into the conditional random field algorithm. This composite approach allows the system to leverage diverse information sources simultaneously, improving identification accuracy without significantly increasing implementation complexity.

Inventive Principle:
Principle #40Composite materials

2Ease of operation

If pre-defined tokens or features are required for search, then the search process is simple, but the adaptability to different concepts and languages deteriorates

Engineering Contradiction:
Improvesimplicity of search processVSAvoidadaptability to different concepts and languages
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The conditional random field algorithm serves as a universal framework that can process multiple feature types (text, tokens, word vectors, layout) and apply to different search concepts and languages. The system doesn't require pre-defined tokens specific to each concept; instead, it learns from training data and adapts to different domains, maintaining operational simplicity through a unified algorithmic approach.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system transitions from static pre-defined token matching to dynamic probabilistic labeling. The conditional random field algorithm dynamically adjusts its predictions based on the input features and training data, enabling adaptability to different concepts and languages while keeping the search process simple through automated learning rather than manual configuration.

Inventive Principle:
Principle #15Dynamics

3Productivity

If traditional search methods are used, then the processing speed is fast, but the ability to handle poor quality text or foreign languages deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidability to handle poor quality text
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces probability thresholds and tolerance parameters in the conditional random field algorithm that allow flexible handling of uncertain text quality. By adjusting these parameters, the system can maintain high processing speed while reliably handling poor quality text, foreign languages, and unresolvable words through probabilistic reasoning rather than strict matching rules.

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If a comprehensive set of features including word vectors and layout is used, then the accuracy of passage identification is improved, but the computational complexity increases

Engineering Contradiction:
Improveaccuracy of passage identificationVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs feature extraction and processing in advance during the training phase. Word vectors, layout information, and other complex features are pre-computed and stored, reducing the computational burden during actual search operations. This preliminary action allows the system to maintain high accuracy through comprehensive features while keeping runtime complexity manageable.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9645988B1System and method for identifying passages in electronic documents
Publication Date: 2017.05.09 KIRA INC
  • US9645988B1 patent drawing
  • US9645988B1 patent drawing
  • US9645988B1 patent drawing

AI summary

The methods proposed here deconstructs training sentences into a stream of features that represent both the sentences and tokens used by the text, their sequence and other ancillary features extracted using natural language processing. Then, we use a conditional random field where we represent the concept we are looking for as state A and the background (everything not concept A) as a state B. The model created by this training phase is then used to locate the concept as a sequence of sentences within a document. This has distinct advantages in accuracy and speed over methods that individually classify each sentence and then use a secondary method to group the classified sentences into passages. Furthermore while previous methods were based on searching for the occurrence of tokens only, the use of a wider set of features enables this method to locate relevant passages even though a different terminology is in use.