Probabilistic Language Models for Discontinuous Text Reading Order
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies fail to accurately identify the reading order of text segments in documents formatted in portable document format (PDF), especially when text is discontinuous or interspersed with images or visual elements, leading to incorrect ordering and limited flexibility across various document types.
Innovation Solution
The development of automated workflows that utilize trained natural language models, such as n-gram models and recurrent neural networks, to iteratively generate and order text segments based on probabilistic language models, incorporating user annotations to update and refine the models for improved accuracy across diverse document formats.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing tagging tools or spatial information methods are used to identify reading order, then the process is simple, but the accuracy is insufficient for discontinuous text segments
Solution Approach 1:
The patent replaces traditional mechanical tagging systems and spatial sorting algorithms with probabilistic language models (n-gram models and recurrent neural networks) that use statistical patterns in language to determine reading order. This substitution enables accurate identification of discontinuous text segments by evaluating the likelihood of text sequences based on learned language patterns rather than relying on document structure tags or simple spatial arrangements.
2Measurement precision
If probabilistic language models are used to order text segments, then the accuracy improves, but the computational complexity and processing time increase
Solution Approach 1:
The patent employs pre-trained probabilistic language models that have been trained on large corpora of text data before deployment. This preliminary training allows the models to quickly evaluate text segment sequences without requiring extensive computation during actual reading order identification. The models can rapidly compute probabilities for different text arrangements based on patterns learned during the pre-training phase, significantly reducing processing time while maintaining high accuracy.
3Adaptability or versatility
If traditional text extraction tools are used, then the implementation is straightforward, but the ability to handle discontinuous text segments is limited
Solution Approach 1:
The patent creates a universal system that can handle various document types and text layouts through probabilistic language modeling. The same n-gram and recurrent neural network models can process different document formats, languages, and text arrangements by evaluating the statistical likelihood of text sequences. This universal approach eliminates the need for document-type-specific processing rules, making the system adaptable to discontinuous text segments across diverse contexts while maintaining a consistent workflow.
Data Source
AI summary
The present invention is directed towards providing automated workflows for the identification of a reading order from text segments extracted from a document. Ordering the text segments is based on trained natural language models. In some embodiments, the workflows are enabled to perform a method for identifying a sequence associated with a portable document. The methods includes iteratively generating a probabilistic language model, receiving the portable document, and selectively extracting features (such as but not limited to text segments) from the document. The method may generate pairs of features (or feature pair from the extracted features). The method may further generate a score for each of the pairs based on the probabilistic language model and determine an order to features based on the scores. The method may provide the extracted features in the determined order.


