Neural Network Table Partition Identification Using Global Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for detecting text fields in unstructured electronic documents rely heavily on manual heuristics, which are inefficient and prone to errors, especially when document layouts are irregular or poorly recognized.

Innovation Solution

The use of neural networks to automatically detect text fields and tables in documents by processing symbol sequences, determining vectors representative of these sequences, and recalculating them based on global document context to establish associations between alphanumeric sequences and table partitions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional manual heuristic methods are used for field detection, then the system is simple to implement, but the detection accuracy and reliability deteriorate

Engineering Contradiction:
Improvedetection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces manual heuristic methods with a neural network-based automated system. The neural network processes document images, performs OCR, extracts features, and detects fields automatically, substituting the mechanical manual process with an intelligent automated system that achieves higher detection accuracy and reliability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If manual heuristic methods are used, then the development time is short, but the productivity and efficiency of field detection deteriorate

Engineering Contradiction:
Improvedetection efficiencyVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-training the neural network on large datasets of documents with various layouts and formats. This pre-training enables the network to quickly and accurately detect fields in new documents without requiring extensive manual configuration or processing time, thereby improving productivity while minimizing time loss.

Inventive Principle:
Principle #10Preliminary action

3Extent of automation

If manual markup is required, then the system has high detection precision, but the ease of operation and automation level deteriorate

Engineering Contradiction:
Improveautomation levelVSAvoidoperational simplicity
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The neural network performs self-service by automatically detecting fields, tables, and their associations without requiring manual markup or human intervention. The system independently processes document images, extracts features, and produces detection results, achieving high automation level while maintaining operational simplicity through a user-friendly interface.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11775746B2Identification of table partitions in documents with neural networks using global document context
Publication Date: 2023.10.03 ABBYY DEVELOPMENT INC
  • US11775746B2 patent drawing
  • US11775746B2 patent drawing
  • US11775746B2 patent drawing

AI summary

Aspects of the disclosure provide for mechanisms for identification of table partitions in documents using neural networks. A method of the disclosure includes obtaining a plurality of symbol sequences of a document having at least one table, determining a plurality of vectors representative of symbol sequences having at least one alphanumeric character or a table graphics element, processing the plurality of vectors using a first neural network to obtain a plurality of recalculated vectors, determining an association between a first recalculated vector and a second recalculated vector, wherein the first recalculated vector is representative of an alphanumeric sequence and the second recalculated vector is associated with a table partition, and determining, based on the association between the first recalculated vector and the second recalculated vector, an association between the alphanumeric sequence and the table partition.