Automated Docket Detection and Data Extraction via Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manually reviewing and processing dockets for information extraction is a time-consuming and error-prone process, especially when dealing with large volumes of documents like invoices, receipts, or credit notes, which requires significant time and resources for accurate data entry.

Innovation Solution

A computer-implemented method using trained neural networks for image processing, including docket detection and data block extraction, which employs deep learning and natural language processing to automatically detect and extract information from images of dockets, such as invoices, receipts, or credit notes, without requiring manual alignment or prior knowledge of the number of dockets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual review and data entry methods are used for processing dockets, then accuracy of information extraction can be maintained through human inspection, but significant time and resources are expended and human error remains a risk

Engineering Contradiction:
Improveaccuracy of information extractionVSAvoidtime for data entry and review
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces the mechanical human review process with an automated computer vision system comprising neural networks for docket detection, OCR for text extraction, and NLP for data block identification. This substitution eliminates manual labor while maintaining extraction accuracy through multiple processing stages including image segmentation, coordinate determination, and validation modules.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service processing where the computer system automatically performs docket detection, text recognition, data block identification, and information extraction without human intervention. The automated pipeline processes images through multiple modules that work independently yet cooperatively to extract and validate data, making the system self-sufficient for high-volume processing.

Inventive Principle:
Principle #25Self-service

2Productivity

If automated image processing methods are used for docket detection, then processing speed and scalability are improved, but handling variations in layout and orientation presents technical challenges

Engineering Contradiction:
Improveprocessing speedVSAvoidhandling layout variations
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic processing where the system adapts to different docket layouts and orientations by using neural networks that learn from training data comprising various formats. The image segmentation and data block detection modules dynamically adjust their processing based on the specific characteristics of each input image, enabling the system to handle diverse document formats while maintaining high processing speed.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes processing parameters based on input characteristics, using trained neural networks that adjust their interpretation based on detected patterns in layout, orientation, and document type. The validation module adjusts acceptance criteria based on the confidence scores generated by previous processing stages, allowing flexible handling of variations while maintaining consistent processing throughput.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If multiple processing modules are used for docket detection and data extraction, then extraction accuracy and validation are improved, but system complexity increases

Engineering Contradiction:
Improveextraction accuracyVSAvoidsystem architecture
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the processing system into distinct functional modules: docket detection module for identifying document locations, OCR module for text recognition, data block detection module for identifying information fields, and validation module for verifying extracted data. Each module performs a specific function with defined inputs and outputs, improving accuracy through specialized processing while managing complexity through modular architecture with clear interfaces between components.

Inventive Principle:
Principle #1Segmentation

4Extent of automation

If trained neural networks are used for docket detection and data block identification, then automation extent and processing efficiency are improved, but training data requirements and computational resources increase

Engineering Contradiction:
Improveautomation of docket processingVSAvoidtraining data volume
Core Design Contradiction:
Extent of automationVSQuantity of substance

Solution Approach 1:

The patent performs preliminary training of neural networks using curated training datasets comprising sample dockets with annotated data blocks and attributes. This preliminary action establishes the automated processing capability before deployment, allowing the system to achieve high automation levels without requiring extensive computational resources during actual processing. The trained models can then process production images efficiently using the learned patterns from the training phase.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20220292861A1Docket Analysis Methods and Systems
Publication Date: 2022.09.15 XERO
  • US20220292861A1 patent drawing
  • US20220292861A1 patent drawing
  • US20220292861A1 patent drawing

AI summary

A computer implemented method for processing images for docket detection and information extraction. The method comprises receiving, at a computer system, an image comprising a representation of a plurality of dockets; and detecting, by a docket detection module of the computer system, a plurality of image segments. Each image segment is associated with one of the plurality of dockets. The method comprises determining, by a character recognition module of the computer system, docket text comprising a set of characters associated with each image segment; and detecting, by a data block detection module of the computer system, based on the docket text, one or more data blocks in each of the plurality of docket segments, wherein each data block is associated with a type of information represented in the docket text.