Receipt OCR Tokenization for Cryptic Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Receipts are often cryptic and difficult for third parties to interpret due to limited space and varying descriptive languages used by vendors, making it challenging to accurately capture and process purchase data from images of receipts.

Innovation Solution

An automated system processes receipt images using optical character recognition (OCR) to generate machine-encoded text, identify tokens, and aggregate potential items of content, which are then stored with the image in a data store, allowing for the extraction of purchase data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If receipts use limited space and vendor-specific descriptive language, then the receipt can be printed on narrow slips with basic transaction information, but the receipt becomes cryptic and difficult for third parties to interpret

Engineering Contradiction:
Improvereceipt printing areaVSAvoidinterpretability of purchase data
Core Design Contradiction:
Area of stationary objectVSLoss of information

Solution Approach 1:

The patent introduces an intermediary processing system that includes OCR technology to convert receipt images into machine-encoded text, token generation to identify meaningful units, and aggregation logic to reconstruct contextual meaning. This intermediary layer bridges the gap between the compact vendor-specific receipt format and the need for third-party interpretability, allowing data extraction without requiring the receipt itself to be spacious or use universal language.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system creates a digital copy of the receipt content through OCR, transforming the physical or image-based receipt into machine-encoded text data. This copy can then be processed, searched, and interpreted without altering the original receipt format, enabling third parties to access and understand purchase information even when the original receipt uses limited space and vendor-specific language.

Inventive Principle:
Principle #26Copying

2Productivity

If receipts use abbreviated and cryptic itemizations, then the receipt fits on narrow paper slips, but accurate capture and processing of purchase data becomes challenging

Engineering Contradiction:
Improvetransaction processing speedVSAvoidaccuracy of data capture
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system employs feedback mechanisms where generated tokens are aggregated and evaluated against known patterns and contexts. The aggregation process refines token interpretations by considering surrounding text, positional information, and vendor-specific patterns, continuously improving accuracy through iterative processing until confident identification of purchase items is achieved.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary OCR conversion and token generation before final data interpretation. By pre-processing the receipt image into machine-encoded text and identifying potential tokens in advance, the system prepares structured data that can be more accurately matched against product databases and vendor catalogs, improving overall data capture accuracy before the final processing stage.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10878232B2Automated processing of receipts and invoices
Publication Date: 2020.12.29 BLINKRECEIPT LLC
  • US10878232B2 patent drawing
  • US10878232B2 patent drawing
  • US10878232B2 patent drawing

AI summary

Systems and methods for automatic processing of receipts to capture data from the receipts are presented. Upon receiving an image of a receipt, an optical character recognition (OCR) of the receipt content embodied in the image is executed. The OCR of the receipt content results in machine-encoded text content of the receipt content embodied in the image. Tokens are generated from the machine-encoded text content and data groups are constructed according to horizontal lines of generated tokens. Potential product items are identified for at least some of the constructed data groups and the potential product items are evaluated for the at least some constructed data groups. The evaluation of the potential product items for the at least some constructed data groups identifies receipt data, such as product items and vendor information, associated with the least some constructed data groups. The identified receipt data is captured and stored with the image of the receipt in a data store.