Document Issuer Identification via Internal String Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for identifying the issuer of documents like receipts and business forms through image processing are complex and inefficient, requiring external server interactions and complicated processing steps to extract and match phone numbers with location information.

Innovation Solution

An image processing apparatus and program that acquires document image data, performs character recognition, and uses specific rules to extract and match character strings within the data to determine the issuer, reducing the need for external server interactions by identifying potential issuers from both URL information and adjacent text.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If external server interaction and XML file analysis are used to specify store names, then location information can be obtained, but the processing complexity increases

Engineering Contradiction:
Improveaccuracy of issuer identificationVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and utilizes issuer name information directly from the document image data itself, removing the dependency on external server interactions and XML file analyses. By focusing on extracting relevant character strings (store names, business names) that are already present in the document images, the system achieves accurate issuer identification while significantly reducing processing complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs self-service by extracting issuer information directly from the document images without requiring external server assistance. The image processing apparatus independently analyzes the document images, extracts relevant text, and determines issuer names using internal processing capabilities, thereby eliminating the need for complex external system interactions.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If multiple processing steps including phone number extraction and server transmission are used, then issuer specification is achieved, but processing time increases

Engineering Contradiction:
Improveaccuracy of issuer identificationVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent removes unnecessary processing steps such as phone number extraction and server transmissions from the workflow. By directly extracting issuer names from document images through character recognition and pattern matching, the system achieves the same identification accuracy with significantly reduced processing time.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary extraction of issuer name candidates during the character recognition phase itself, rather than requiring subsequent separate processing steps. By identifying and extracting potential issuer names early in the processing pipeline, the system eliminates the need for time-consuming sequential operations.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If character recognition and rule-based extraction are performed locally, then external server dependency is reduced, but computational requirements increase

Engineering Contradiction:
Improveindependence from external serversVSAvoidcomputational energy consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the essential information needed for issuer identification directly from document images, avoiding the need for comprehensive server-based processing. By focusing extraction efforts on specific patterns and characteristics relevant to issuer names, the system achieves server independence with optimized computational energy consumption.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10832081B2Image processing apparatus and non-transitory computer-readable computer medium storing an image processing program
Publication Date: 2020.11.10 SEIKO EPSON CORP
  • US10832081B2 patent drawing
  • US10832081B2 patent drawing
  • US10832081B2 patent drawing

AI summary

There is provided an image processing apparatus including a control unit that acquires document image data generated by reading a document and recognizes character strings included in the document image data by character recognition and a storage unit that stores a specific rule for extracting an issuer of the document, in which the control unit extracts a first character string from the character strings included in the document image data based on the specific rule, extracts a second character string which matches at least a part of the first character string from a portion other than the first character string among the character strings included in the document image data, and determines the first character string or the second character string as the issuer.