Form Identification via Spatial-Semantic Coordinate Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document and form analysis techniques face challenges in efficiently matching and registering forms due to issues like scan noise, OCR errors, scaling, and rotation, which complicate the direct template matching process.

Innovation Solution

The approach maintains spatial information from a 2D image space while injecting semantic information from optical or image character recognition (OICR) data, allowing for efficient keyword matching and form registration regardless of image quality or transformations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multi-scaling and rotation techniques are used to match keywords, then the system can handle scaling and rotation transformations, but direct template matching becomes difficult

Engineering Contradiction:
Improvehandling of scaling and rotation transformationsVSAvoidtemplate matching process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a coordinate transformation intermediary that maps form fields between image space coordinates and document space coordinates. This intermediary layer enables matching without direct template comparison by transforming query coordinates through the same field mapping relationships used in multi-scaling and rotation, thus resolving the contradiction between handling transformations and maintaining simple matching.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If multi-modality technique using semantic information is used, then semantic relationships can be captured, but the process becomes complicated with multiple stages

Engineering Contradiction:
Improvesemantic information preservationVSAvoidprocessing stages
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges spatial coordinate information and semantic information into a unified field mapping relationship. Instead of using separate multi-stage processing, the invention combines both types of information in a single coordinate transformation framework, where the mapping from image coordinates to document coordinates inherently preserves both spatial relationships and semantic meaning, thus reducing complexity while maintaining information integrity.

Inventive Principle:
Principle #5Merging (Combining)

3Device complexity

If traditional template matching is used, then the process is simple, but it is sensitive to scan noise, color degradation, and OCR errors

Engineering Contradiction:
Improvematching process simplicityVSAvoidrobustness to image quality degradation
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent replaces the mechanical template matching process (direct pixel-by-pixel comparison) with a coordinate transformation-based system. Instead of mechanically comparing image patterns that are sensitive to noise and degradation, the invention uses a mathematical mapping model that transforms coordinates through the field relationships, which is inherently more robust to image quality issues while maintaining process simplicity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If image-based matching is used, then spatial information is preserved, but the system is affected by image quality and transformations

Engineering Contradiction:
Improvespatial information accuracyVSAvoidefficacy under image degradation
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent transitions from operating solely in image space to operating in document space through coordinate transformation. By mapping form fields to a standardized document coordinate system, the invention preserves spatial information accuracy while eliminating sensitivity to image-space transformations and degradation, as the document space representation is invariant to these issues.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12340607B2Method and apparatus for form identification and registration
Publication Date: 2025.06.24 KONICA MINOLTA BUSINESS SOLUTIONS USA INC
  • US12340607B2 patent drawing
  • US12340607B2 patent drawing
  • US12340607B2 patent drawing

AI summary

Aspects of the present invention relate to a machine learning system that is trained to identify forms, performing a method that includes receiving a form as an input image; identifying a field in the input image; identifying boundaries of the field; identifying locations of characters in the field; creating a two-dimensional space containing special characters; replacing the special characters with the characters in the field; identifying one or more keywords in the field based on identification of words and/or location of words; and responsive to an indication that the identifying one or more keywords yielded an incorrect result, updating the machine learning system. In another aspect, the machine learning system is used to identify forms, and can identify whether a form requires registration and, if registration is required, performing the registration.