Form Data Acquisition via Dynamic Attribute Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing form data acquisition systems rely on static relationship determination rules and attribute specifications, which are inadequate for accurately processing forms with varying formats and positional attributes, leading to inefficiencies in data recognition and extraction.

Innovation Solution

A system that includes a character string attribute learning unit, attribute-positional relation learning unit, attribute probability acquisition unit, and attribute probability correction unit to create and apply models and rules for determining the probability of attributes and their positional relations in form data, enhancing the accuracy of form data acquisition from images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If static relationship determination rules are used to determine word relationships based on positional relations and dictionary attributes, then the system structure remains simple, but the accuracy of form data acquisition deteriorates when processing forms with varying formats and positional attributes

Engineering Contradiction:
Improveaccuracy of form data acquisitionVSAvoidsystem structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms static relationship determination rules into dynamic learning models. The character string attribute learning unit and attribute-positional relation learning unit dynamically learn from training data to create adaptive models that adjust to varying form formats, replacing fixed rules with flexible, data-driven approaches that improve accuracy while managing complexity through automated learning processes

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes parameters by learning optimal attribute determination thresholds and positional relation weights from training data. The attribute probability acquirement unit calculates probabilities based on learned models rather than fixed parameters, allowing the system to adapt to different form types by adjusting internal parameters through the learning process

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If dynamic learning models are implemented to improve accuracy of attribute determination, then the accuracy of form data acquisition improves, but the device complexity increases due to multiple learning units and models

Engineering Contradiction:
Improveaccuracy of attribute probabilityVSAvoidnumber of learning units and models
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex learning task into distinct functional units: character string attribute learning unit for learning character attributes, attribute-positional relation learning unit for learning positional relationships, attribute probability acquirement unit for calculating probabilities, and attribute probability correction unit for refining results. This segmentation manages complexity by creating specialized, modular components that can be developed and maintained independently

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The attribute probability acquirement unit acts as an intermediary between the learning models and the final attribute determination. It applies the learned character string attribute model to recognize character attributes and calculates probabilities, which are then refined by the correction unit using positional relations, providing a structured intermediate processing step that manages system complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If character recognition is performed on form images followed by attribute determination, then the system can process diverse form formats, but the time required for data acquisition increases due to multiple processing steps

Engineering Contradiction:
Improveability to process diverse form formatsVSAvoidprocessing time for form data acquisition
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary learning actions by training the character string attribute model and attribute-positional relation model on training data before actual form processing. This preliminary training phase enables the system to quickly process diverse forms during operation, as the heavy learning work has already been completed in advance, reducing processing time for actual data acquisition tasks

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system maintains continuous useful action by implementing an iterative learning process where the learning units continuously refine models based on training data, and the correction unit continuously adjusts attribute probabilities based on positional relations. This continuous refinement process improves accuracy over time without requiring complete reprocessing of forms

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS11676409B2Form data acquirement system and non-transitory computer readable recording medium storing form data acquiring program
Publication Date: 2023.06.13 KYOCERA DOCUMENT SOLUTIONS INC
  • US11676409B2 patent drawing
  • US11676409B2 patent drawing
  • US11676409B2 patent drawing

AI summary

An information processing apparatus learns an attribute-positional relation in a form for learning to create an attribute-positional relation rule for a character string in a form, determines correspondence between a character string in a result of character recognition executed on an image of the form for learning and an attribute in form data for learning based on the form for learning to create a character string attribute model for acquiring a probability of an attribute of the character string in the form, applies the character string attribute model to a character string in a result of character recognition executed on an image of the form to acquire the probability of an attribute, and corrects the probability based on a position in the form of the character string in the result of the character recognition executed on the image of the form and the attribute-positional relation rule.