Neural Network Object Detection via Synthetic Data Blending

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods, such as OCR and classical computer vision, are ineffective for detecting and extracting signatures from digitized documents, especially in unstructured documents where the signature's location is not predefined and may overlap with other elements, and deep learning models require significant training data that is challenging to obtain for unique objects like signatures.

Innovation Solution

A system and method that uses a processor coupled with a data lake to generate a training dataset by blending pre-defined representations of objects with pruned original documents, allowing for the training of a neural network model to automate the detection of objects in test documents, including signatures, regardless of their location.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If deep learning neural networks are used for object detection, then detection accuracy can be improved, but the requirement for significant training data increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidtraining data requirement
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by pre-defining object representations and preparing training data in advance through automated blending techniques. The system pre-processes documents by removing objects, then blends them with synthesized object representations to create ready-to-use training datasets before model training begins, reducing the need for extensive manual data collection.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating synthetic training examples through blending pre-defined object representations with pruned document images. Instead of collecting diverse real-world examples manually, the system generates multiple copies and variations of training data by combining standardized object representations with different document backgrounds, effectively multiplying the utility of limited training samples.

Inventive Principle:
Principle #26Copying

2Device complexity

If conventional computer vision methods are used for signature extraction, then the process can be simplified, but detection accuracy deteriorates when objects overlap or are in unstructured locations

Engineering Contradiction:
Improveprocess simplicityVSAvoiddetection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent applies segmentation by dividing the document processing into distinct stages: object removal, background preparation, and selective blending. The system segments the training data generation process into pruned image creation (removing original objects) and controlled object reinsertion, allowing simple processing steps to be combined for sophisticated detection capabilities.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses an intermediary approach by introducing blended images as intermediate training data between simple document scans and complex detection requirements. These blended images serve as mediators that teach the model to recognize objects in various contexts without requiring complex processing of real-world variability.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If training datasets are made customizable for any object location, then detection versatility improves, but dataset generation complexity increases

Engineering Contradiction:
Improvedetection versatilityVSAvoiddataset generation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies universality by creating a multi-functional training data generation system that can handle various object types (signatures, stamps, logos) and document formats through a single blended image generation process. The same blending technique works for different objects and document structures, making the system versatile without requiring separate processing pipelines for each case.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses parameter changes by varying the blending parameters, object positions, and document characteristics in the training data generation process. By changing these parameters systematically, the system generates diverse training examples that teach the model to detect objects in various locations and contexts, achieving versatility through controlled parameter variation rather than complex manual dataset curation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12073642B2Object detection using neural networks
Publication Date: 2024.08.27 ACCENTURE GLOBAL SOLUTIONS LTD
  • US12073642B2 patent drawing
  • US12073642B2 patent drawing
  • US12073642B2 patent drawing

AI summary

Systems and methods for facilitating an automated detection of an object in a test document are disclosed. A system may include a processor including a dataset generator. The dataset generator may obtain a first input image and a first original document from a data lake. The dataset generator may prune a portion of the first original document to obtain a pruned image. The dataset generator may blend the first input image with the pruned image to generate a modified image. The modified image may include the pruned image bearing the first pre-defined representation. The modified image may be combined with the first original document to generate a training dataset. The training dataset may be utilized to train a neural network based model to obtain a trained model for the automated detection of the object in the test document.