Neural Network Data Extraction for Structured Graphic Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for data extraction from structured graphic documents, such as packaging and technical brochures, require significant human intervention, leading to errors, time-consuming processing, and high costs.

Innovation Solution

A data extraction process utilizing a network of neurons trained on structured graphic documents to identify and extract different types of graphic data (texts, tables, signs, graphic codes, and illustrations) with minimal human intervention, employing specific extraction modules for each data type.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If character recognition is used to extract data from structured graphic documents, then text extraction is attempted, but the recognition becomes ineffective or incorrect due to illustrations, different fonts, and text colors

Engineering Contradiction:
Improvecharacter recognition accuracyVSAvoiddata extraction complexity
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments the structured graphic document into multiple distinct data types (text, tables, signs, graphic codes, illustrations) and processes each segment with specialized extraction modules. This segmentation allows the system to handle each data type appropriately rather than attempting uniform character recognition on the entire document, thereby improving extraction accuracy while managing complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary classification step that identifies and categorizes different data types before extraction. This intermediary layer acts as a mediator between the raw graphic document and the extraction processes, routing different data types to appropriate modules (OCR for text, table structure analysis for tables, sign recognition for logos, etc.), thus improving overall extraction precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If human intervention is used for data extraction, then accuracy can be maintained, but the process becomes time-consuming and expensive

Engineering Contradiction:
Improveextraction accuracyVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements self-service extraction modules that autonomously process different data types without requiring manual intervention. Each module (text extraction, table extraction, sign recognition, graphic code reading) is designed to independently identify and extract its corresponding data type with high accuracy, enabling automated processing that maintains reliability while dramatically improving productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent creates a universal extraction system that handles multiple data types through a single integrated platform. The system uses a common architecture with specialized modules that can process text, tables, signs, and graphic codes uniformly, eliminating the need for separate manual processing workflows and enabling scalable automated extraction across diverse document types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If manual data extraction is performed, then human judgment can be applied, but input errors occur and the process is costly

Engineering Contradiction:
Improveextraction flexibilityVSAvoidinput accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent replaces the mechanical human operation of manual data extraction with automated computational systems. Specialized algorithms and machine learning models substitute for human eyes and hands in identifying and extracting different data types, eliminating human-induced input errors while maintaining the flexibility to handle various document formats and layouts through adaptive processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentEP4531005A1Method for extracting data from a structured graphic document, program product and recording medium for implementing such a method
Publication Date: 2025.04.02 2JDB
  • EP4531005A1 patent drawingFigure 1~2
  • EP4531005A1 patent drawingFigure 3A~3B
  • EP4531005A1 patent drawingFigure 4A~4D

AI summary

The invention relates to a method for extracting data from a structured graphic document capable of comprising at least two distinct types of graphical data selected from among texts, tables, symbols, graphic codes, and illustrations. The method implements such data extraction using at least one neural network previously trained using a training method and a plurality of data extraction modules, each adapted to a respective type of graphical data selected from among texts, tables, symbols, graphic codes, and illustrations. The method includes identifying the areas of the structured graphic document using at least one neural network and selecting the extraction module corresponding to the identified data type.