Machine-Readable Code Coordinates for Document Metadata Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing information retrieval techniques are inefficient in extracting only the required information from unstructured or structured data sources, leading to process overhead and increased time consumption due to the need for extensive preprocessing.

Innovation Solution

A computer-implemented method and system that scans machine-readable codes from documents to determine metadata coordinates, extracts relevant information, identifies the type of information, and sends it to devices for specific actions, using image processing techniques to classify and recognize the information as pictures, numbers, or symbols.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If existing information retrieval techniques are used to extract all information from data sources, then complete information is obtained, but process overhead and time consumption increase due to extensive preprocessing

Engineering Contradiction:
Improveinformation completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent extracts only the required specific information from documents using machine-readable codes as guides, rather than extracting all information. The machine-readable code contains coordinates that directly point to the location of needed metadata, allowing the system to extract only relevant data elements and skip unnecessary preprocessing of unrelated content

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The machine-readable code is embedded in the document in advance with pre-calculated coordinates of the metadata. This preliminary action allows the extraction system to immediately locate required information without performing extensive search and preprocessing operations, thereby reducing processing time while maintaining information completeness

Inventive Principle:
Principle #10Preliminary action

2Productivity

If machine-readable code with coordinates is used to extract specific metadata, then processing efficiency improves, but the document structure becomes more complex

Engineering Contradiction:
Improvedata extraction efficiencyVSAvoiddocument structure complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The machine-readable code serves as an intermediary element that bridges the document content and the extraction system. It contains coordinate information that mediates between the physical document structure and the digital extraction process, enabling efficient location of metadata without requiring the extraction system to parse the entire document structure

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent uses machine-readable codes that encode coordinate information as a simplified copy or representation of the metadata locations. Instead of implementing complex structural markers throughout the document, the system creates a compact coordinate copy that points to the actual data locations, reducing overall document structure complexity while maintaining extraction efficiency

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11216798B2System and computer implemented method for extracting information in a document using machine readable code
Publication Date: 2022.01.04 VISA INTERNATIONAL SERVICE ASSOCIATION
  • US11216798B2 patent drawing
  • US11216798B2 patent drawing
  • US11216798B2 patent drawing

AI summary

The present disclosure relates to a system and computer-implemented method for extracting information in a document using machine-readable code. The machine-readable code is scanned from a document for determining coordinates of one or more metadata among a plurality of metadata in the document. Information corresponding to the coordinates of the one or more metadata is extracted from the document. The type of the information extracted from the one or more metadata is identified. Finally, the identified information is sent to the one or more devices for performing one or more actions using the identified information.