Machine-Readable Code Coordinates for Document Metadata Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing information retrieval techniques are inefficient in extracting only the required information from unstructured or structured data sources, leading to process overhead and increased time consumption due to the need for extensive preprocessing.
Innovation Solution
A computer-implemented method and system that scans machine-readable codes from documents to determine metadata coordinates, extracts relevant information, identifies the type of information, and sends it to devices for specific actions, using image processing techniques to classify and recognize the information as pictures, numbers, or symbols.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If existing information retrieval techniques are used to extract all information from data sources, then complete information is obtained, but process overhead and time consumption increase due to extensive preprocessing
Solution Approach 1:
The patent extracts only the required specific information from documents using machine-readable codes as guides, rather than extracting all information. The machine-readable code contains coordinates that directly point to the location of needed metadata, allowing the system to extract only relevant data elements and skip unnecessary preprocessing of unrelated content
Solution Approach 2:
The machine-readable code is embedded in the document in advance with pre-calculated coordinates of the metadata. This preliminary action allows the extraction system to immediately locate required information without performing extensive search and preprocessing operations, thereby reducing processing time while maintaining information completeness
2Productivity
If machine-readable code with coordinates is used to extract specific metadata, then processing efficiency improves, but the document structure becomes more complex
Solution Approach 1:
The machine-readable code serves as an intermediary element that bridges the document content and the extraction system. It contains coordinate information that mediates between the physical document structure and the digital extraction process, enabling efficient location of metadata without requiring the extraction system to parse the entire document structure
Solution Approach 2:
The patent uses machine-readable codes that encode coordinate information as a simplified copy or representation of the metadata locations. Instead of implementing complex structural markers throughout the document, the system creates a compact coordinate copy that points to the actual data locations, reducing overall document structure complexity while maintaining extraction efficiency
Data Source
AI summary
The present disclosure relates to a system and computer-implemented method for extracting information in a document using machine-readable code. The machine-readable code is scanned from a document for determining coordinates of one or more metadata among a plurality of metadata in the document. Information corresponding to the coordinates of the one or more metadata is extracted from the document. The type of the information extracted from the one or more metadata is identified. Finally, the identified information is sent to the one or more devices for performing one or more actions using the identified information.


