Optical Structure Recognition Module for Heterogeneous Document Chemical Structures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies are inadequate in systematically processing multiple heterogeneous types of documents to identify chemical structures, often missing molecule images due to limited capabilities in handling unstructured information and specific image types.
Innovation Solution
The development of an apparatus and method that uses an optical structure recognition module to identify and compile chemical structures from various electronic file types, including non-embedded images, by applying filters to eliminate false positives and associating confidence factors with derived chemical structure objects, enabling robust processing and searching across diverse document formats.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing technologies search for specific image types (e.g., OLE images) in specific document types, then they can identify some chemical structures, but they miss many molecule images and cannot systematically process multiple heterogeneous document types
Solution Approach 1:
The patent implements a universal document processing system that can handle multiple heterogeneous document types (Word, PDF, PowerPoint, etc.) through a unified architecture. The system uses multiple image extraction modules that can process different document formats, and a comprehensive filtering system that validates chemical structures across all document types, thereby achieving both versatility in document handling and reliability in structure identification
Solution Approach 2:
The system employs feedback mechanisms through multiple filtering stages (size filter, shape filter, chemical validity filter) that continuously validate and refine the extracted chemical structures. The confidence scoring system provides feedback on the reliability of each identified structure, allowing the system to adjust its processing and ensure accurate identification across diverse document formats
2Productivity
If the system processes all graphical images to identify chemical structures, then it can find more structures, but it generates many false positives from non-chemical graphical elements
Solution Approach 1:
The patent segments the image processing into distinct functional modules: extraction module, size filtering module, shape filtering module, and chemical validity filtering module. Each module performs a specific function in the identification pipeline, allowing the system to process large volumes of images while maintaining precision through progressive filtering that eliminates false positives at each stage
Solution Approach 2:
The patent introduces intermediary filtering mechanisms between image extraction and final identification. The size filter, shape filter, and chemical validity filter act as intermediaries that mediate between the raw extracted images and the final chemical structure identification, ensuring that only valid chemical structures are identified while maintaining high productivity
3Measurement precision
If the system applies multiple filtering criteria to eliminate false positives, then identification accuracy improves, but processing complexity increases
Solution Approach 1:
The patent divides the complex filtering system into separate, manageable modules: size filter, shape filter, and chemical validity filter. Each module handles a specific aspect of validation, making the overall complex system easier to implement, maintain, and optimize while achieving high identification accuracy through the coordinated operation of these segmented filtering components
Data Source
Figure 1
Figure 2
Figure 3
AI summary
In various embodiments, multiple heterogeneous documents are processed to identify structures, such as chemical structures, contained therein, including non-embedded structures. Also described is a graphical user interface that permits a user to search for a structure or substructure within a set of electronic documents, then displays the matching structures as well as the actual pages of the documents on which the matching structures are found. Display of the actual pages allows the user to verify the matches and provides helpful context for the user.