3D Molecular Coordinate Extraction from Non-Editable PDFs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods are inefficient in extracting accurate and computable molecular data from complex non-readable documents like PDFs, which are common in scientific literature, due to varied data formats and lack of standardization, leading to difficulties in computational studies.
Innovation Solution
A method that uses pattern recognition to separate and encode 3D molecular coordinates from non-molecular text in PDFs, generating a bond matrix for accurate conversion into standard interoperability formats like SDF and MOL, ensuring reusability and accuracy by comparing with standardized data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If molecular data is stored in PDF format for printing and reading convenience, then document readability and printability are improved, but molecular data is lost or buried and cannot be used for computational studies
Solution Approach 1:
The patent extracts molecular data from PDF documents by identifying and separating atomic coordinate information from other document content. The system locates coordinate data in supplementary information files, extracts the relevant molecular information, and converts it to computable formats while leaving the original PDF intact for reading purposes.
Solution Approach 2:
The patent introduces an intermediary conversion process that transforms molecular coordinate data from PDF format into intermediate computable formats. This mediator system enables the molecular data to be used for computational studies without compromising the original document's readability and printability.
2Productivity
If existing methods are used to extract molecular data from PDFs, then data extraction is attempted, but accuracy and computational usability are compromised due to varied data formats and lack of standardization
Solution Approach 1:
The patent changes the parameters of data extraction by specifically targeting atomic coordinate information in standardized formats within PDFs. The system adjusts extraction parameters to recognize and parse coordinate data according to specific patterns, improving both extraction efficiency and data accuracy for computational use.
Data Source
AI summary
A method and system for harvesting molecular structures from non-editable documents is disclosed herein. A non-editable storage document is fed by a feeder which is received by a receiver. The molecular and non-molecular data contained in the non-editable storage document is recognized. The three-dimensional coordinates of the molecular data is separated using a pattern recognition. The molecular coordinates are encoded by a pattern sequence. A bond matrix data of the encoded data is generated. Subsequently the bond matrix data for accuracy is verified by comparing with a stored standardized data into a library.


