3D Molecular Coordinate Extraction from Non-Editable PDFs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods are inefficient in extracting accurate and computable molecular data from complex non-readable documents like PDFs, which are common in scientific literature, due to varied data formats and lack of standardization, leading to difficulties in computational studies.

Innovation Solution

A method that uses pattern recognition to separate and encode 3D molecular coordinates from non-molecular text in PDFs, generating a bond matrix for accurate conversion into standard interoperability formats like SDF and MOL, ensuring reusability and accuracy by comparing with standardized data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If molecular data is stored in PDF format for printing and reading convenience, then document readability and printability are improved, but molecular data is lost or buried and cannot be used for computational studies

Engineering Contradiction:
Improvedocument readabilityVSAvoidmolecular data availability
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent extracts molecular data from PDF documents by identifying and separating atomic coordinate information from other document content. The system locates coordinate data in supplementary information files, extracts the relevant molecular information, and converts it to computable formats while leaving the original PDF intact for reading purposes.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary conversion process that transforms molecular coordinate data from PDF format into intermediate computable formats. This mediator system enables the molecular data to be used for computational studies without compromising the original document's readability and printability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If existing methods are used to extract molecular data from PDFs, then data extraction is attempted, but accuracy and computational usability are compromised due to varied data formats and lack of standardization

Engineering Contradiction:
Improvedata extraction efficiencyVSAvoidmolecular data accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent changes the parameters of data extraction by specifically targeting atomic coordinate information in standardized formats within PDFs. The system adjusts extraction parameters to recognize and parse coordinate data according to specific patterns, improving both extraction efficiency and data accuracy for computational use.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11456061B2Method for harvesting 3D chemical structures from file formats
Publication Date: 2022.09.27 COUNCIL OF SCI & IND RES
  • US11456061B2 patent drawing
  • US11456061B2 patent drawing
  • US11456061B2 patent drawing

AI summary

A method and system for harvesting molecular structures from non-editable documents is disclosed herein. A non-editable storage document is fed by a feeder which is received by a receiver. The molecular and non-molecular data contained in the non-editable storage document is recognized. The three-dimensional coordinates of the molecular data is separated using a pattern recognition. The molecular coordinates are encoded by a pattern sequence. A bond matrix data of the encoded data is generated. Subsequently the bond matrix data for accuracy is verified by comparing with a stored standardized data into a library.