OCSR tools convert images of chemical structures into machine-readable formats such as SMILES, InChIKey and MOL files. In practice you have three options: redraw structures by hand in ChemDraw or Ketcher, run an open-source model such as OSRA, DECIMER or MolScribe, or use a hosted service like Mollens that takes a PDF in and returns a reviewed, source-linked molecule package. The right fit depends on volume, how much setup you can absorb and how much provenance you need.
Key Takeaways
- Manual redrawing is accurate for a handful of structures but scales poorly, and provenance lives only in your own notes.
- Open-source OCSR models are free and scriptable, which suits research and custom pipelines, but you own setup, PDF handling, segmentation and review.
- Mollens is a hosted workflow: upload a PDF, extract, review each structure against its highlighted source region, then export SDF, SMILES, CSV or PNG.
- Published accuracy figures for any tool vary by dataset, so test candidates on your own documents.
- Every approach still needs a human check against the source image.
What should chemical structure recognition software actually do?
Optical chemical structure recognition (OCSR) is the automated conversion of a 2D chemical structure drawing into a machine-readable molecular representation. For a deeper primer, see our explainer on what OCSR is and how it works.
Recognizing a single cropped image is only one step. Extracting structures from a real patent or paper involves a longer chain:
- Page handling: rendering PDF pages, including scans, at usable resolution.
- Layout and segmentation: finding each structure on a crowded page and separating it from text and tables.
- Recognition: turning each drawing into atoms and bonds, then SMILES or a MOL block.
- Identifiers and validation: generating a formula and InChIKey, and checking the result is chemically sensible.
- Review and provenance: linking each structure back to its page and position.
- Export: writing SDF, SMILES or CSV for downstream tools.
When you compare OCSR tools or any chemical structure recognition software, ask which of these steps each option covers and which ones you would have to build or do by hand.
Option 1: Manual redrawing in ChemDraw or Ketcher
Manual redrawing means a chemist looks at each structure and rebuilds it in a sketcher such as ChemDraw or the open-source Ketcher, then exports SMILES or MOL.
Strengths
- High control. An experienced chemist interprets ambiguous drawings, abbreviations and stereochemistry with judgment.
- No new software if your team already uses a sketcher.
Trade-offs
- Time. Medicinal chemistry patents often contain dozens to hundreds of example compounds.
- Transcription errors. A missed methyl or a flipped wedge is easy to introduce and hard to spot later.
- Manual provenance. Page and position references exist only if someone records them.
Manual work remains a sensible choice for a few structures, or as the final check on automated output.
Option 2: Open-source OCSR models (OSRA, DECIMER, MolScribe)
Three widely cited open-source projects:
- OSRA, developed at the US National Cancer Institute, is a rule-based recognizer that vectorizes structure images and rebuilds the molecular graph (Filippov and Nicklaus, J Chem Inf Model, 2009).
- DECIMER is a deep learning project that translates structure images into SMILES and includes a segmentation component for locating structures on pages (Rajan et al., J Cheminform).
- MolScribe is an image-to-graph model that predicts atoms, bonds and their coordinates, then assembles the molecule (Qian et al., J Chem Inf Model, 2023).
Strengths
- Free to use under open-source licenses. Check each project’s terms before commercial use.
- Scriptable. You can call them from Python or the command line and drop them into a cheminformatics pipeline.
- Transparent and self-hostable. Code and methods are public, which helps research and method comparison.
Trade-offs
- Setup. Expect to manage environments, dependencies, model weights and sometimes a GPU.
- Page-level work. Several models expect a cropped structure image. Rendering PDFs, segmenting pages and filtering non-structure regions may be separate steps you assemble yourself.
- DIY review. A side-by-side view of each prediction and its source region, plus page-level provenance, usually needs your own tooling.
- Data-dependent accuracy. Published results vary by benchmark dataset, image quality and drawing style, so numbers from one paper rarely transfer directly to your patents.
For a cheminformatics group with engineering time, these models are a strong foundation. For an analyst who needs a clean SDF this week, the assembly work can outweigh the savings.
Option 3: Mollens, a hosted PDF-to-package workflow
Mollens is Patsnap’s AI chemical structure extraction tool that turns chemical PDFs into source-linked molecule packages. It runs on an in-house OCSR model (V1.0) and is available through self-service sign-up in Eureka Life Science.
The workflow covers the full chain described above:
- Upload a PDF, PNG, JPG or JPEG, as a single file or a batch, up to 50 MB per file.
- Extract. Mollens finds and recognizes the structures on each page.
- Review. Click any structure to jump to its highlighted source region on the original page.
- Save to Library or export as SDF, SMILES, CSV or PNG.
Each structure comes with a 2D image, molecular formula, SMILES, InChIKey, MOL/SDF data, source page number and highlighted source region, plus built-in PubChem lookup. The “All” export is a zip with the source PDF, large PNG structure images, CSV, SDF, SMILES and an offline reader.html that opens without Mollens.
Optional fragment extraction returns scaffolds or cores, substituents and linkers. These are partial structures, not complete compounds. Full Markush interpretation is not supported, so generic claims still need expert reading.
In Patsnap’s internal evaluation (June 2026) on 800 patent pages with 8,816 source structures, Mollens reached recall above 95% and precision above 99%, matched by page, InChIKey and IoU localization. Results should still be checked against the source image. This evaluation was not run on the same set as third-party tools, so it is not a head-to-head comparison.
Pricing is pay-per-structure: 10,000 free Credits at sign-up, 10 Credits per successfully extracted structure (about 1,000 free structures), and top-ups at $1 for 100 Credits (about $0.10 per structure). Nothing is charged when no structure is recognized. Uploads stay in an authenticated workspace, and shared pages exist only when you explicitly share.
The trade-off: Mollens is a web workspace, not a code library. The Mollens product overview and FAQ covers features in more detail, or you can try Mollens on your own PDF.
See it on a real patent
The live demo extracts 15 structures from WO2023023255A1 with page-level highlights. Upload your own PDF and the first ~1,000 structures are covered by free Credits.
How do OCSR tools compare side by side?
| Criterion | Manual redrawing | Open-source OCSR models | Mollens |
|---|---|---|---|
| Setup | Existing sketcher (ChemDraw license or free Ketcher) | Install environments, dependencies and model weights; a GPU may help | Browser sign-up; no installation |
| PDF and layout handling | A human reads the page | Varies by project; rendering and segmentation are often separate steps | Built in: PDF and image upload, page-level detection, batch processing |
| Source traceability | Only if recorded manually | Usually DIY; coordinates may be available, the review UI is up to you | Page number and highlighted source region for every structure |
| Export formats | Whatever the sketcher supports, such as SMILES or MOL | Typically SMILES or MOL from code; packaging is up to you | SDF, SMILES, CSV, PNG, plus an “All” zip with an offline reader.html |
| Review interface | The sketcher itself | Build your own | Click-to-source review, Library, similarity search |
| Cost model | Chemist time | Free software; compute and engineering time | Pay per successfully extracted structure; 10,000 free Credits |
| Best for | A few structures, or final QC | Research, method development, custom pipelines | Patent and paper extraction where analysts need traceable output quickly |
How to choose the best way to extract structures from patents
Use this checklist before you commit:
- Count the structures. Under a dozen, redrawing may be fastest. Hundreds across many documents favors automation.
- Check who will run it. With Python and DevOps support, open-source models are viable. If the users are chemists or analysts, a hosted tool avoids setup.
- Decide how much provenance you need. For SAR, prior art or competitive work, source traceability matters.
- Look at your documents. Scans and dense example tables stress page handling more than recognition.
- Define the output: SDF, SMILES, CSV or images, and whether colleagues without the tool must open it.
- Test on your own pages. Published results vary by dataset, so run a representative patent through each candidate and check it against the source.
- Price the whole job. Include chemist and engineering hours, not only license or Credit costs.
Many teams combine approaches: an automated image to SMILES tool for the bulk and manual redrawing for the few structures that need judgment. Our guide to chemical structure extraction from patents walks through the end-to-end workflow, and the PDF to SDF and SMILES how-to covers export in detail.
Note: structure extraction supports, but does not replace, review by a qualified patent professional.
Frequently Asked Questions
What are OCSR tools used for?
OCSR tools convert chemical structure drawings in patents, papers and scanned documents into machine-readable data such as SMILES, InChIKey and MOL or SDF files. Medicinal chemists use them to build SAR tables, IP analysts use them to prepare prior art and freedom-to-operate searches, and BD teams use them to map competitor chemical matter. Output should always be reviewed against the source image.
Are DECIMER, MolScribe and OSRA free?
DECIMER, MolScribe and OSRA are open-source projects you can download and run without license fees, subject to each project’s license terms. The real costs are setup, compute and the engineering time needed for PDF rendering, page segmentation, validation and review. For research groups with that capacity they are a good fit. Teams without it often prefer a hosted workflow.
Which OCSR tool is the most accurate?
There is no single answer. Published accuracy for OCSR models varies by benchmark dataset, image quality and drawing style, and tools are rarely evaluated on the same documents. Patsnap’s internal evaluation (June 2026) reported Mollens recall above 95% and precision above 99% on 800 patent pages. The reliable approach is to test candidates on your own documents and check results against the source.
Can I convert a chemical structure image to SMILES for free?
Yes. Open-source models such as DECIMER, MolScribe and OSRA can convert a structure image to SMILES at no license cost if you set them up locally. Mollens gives 10,000 free Credits at sign-up, which covers about 1,000 successfully extracted structures at 10 Credits each, and no Credits are charged when no structure is recognized.
Does Mollens handle Markush structures?
Mollens extracts complete discrete structures and can optionally extract fragments such as scaffolds, substituents and linkers. Fragments are partial structures, not complete compounds. Full Markush interpretation, including enumeration of generic claims, is not supported. For Markush-heavy claims, use Mollens to capture the drawn cores and example compounds, then have an expert interpret the variable groups.
Do I still need to review automated OCSR output?
Yes. Even strong recognition models can misread stereochemistry, abbreviations, overlapping labels or low-quality scans. Review is faster when each structure links back to its source. In Mollens you click a structure to jump to its highlighted region on the original page. With open-source models, plan to build a comparable side-by-side view so reviewers can confirm each result.
The quickest way to compare OCSR tools is on your own pages. Run one representative patent through Mollens alongside your current method and check each structure against its source. Try Mollens free →