Small Molecule Fragmentation via pBRICS and Murcko Scaffolds
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing fragmentation methods in drug discovery struggle to effectively handle small molecules and all substituents of a core scaffold, often providing partial explanations and generating a large number of unique fragments, which hinders the prediction of drug-like properties.
Innovation Solution
A processor-implemented method and system that uses a post-processing Breaking of Retrosynthetically Interesting Chemical Substructures (pBRICS) technique for fine-grained fragmentation, combined with a deep learning model trained on domain-aware graphs generated from Murcko scaffolds and Gradient Class Activation Maps (GradCAM) for node-level contribution analysis, to identify significant fragments and optimize molecules.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If existing fragmentation methods (BRICS, RECAP, SynDiR) are used to fragment molecules, then the molecules can be divided into substituents and core scaffolds, but these methods produce a large number of unique fragments and cannot handle very small molecules effectively
Solution Approach 1:
The patent applies segmentation by dividing molecules into core scaffolds and substituents using Murcko scaffold extraction, then further segmenting substituents into smaller meaningful fragments. This hierarchical segmentation approach reduces the total number of unique fragments while maintaining chemically relevant boundaries, directly addressing the contradiction between fragment quantity and fragmentation accuracy.
Solution Approach 2:
Instead of starting with whole molecules and fragmenting them (which produces too many unique fragments), the patent inverts the approach by first identifying core scaffolds and then extracting substituents relative to these scaffolds. This inverted workflow ensures that small molecules are handled appropriately and reduces the number of unique fragments generated.
2Loss of information
If atom-based explanation methods are used, then explanations can be provided for model predictions, but only a fraction of the fragment is shown to be significant rather than the complete fragment
Solution Approach 1:
The patent extracts complete substituent fragments from molecules relative to identified core scaffolds, rather than relying on atom-based gradient methods that only highlight portions of fragments. This extraction approach ensures that entire chemically meaningful fragments are identified as significant, preserving explanation completeness while maintaining fragment identification accuracy.
Solution Approach 2:
The patent introduces core scaffold identification as an intermediary step between molecular representation and fragment analysis. By first identifying the core scaffold and then defining substituents relative to it, the method provides a structured framework that enables complete fragment identification while maintaining connection to the original molecular structure for accurate explanation.
3Adaptability or versatility
If conventional fragmentation methods are used, then molecules can be divided into components, but they fail to fragment all substituents of a core scaffold including small groups like halogen atoms, methyl and hydroxyl groups
Solution Approach 1:
The patent employs a dynamic and flexible fragmentation approach that adapts to different molecule types and sizes. By using graph convolutional networks to identify core scaffolds and then extracting substituents based on connectivity to these scaffolds, the method can handle anything from very small molecules to large complex molecules, achieving high fragmentation coverage without requiring overly complex specialized procedures for different molecule classes.
Data Source
AI summary
The embodiments of present disclosure herein address the inability of existing techniques to fragment both small molecules and substituents of a core scaffold. It addresses generation of lesser number of unique fragments which hinders application of graph propagation approaches to predict properties from molecular datasets. The method and system for extraction of small molecule fragments and their explanation for drug-like properties. A molecular graph representation is used to train graph convolution network (GCN) models for prediction of various absorption, distribution, metabolism, excretion, and toxicity (ADMET) properties. The models developed are compared with an existing atom-level graph model trained using a similar architecture. Further, the explanations obtained from the predictive models are validated based on their relevance to the existing knowledgebase of substructure contributions using matched molecular pairs (MMP) analysis.


