Deep Learning Framework for Small Molecule Bioactivity Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for predicting the chemical bioactivity of small molecules towards miRNA targets are limited by sparse biological data, leading to ineffective generalization and unreliable predictions, especially when dealing with small datasets.
Innovation Solution
A deep learning framework is developed to predict the biological activity of small molecules on miRNA targets by using a novel objective function that learns chemical space from large chemical structures and expands the size of training sets using unlabeled data, scaled by a parameter α.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If machine learning models are trained on small biological datasets, then the models can be developed for disease targets with limited data, but the models suffer from overfitting, insufficient representation, and poor generalization
Solution Approach 1:
The patent introduces an intermediary approach by using large chemical datasets with unknown biological activity as a bridge between available small labeled datasets and the target prediction task. This intermediary data allows the model to learn general chemical patterns without requiring extensive labeled examples for each specific disease target, thereby improving reliability while maintaining adaptability to small datasets.
Solution Approach 2:
The training process is segmented into two distinct phases: first training on large chemical datasets with unknown biological activity to learn general chemical space representations, then fine-tuning on small labeled datasets for specific disease targets. This segmentation allows the model to acquire general knowledge from abundant data while adapting to specific tasks with limited data, resolving the contradiction between adaptability and reliability.
2Measurement precision
If deep learning models are used to extract complex features from molecular structures, then the models can capture intricate patterns, but the models require large datasets to avoid overfitting and ensure effective generalization
Solution Approach 1:
The patent applies preliminary action by pre-training the deep learning model on large chemical datasets to establish robust feature extraction capabilities before encountering the limitation of small labeled datasets. This preliminary training phase allows the model to learn general molecular patterns and representations that can be transferred to specific disease targets, reducing the amount of task-specific data needed while maintaining high measurement precision.
Solution Approach 2:
The patent changes the training parameters by transitioning from task-specific training on small datasets to general chemical space training on large datasets. This parameter change in the training objective and data distribution allows the model to develop strong feature extraction abilities without requiring extensive task-specific labeled data, thereby achieving high measurement precision with limited quantity of labeled substance.
3Adaptability or versatility
If generative models are trained on large chemical databases to produce novel drug candidates, then the models can explore chemical space, but the models are limited to producing compounds similar to the training set and often generate non-synthesizable compounds
Solution Approach 1:
The patent uses chemical datasets with unknown biological activity as an intermediary training resource that provides diverse chemical structures without the constraints of known drug-like properties. This intermediary training data allows the model to explore broader chemical space and generate more novel structures that are not merely variations of existing drugs, while still maintaining reasonable synthesizability by learning from real chemical data rather than purely synthetic generation.
Data Source
AI summary
A computer implemented method may collect a set of known drug compounds from a database, and create augmented datasets with labeled and unlabeled chemical data and biological data. A computer implemented method may train a neural network using the expanded training set to generate a prediction score of bioactivity of candidate drug compounds towards the biological target, wherein the contribution of the second plurality of unlabeled samples in training the neural network is scaled by a parameter. A computer implemented method may output one or more refined neural network models capable of generating candidate drug compounds with predicted activity. A computer implemented method may generate a candidate drug compound by inputting a candidate drug compound with chemical data into the trained neural network model.


