Deep Learning Framework for Small Molecule Bioactivity Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for predicting the chemical bioactivity of small molecules towards miRNA targets are limited by sparse biological data, leading to ineffective generalization and unreliable predictions, especially when dealing with small datasets.

Innovation Solution

A deep learning framework is developed to predict the biological activity of small molecules on miRNA targets by using a novel objective function that learns chemical space from large chemical structures and expands the size of training sets using unlabeled data, scaled by a parameter α.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If machine learning models are trained on small biological datasets, then the models can be developed for disease targets with limited data, but the models suffer from overfitting, insufficient representation, and poor generalization

Engineering Contradiction:
ImproveAbility to train on small datasetsVSAvoidPrediction reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces an intermediary approach by using large chemical datasets with unknown biological activity as a bridge between available small labeled datasets and the target prediction task. This intermediary data allows the model to learn general chemical patterns without requiring extensive labeled examples for each specific disease target, thereby improving reliability while maintaining adaptability to small datasets.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The training process is segmented into two distinct phases: first training on large chemical datasets with unknown biological activity to learn general chemical space representations, then fine-tuning on small labeled datasets for specific disease targets. This segmentation allows the model to acquire general knowledge from abundant data while adapting to specific tasks with limited data, resolving the contradiction between adaptability and reliability.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If deep learning models are used to extract complex features from molecular structures, then the models can capture intricate patterns, but the models require large datasets to avoid overfitting and ensure effective generalization

Engineering Contradiction:
ImproveFeature extraction accuracyVSAvoidDataset size requirement
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by pre-training the deep learning model on large chemical datasets to establish robust feature extraction capabilities before encountering the limitation of small labeled datasets. This preliminary training phase allows the model to learn general molecular patterns and representations that can be transferred to specific disease targets, reducing the amount of task-specific data needed while maintaining high measurement precision.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the training parameters by transitioning from task-specific training on small datasets to general chemical space training on large datasets. This parameter change in the training objective and data distribution allows the model to develop strong feature extraction abilities without requiring extensive task-specific labeled data, thereby achieving high measurement precision with limited quantity of labeled substance.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If generative models are trained on large chemical databases to produce novel drug candidates, then the models can explore chemical space, but the models are limited to producing compounds similar to the training set and often generate non-synthesizable compounds

Engineering Contradiction:
ImproveChemical space exploration capabilityVSAvoidCompound synthesizability
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent uses chemical datasets with unknown biological activity as an intermediary training resource that provides diverse chemical structures without the constraints of known drug-like properties. This intermediary training data allows the model to explore broader chemical space and generate more novel structures that are not merely variations of existing drugs, while still maintaining reasonable synthesizability by learning from real chemical data rather than purely synthetic generation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250046404A1Systems and methods relating to a machine learning prediction model for predicting molecular targets of small molecules
Publication Date: 2025.02.06 KBR WYLE SERVICES LLC
  • US20250046404A1 patent drawing
  • US20250046404A1 patent drawing
  • US20250046404A1 patent drawing

AI summary

A computer implemented method may collect a set of known drug compounds from a database, and create augmented datasets with labeled and unlabeled chemical data and biological data. A computer implemented method may train a neural network using the expanded training set to generate a prediction score of bioactivity of candidate drug compounds towards the biological target, wherein the contribution of the second plurality of unlabeled samples in training the neural network is scaled by a parameter. A computer implemented method may output one or more refined neural network models capable of generating candidate drug compounds with predicted activity. A computer implemented method may generate a candidate drug compound by inputting a candidate drug compound with chemical data into the trained neural network model.