Compound Property Prediction Model Using Spatial Structure Pre-training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current AI-driven drug design methods face challenges in accurately predicting compound properties due to the imbalance in labeled and unlabeled sample compounds, requiring efficient methods to leverage large amounts of unlabeled data for training robust prediction models.
Innovation Solution
A method is proposed that involves training a compound property prediction model by first acquiring spatial structure information of sample compounds and using it to develop a spatial structure prediction model, followed by continuing training with labeled compounds to incorporate property information, effectively utilizing a large number of unlabeled samples to enhance prediction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional deep learning methods are used to train compound property prediction models, then the models can be trained with available labeled data, but the prediction accuracy is limited due to the small number of labeled samples and the inability to effectively leverage unlabeled data
Solution Approach 1:
The patent applies preliminary action by first pre-training the neural network model on a large number of unlabeled compound structures to learn fundamental molecular representation patterns. This preliminary training phase prepares the model with general structural knowledge before it is fine-tuned on the limited labeled property data, thereby improving prediction accuracy despite the small number of labeled samples.
Solution Approach 2:
The patent introduces an intermediary self-supervised learning task (such as predicting masked atomic properties or molecular properties from structural information) that acts as a bridge between unlabeled structural data and the final property prediction task. This intermediary task enables the model to effectively learn from unlabeled data and transfer the learned representations to the target property prediction task.
2Measurement precision
If more labeled sample compounds are used to improve prediction accuracy, then the model performance improves, but the cost of data labeling and computing resources increases significantly
Solution Approach 1:
The patent extracts and utilizes the valuable structural information contained in unlabeled compound data through self-supervised pre-training. By extracting general molecular representation patterns from unlabeled data, the model achieves better performance without requiring proportional increases in labeled data or computing resources for full supervised training.
Solution Approach 2:
The pre-training phase serves as a preliminary action that prepares the model with general molecular knowledge before the actual property prediction task. This separation of learning stages allows efficient use of computing resources by performing the computationally intensive learning on abundant unlabeled data first, then requiring minimal resources for fine-tuning on labeled data.
3Measurement precision
If a large number of unlabeled compounds are utilized for training, then the model can achieve high accuracy with fewer labeled samples, but the training process becomes more complex
Solution Approach 1:
The patent segments the training process into distinct phases: a pre-training phase on unlabeled structural data and a fine-tuning phase on labeled property data. This segmentation simplifies the overall training complexity by breaking down the challenging task of learning from mixed labeled and unlabeled data into manageable stages with clear objectives for each phase.
Data Source
AI summary
A method for predicting a compound property, apparatuses, an electronic device, a computer readable storage medium, and a computer program product are provided. The method includes: for each first sample compound of first sample compounds, acquiring spatial structure information of a spatial structure formed by atoms and chemical bonds that constitute the first sample compound; training, using the first sample compounds as input samples and pieces of corresponding spatial structure information as output samples, to obtain a spatial structure prediction model; and continuing training, using second sample compounds as input samples and pieces of corresponding property information as output samples, to obtain the compound property prediction model on the basis of the spatial structure prediction model.


