Compound Property Prediction Model Using Spatial Structure Pre-training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current AI-driven drug design methods face challenges in accurately predicting compound properties due to the imbalance in labeled and unlabeled sample compounds, requiring efficient methods to leverage large amounts of unlabeled data for training robust prediction models.

Innovation Solution

A method is proposed that involves training a compound property prediction model by first acquiring spatial structure information of sample compounds and using it to develop a spatial structure prediction model, followed by continuing training with labeled compounds to incorporate property information, effectively utilizing a large number of unlabeled samples to enhance prediction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional deep learning methods are used to train compound property prediction models, then the models can be trained with available labeled data, but the prediction accuracy is limited due to the small number of labeled samples and the inability to effectively leverage unlabeled data

Engineering Contradiction:
Improveprediction accuracyVSAvoidnumber of labeled samples
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by first pre-training the neural network model on a large number of unlabeled compound structures to learn fundamental molecular representation patterns. This preliminary training phase prepares the model with general structural knowledge before it is fine-tuned on the limited labeled property data, thereby improving prediction accuracy despite the small number of labeled samples.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary self-supervised learning task (such as predicting masked atomic properties or molecular properties from structural information) that acts as a bridge between unlabeled structural data and the final property prediction task. This intermediary task enables the model to effectively learn from unlabeled data and transfer the learned representations to the target property prediction task.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If more labeled sample compounds are used to improve prediction accuracy, then the model performance improves, but the cost of data labeling and computing resources increases significantly

Engineering Contradiction:
Improveprediction accuracyVSAvoidcomputing resources
Core Design Contradiction:
Measurement precisionVSUse of energy by stationary object

Solution Approach 1:

The patent extracts and utilizes the valuable structural information contained in unlabeled compound data through self-supervised pre-training. By extracting general molecular representation patterns from unlabeled data, the model achieves better performance without requiring proportional increases in labeled data or computing resources for full supervised training.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The pre-training phase serves as a preliminary action that prepares the model with general molecular knowledge before the actual property prediction task. This separation of learning stages allows efficient use of computing resources by performing the computationally intensive learning on abundant unlabeled data first, then requiring minimal resources for fine-tuning on labeled data.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If a large number of unlabeled compounds are utilized for training, then the model can achieve high accuracy with fewer labeled samples, but the training process becomes more complex

Engineering Contradiction:
Improveprediction accuracyVSAvoidtraining process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the training process into distinct phases: a pre-training phase on unlabeled structural data and a fine-tuning phase on labeled property data. This segmentation simplifies the overall training complexity by breaking down the challenging task of learning from mixed labeled and unlabeled data into manageable stages with clear objectives for each phase.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20220122697A1Method for training compound property prediction model and method for predicting compound property
Publication Date: 2022.04.21 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20220122697A1 patent drawing
  • US20220122697A1 patent drawing
  • US20220122697A1 patent drawing

AI summary

A method for predicting a compound property, apparatuses, an electronic device, a computer readable storage medium, and a computer program product are provided. The method includes: for each first sample compound of first sample compounds, acquiring spatial structure information of a spatial structure formed by atoms and chemical bonds that constitute the first sample compound; training, using the first sample compounds as input samples and pieces of corresponding spatial structure information as output samples, to obtain a spatial structure prediction model; and continuing training, using second sample compounds as input samples and pieces of corresponding property information as output samples, to obtain the compound property prediction model on the basis of the spatial structure prediction model.