Compound activity prediction based on Boltzmann noise polymorphism SMILE type training large model
Through chemical rules-driven SMILE polymorphism expansion and Boltzmann noise injection, combined with dynamic low-rank fine-tuning technology, the problems of data scarcity and chemical polymorphism neglect in compound activity prediction are solved, and high-precision and efficient pEC50 value prediction are achieved.
Patent Information
- Application Number
- CN202510185355.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-27
AI Technical Summary
Existing compound activity prediction methods face the problems of data scarcity and neglect of chemical polymorphisms, which leads to the model being easily overfitted and unable to effectively characterize the real form of the compound.
Through SMILE polymorphism expansion based on chemical rules, semantic equivalent SMILE variants are generated, the data volume is increased and the polymorphic characteristics of the compound are considered; at the same time, Boltzmann distributed noise is introduced to simulate physiological thermodynamic fluctuations, and dynamic low-rank fine-tuning technology is combined to achieve efficient adaptation of the large model.
High-precision pEC50 value prediction is achieved, which improves the stability and physiological correlation of the model, reduces resource consumption, and provides reliability evaluation through polymorphic prediction distribution.
Smart Images

Figure SMS_2
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of computational chemistry and artificial intelligence, and particularly relates to a method for predicting compound activity based on a pre - trained large - language model, and more particularly to a technical solution for achieving high - precision pEC50 value prediction through Boltzmann noise injection and SMILE - style polymorphism expansion. Background Art
[0002] Compound activity prediction is a core part of compound property research, especially in drug discovery. Traditional QSAR models rely on manual feature engineering and have limited generalization ability. In recent years, although deep - learning models based on SMILE - style (such as ChemBERTa) have made progress, they still face two major challenges: Data scarcity: The experimentally measured pEC50 data is limited, resulting in the model being prone to overfitting.
[0003] Ignoring chemical polymorphism: A single SMILE - style cannot represent the true existing forms of compounds such as stereoisomerism and tautomerism.
[0004] Although existing methods (such as data augmentation and transfer learning) partially alleviate the above problems, they have the following defects: The SMILE - styles generated by random augmentation may violate chemical rationality; The influence of molecular thermal motion on binding free energy in the physiological environment is not considered.
[0005] The present invention solves the above problems through the following innovations: (1) SMILE - style polymorphism expansion based on chemical rules to ensure the rationality of data augmentation; (2) Introducing Boltzmann distribution noise to simulate physiological thermodynamic fluctuations; (3) Combining dynamic low - rank fine - tuning technology to achieve efficient adaptation of large models. Summary of the Invention
[0006] The present invention proposes a method for predicting compound activity by integrating chemical polymorphism and thermodynamic noise. Through data - driven polymorphism expansion, physically inspired noise injection, and parameter - efficient model fine - tuning, high - precision pEC50 value prediction is achieved. The following elaborates in detail from three aspects: core technical solutions, implementation details, and innovation mechanisms:
[0007] The data augmentation strategy of the present invention is based on generating semantically equivalent SMILE - style variants through systematic structural transformation to greatly increase the data volume. By considering the polymorphism characteristics of compounds in the real environment, on the one hand, overfitting caused by exactly the same pEC50 is avoided, and on the other hand, the true numerical situation of compounds is reproduced. The specific implementation is as follows:
[0008] Using the EnumerateStereoisomers module of RDKit, traverse all chiral centers and double bond geometric isomerism sites in the molecule to generate all possible stereoisomers.
[0009] Retain the key chiral centers of the pharmacophore (such as the drug-target binding site) through the Tetrahedral Stereochemistry algorithm, randomly flip the non-key chiral centers, and generate a large amount of data to meet the needs of large model training.
[0010] Based on a predefined SMARTS pattern library (covering 30 common tautomers such as keto-enol and amine-imine), use the TautomerEnumerator module for structure enumeration.
[0011] Retain stable configurations with ΔE < 2 kcal / mol through energy screening to avoid generating high-energy unstable forms.
[0012] For cyclic compounds (such as benzene rings and piperidine rings), apply the RingConformationEnumerator module to generate ring plane flip conformations and retain low-strain conformations with RMSD < 0.5 Å.
[0013] Calculate the ECFP4 fingerprints (radius = 2, 2048 bits) of the original SMILE formula and all variants, and use the Tanimoto coefficient to evaluate the pharmacophore similarity.
[0014] Eliminate variants with a Tanimoto coefficient less than 0.85 to ensure that the extended isomers retain the core pharmacophore features.
[0015] Technical effect: Through the above steps and controlling the maximum value, a single SMILE formula can be extended to more than 500 chemically reasonable variants, the data volume is increased by 500 times, and the pharmacophore consistency is maintained.
[0016] To simulate the influence of molecular thermal motion on the binding free energy in a physiological environment, the present invention designs a noise injection mechanism based on statistical physics:
[0017] Based on molecular dynamics (MD) simulation data (AMBER force field, 310K, 100 ns sampling), statistically obtain the distribution characteristics of ΔG, and its probability density function is:
[0018]
[0019] Among them, μ = 0, σ = 0.75 kcal / mol, and it is verified to be consistent with the simulation data through the K-S test (p = 0.62 > 0.05).
[0020] According to the linear response theory, combined with the free energy change Δ G The relationship with pEC50 is as follows: Substitute the physiological conditions ( R = 1.987 cal / (mol·K), T = 310 K), and the conversion coefficient 0.000733 mol / cal is obtained The actual correction formula simplifies the corrected pEC50 to be equal to the original pEC50 + 0.733Δ G。
[0021] After noise injection into the Tox21 test set, the coefficient of variation (CV) of the pEC50 value increased from 0.12 to 0.35, approaching the experimental measurement error range (CV = 0.28 - 0.42), verifying the rationality of the noise level.
[0022] To achieve efficient adaptation of large language models, the present invention proposes an improved low-rank fine-tuning technique, and the key innovations are as follows:
[0023] The recommended ranks for each layer in the Transformer are {"q_proj": 32, "k_proj": 16, "v_proj": 16, "o_proj": 32, "gate_proj": 8, "up_proj": 4, "down_proj": 4}.
[0024] The main loss function is MAE, and L2 regularization is introduced to constrain the low-rank matrix to prevent overfitting.
[0025] Adopt a staged learning rate schedule: The first 10 epochs: fix the learning rate at 2e-5 to warm up the model; Subsequent epochs: cosine annealing schedule with a minimum learning rate of 1e-6.
[0026] Gradient clipping (max_norm = 1.0) and mixed precision training (FP16) accelerate convergence.
[0027] Technical effect: On the LLaMA-7B model, DORA fine-tuning only requires training 0.3% of the parameters (about 2.3M), the training speed is 5.2 times faster than full-parameter fine-tuning, and the memory occupancy is reduced by 73%.
[0028] The fine-tuned model can predict compound activity through the following steps:
[0029] Input a single SMILE formula, automatically perform stereoisomerism, tautomerism, and ring flipping operations to generate 500 variants.
[0030] Input 500 variants into the model and batch-calculate the predicted pEC50 values of each variant, which takes less than 0.5 seconds under GPU acceleration.
[0031] Calculate the mean and standard deviation of the predicted values of all variants; The final output is in the form of "mean ± standard deviation" (such as 7.2 ± 0.3), reflecting the uncertainty of the prediction results.
[0032] Key conclusions: The MAE of the present invention is reduced by 76% compared with the traditional QSAR model; The polymorphism expansion improves the prediction stability; The DORA fine-tuning, while maintaining the accuracy, only consumes 1 / 300 of the resources of the full fine-tuning.
[0033] Chemistry-driven data augmentation: Break through the bottleneck of chemical irrationality of traditional random amplification through rule-constrained polymorphism expansion; Physics-inspired noise model: Quantify the thermodynamic fluctuations into computable ΔG noise to improve the physiological relevance of the prediction; Parameter-efficient fine-tuning architecture: Achieve the lightweight of the large model through dynamic low-rank adaptation and solve the resource limitation problem in the drug discovery scenario; Uncertainty quantification output: Provide reliability assessment through the polymorphism prediction distribution to assist drug chemists in decision-making.
Claims
1. A method for predicting compound activity based on a Boltzmann noise multi-state SMILE training model, characterized in that: The following steps are involved: Step 1: Data collection and preprocessing ○ The activity data of compounds with R² values greater than 0.8 were screened from the Tox21 data set, and a corresponding relationship data set between the SMILE formula and the pEC50 value was established; ○ Taking advantage of the chemical polymorphism of SMILE formulas, each original SMILE formula is expanded into 500 semantically equivalent but structurally isomeric SMILE formulas through stereoisomerism, tautomerism, and ring flipping operations; Step 2: Boltzmann noise generation and pEC50 value correction ○ Based on the basic molecular internal energy of water molecules at physiological body temperature (1.5 kcal / mol), a random noise value ΔG that conforms to the Boltzmann distribution is generated, with a range of 0 ± 1.5 kcal / mol; ○ Apply ΔG correction to the pEC50 value corresponding to each expanded SMILE formula; Step 3: Polymorphism dataset construction ○ The expanded SMILE formula and its corrected pEC50 values are used as a training dataset to ensure that the 500 isomers corresponding to the same original SMILE formula have differentiated pEC50 labels; Step 4: Fine-tune the large language model ○ Use a pre-trained large language model (LLM) as the infrastructure, with SMILE-style strings as input and predicted pEC50 values as output; ○ Apply the dynamic optimization rank adaptation (DORA) fine-tuning technique to dynamically adjust the model parameters through low-rank decomposition, and the objective function is to minimize the mean absolute error (MAE) between the predicted pEC50 and the corrected pEC50; ○ During the fine-tuning process, the underlying parameters of the model are frozen, and only the top adaptation layer is fine-tuned. The fine-tuning parameters include: r=16 of the low-rank matrix, LoRA_alpha=64, learning rate lr=2e-5, and training batch size batch_size=16; Step 5: Model validation and iterative optimization ○ The data set is divided into training set, validation set and test set according to the ratio of 8:1:
1. The training is terminated when the MAE of the test set is less than 1.
5. ○ After training, the model can be used directly to predict compound activity.
2. The method according to claim 1, characterized in that The SMILE expansion method includes: ○ Adopt the chemical equivalent transformation rules in the RDKit toolkit and use SMARTS pattern matching to ensure that the expanded SMILE formula has the same pharmacophore core structure as the original formula; ○ Eliminate non-essential structural differences through Kekulé form unification and chiral center standardization.
3. The method according to claim 1, characterized in that The Boltzmann noise generation satisfies: ○ The noise value ΔG obeys the Boltzmann distribution function, and its consistency with the thermodynamic distribution is verified by the Kolmogorov-Smirnov test (p>0.05).
4. The method according to claim 1, characterized in that The DORA fine-tuning technology includes: ○ Insert low-rank adaptation matrix in Transformer layer; ○ The cosine annealing learning rate scheduling strategy is adopted, with an initial learning rate of 2e-5 and a period of 50 epochs.
5. The method according to claim 1, characterized in that When the model is inferring: ○ After inputting a single SMILE formula, 500 polymorphic variants are automatically generated and the pEC50 values are predicted for each; ○ The final activity prediction result is pEC50.