Beta-secretase inhibitor molecular activity prediction method based on SMILES sequence and graph information fusion learning thereof

By using the fusion learning method of SMILES sequence and its graph information in the prediction of molecular activity of β secretase inhibitors, the problems of prediction accuracy and inefficiency in the prior art are solved, and fast and accurate molecular activity prediction is achieved, and the efficiency of compound structure optimization and virtual screening is improved.

CN120072040APending Publication Date: 2025-05-30FUJIAN UNIV OF TRADITIONAL CHINESE MEDICINE

Patent Information

Application Number
CN202510140149.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has difficulties in predicting molecular activity of beta secretase inhibitors, including relying on complex QSAR models and manually designed molecular descriptors, resulting in prediction accuracy and inefficiency.

Method used

Using a fusion learning method based on SMILES sequence and its graph information, the SMILES sequence and molecular graph information are fusion learning to predict the molecular activity of β-secretase inhibitors by designing a BiLSTM-Transformer framework with enhanced hierarchical attention mechanism and a graph coding model with atomic-level feature attention mechanism.

Benefits of technology

It realizes rapid and accurate prediction of the molecular activity of β-secretase inhibitors, improves the efficiency of compound structure optimization and virtual screening, and reduces the cost and time of experimental testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072040A_ABST
    Figure CN120072040A_ABST
Patent Text Reader

Abstract

The invention relates to a beta secretase inhibitor molecular activity prediction method and system based on SMILES sequence and graph information fusion learning thereof, and belongs to the field of inhibitor molecular activity prediction. The method comprises the following steps: firstly, designing a BiLSTM-Transformer framework with an enhanced layered attention mechanism, and coding a functional subsequence derived from an SMILES character string of a beta secretase inhibitor molecule; functional subsequences are extracted from SMILES character strings and are divided into multiple levels on the functional group level, the ion level and the atom level, and understanding of molecular structure details can be enhanced. Secondly, a graph coding model with an atomic-scale feature attention mechanism is designed based on a molecular graph structure generated from an SMILES character string, so that key atoms and chemical bonds in the molecular graph structure are highlighted, and the feature representation capability of the molecular graph is further enhanced. And finally, through a global attention mechanism and a weighting module, carrying out fusion learning on molecular information of different modes to obtain a prediction result of the molecular activity of the beta-secretase inhibitor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of predicting the activity of inhibitor molecules, and particularly relates to a method and system for predicting the activity of β-secretase inhibitor molecules based on the fusion learning of SMILES sequences and their graph information. Only by providing the SIMLES string of the β-secretase inhibitor to be tested, the molecular activity can be quickly and accurately predicted, greatly improving the efficiency of compound structure optimization and virtual screening. Background Art

[0002] β-secretase is an enzyme that is very important in the research of Alzheimer's disease (AD), and its inhibitor is a potential drug target for the treatment of Alzheimer's disease. Over the years, computer-aided drug design has been widely used in the prediction of the activity of β-secretase inhibitor molecules, such as Ponzoni et al. [1] used a QSAR classification model to predict the activity of β-secretase inhibitors; Aswathy et al. [2] utilized R-group search and molecular docking methods, integrating 3D-QSAR and binding modes to predict the molecular activity of β-secretase inhibitors, significantly improving the efficiency of related drug screening. Such methods largely rely on virtual screening technologies such as QSAR models or molecular docking. However, designing an effective QSAR model is a relatively difficult task. Also, traditional machine learning algorithms, such as support vector machine (SVM), random forest (RF), K-nearest neighbor (KNN), etc. have also been applied in the prediction and classification of the activity of β-secretase inhibitor molecules [3] . However, these methods usually rely on manually designed molecular descriptors or extracted molecular fingerprint features, and still face major limitations in capturing the complex relationship features between molecules, which also directly affects the accuracy of prediction.

[0003] In recent years, methods based on graph neural networks (GNNs) have been widely studied and applied in the field of molecular property prediction and drug discovery. Such as recently Deng et al. [4] implemented a classification method XGraphBoost that integrates GNN with the XGBoost machine learning algorithm, where GNN extracts features and XGBoost is used for classification; Tan et al. [5] proposed a method for predicting molecular bioactivity with a multi-characteristic fusion graph convolution method; Zhang et al. [6] proposed a fragment-oriented multi-scale graph attention model, which is applied to the prediction of molecular properties; Choo et al. [7] implemented a fingerprint-enhanced graph attention network (FinGAN) for antibiotic discovery; Bongini et al. [8]A composite GNN model is introduced, which processes heterogeneous molecular graphs by using a multi-state update mechanism and is more efficient than classical GNNs in extracting information and improving molecular property prediction tasks. On the other hand, significant progress has also been made in molecular representation learning using SMILES (Simplified Molecular Input Line Entry System) sequence information. For example, Tran et al. [9] Proposed a semi-supervised learning model based on the Transformer model, which uses pre-training and fine-tuning strategies to learn SMILES character sequences for predicting antimalarial drug candidates. Zheng et al.

[10] Implemented a BERT model based on SMILES character sequences for predicting drug molecular properties. These methods usually emphasize molecular graph structures or semantic information derived from SMILES sequences, but fail to fully utilize the global information within molecules and their inherent complex relationships, and still need to enhance the representation of the overall molecular features.

[0004] In recent years, the national invention patent documents of the technical solutions closest to the proposal of this application retrieved and their brief contents are as follows:

[0005] CN202211055546.8 - A method and system for predicting the activity of new target compounds based on deep learning. The steps of the method of this invention include: 1) converting the molecules of the compounds to be screened into a set number of SMILES formulas; 2) performing data preprocessing on the SMILES formula data to obtain a one-hot vector set of the compounds to be screened; 3) constructing and training a molecular activity prediction model, and predicting the activity of the compounds to be screened through the trained molecular activity prediction model based on the one-hot vector set of the compounds to be screened.

[0006] CN202211562456.8 - Pan Xiaolin, Zhang Yueqing, Zhang Zenghui, Ji Changge. A method for predicting the solvation energy of small molecule compounds based on graph convolutional neural network. The method of this invention introduces atomic feature descriptors containing three-dimensional structure information and two-dimensional topological information, and automatically extracts chemical patterns related to solvation energy in compounds based on graph convolutional neural network to construct a prediction model. Among them, in order to improve the robustness of the prediction model, this method does not directly predict the solvation energy of the compound, but predicts the energy contribution of each atom in the compound.

[0007] CN202410769513. - A method for predicting the activity of KRAS inhibitors based on machine learning. The steps of the method of this invention include: 1) Collect KRAS inhibitor data from ChEMBL, BindingDB, and PubChem databases, and then preprocess the collected data, including data cleaning, data deduplication, and label extraction. The label refers to the activity label of the inhibitor molecule; 2) Calculate features for the preprocessed data, including MACCS fingerprints, ECFP4 fingerprints, and Mordred descriptors, and screen out invalid and redundant features. Then randomly divide the obtained data into a training set and a test set; 3) Use mutual information feature selection for feature screening, and select the feature set that contributes the most to the model prediction performance on the training set; 4) Construct a support vector machine (SVM) classification model, input the training set data after feature screening into the model for training, and use the test set to evaluate the model; 5) Apply the classification model to an external validation set for further verification and evaluation, predict the activity of unknown molecules, and output the prediction results of whether each molecule has KRAS inhibitory activity.

[0008] The existing technical solutions have the following disadvantages:

[0009] 1) Based on virtual screening technologies such as QSAR (quantitative structure-activity relationship) models or molecular docking, it is a relatively difficult task. Constructing an efficient QSAR model requires a deep understanding of the structural characteristics, action mechanisms, and complex biological environments of compounds. In addition, the success of the model also depends on the acquisition and processing of large-scale high-quality data sets, including steps such as data cleaning, standardization, and feature extraction. Molecular docking, which simulates the process of molecules binding to receptor proteins, is also a major challenge to accurately predict the binding mode and binding free energy.

[0010] 2) In the classification task of molecular activity prediction, classical machine learning methods, such as support vector machine (SVM), random forest (RF), K-nearest neighbor (KNN), and extreme gradient boosting (XGBoost), rely on manually designed molecular descriptors or molecular fingerprint features extracted from compound structures for prediction. These manually designed features still have significant limitations in capturing the complex relationships and chemical properties between molecules, not only restricting the performance of the model in dealing with non-linear and high-dimensional data, but also possibly leading to information loss or noise introduction, thus directly affecting the accuracy and reliability of the prediction.

[0011] 3) Currently, deep learning-based methods have been widely used in molecular activity prediction research, including the use of advanced models such as graph neural networks (GNNs), Transformers, or BERT. These methods usually focus on analyzing the graph structure of molecules or the semantic information extracted from SMILES sequences, but they ignore the integration of global information within molecules, resulting in limitations in characterizing the overall properties of molecules. Summary of the Invention

[0012] The object of the present invention is to solve the limitations existing in the prior art and provide a method and system for predicting the molecular activity of β-secretase inhibitors based on the fusion learning of SMILES sequences and their graph information. The method, firstly, designs a BiLSTM-Transformer framework with an enhanced hierarchical attention mechanism to encode the functional subsequences derived from the SMILES strings of β-secretase inhibitor molecules. The functional subsequences are extracted from the SMILES strings and divided into multiple levels at the functional group, ion, and atomic levels, which can enhance the understanding of the molecular structure details. Secondly, based on the molecular graph structure generated from the SMILES strings, a graph encoding model with an atomic-level feature attention mechanism is designed, which not only highlights the key atoms and chemical bonds in the molecular graph structure but also further enhances the feature representation ability of the molecular graph. Finally, through the global attention mechanism and the weighting module, the molecular information of different modalities is fused and learned to obtain the prediction result of the molecular activity of β-secretase inhibitors.

[0013] To achieve the above object, the technical solution of the present invention is: A method for predicting the molecular activity of β-secretase inhibitors based on the fusion learning of SMILES sequences and their graph information, including:

[0014] Design a BiLSTM-Transformer framework with an enhanced hierarchical attention mechanism to encode the functional subsequences derived from the SMILES strings of β-secretase inhibitor molecules;

[0015] Based on the molecular graph structure generated from the SMILES strings, design a graph encoding model with an atomic-level feature attention mechanism to encode the molecular graph information of β-secretase inhibitor molecules;

[0016] Through the global attention mechanism and the weighting module, fuse and learn the molecular information of different modalities to obtain the prediction result of the molecular activity of β-secretase inhibitors.

[0017] In an embodiment of the present invention, the functional subsequences are extracted from the SMILES strings and divided into functional group subsequences, ion subsequences, and atomic subsequences.

[0018] In an embodiment of the present invention, the specific method for extracting the functional subsequences is as follows:

[0019] First, analyze the SMILES character of the β-secretase inhibitor molecule, discover and extract the frequently occurring functional group subsequences therein, and perform annotation to enhance the readability and chemical interpretability of these functional group subsequences;

[0020] Next, based on the encoding rules of SMILES, extract the structure between the "[" and "]" symbols in the remaining part of the molecule, and process it as an ionic group to obtain an ionic subsequence;

[0021] Finally, for the remaining part, each character is separately parsed as an atomic feature for separate encoding to obtain an atomic subsequence.

[0022] In an embodiment of the present invention, the BiLSTM-Transformer framework with an enhanced hierarchical attention mechanism includes three sub-modules: word embedding Embedding, hierarchical attention mechanism, and BiLSTM and Transformer encoding; the implementation of each sub-module is as follows:

[0023] (1) Embedding

[0024] After the data preprocessing of the SMILES character sequence of a molecule, it is composed of multiple subsequences, expressed as:

[0025] S = {s i | i = 1, 2, …, s T}, s i ∈D

[0026] where s i represents subsequences from three different levels, T represents the number of subsequences after the SMILES is split, and D represents the total dictionary of the split labeled subsequences; one-hot encoding is used for each SMILES molecule to map the sequence S to a feature vector:

[0027]

[0028] (2) Hierarchical attention mechanism

[0029] When the hierarchical attention mechanism sub-module processes the SMILES sequence feature encoding, by dynamically weighting the encoding vectors of different levels, it makes full use of the attention weights to calculate the correlation of features, effectively captures the context information between each level, realizes the fusion of multi-level features, thereby enhancing the learning effect of the subsequent BiLSTM model and improving the expression ability of molecular features; the specific steps are as follows:

[0030] 1) Perform a linear transformation on the input features; assume the input is an element vector of where B is the batch size, T is the sequence length, and H is the input feature dimension; the linear transformation is defined as:

[0031]

[0032] Among them, MLP is a multi-layer perceptron, is the weight matrix, is the bias term, A is the dimension of the attention feature, and tanh is a hyperbolic tangent activation function;

[0033] 2) Calculate the context vector for the output after linear transformation:

[0034]

[0035] Among them, context represents the linear transformation of the context vector, and SM is the softmax function, which is used to normalize the scores;

[0036] 3) Calculate the weighted value of each input, and ⊙ represents element-wise multiplication:

[0037]

[0038] 4) By aggregating the weighted inputs and combining with the original inputs:

[0039]

[0040] 5) Update the output to:

[0041]

[0042] Among them, α is the attention weight, and the default value is 0.9; the dimension of the final output data of the formula is B×T×H.

[0043] (3) BiLSTM and Transformer Encoding

[0044] For the input β-secretase inhibitor molecule SMILES sequence The output of BiLSTM is expressed as:

[0045]

[0046] Among them: is the forward s i hidden state, is the backward s i hidden state, is the s i th input subsequence in the input sequence; the final output of BiLSTM combines the forward and backward hidden states and is expressed as:

[0047]

[0048] Among them represents the splicing operation;

[0049] The output result of the BiLSTM is passed to the Transformer, and the self-attention mechanism is used to extract the potential representation of the SMILES character sequence.

[0050] In an embodiment of the present invention, before encoding the β-secretase inhibitor molecular graph information using the graph encoding model, the β-secretase inhibitor molecular SMILES string needs to be converted into molecular graph information using the deep graph learning framework DGL and the cheminformatics tool RDKit. A molecule is represented by an undirected graph G(v, e), where the atoms in the molecule correspond to the nodes v, and the chemical bonds correspond to the edges e. The extracted atomic features include 26-dimensional information on element type, implicit valence, valence electrons, bonding, charge, and hybridization type. The extracted edge features include 6-dimensional information on single bonds, double bonds, triple bonds, ring formation, aromatic rings, and conjugation.

[0051] In an embodiment of the present invention, the specific implementation manner of the graph encoding model for encoding the β-secretase inhibitor molecular graph information is as follows:

[0052] The atomic and chemical bond features extracted from the β-secretase inhibitor molecular SMILES string are passed to a GNN module composed of multiple communication message passing mechanisms. Each GNN module has an atomic-level attention mechanism to construct an atomic interaction layer, enabling the model to better capture local atomic interactions and emphasize the key information within the molecule. The core module of the GNN consists of two steps: aggregation and communication, as follows:

[0053] Let the input graph: G = (V, E), including node attributes X V and edge attributes X E ; the initial node hidden representation is The hidden representation of the edge is They are propagated in the k-th iteration as follows:

[0054] 1) Aggregate:

[0055]

[0056] Among them, is the message generated by node v, is the hidden representation of the edge in the (k - 1)-th iteration;

[0057] 2) Communicate:

[0058]

[0059] Among them, is the hidden representation of the node; in the transfer function, a feature attention mechanism is introduced. During the message passing process, the importance of weighted node features is expressed as:

[0060]

[0061] Among them, W h , W m and W x are learnable weight matrices, σ is the activation function, is the attention score of the influence degree of the feature of node v on its molecular representation;

[0062] 3) Edge representation update:

[0063]

[0064] Among them, is the hidden representation of the node, is the hidden representation of the edge, W is a learnable weight matrix, and σ is the activation function;

[0065] 4) After L iterations, the final message aggregation:

[0066]

[0067] Among them, is the hidden representation of the edge in the L-th iteration;

[0068] 5) After L iterations, the final node representation update:

[0069]

[0070] Among them, is the node hidden representation in the L-th iteration.

[0071] In an embodiment of the present invention, the Aggregate function includes a message enhancer, which generates the maximum pooling result of the sum of messages of the edge hidden state h e and calculates the sum of the element products of h e ; the Communicate function represents a multi-layer perceptron.

[0072] In an embodiment of the present invention, through the global attention mechanism and the weighted module, the molecular information of different modalities is fused and learned to obtain the prediction result of the molecular activity of β-secretase inhibitor, specifically:

[0073] First, using the global attention mechanism, the different modality representations of each atom in the molecule are integrated to obtain the graph-level structure representation of the entire SMILES molecule, thereby obtaining a unified feature vector and the feature vectors corresponding to two modalities;

[0074] Next, a fusion layer is designed to generate a comprehensive representation, that is, the features of different patterns are weighted and combined through a non - linear function, and a bias term is added. The formula is as follows:

[0075] H = f(H s , H g ) + b = W s ·H s + W g ·H g + b

[0076] Where f is a non - linear function, H s , H g represent the feature representations based on SMILES sequence encoding and GNN encoding respectively, W s and W g represent learnable weights, and b represents the bias term.

[0077] In an embodiment of the present invention, the method adopts a loss function based on similarity calculation. The loss function is as follows:

[0078]

[0079] Where Z g (i) and Z s (i) represent the feature vectors of s and g of the i - th sample respectively, cos represents the cosine similarity, T represents the contrast loss rate, and N is the number of sample information; r ∈ {g, s} is used to represent the identifier of the feature space, s represents sequence information, and g represents graph structure; the contrast loss is used to enhance the model's ability to recognize the differences between different feature vectors, and at the same time, the label loss is introduced to ensure the prediction accuracy of the model for labels.

[0080] The present invention also provides a β - secretase inhibitor molecular activity prediction system based on the fusion learning of SMILES sequences and their graph information, including:

[0081] A data pre - processing module that extracts functional subsequences from the SMILES strings of β - secretase inhibitor molecules; at the same time, converts the SMILES strings of β - secretase inhibitor molecules into molecular graph information;

[0082] An SMILES character subsequence encoding module that designs a BiLSTM - Transformer framework with an enhanced hierarchical attention mechanism to encode the extracted functional subsequences;

[0083] A graph information encoding module that designs a graph encoding model with an atomic - level feature attention mechanism to encode the β - secretase inhibitor molecular graph information;

[0084] The feature fusion learning prediction module fuses and learns molecular information of different modalities to obtain the prediction result of the molecular activity of β-secretase inhibitors.

[0085] Compared with the prior art, the present invention has the following beneficial effects: For the method and system of the present invention, only the SMILES string of the β-secretase inhibitor compound to be tested needs to be provided, and its molecular activity can be quickly and accurately predicted, greatly improving the efficiency of compound structure optimization and virtual screening. The method and system of the present invention have low application costs, are simple and fast, can save a large amount of manpower, costs and time required for experimental testing, and can also provide new ideas for predicting the properties of other chemical molecules, having important practical significance and theoretical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] Figure 1 It is a block diagram of the overall technical solution of the present invention.

[0087] Figure 2 It is a graph of the index values during the training process of the solution of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0088] The technical solution of the present invention will be specifically described below with reference to the accompanying drawings.

[0089] The present invention provides a method for predicting the molecular activity of β-secretase inhibitors based on the fusion learning of SMILES sequences and their graph information, including:

[0090] Design a BiLSTM-Transformer framework with an enhanced hierarchical attention mechanism to encode the functional subsequences derived from the SMILES string of β-secretase inhibitor molecules;

[0091] Based on the molecular graph structure generated from the SMILES string, design a graph encoding model with an atomic-level feature attention mechanism to encode the molecular graph information of β-secretase inhibitor molecules;

[0092] Through the global attention mechanism and the weighting module, fuse and learn molecular information of different modalities to obtain the prediction result of the molecular activity of β-secretase inhibitors.

[0093] The present invention also provides a system for predicting the molecular activity of β-secretase inhibitors based on the fusion learning of SMILES sequences and their graph information, including:

[0094] The data preprocessing module extracts the functional subsequences from the SMILES string of β-secretase inhibitor molecules; at the same time, converts the SMILES string of β-secretase inhibitor molecules into molecular graph information;

[0095] SMILES character subsequence encoding module, design a BiLSTM-Transformer framework with an enhanced hierarchical attention mechanism to encode the extracted functional subsequences;

[0096] Graph information encoding module, design a graph encoding model with an atomic-level feature attention mechanism to encode the molecular graph information of β-secretase inhibitors;

[0097] Feature fusion learning and prediction module, fuse and learn different modalities of molecular information to obtain the prediction result of the activity of β-secretase inhibitor molecules.

[0098] The following is the specific implementation process of the present invention.

[0099] The present invention proposes a method for predicting the activity of β-secretase inhibitor molecules based on the fusion learning of SMILES sequences and their graph information. First, a BiLSTM-Transformer framework with an enhanced hierarchical attention mechanism is designed to encode the functional subsequences derived from the SMILES strings of β-secretase inhibitor molecules. The functional subsequences are extracted from the SMILES strings and divided into multiple levels at the functional group, ion, and atomic levels, which can enhance the understanding of the molecular structure details. Second, based on the molecular graph structure generated from the SMILES strings, a graph encoding model with an atomic-level feature attention mechanism is designed, which not only highlights the key atoms and chemical bonds in the molecular graph structure but also further enhances the feature representation ability of the molecular graph. Finally, through the global attention mechanism and the weighting module, different modalities of molecular information are fused and learned to obtain the prediction result of the activity of β-secretase inhibitor molecules.

[0100] The overall technical solution of the present invention includes the content of four modules: data preprocessing, SMILES character subsequence encoding, graph information encoding, and feature fusion learning and prediction. The data preprocessing module is a SMILES molecular string parsing module. First, it extracts the subsequences at the functional group, ion, and atomic levels from the SMILES character sequence as the input of the SMILES character subsequence encoding module. Second, it converts the SMILES character sequence into molecular graph information as the input of the graph information encoding module. The SMILES character subsequence encoding module is a BiLSTM-Transformer model with an enhanced hierarchical attention mechanism, which can encode the functional substrings of the SMILES characters of β-secretase inhibitor molecules. The graph information encoding module is a graph encoding model with an atomic-level feature attention mechanism, which encodes the molecular graph information of β-secretase inhibitor molecules. The feature fusion learning and prediction module fuses and learns different modalities of molecular information through the global attention mechanism and the weight module to obtain the prediction result of the activity of β-secretase inhibitor molecules. The overall technical solution is as Figure 1 shown. The following are the specific descriptions of each module.

[0101] 1. Data preprocessing module

[0102] The experimental data of the present invention come from publicly reported literature data. There are 1,548 β-secretase inhibitor molecular compounds, presented in SMILES format, and each compound has an IC50 activity measurement value. SMILES is a specification that clearly describes the molecular structure with ASCII strings, that is, a simplified text format for representing chemical molecular structures. The data preprocessing module is a process of parsing SMILES strings.

[0103] 1.1 Extraction of SMILES character subsequences

[0104] Currently, some models based on learning SMILES character sequences generally directly input atomic-level elements and chemical bonds from SMILES for training, often ignoring important substructural information within the molecule. The present invention analyzed the SMILES character sequences of β-secretase inhibitor molecules and found many functional groups with unique chemical characteristics. For this reason, a hierarchical feature parsing method is proposed, that is, extracting subsequences at the functional group, ion, and atomic levels from the SMILES character sequence as the input of the SMILES character subsequence encoding module. The specific steps are as follows:

[0105] 1) First, analyze the SMILES characters of the dataset molecules, find and extract the frequently occurring functional group subsequences therein, and label them to enhance the readability and chemical interpretability of these functional group subsequences to better support the training of deep learning models.

[0106] 2) Then, based on the encoding rules of SMILES, extract the structure between the "[" and "]" symbols in the remaining part of the molecule and treat it as an ionic group.

[0107] 3) Finally, for the remaining part, each character is separately parsed into atomic features for separate encoding. An example of the parsing result of a SMILES character sequence of a β-secretase inhibitor molecule is shown in Table 1.

[0108] Table 1 Example of parsing a SMILES character sequence of a β-secretase inhibitor molecule

[0109]

[0110] 1.2 Conversion of SMILES sequence into graph information

[0111] The present invention utilizes the deep graph learning framework DGL (Deep Graph Library) and the cheminformatics tool RDKit to convert the SMILES string of the compound to be tested into corresponding graph data as the input of the graph information encoding model. A molecule can be represented by an undirected graph G(v, e), where the atoms in the molecule correspond to nodes v and the chemical bonds correspond to edges e. The present invention extracts 26-dimensional information of atomic features including element type, implicit valence, valence electrons, bonding, charge, hybridization type, etc., and 6-dimensional information of edge features including single bond, double bond, triple bond, ring formation, aromatic ring, conjugation, etc.

[0112] 2. SMILES Character Subsequence Encoding Module

[0113] As described above, the present invention converts β-secretase inhibitors into three levels of sequence information, including functional group, ion, and atom levels. For this purpose, the present invention designs a hierarchical attention mechanism to perform hierarchical weighted learning on features at different levels, enabling the model to adaptively allocate attention weights between features at different levels, thereby effectively extracting detailed information at each level. This module mainly consists of three sub-modules: word embedding, hierarchical attention mechanism, and BiLSTM and Transformer encoding.

[0114] 2.1 Word Embedding

[0115] After preprocessing the SMILES character sequence data of a molecule, it consists of multiple subsequences and can be expressed as:

[0116] S = {s i | i = 1, 2, …, s T}, s i ∈D (1)

[0117] where s i represents subsequences from three different levels, T represents the number of subsequences after splitting the SMILES, and D represents the total dictionary of the split labeled subsequences. Then, one-hot encoding is used for each SMILES molecule to map the sequence S to a feature vector:

[0118]

[0119] 2.2 Hierarchical Attention Mechanism

[0120] When the hierarchical attention mechanism sub-module processes the feature encoding of the SMILES sequence, by dynamically weighting the encoding vectors of different levels (functional groups, ionic groups, and atoms), it makes full use of the attention weights to calculate the correlation of features, can effectively capture the context information between each level, realize the fusion of multi-level features, thereby enhancing the learning effect of the subsequent BiLSTM model and improving the expression ability of molecular features. The specific steps are as follows:

[0121] 1) First, perform a linear transformation on the input features. Let the input be an element vector where B is the batch size, T is the sequence length, and H is the input feature dimension. The linear transformation is defined as:

[0122]

[0123] where MLP is a multi-layer perceptron, is the weight matrix, is the bias term, A is the dimension of the attention feature, and tanh is a hyperbolic tangent activation function.

[0124] 2) Then, calculate the context vector for the output after the linear transformation:

[0125]

[0126] where context represents the linear transformation of the context vector, and SM is the softmax function, which is used to normalize the scores.

[0127] 3) Then, calculate the weighted value of each input, and ⊙ represents element-wise multiplication:

[0128]

[0129] 4) Then, by aggregating the weighted inputs and combining with the original inputs:

[0130]

[0131] 5) Finally, update the output as:

[0132]

[0133] where α is the attention weight, and the default value is 0.9; the dimension of the data finally output by the formula is B×T×H.

[0134] 2.3. BiLSTM and Transformer Encoding

[0135] BiLSTM (Bidirectional Long Short-Term Memory), that is, bidirectional long short-term memory network. It contains two LSTM networks, a forward LSTM and a backward LSTM. The forward LSTM processes the input data sequentially from the beginning to the end of the sequence, and can capture the influence of the information in the front of the sequence on the back; the backward LSTM processes the input data in reverse from the end to the beginning of the sequence, and can capture the influence of the information in the back of the sequence on the front. Therefore, BiLSTM can utilize both the forward and backward information of the SMILES sequence of β-secretase inhibitor molecules to better understand the overall characteristics of the sequence data.

[0136] Specifically, for the input SMILES sequence of β-secretase inhibitor molecules The output of BiLSTM can be expressed as:

[0137]

[0138] Where: is the forward s i hidden state, is the backward s i hidden state, is the s-th input subsequence in the input sequence. The final output of BiLSTM combines the forward and backward hidden states, and can be expressed as: i

[0139]

[0140] Where represents the concatenation operation. Finally, the output result of BiLSTM is passed to a standard Transformer, which uses self-attention mechanism to extract the latent representation of the SMILES character sequence.

[0141] 3. Graph Information Encoding Module

[0142] A molecule can be represented by an undirected graph G(v, e), where the atoms in the molecule correspond to nodes v, and the chemical bonds correspond to edges e. The present invention transfers the atom and chemical bond features extracted by data preprocessing into a GNN module composed of multiple communication message passing mechanisms. Each GNN module has an atomic-level attention mechanism to construct an inter-atomic interaction layer, enabling the model to better capture local atomic interactions and emphasize the key information within the molecule. The core module of GNN consists of two steps: Aggregate and Communicate. Given the input graph: G = (V, E), including node attributes X V and edge attributes X E . The initial node hidden representation is The hidden representation of the edge is They are propagated in the k-th iteration as follows:

[0143] 1) Aggregate:

[0144]

[0145] where, is the message generated by node v, is the hidden representation of the edge in the (k - 1)-th iteration;

[0146] 2) Communicate:

[0147]

[0148] where, is the hidden representation of the node. In the communication function, a feature attention mechanism is introduced. During the message passing process, the importance of weighted node features can be expressed as:

[0149]

[0150] where, W h 、W m and W x are learnable weight matrices, σ is the activation function, is the attention score of the influence of the feature of node v on its molecular representation.

[0151] 3) Edge representation update:

[0152]

[0153] where, is the hidden representation of the node, is the hidden representation of the edge, W is a learnable weight matrix, and σ is the activation function. After L iterations, the following operations are performed to obtain the final message and node representation.

[0154] 4) Final message aggregation:

[0155]

[0156] where, is the hidden representation of the edge in the L-th iteration;

[0157] 5) Final node representation update:

[0158]

[0159] where, is the node hidden representation in the L-th iteration.

[0160] The Aggregate function contains a message enhancer that generates the max-pooling result of the message sum of the edge-hidden state he and calculates the sum of the element products he; the Communicate function represents a multi-layer perceptron. This process enhances atomic-level information exchange by optimizing node representations using message aggregation and attention mechanisms, focusing on the most relevant connections in the graph structure. Then, through the feature attention module, the molecular feature representation vector is used to adaptively focus on the feature dimensions of each atom within the same molecule, re-evaluating and adjusting the weights of the atomic feature vectors, and finally generating a reconstructed feature vector.

[0161] 4. Fusion Learning Prediction Module

[0162] To effectively fuse information from different modalities, the present invention first uses a global attention mechanism to integrate the different modality representations of each atom within the molecule, obtaining a graph-level structure representation of the entire SMILES molecule, thereby obtaining a unified feature vector and the feature vectors corresponding to the two modalities. Then, a fusion layer is designed to generate a comprehensive representation, that is, the features of different modalities are weighted and combined through a non-linear function, and a bias term is added, as shown in the following formula:

[0163] H = f(H s , H g ) + b = W s ·H s + W g ·H g + b (17)

[0164] where f is a non-linear function, H s , H g represent the feature representations based on SMILES sequence encoding and GNN encoding respectively, W s and W g represent learnable weights, and b represents the bias term.

[0165] 5. Loss Function

[0166] To effectively reduce the differences in feature representations among different methods, the present invention adopts a loss function based on similarity calculation, and the loss function is as follows:

[0167]

[0168] where, Z g (i) and Z s(i) Respectively represent the feature vectors of s and g for the i-th sample, cos represents the cosine similarity, T represents the contrast loss rate, and N is the number of sample information. r ∈ {g, s} is used to represent the identifier of the feature space, s represents sequence information, and g represents graph structure. The contrast loss is used to enhance the model's ability to recognize the differences between different feature vectors, and at the same time, the label loss is introduced to ensure the prediction accuracy of the model for labels.

[0169] 6. Experimental Results

[0170] The experimental dataset of the present invention is randomly divided into a training set, a validation set, and a test set at a ratio of 8:1:1. The experimental results of using the method proposed in the present invention are compared with the three recent β-secretase inhibitor molecular activity prediction models QSARBioPred[1], XGraphBoost[4], and FraGAT[6]. Among them, the QSARBioPred method uses molecular descriptor feature extraction and traditional machine learning classification techniques; the XGraphBoost method is a classification method that integrates GNN and the XGBoost machine learning algorithm; the FraGAT method is a method based on graph neural network (GNN). The comparison results are shown in Table 2. The definitions of each evaluation index are as follows.

[0171] Table 2 Comparison of Experimental Results

[0172]

[0173]

[0174] The evaluation indexes include: accuracy (ACC), sensitivity (SE), specificity (SP), Matthews correlation coefficient (MCC), F1-score, precision-recall curve (PRC), area under the curve (AUC), etc., a total of 7 indexes. The definitions of each index are as follows, where TP is true positive, FP is false positive, TN is true negative, and FN is false negative.

[0175]

[0176] Figure 2 The changes of seven classification metric values in the training iteration of the present solution are given. Each metric is represented by a curve of a different color. The x-axis represents the number of iterations, and the y-axis represents the values of each evaluation index.

[0177] It can be seen that compared with other methods, the experimental results of the present invention show higher reliability in multiple indexes such as ACC, F1-score, SP, MCC, and PRC, indicating the effectiveness of the method proposed in the present invention in accurately predicting the molecular activity of β-secretase inhibitors and highlighting its potential advantages in this field.

[0178] References:

[0179] [1] Ponzoni I., Víctor Sebastián-Pérez, María J. Martínez, et al. QSAR Classification Models for Predicting the Activity of Inhibitors of Beta-Secretase (BACE1) Associated with Alzheimer's Disease[J]. Scientific Reports, 2019, 9(1): 10341.

[0180] [2] Huang D., Liu Y., Shi B., et al. Comprehensive 3D-QSAR and binding mode of BACE-1 inhibitors using R-group search and molecular docking[J]. Journal of molecular graphics & modelling, 2013, 45C(18): 65 - 83.

[0181] [3] Singh R., Ganeshpurkar A., Ghosh P., et al. Classification of Beta-site Amyloid Precursor Protein Cleaving Enzyme 1 Inhibitors by Using Machine Learning Methods[J]. Chemical Biology & Drug Design, 2021, 98(6): 1079 - 1097.

[0182] [4] Daiguo Deng, Xiaowei Chen, Ruochi Zhang, et al. XGraphBoost: Extracting Graph Neural Network-Based Features for a Better Prediction of Molecular Properties[J]. Journal of Chemical Information and Modeling, 2021, 61(6), 2697 - 2705.

[0183] [5] Tan Lulu, Zhang Xinxin, Zhou Yinzuo. Prediction of Molecular Bioactivity by Multi-Feature Fusion Graph Convolution Method [J]. Journal of University of Electronic Science and Technology of China, 2021, 50(6): 921-929.

[0184] [6] Zhang Z., Guan J., Zhou S. FraGAT: A Fragment-Oriented Multi-Scale Graph Attention Model for Molecular Property Prediction [J]. Bioinformatics, 2021, 37(18), 2981-2987.

[0185] [7] Choo H Y, Wee J J, Shen C., et al. Fingerprint-Enhanced Graph Attention Network(FinGAT) Model for Antibiotic Discovery [J]. J Chem Inf Model., 2023, 63(10): 2928-2935.

[0186] [8] Bongini P., Pancino N., Bendjeddou A., et al. Composite Graph Neural Networks for Molecular Property Prediction [J]. Int.J.Mol.Sci., 2024, 25(12): 6583.

[0187] [9] Tran T, Ekenna C. Molecular Descriptors Property Prediction Using Transformer-Based Approach [J]. Int.J.Mol.Sci., 2023, 24(15): 11948.

[0189]

[10] Zheng X., Tomiura Y. A BERT-Based Pretraining Model for Extracting Molecular Structural Information from a SMILES Sequence [J]. J.Cheminform., 2024, 16(1): 71-80.

[0190] The above are the preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention in terms of the functions and effects produced shall fall within the protection scope of the present invention.

Claims

1. A method for predicting the molecular activity of β-secretase inhibitors based on SMILES sequence and its graph information fusion learning, characterized in that: include: Design a BiLSTM-Transformer framework with enhanced hierarchical attention mechanism to encode functional subsequences derived from SMILES strings of β-secretase inhibitor molecules; Based on the molecular graph structure generated from SMILES strings, a graph encoding model with atomic-level feature attention mechanism is designed to encode the molecular graph information of β-secretase inhibitors. Through the global attention mechanism and weighted module, the molecular information of different modalities is fused and learned to obtain the prediction results of the molecular activity of β-secretase inhibitors.

2. A method for predicting the molecular activity of β-secretase inhibitors based on SMILES sequence and its graph information fusion learning according to claim 1, characterized in that: Functional subsequences are extracted from SMILES strings and divided into functional group subsequences, ion subsequences, and atom subsequences.

3. A method for predicting the molecular activity of β-secretase inhibitors based on SMILES sequence and its graph information fusion learning according to claim 1 or 2, characterized in that: The functional subsequence extraction method is as follows: First, the SMILES characters of β-secretase inhibitor molecules were analyzed to find and extract the frequently appearing functional group subsequences, and annotate them to enhance the readability and chemical interpretability of these functional group subsequences. Next, based on the encoding rules of SMILES, the structure between the "[" and "]" symbols in the rest of the molecule is extracted and treated as an ionic group to obtain an ionic subsequence; Finally, for the remaining part, each character is parsed into atomic features and encoded separately to obtain atomic subsequences.

4. The method for predicting the molecular activity of β-secretase inhibitors based on SMILES sequence and its graph information fusion learning according to claim 2, characterized in that: The BiLSTM-Transformer framework with enhanced hierarchical attention mechanism includes three submodules: word embedding, hierarchical attention mechanism, BiLSTM and Transformer encoding; each submodule is implemented as follows: (1) Embedding After preprocessing, the SMILES character sequence data of a molecule consists of multiple subsequences, which are expressed as: S={s i |i=1,2,…,s T },s i ∈D where s i Represents subsequences from three different levels, T represents the number of subsequences after SMILES is split, and D represents the total dictionary of the split labeled subsequences; use one-hot encoding for each SMILES molecule to map the sequence S into a feature vector: (2) Hierarchical Attention Mechanism When processing SMILES sequence feature encoding, the hierarchical attention mechanism submodule dynamically weights the encoding vectors of different levels and makes full use of the attention weights to calculate the relevance of features, so as to effectively capture the contextual information between each level and realize the fusion of multi-level features, thereby enhancing the learning effect of the subsequent BiLSTM model and improving the expression ability of molecular features. The specific steps are as follows: 1) Perform linear transformation on the input features; let the input be the element vector Where B is the batch size, T is the sequence length, and H is the input feature dimension; The linear transformation is defined as: Among them, MLP is a multi-layer perceptron. is the weight matrix, is a bias term, A is the dimension of the attention feature, and tanh is a hyperbolic tangent activation function; 2) Calculate the context vector of the linearly transformed output: Among them, context represents the linear transformation through the context vector, and SM is the softmax function used to normalize the score; 3) Calculate the weighted value of each input, ⊙ represents element-by-element multiplication: 4) By aggregating the weighted inputs, and combining the original inputs: 5) Update the output to: Among them, α is the attention weight; the dimension of the final output data of the formula is B×T×H. (3) BiLSTM and Transformer encoding For the input β-secretase inhibitor molecule SMILES sequence The output of BiLSTM is represented as: in: is positive i Hidden state, It is the reverse i Hidden state, is the sth in the input sequence i The final output of BiLSTM combines the forward and reverse hidden states and is expressed as: in Represents a splicing operation; The output of the BiLSTM is passed to the Transformer, which uses a self-attention mechanism to extract the latent representation of the SMILES character sequence.

5. The method for predicting the molecular activity of β-secretase inhibitors based on SMILES sequence and its graph information fusion learning according to claim 1, characterized in that: Before using the graph encoding model to encode the β-secretase inhibitor molecular graph information, the β-secretase inhibitor molecule SMILES string needs to be converted into molecular graph information using the deep graph learning framework DGL and the chemical informatics tool RDKit. A molecule is represented by an undirected graph G(v,e), where the atoms in the molecule correspond to nodes v and the chemical bonds correspond to edges e. The extracted atomic features include 26-dimensional information such as element type, implicit valence, valence electrons, bonding, charge, and hybridization type. The extracted edge features include 6-dimensional information such as single bond, double bond, triple bond, cyclization, aromatic ring, and conjugation.

6. A method for predicting the molecular activity of β-secretase inhibitors based on SMILES sequence and its graph information fusion learning according to claim 5, characterized in that: The specific implementation method of the graph encoding model to encode the β-secretase inhibitor molecular graph information is as follows: The atomic and chemical bond features extracted from the SMILES string of the β-secretase inhibitor molecule are passed to a GNN module composed of multiple communication message passing mechanisms. Each GNN module has an atomic-level attention mechanism to build an atomic interaction layer, which enables the model to better capture local atomic interactions and emphasize key information within the molecule. The core module of the GNN consists of two steps: aggregation and transmission, as follows: Assume the input graph: G = (V, E), including node attributes X V and edge attribute X E ; The initial node hidden representation is The hidden representation of the edge is They are propagated in the kth iteration as follows: 1) Aggregate: in, is the message generated by node v, is the hidden representation of the edge at the k-1th iteration; 2) Communicate: in, is the hidden representation of the node; in the transfer function, the feature attention mechanism is introduced. In the process of message passing, the importance of weighted node features is expressed as: Among them, W h , W m and W x is the learnable weight matrix, σ is the activation function, is the attention score of the influence of the feature of node v on its molecular representation; 3) Edge representation update: in, is the hidden representation of the node, is the hidden representation of the edge, W is a learnable weight matrix, and σ is the activation function; 4) After L iterations, the final message aggregation: in, is the hidden representation of the edge at the Lth iteration; 5) After L iterations, the final node representation is updated: in, Hidden representation for the Lth iteration node.

7. A method for predicting the molecular activity of β-secretase inhibitors based on SMILES sequence and graph information fusion learning according to claim 6, characterized in that: The Aggregate function contains a message enhancer, which generates the edge hidden state h e The maximum pooling result of the sum of messages and calculate its element product h e The communicate function represents a multi-layer perceptron.

8. The method for predicting the molecular activity of β-secretase inhibitors based on SMILES sequence and its graph information fusion learning according to claim 1, characterized in that: Through the global attention mechanism and weighted module, the molecular information of different modalities is fused and learned to obtain the prediction results of the molecular activity of β-secretase inhibitors, specifically: First, the global attention mechanism is used to integrate the different modal representations of each atom in the molecule to obtain the graph-level structural representation of the entire SMILES molecule, thereby obtaining a unified feature vector and feature vectors corresponding to the two modes; Next, a fusion layer is designed to generate a comprehensive representation, that is, the features of different modes are weighted and combined through a nonlinear function, and a bias term is added. The formula is as follows: H=f(H s ,H g )+b=W s ·H s +W g ·H g +b Among them, f is a nonlinear function, H s ,H g They represent the feature representation based on SMILES sequence encoding and GNN encoding respectively, and W s and W g represents the learnable weight and b represents the bias term.

9. The method for predicting the molecular activity of β-secretase inhibitors based on SMILES sequence and its graph information fusion learning according to claim 1, characterized in that: The method adopts a loss function based on similarity calculation, and the loss function is as follows: Among them, Z g (i) and Z s (i) represents the feature vectors of s and g of the i-th sample, cos represents cosine similarity, T represents contrast loss rate, and N is the number of sample information; r∈{g,s} is used to represent the identifier of the feature space, s represents sequence information, and g represents graph structure; contrast loss is used to enhance the model's ability to recognize the differences between different feature vectors, and label loss is introduced to ensure the model's accuracy in predicting labels.

10. A β-secretase inhibitor molecular activity prediction system based on SMILES sequence and its graph information fusion learning, characterized in that: include: The data preprocessing module extracts functional subsequences from the SMILES string of the β-secretase inhibitor molecule; at the same time, the SMILES string of the β-secretase inhibitor molecule is converted into molecular graph information; SMILES character subsequence encoding module, a BiLSTM-Transformer framework with enhanced hierarchical attention mechanism is designed to encode the extracted functional subsequences; Graph information encoding module: a graph encoding model with atomic-level feature attention mechanism is designed to encode the molecular graph information of β-secretase inhibitors; The feature fusion learning prediction module fuses the molecular information of different modalities to obtain the prediction results of the molecular activity of β-secretase inhibitors.

Citation Information

Patent Citations

  • New target compound activity prediction method and system based on deep learning

    CN115331750A

  • Method for predicting solvation energy of small molecule compound based on graph convolutional neural network

    CN115938501A

  • KRAS inhibitor activity prediction method based on machine learning

    CN118800344A

Cited By

  • Enzyme EC number prediction method

    CN121583339A

  • Transform network-based biological metabolic pathway inverse synthesis design method and system

    CN121617479A

  • Method for generating multi-mode synthesizable molecules perceived by chemical reaction

    CN121789834A

  • A chemical reaction-sensing multimodal synthetic molecule generation method

    CN121789834B