A method for generating a targeting ligand based on an active fragment

By using a targeted ligand generation method based on active fragments, the problem of generating drug ligand molecules that are not synthesizable or difficult to synthesize in existing technologies has been solved. This method achieves efficient and synthetic targeted ligand molecule generation, applicable to proteins with unknown three-dimensional structures, thus improving the generation efficiency and applicability of the model.

CN116758977BActive Publication Date: 2026-01-02XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310409171.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-17
Publication Date
2026-01-02
Estimated Expiration
2043-04-17

AI Technical Summary

Technical Problem

Existing technologies fail to adequately consider the syntheticity of molecules when generating drug ligand molecules that can target specific proteins. They are not applicable to proteins with unknown three-dimensional structures, and the models suffer from low training efficiency, low fault tolerance, high complexity, and poor generalization ability.

Method used

A targeted ligand generation method based on active fragments is adopted. Through fragment sequencing, model training and targeted ligand molecule generation process, features are extracted by fragment sequence encoder and target encoder, and combined with prior network to predict complex information to generate synthetic targeted ligand molecules.

Benefits of technology

It improves the syntheticity and targeting of generated molecules, expands the applicability of the model, enhances generation efficiency and fault tolerance, reduces model complexity, and achieves more efficient generation of drug ligand molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116758977B_ABST
    Figure CN116758977B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on active fragment's targeted ligand generation method, including fragment sequencing process: based on fragmentation method and sequencing method processing ligand molecule for training obtains fragment sequence;Model training process: ligand molecule fragment sequence and target protein amino acid sequence for training are input into targeted ligand molecule generation model, the molecular representation of ligand molecule fragment sequence is extracted by fragment sequence encoder, and the target representation of target protein amino acid sequence is extracted by target encoder, feature fusion is used after fragment sequence decoder output fragment sequence, and optimization model;Targeted ligand molecule generation process: target protein amino acid sequence is input into the model after training, the target representation is extracted by target encoder, and compound information is predicted by priori network and is input into fragment sequence decoder, and output fragment sequence, again based on molecular reconstruction method fragment sequence is recombined into targeted ligand molecule.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of molecule generation and natural language processing, and particularly relates to a method for generating a targeting ligand based on an active fragment. BACKGROUND

[0002] Designing a ligand molecule that can bind to a given protein target is a particularly important step in drug development. In recent years, deep generative models have made significant progress in the field of molecule generation. Various different models such as variational autoencoder, generative adversarial network, diffusion model, etc. are used to generate molecule SMILES (Simplified Molecular Input Line Entry System) characters, molecule graphs or molecule three-dimensional conformations. These methods aim to learn the distribution of real molecules and sample novel and effective chemical molecules from the distribution, but cannot be directly used to generate ligand molecules that can target specific protein targets, and further screening is required through docking or other affinity prediction methods.

[0003] Recently, researchers have begun to explore how to use molecule and protein target interaction information to generate compounds that can bind to the target. Two technical solutions are most similar to the present application. The first method LiGAN first encodes the protein pocket using a variational autoencoder composed of a convolutional neural network, decodes the generated molecule image from the encoded features, and then decodes the molecule SMILES characters from the molecule image. The second method 3D-SBDD proposes an autoregressive three-dimensional molecule generation method, which models the probability of atom occurrence in the three-dimensional space of the protein pocket, and then continuously samples atoms to generate the final molecule. Both methods generate ligand molecules based on protein pockets, and the design of the algorithm does not fully consider the synthesizability of the generated molecules.

[0004] There are three major limitations in the prior art for designing an algorithm to generate a synthetic drug ligand molecule that can bind to a given protein target:

[0005] (1) The prior art does not consider the synthesizability of the molecule, resulting in unreasonable or difficult-to-synthesize molecules;

[0006] (2) Most of the prior art is based on three-dimensional information of the protein binding pocket to predict and generate molecules, which cannot be applied to proteins with unknown three-dimensional structures or proteins with unknown targets. The prediction results of the entire protein ligand molecule prediction scheme lack targeting, and an additional screening step is required.

[0007] (3) The prior art generates atoms one by one, and the model training process involves predicting three-dimensional coordinates, which is low in efficiency, fault tolerance, model complexity, and generalization ability. Summary of the Invention

[0008] To address the aforementioned problems, this invention proposes a method for generating targeted ligands based on active fragments, which can effectively overcome three limitations in the prior art.

[0009] The technical solution adopted in this invention is as follows:

[0010] A method for generating targeted ligands based on active fragments, comprising:

[0011] Fragment serialization process: Based on fragmentation methods and serialization methods, the ligand molecules used for training are processed to obtain fragment sequences;

[0012] Model training process: The ligand molecule fragment sequence and the target protein amino acid sequence for training are input into the target ligand molecule generation model. The molecular characterization of the ligand molecule fragment sequence is extracted by the fragment sequence encoder, and the target characterization of the target protein amino acid sequence is extracted by the target encoder. After feature fusion, the fragment sequence is output by the fragment sequence decoder, and the target ligand molecule generation model is optimized.

[0013] The process of generating targeted ligand molecules is as follows: The amino acid sequence of the target protein is input into the trained targeted ligand molecule generation model. The target characterization is extracted by the target encoder, and the complex information is predicted by the prior network and input into the fragment sequence decoder. The fragment sequence decoder outputs the fragment sequence, and then the fragment sequence is recombined into the targeted ligand molecule based on the molecular reconstruction method.

[0014] Furthermore, the fragmentation method includes the following steps:

[0015] Based on the rules of reversible synthetic chemical substructure decomposition, all breakable chemical bonds in the ligand molecule are matched. All breakable chemical bonds are broken, and virtual atoms are connected to the atoms at both ends of all breakable chemical bonds as breakpoints and marked. This is used to query the fragments at both ends of any breakable chemical bond and the breakpoints within the fragments.

[0016] Furthermore, the serialization method includes the following steps:

[0017] Step 101: Construct a fragmentation method to obtain all fragments with breakpoint markers, treating each fragment as a node; sort the breakpoints within each fragment in ascending order according to their RDKit atomic index numbers, and assign each breakpoint a breakpoint number x = 1, 2, ..., P. N P N Let A be the number of breakpoints on a node; construct a quadruple [N1,,2,] for each of the original breakable chemical bonds, indicating that there is an edge between breakpoint x on node N1 and breakpoint y on node N2; if a node Ax There is an edge between the current node N, then A x is called the adjacent node of the current node N;

[0018] Step 102: Take the fragment with the largest relative mass as the initial node, denoted as node N, at this time node N as the root node, parent node N par is empty;

[0019] Step 103: Add node N to the initial sequence L with length 0, and mark the breakpoint number in the order x = 1, 2, …, P N , according to the four-tuple undirected edge, traverse the adjacent node A x of node N;

[0020] Step 104: If node N has a parent node N par , and the adjacent node A x is the parent node N par of node N, then add a special mark "*" to the sequence L; otherwise, N par =, N = A x , return to step 103 for recursive iteration, and convert the nodes in sequence L into fragment representations in turn, output a fragment sequence containing a special mark "*".

[0021] Further, in step 101, for each fragment, the equivalence relationship between the virtual atoms needs to be found: combine the virtual atoms two by two as [D x D y ], and replace the virtual atom D x with a virtual atom with neutron number 1, and replace the other virtual atoms with virtual atoms with neutron number 0, to obtain the SMILES string expression S x of the fragment at this time; replace the virtual atom Y with a virtual atom with neutron number 1, and replace the other virtual atoms with virtual atoms with neutron number 0, to obtain the SMILES string expression S y of the fragment at this time; if the SMILES string expressions S x and S y are the same, then it is determined that there is an equivalence relationship between the virtual atoms in this group;

[0022] Further, according to the equivalence relationship, the virtual atom sequence is sorted, so that the virtual atoms with equivalence relationship are symmetrically distributed at the beginning and end of the sequence; if such a sequence exists, it is determined that the fragment has symmetry, and the virtual atoms are marked with breakpoint numbers according to the sequence; otherwise, the initial order is marked.

[0023] Further, in step 104, if If A1 is the parent node of node N, and the fragment has symmetry, then traverse the breakpoints in reverse and no special mark is needed.

[0024] Further, the molecular reconstruction method comprises the following steps:

[0025] Step 201: For each valid node N in sequence L, obtain the fragment representation, and determine the maximum number of adjacent nodes P according to the number of breakpoints in each node fragment N ; let the index c on sequence L at this time be 0, the parent node N par be empty, and the breakpoint number x on the current parent node N par be empty;

[0026] Step 202: If node L[c] is not a special mark “*”, then extract node N = L[c], and the adjacent node sequence of node N is L N , with an initial length of 0, and proceed to step 203; if node L[c] is a special mark “*”, then directly return the index c and the current node N;

[0027] Step 203: If sequence L N is not full, then set c = c + 1, N par = N, and x = Length(L N ), and return to step 202 to obtain the next index c and adjacent node; add the adjacent node to sequence L N ; if sequence L N is full, and the parent node N par is not empty, then find the order of the special mark “*” in sequence L N , denoted as y, as the breakpoint number of the current node N, record the four-tuple [N par , x, N, y] as an edge, and return the index c and the current node N; recursively iterate to obtain a set of fragment representations and a set of four-tuple undirected edges;

[0028] Step 204: According to the fragment representations and four-tuple undirected edges obtained in step 203, re-establish the chemical bonds between the fragments to form a complete molecule.

[0029] Further, the fragment sequence encoder is configured to:

[0030] Encode the fragment sequence of the ligand molecule based on a multi-layer gated recurrent unit to obtain a molecular representation H g : the feature representation of a given fragment sequence where n is the length of the fragment sequence, and d is the fragment feature dimension, i.e., the dictionary size; input the fragments in the sequence one by one into the encoder, and take the hidden layer of the gated recurrent unit in the final state as the molecular representation H g .

[0031] Further, the target point encoder is configured to:

[0032] The multi-layer Transformer model based on the multi-head attention mechanism encodes the protein amino acid sequence to obtain an amino acid feature sequence: the protein amino acid sequence is represented as where m is the length of the amino acid sequence, and d is the feature dimension, i.e., the amino acid type; the Transformer layers are stacked and connected in residual, to obtain an amino acid feature sequence S' = [s'1,..., s'm]. m ];

[0033] The target point contribution degree perceptron based on the multi-layer perceptron obtains a target point representation: the amino acid feature sequence is analyzed and estimated by the target point contribution degree perceptron of the multi-layer perceptron to obtain a contribution degree a i = Softmax(MLP α (s′ i )), where Softmax represents a normalized exponential function, and MLP α represents a multi-layer perceptron. The target point representation is obtained by weighted sum of the target point features

[0034] The actual contribution degree is expressed based on a coefficient negatively correlated with the distance from the target point center to each residue:

[0035]

[0036] where d i is the distance from the i-th residue centroid to the binding site center, and k and τ are hyperparameters;

[0037] The KL divergence between the contribution degree a i and the actual contribution degree b i is minimized based on the contribution degree loss KL(a||b) to guide the target ligand molecule generation model to perceive the target point information and focus on predicting the ligand molecules adapted to the residues in the pocket.

[0038] Further, the feature fusion includes the following steps:

[0039] Step 301: splice the molecule representation H g and the target point representation S g to obtain a vector T, obtain the mean value m and the variance s in the complex feature hidden space containing the molecule-protein interaction information through the recognition network, and obtain the hidden vector z through the resampling technique, i.e.:

[0040] m = MLP μ (T), s = MLP σ (T)

[0041]

[0042] wherein MLP μ and MLP σ are multi-layer perceptrons;

[0043] Step 302: concatenate the latent vector z with the target point representation S g and obtain the sampled feature representation C.

[0044] Further, the fragment sequence decoder is configured to:

[0045] receive the sampled feature representation C as an initial hidden layer parameter based on a multi-layer gated recurrent unit; use a known sequence in a model training process for training, and iteratively output a prediction result in a target ligand molecule generation process.

[0046] The present application has the following beneficial effects:

[0047] (1) The present application designs a generation algorithm based on molecular fragments, which can ensure the synthesizability of the generated molecules, and improves the overall target chemical property expression of the molecules by utilizing the individual chemical properties of the fragments and the similarity between individuals;

[0048] (2) The present application provides a generation method based on protein sequences, and simultaneously predicts a binding site to guide ligand molecule generation, thereby expanding the application scope of the model;

[0049] (3) The present application uses sequences as the output of the model for post-processing, which can improve the fault tolerance of the model, and provides convenience for utilizing the excellent algorithms in the current natural language processing field, thereby further improving the performance of the model and realizing interdisciplinary and industry empowerment. BRIEF DESCRIPTION OF DRAWINGS

[0050] Figure 1 is a flow chart of a target ligand generation method based on active fragments according to Embodiment 1 of the present application.

[0051] Figure 2 is a contribution degree perceptron prediction result comparison chart according to Embodiment 2 of the present application. DETAILED DESCRIPTION

[0052] In order to have a more clear understanding of the technical features, objectives and effects of the present application, the specific embodiments of the present application will now be described. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application, i.e., the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative efforts fall within the scope of protection of the present application.

[0053] Example 1

[0054] like Figure 1 As shown, this embodiment provides a method for generating targeted ligands based on active fragments, including a fragment serialization process, a model training process, and a targeted ligand molecule generation process. The fragment serialization process involves processing the ligand molecules used for training using fragmentation and serialization methods to obtain fragment sequences. The model training process involves inputting the ligand molecule fragment sequences and the target protein amino acid sequences used for training into a targeted ligand molecule generation model. A fragment sequence encoder extracts the molecular characterization of the ligand molecule fragment sequences, and a target encoder extracts the target characterization of the target protein amino acid sequences. After feature fusion, a fragment sequence decoder outputs the fragment sequences, and the targeted ligand molecule generation model is optimized. The targeted ligand molecule generation process involves inputting the target protein amino acid sequences into the trained targeted ligand molecule generation model. A target encoder extracts the target characterization, and a priori network predicts complex information, which is then input into the fragment sequence decoder. The fragment sequence decoder outputs the fragment sequences, and the fragment sequences are then reconstructed into targeted ligand molecules based on a molecular reconstruction method.

[0055] In this embodiment, the fragmentation method, the serialization method, and the molecular reconstruction method together constitute Algorithm 1: Active Fragment Generation, Serialization, and Molecular Reconstruction Algorithm; the fragment sequence encoder, the target encoder, and the fragment sequence decoder together constitute Algorithm 2: Targeted Active Fragment Sequence Production Algorithm, as detailed below.

[0056] I. Algorithm 1: Active Fragment Generation, Serialization, and Molecular Reconstruction Algorithm

[0057] (1) Fragmentation method

[0058] This embodiment fragments molecules based on the Breaking of Retrosynthetically Interesting Chemical Substructures (BRICS) principle. BRICS breaks chemical bonds in a molecule that can undergo chemical reactions, and then attaches "virtual" atoms (atomic number 0) to each end of the cleavage site. BRICS preserves chemically significant molecular components, ensuring their integrity, and only breaks single bonds between aromatic rings and side chains.

[0059] Based on the decomposition rules of reversible synthetic chemical substructures, all breakable chemical bonds in the ligand molecule are matched, all breakable chemical bonds are broken, and virtual atoms are connected to the atoms at both ends of all breakable chemical bonds as breakpoints and marked. This is used to query the fragments at both ends of any breakable chemical bond and the breakpoints within the fragments.

[0060] (2) Serialization method

[0061] The serialization method of the present embodiment comprises the following steps:

[0062] Step 101: Construct a node using the above fragmentation method to obtain all fragments marked with breakpoints; sort the breakpoints in each fragment according to the RDKit atomic index sequence number from small to large, and assign a breakpoint number x = 1, 2, …, P to each breakpoint N , P N is the number of breakpoints on the fragment, i.e. the node; construct a quadruple [N1, x, N2, y] for each original breakable chemical bond, indicating that there is an edge between the xth breakpoint on node N1 and the yth breakpoint on node N2; if there is an edge between a node A x and the current node N, then A x is called the adjacent node of the current node N;

[0063] Step 102: Take the fragment with the largest relative mass as the initial node, denoted as node N, at this time node N as the root node, the parent node N par is empty;

[0064] Step 103: Add node N to the initial sequence L with length 0, and traverse the adjacent nodes A N of node N according to the breakpoint number order x = 1, 2, …, P x and the undirected edge according to the quadruple;

[0065] Step 104: If node N has a parent node N par , and the adjacent node A x is the parent node N par of node N, then add a special mark "*" to the sequence L; otherwise, N par = N, N = A x , return to step 103 for recursive iteration, and convert the nodes in the sequence L into fragment representations one by one, and output a fragment sequence containing a special mark "*".

[0066] Preferably, after the serialization method ends, the consistency of the molecular SMILES (Simplified Molecular Input Line Entry System) string is verified using a molecular reconstruction method.

[0067] Preferably, in step 101, for each fragment, the equivalence relationship between its virtual atoms needs to be found: combine the virtual atoms two by two as [D x , D y ], and replace the virtual atoms D xReplace each virtual atom with a neutron count of 1, and replace all other virtual atoms with virtual atoms with a neutron count of 0. Obtain the SMILES string representation S of this segment. x Replace the virtual atom Y with a virtual atom with 1 neutron, and replace the other virtual atoms with virtual atoms with 0 neutrons, to obtain the SMILES string representation S of this segment. y If the SMILES string represents S x With S y If they are the same, then the virtual atoms in this group are considered to be equivalent.

[0068] Next, according to the equivalence relation, the virtual atom sequence is sorted so that virtual atoms with equivalence relation are symmetrically distributed at the beginning and end of the sequence; if such a sequence exists, the segment is determined to be symmetrical, and the virtual atoms are marked with breakpoint numbers according to the sequence; otherwise, they are marked according to the initial order.

[0069] Preferably, in step 104, if If A1 is the parent node of node N, no special marker is needed; if A1 is the parent node of node N and the segment has symmetry, then the breakpoints are traversed in reverse, and no special marker is needed. This operation aims to shorten the sequence length. Correspondingly, in step 203 of the molecular reconstruction method, if Length(L N )= N -1, then it needs to be in L N Add a special symbol "*" to the middle.

[0070] (3) Molecular reconstruction methods

[0071] The molecular reconstruction method in this embodiment includes the following steps:

[0072] Step 201: For each valid node N in sequence L, obtain a segment representation, and determine the maximum number of adjacent nodes P based on the number of breakpoints in each node segment. N Let the index c = 0 on sequence L at this time, and the parent node N par If empty, the current parent node N par The breakpoint number x is empty;

[0073] Step 202: If node L[c] is not marked with a special symbol "*", then extract node N = L[c]. The sequence of adjacent nodes of node N is L. N The initial length is 0, proceed to step 203; if node L[c] is marked with a special symbol "*", then directly return the index c and the current node N;

[0074] Step 203: If sequence L N If not full, then set c = c + 1, C par =,x=Length(LN ), return step 202 to get the next index c and the adjacent node, and add the adjacent node to the sequence L N ; if the sequence L N is full, and the parent node N par is not empty, find the order of the special mark "*" in the sequence L N , denoted as y, as the breakpoint number of the current node N, record the four-tuple [ par ,,,] as an edge, and return the index c and the current node N; recursively iterate to obtain a set of fragment representations and a set of four-tuple undirected edges;

[0075] Step 204: According to the fragment representations and four-tuple undirected edges obtained in step 203, re-establish the chemical bonds between the fragments to form a complete molecule, and finally output the SMILES string of the molecule.

[0076] Preferably, the active fragment generation and sequencing and the fragment sequence obtained by the molecular reconstruction algorithm can be applied to other molecular generation tasks.

[0077] II. Algorithm 2: Targeted active fragment sequence production algorithm

[0078] As shown in Figure 1 , the targeted active fragment sequence production algorithm consists of a fragment sequence encoder, a target encoder, and a fragment sequence decoder. Two pre-training schemes are used in the model training process: a large-scale small molecule dataset for improving the performance of the fragment encoder and decoder, and a large-scale protein pre-training model for feature extraction of the target encoder; two self-supervised schemes: a target contribution sensor for target encoder pooling operation, and a Bag-of-words (BOW) decoder for improving the amount of information in the hidden space.

[0079] (1) Fragment sequence encoder

[0080] The fragment sequence of the ligand molecule is encoded based on a multi-layer Gated Recurrent Unit (GRU) to obtain a molecular representation H g : feature representation of a given fragment sequence where n is the length of the fragment sequence, and d is the fragment feature dimension, i.e., the dictionary size; the fragments in the sequence are input into the encoder one by one, and the hidden layer of the gated recurrent unit in the final state is taken as the molecular representation H g .

[0081] (2) Target encoder

[0082] The multi-head attention mechanism-based multi-layer Transformer model encodes the protein amino acid sequence to obtain an amino acid feature sequence: The Transformer model uses a self-attention mechanism instead of a recurrent structure, has parallel computing capability and better long sequence modeling capability. The protein amino acid sequence can be understood as a sequence composed of more than twenty kinds of words, and the protein amino acid sequence can be represented as where m is the length of the amino acid sequence, and d is the feature dimension, i.e., the amino acid type; the Transformer layers are stacked and connected in residual, to obtain an amino acid feature sequence S' = [s'1,..., s'm]. m ].

[0083] Preferably, in order to improve the quality of protein sequence feature extraction and reduce the training cost, the embodiment adopts a large-scale protein pre-training model ESM-2 weight parameter for the multi-layer Transformer model, which has great advantages in protein sequence, structure prediction and property prediction. In actual use, the embodiment designs a caching strategy to extract all amino acid sequence features, and splices and aligns them on demand in subsequent training.

[0084] The target contribution perceptron based on the multi-layer perceptron obtains a target representation: The target contribution perceptron based on the multi-layer perceptron analyzes and estimates the contribution of the amino acid feature sequence and the target point α i = Softmax(MLP α (s' i )) one by one, where Softmax represents a normalized exponential function, MLP α represents a multi-layer perceptron, and the target representation is obtained by weighted sum of the target feature

[0085] The embodiment designs a coefficient negatively related to the distance from the target center to each residue to express the actual contribution:

[0086]

[0087] where d i is the distance from the i-th residue centroid to the binding site center, and k and τ are hyperparameters, and the specific values are, for example, k = 20 and τ = 2.5.

[0088] The embodiment adds a contribution loss KL(α||β) to minimize the KL divergence of the contribution α i and the actual contribution β i In this way, the target ligand molecule generation model perceives the target information and focuses on predicting the ligand molecules adapted by the residues in the pocket.

[0089] Since the guidance of this target only relies on external loss function, the model can be applied to datasets that do not contain three-dimensional structures, broadening the data sources and application scope. In addition, the weight of the perception output can be analyzed for interpretability and visualization, and the target position of novel proteins can be predicted.

[0090] (3) Feature fusion

[0091] The feature fusion of the embodiment includes the following steps:

[0092] Step 301: splice the molecular representation H g and the target representation S g to obtain the vector T, and obtain the mean μ and variance σ in the complex feature hidden space containing the molecular and protein interaction information through the recognition network, and obtain the hidden vector z through the resampling technology, that is:

[0093] μ = MLP μ (T), σ = MLP σ (T)

[0094]

[0095] wherein MLP μ and MLP σ are both multilayer perceptrons;

[0096] Step 302: splice the hidden vector z and the target representation S g again to obtain the sampled feature representation C.

[0097] In the prediction stage, the model will directly predict the hidden space mean μ ′ and variance σ ′ of the fused molecular feature through the prior network, and obtain the hidden vector z ′ through the resampling technology to complete the subsequent process.

[0098] (4) Fragment sequence decoder

[0099] Based on the multilayer gated recurrent unit, the sampled feature representation C is received as the initial hidden layer parameter. In the model training process, known sequences are used for training, and in the process of generating targeted ligand molecules, the prediction results are iteratively output.

[0100] In addition, the embodiment parallelly connects a Bag-of-words (BOW) decoder similar to kgCVAE in the fragment sequence decoder, which is essentially a multi-classifier for receiving the feature representation C and calculating the probability of the occurrence of each fragment in the sequence through a multi-layer perceptron. By minimizing the error between the classification result and the unordered fragment set, the feature representation C can contain more category information of the fragments behind the sequence, reducing the dependence of the decoding process on the sequence of fragments and the starting fragment. In addition, this method can alleviate the problem that the KL divergence approaches 0 during the training process but the reconstruction error does not decrease, guiding the model to generate hidden vectors with more information.

[0101] Preferably, the encoder and decoder in the active fragment sequence production algorithm can be implemented by other models, such as other types of recurrent neural networks, convolutional neural networks, and other deep learning models; there are many alternatives for the generation method, such as diffusion models, generative adversarial networks, etc.

[0102] Preferably, the target contribution perceptron self-supervised scheme used in the embodiment can be applied to other protein-related tasks.

[0103] Preferably, the data set used to train the model can be replaced with other protein-ligand interaction data sets.

[0104] Embodiment 2

[0105] Based on embodiment 1, the embodiment of the present application provides a method for generating active fragment-based targeting ligands, which includes a model training process and a targeting ligand molecule generation process, and the specific description is as follows.

[0106] The embodiment provides an active fragment-based targeting ligand generation method, which includes a model training process and a targeting ligand molecule generation process, and the specific description is as follows.

[0107] I. Process 1: Model training process

[0108] (1) Data processing and fragment dictionary construction

[0109] The embodiment involves two data sets in the training process. In the pre-training process, a large-scale drug-like molecule data set ChEMBL is mainly used, which contains about 2.2 million small molecules. The active fragment generation and sequencing and molecule reconstruction algorithm (algorithm 1) of embodiment 1 is used to process the molecules in the data set one by one; all effective fragments in the sequence are extracted, only fragments with an occurrence frequency greater than or equal to 100 are retained, special reserved characters such as sequence start and end characters are added, and a fragment dictionary containing 5113 words is obtained. Among the 1 million fragment sequences covered by the dictionary, 80% are used as the training set and 20% are used as the validation set.

[0110] In the formal training, this embodiment selected the CrossDocked2020 dataset and adopted a similar processing method to 3D-SBDD for the dataset. Specifically, this dataset initially contained 225,000 protein-ligand pairings at different quality levels. First, data with a binding conformation RMSD greater than [value missing] were filtered out. The data yielded 184,057 data points. The mmseqs2 method was then used to cluster the data based on 30% sequence identity, and 100,000 protein-ligand pairs were randomly selected to form the training dataset. 80% of these pairs were used as the training set, 20% as the validation set, and the remaining 100 proteins were used as the test set. Fragment sequences were calculated using the same method, and all fragments were retained, resulting in a dictionary of length 7487, which was then merged into the dictionary from the pre-training stage.

[0111] The final fragment dictionary is 10701 in length. This dictionary and the processed dataset are then used for subsequent training and testing.

[0112] (2) Pre-training of segment sequence encoder and decoder

[0113] Using the aforementioned large-scale drug-like molecule dataset ChEMBL and its integrated dictionary, the model's fragment sequence encoder and decoder were pre-trained. In this stage, the input fragments are encoded into latent vectors solely by the fragment sequence encoder, while the target representation H is incorporated into the feature fusion process. g Set to zero, and predict the sequence using the fragment sequence decoder.

[0114] The purpose of pre-training is to expand its chemical space and enhance its ability to perceive commonalities among fragments. This effectively alleviates the overfitting phenomenon that may occur in fragment sequence encoders and decoders when the complex dataset is too small, and it is hoped that this measure can improve the legitimacy and diversity of the output molecules in the prediction stage.

[0115] This stage involves three loss functions: divergence loss, reconstruction error, and BOW loss. Divergence loss is the KL divergence between the probability distribution of the recognition network output and the normal distribution; reconstruction error is the cross-entropy between the decoder output sequence and the input sequence; and BOW loss is the cross-entropy between the predicted segment and the actual segment from the bag-of-words classifier. The Adam optimizer is used for optimization, selecting the weights with the smallest error on the validation set as the subsequent starting weights.

[0116] (3) Complete network training

[0117] The input fragment is encoded into molecular characterization S using a fragment sequence encoder. g And compared with the target characterization H extracted by the target encoder gAfter feature fusion and sampling, the input is input to the fragment sequence decoder to output the fragment sequence. In addition to the reconstruction error and BOW loss in the pre-training stage, the divergence loss is changed to the KL divergence between the z distribution output by the recognition network and the z distribution output by the prior network ′ The discrete KL divergence between the predicted value and the actual value of the target point contribution degree sensor contribution degree is calculated, and the value is optimized.

[0118] The Adam optimizer is used for optimization, and the weights after 200 rounds of training are selected as the final test benchmark, so that the target point sensor can capture the binding site position, and the prior network can predict the ligand molecule information from the target point information.

[0119] II. Process 2: Targeted molecule generation process

[0120] The input protein sequence or three-dimensional structure file is preprocessed in the same way as in process 1, and is input into the target point encoder to obtain the target point information representation; the complex information is directly predicted by the prior network and input into the fragment sequence decoder to output the fragment sequence. The molecule is reconstructed by the molecular reconstruction method in algorithm 1 to obtain the prediction result.

[0121] III. Test comparison

[0122] In order to reflect the advantages of the method of the present application, and to compare with the two best existing technologies, in this embodiment, 100 non-repeating molecules are generated for each of the 100 proteins in the test set, and the 10000 molecules are tested. The data of LiGAN and 3D-SBDD in Table 1 are from the article “A 3D generative model for structure-based drug design” published by Luo S et al. in 2021, and the names of the 100 proteins are known and the same as the test in this embodiment, the test method is similar to that of Luo S et al., and the time consumption data is measured by this embodiment. Due to the influence of random seeds, test equipment and environment, the data in Table 1 will have certain fluctuations when reproduced later.

[0123] Table 1-Comparison of ligand generation effects of three methods

[0124]

[0125] (1) The present application can generate drug-like molecules with better binding affinity

[0126] The ability of the generated molecules to effectively bind to the target is the most important indicator for evaluating the generated model. In this embodiment, the docking software AutoDock Vinav1.2.3 is used to dock the molecules generated in this embodiment with the corresponding target molecules, and the docking score is calculated. The lower the score, the easier it is to bind to the target. As can be seen from Table 1, the molecules generated by the present application have better docking scores with protein targets, indicating better binding ability with the target. In addition, the drugability of the molecules generated by the present application is better than the best existing technologies LiGAN and 3D-SBDD, while having a certain diversity, indicating that the present application can generate diverse drug-like molecules.

[0127] (2) The molecules generated by the present application are easier to synthesize

[0128] In order to evaluate whether the generated molecules are easy to synthesize, the synthesizability score of the generated molecules is calculated using RDKit. The score is between 0 and 1, and the higher the score, the easier it is to synthesize. Table 1 shows that the synthesizability score of the molecules generated by the present application is much lower than that of the two existing best technologies, indicating that the molecules generated by the present application are easier to synthesize.

[0129] (3) The molecules generated by the present application have higher efficiency

[0130] This embodiment counts the time for generating 100 molecules by various methods, as shown in Table 1. 3D-SBDD is the longest in time due to its self-regressive generation method, while the present application takes much less time than existing methods, having better generation efficiency.

[0131] (4) The present application has precise binding site perception ability

[0132] Figure 2 is the binding site of a protein sequence (2RMAChain A) in the test set predicted by the target contribution perception device of the present application. Overall, the trend of the predicted curve is basically the same as that of the actual curve, and the prediction of the 100th residue with strong contribution to the ligand is very accurate. This shows that the scheme has precise binding site perception ability and has obvious advantages in practicality and scope of application.

[0133] It should be noted that, for the foregoing method embodiments, in order to facilitate description, they are expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the order of the described actions, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

Claims

1. A method for the generation of active fragment-based targeting ligands, characterized in that, The method comprises the following steps: The fragment serialization process comprises the following steps: processing the ligand molecules for training based on a fragmentation method and a serialization method to obtain fragment sequences; The model training process comprises the following steps: inputting the ligand molecule fragment sequences and the target protein amino acid sequences for training into a targeted ligand molecule generation model, extracting the molecular representation of the ligand molecule fragment sequences by a fragment sequence encoder, extracting the target representation of the target protein amino acid sequences by a target encoder, fusing the features, outputting the fragment sequences by a fragment sequence decoder, and optimizing the targeted ligand molecule generation model; The targeted ligand molecule generation process comprises the following steps: inputting the target protein amino acid sequences into the trained targeted ligand molecule generation model, extracting the target representation by the target encoder, predicting the complex information by the prior network and inputting the complex information into the fragment sequence decoder, outputting the fragment sequences by the fragment sequence decoder, and recombining the fragment sequences into the targeted ligand molecule based on a molecular reconstruction method; The fragmentation method comprises the following steps: based on the reversible synthetic chemical substructure decomposition rule, matching all breakable chemical bonds in the ligand molecules, breaking all the breakable chemical bonds, connecting virtual atoms to the atoms at both ends of all the breakable chemical bonds as break points and marking the virtual atoms, and querying the fragments at both ends of any breakable chemical bond and the break points in the fragments; The target encoder is configured to: The multi-layer Transformer model based on the multi-head attention mechanism encodes the protein amino acid sequence to obtain an amino acid feature sequence: the protein amino acid sequence is represented as , where m is the length of the amino acid sequence, and d is the feature dimension, i.e., the amino acid type; the Transformer layers are stacked and connected in a residual manner to obtain the amino acid feature sequence . The target point contribution degree perceptron based on the multilayer perceptron obtains a target point representation: the target point contribution degree perceptron based on the multilayer perceptron analyzes the contribution degree of the estimated amino acid feature sequence and the target point one by one wherein Softmax represents a normalized exponential function, represents a multilayer perceptron, and a target point representation is obtained by weighted summation of target point features ; Express the actual contribution degree based on a coefficient negatively correlated with the distance from the target center to each residue: wherein is the distance from the center of the residue to the center of the binding site, is the distance from the center of the residue to the center of the binding site, and is a hyperparameter; Based on the contribution loss To minimize the contribution And the actual contribution Of the KL divergence to guide the target ligand molecule generation model to perceive the target point information and focus on predicting the ligand molecules adapted to the residues in the pocket.

2. The method for generating an active fragment-based targeting ligand according to claim 1, wherein, The serialization method comprises the following steps: Step 101: Construct a fragmentation method to obtain all fragments with breakpoint markers, treating each fragment as a node; sort the breakpoints within each fragment in ascending order according to their RDKit atomic index numbers, and assign a breakpoint number to each breakpoint. , This represents the number of breakpoints on a fragment, i.e., a node; a quadruple is constructed for each of the original breakable chemical bonds. , indicating the node On Breakpoint and Node On There are edges between the breakpoints; if a node With the current node If there is an edge between them, then it is called For the current node The adjacent nodes; Step 102: take the fragment with the largest relative mass as the initial node, denoted as node At this time, node As the root node, the parent node Is empty; Step 103: add the node N to the sequence of nodes with initial length 0 , in the order of the breakpoint numbers , according to the quadruple undirected edges, traverse the adjacent nodes of the node N ; Step 104: If node has a parent node and the adjacent node is the parent node of node , then a special mark "*" is added to the sequence ; otherwise, set , , return to step 103 for recursive iteration, and convert the nodes in the sequence into fragment representations in turn, and output a fragment sequence containing the special mark "*".

3. The method for generating a target ligand based on an active fragment according to claim 2, characterized in that, In step 101, for each fragment, the equivalence relationship between its virtual atoms needs to be found: combining the virtual atoms two by two into , and replacing the virtual atom with a virtual atom of neutron number 1, and replacing the other virtual atoms with virtual atoms of neutron number 0, to obtain the SMILES string expression of the fragment at this time ; replacing the virtual atom Y with a virtual atom of neutron number 1, and replacing the other virtual atoms with virtual atoms of neutron number 0, to obtain the SMILES string expression of the fragment at this time ; if the SMILES string expressions are the same as , it is judged that the current group of virtual atoms has an equivalence relationship; According to the equivalence relation, the virtual atom sequences are sorted, the virtual atoms with the equivalence relation are symmetrically distributed at the beginning and end of the sequence, if such a sequence exists, it is determined that the fragment has symmetry, and the virtual atoms are marked with break point numbers according to the sequence; otherwise, the initial order is marked.

4. The method for generating a target ligand based on an active fragment according to claim 2, wherein, In step 104, if is the parent node of node N, no special mark is needed; if is the parent node of node N and the segment has symmetry, the breakpoint is traversed reversely and no special mark is needed.

5. The active fragment-based targeted ligand generation method of claim 1, wherein, The molecular reconstruction method comprises the following steps: Step 201: For the sequence Each valid node in Obtain fragment representations and determine the maximum number of adjacent nodes based on the number of breakpoints in each node fragment. Let the sequence at this time be Index on Parent node Empty, current parent node Breakpoint number Empty; Step 202: If the node is not the special marker "*", the node is extracted, the sequence of adjacent nodes of the node N is with an initial length of 0, and step 203 is entered; if the node is the special marker "*", the index and the current node are returned directly; Step 203: if the sequence is not full, set , , , return to step 202 to get the next index and the adjacent node, add the adjacent node to the sequence ; if the sequence is full and the parent node is not empty, find the order of the special mark "*" in the sequence , record it as , take it as the breakpoint number of the current node , record the quadruple as an edge, return the index and the current node ; recursively iterate to obtain a group of fragment representations and a group of quadruple undirected edges; Step 204: according to the fragment representation and the four-tuple undirected edge obtained in step 203, reestablishing chemical bonds between the fragments to form a complete molecule.

6. The method for the production of a target ligand based on active fragments according to any one of claims 1 to 5, characterized in that, The fragment sequence encoder is configured to: The fragment sequence of the ligand molecule is encoded based on a multi-layer gated recurrent unit to obtain a molecular representation : the feature representation of a given fragment sequence , where n is the length of the fragment sequence and d is the fragment feature dimension, i.e., the dictionary size; the fragments in the sequence are input into the encoder one by one, and the hidden layer of the gated recurrent unit in the final state is taken as the molecular representation .

7. The active-fragment based targeted ligand generation method according to claim 1, wherein, The feature fusion comprises the following steps: Step 301: Characterize molecules and target points Concatenate to get vector Get mean and variance in complex feature latent space containing molecule-protein interaction information by identifying network and variance Get latent vector by resampling technique i.e.: wherein and are both multilayer perceptrons; Step 302: obtaining the latent vector With target characterization The sampled feature representation C is obtained by splicing again.

8. The method for generating a target ligand based on an active fragment according to claim 7, wherein, The fragment sequence decoder is configured to: Receiving a sampled feature representation based on a multi-layer gated recurrent unit , as initial hidden layer parameters; using known sequences for training during a model training process, and iteratively outputting predicted results during a targeted ligand molecule generation process.

Citation Information

Patent Citations

  • Universal method for drug design aiming at different target proteins

    CN114464270A

  • System and method for prediction of protein-ligand bioactivity and pose propriety

    US11256994B1