Method and device for training ligand information generation model, and method and device for generating ligand information
The ligand information generation model is trained through neural network models, which solves the problem of generating ligands with high binding affinity with receptors and improves drug development efficiency.
Patent Information
- Application Number
- PCT/CN2025/077641
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-20
- Filing Date
- 2025-02-17
- Publication Date
- 2025-08-28
AI Technical Summary
How to produce ligands that bind to receptors has become an urgent problem. It is difficult for the prior art to efficiently generate ligands with high binding affinity to receptors.
By obtaining sample receptor information and sample ligand information, using neural network model to denoise the reference noise data, the ligand information generation model is obtained, and the reference ligand information is generated based on the reference receptor information to ensure high binding affinity.
It realizes the generation of ligand information with high binding affinity, improves drug research and development efficiency, and can accurately output ligand information.
Smart Images

Figure CN2025077641_28082025_PF_FP_ABST
Abstract
Description
Training of ligand information generation model, method and device for generating ligand information
[0001] This application claims priority to Chinese patent application No. 202410192333.2, filed on February 20, 2024, entitled “Training of ligand information generation model, method and device for generating ligand information”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a method and device for training a ligand information generation model, and generating ligand information. Background Art
[0003] In the drug design industry, artificial intelligence (AI) technology can be used to design ligands. Through ligand experimental verification, therapeutically effective ligand-based drugs can be obtained. Generally, ligands bind to receptors, triggering signaling and biochemical reactions, thereby influencing receptor function and exerting therapeutic effects. Therefore, generating ligands that can bind to receptors has become a pressing issue. Summary of the Invention
[0004] The present application provides a method and device for training a ligand information generation model and generating ligand information, which can train a ligand information generation model, generate reference ligand information based on reference receptor information through the ligand information generation model, and the binding affinity between the ligand described by the reference ligand information and the receptor described by the reference receptor information is high. The technical solution includes the following contents.
[0005] In a first aspect, a method for training a ligand information generation model is provided, the method being executed by an electronic device and comprising:
[0006] Acquiring sample receptor information and sample ligand information, wherein the binding affinity between the ligand described by the sample ligand information and the receptor described by the sample receptor information is not less than the set affinity;
[0007] Denoising the reference noise data based on the sample receptor information using a neural network model to be trained to obtain predicted ligand information;
[0008] determining a first loss for characterizing a difference between the sample ligand information and the predicted ligand information;
[0009] The neural network model is trained based on the first loss to obtain a ligand information generation model, and the ligand information generation model is used to generate reference ligand information based on reference receptor information.
[0010] In a second aspect, a method for generating ligand information is provided, the method being executed by an electronic device, the method comprising:
[0011] Obtain reference receptor information;
[0012] The reference noise data is denoised based on the reference receptor information by a ligand information generation model to obtain reference ligand information. The ligand information generation model is trained according to the training method of the ligand information generation model described in the first aspect.
[0013] In a third aspect, a training device for a ligand information generation model is provided, the device comprising:
[0014] an acquisition module, configured to acquire sample receptor information and sample ligand information, wherein the binding affinity between the ligand described by the sample ligand information and the receptor described by the sample receptor information is not less than a set affinity;
[0015] a denoising module, configured to denoise the reference noise data based on the sample receptor information using a neural network model to be trained to obtain predicted ligand information;
[0016] a determining module, configured to determine a first loss for characterizing a difference between the sample ligand information and the predicted ligand information;
[0017] A training module is used to train the neural network model based on the first loss to obtain a ligand information generation model, wherein the ligand information generation model is used to generate reference ligand information based on reference receptor information.
[0018] In a fourth aspect, a device for generating ligand information is provided, the device comprising:
[0019] An acquisition module, used for acquiring reference receptor information;
[0020] A denoising module is used to denoise the reference noise data based on the reference receptor information through a ligand information generation model to obtain reference ligand information. The ligand information generation model is trained according to the training method of the ligand information generation model described in the first aspect.
[0021] In a fifth aspect, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor so that the electronic device implements the training method for the ligand information generation model described in the first aspect or the ligand information generation method described in the second aspect.
[0022] In the sixth aspect, a non-volatile computer-readable storage medium is also provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to enable the electronic device to implement the training method of the ligand information generation model described in the first aspect or the ligand information generation method described in the second aspect.
[0023] In the seventh aspect, a computer program is also provided, and the computer program is at least one, and at least one computer program is loaded and executed by a processor to enable the electronic device to implement the training method of the ligand information generation model described in the first aspect or the ligand information generation method described in the second aspect.
[0024] In the eighth aspect, a computer program product is also provided, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor so that the electronic device can implement the training method of the ligand information generation model described in the first aspect or the ligand information generation method described in the second aspect.
[0025] The technical solution provided in this application first uses a neural network model to denoise reference noise data based on sample receptor information to obtain predicted ligand information, then determines a first loss based on the sample ligand information and the predicted ligand information, and trains the neural network model based on the first loss, so that the neural network model is optimized in the direction of making the predicted ligand information close to the sample ligand information, so that the model can accurately output the ligand information. In addition, because the binding affinity between the ligand described by the sample ligand information and the receptor described by the sample receptor information is not less than the set affinity, the trained ligand information generation model can generate reference ligand information based on the reference receptor information, and the binding affinity between the ligand described by the reference ligand information and the receptor described by the reference receptor information is high. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] FIG1 is a schematic diagram of an implementation environment for a training method for a ligand information generation model or a method for generating ligand information provided in an embodiment of the present application;
[0027] FIG2 is a flow chart of a training method for a ligand information generation model provided in an embodiment of the present application;
[0028] FIG3 is a schematic diagram of an attention processing method provided by an embodiment of the present application;
[0029] FIG4 is a schematic diagram of a training method for a ligand information generation model provided in an embodiment of the present application;
[0030] FIG5 is a flow chart of a method for generating ligand information provided in an embodiment of the present application;
[0031] FIG6 is a schematic diagram of a training method for a polypeptide sequence denoising diffusion model using a sample target protein as a condition, provided in an embodiment of the present application;
[0032] FIG7 is a schematic diagram comparing a natural polypeptide and a produced polypeptide provided in an example of the present application;
[0033] FIG8 is a schematic diagram comparing another natural polypeptide and a produced polypeptide provided in the Examples of the present application;
[0034] FIG9 is a schematic diagram of a docking score and physicochemical similarity provided in an embodiment of the present application;
[0035] FIG10 is a schematic diagram comparing another natural polypeptide and a produced polypeptide provided in the Examples of the present application;
[0036] FIG11 is a schematic structural diagram of a training device for a ligand information generation model provided in an embodiment of the present application;
[0037] FIG12 is a schematic structural diagram of a device for generating ligand information provided in an embodiment of the present application;
[0038] FIG13 is a schematic structural diagram of a terminal device provided in an embodiment of the present application;
[0039] FIG14 is a schematic diagram of the structure of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0040] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0041] In the drug design industry, ligands can be designed based on artificial intelligence technology. Through experimental validation of these ligands, therapeutically effective ligand-based drugs can be obtained. For example, a ligand is a polypeptide sequence, a compound formed by the dehydration condensation of at least three amino acids. Generally, ligands bind to receptors, influencing their function and thus exerting therapeutic effects. Therefore, generating ligands that can bind to receptors has become a pressing issue.
[0042] The present application provides a method for training a ligand information generation model or a method for generating ligand information, which can be trained to obtain a ligand information generation model. The ligand information generation model generates reference ligand information based on reference receptor information, and the binding affinity between the ligand described by the reference ligand information and the receptor described by the reference receptor information is high, which is conducive to accelerating the development of ligand drugs.
[0043] As shown in Figure 1, Figure 1 is a schematic diagram of an implementation environment for a method for training a ligand information generation model or a method for generating ligand information provided in an embodiment of the present application, and the implementation environment includes a terminal device 101 and a server 102. The method for training a ligand information generation model or a method for generating ligand information in the embodiment of the present application can be executed by the terminal device 101, can be executed by the server 102, or can be executed jointly by the terminal device 101 and the server 102.
[0044] The terminal device 101 can be a mobile phone or a computer. Mobile phones include but are not limited to smart phones, foldable phones, and flip phones. Computers include but are not limited to game consoles, desktop computers, tablet computers, and laptop portable computers. In actual applications, the terminal device 101 can also include smart TVs, smart car-mounted devices, intelligent voice interaction devices, smart home appliances, etc. The server 102 can be a single server, or a server cluster consisting of multiple servers, or any one of a cloud computing platform and a virtualization center, which is not limited in the embodiments of the present application. The server 102 can be connected to the terminal device 101 through a communication network, which is a wired network or a wireless network. The server 102 can have functions such as data processing, data storage, and data transmission and reception, which are not limited in the embodiments of the present application. The number of terminal devices 101 and servers 102 is not limited and can be one or more.
[0045] In an exemplary embodiment, the terminal device 101 or the server 102 pre-trains a ligand information generation model based on sample receptor information and sample ligand information. When reference receptor information is acquired, the reference receptor information is input into the ligand information generation model, which then generates and outputs reference ligand information based on the reference receptor information.
[0046] In an exemplary embodiment, server 102 pre-trains a ligand information generation model based on sample receptor information and sample ligand information. When terminal device 101 obtains reference receptor information, it transmits the reference receptor information to server 102 via the communication network. After receiving the reference receptor information, server 102 inputs the reference receptor information into the ligand information generation model. The ligand information generation model then generates and outputs reference ligand information based on the reference receptor information, and transmits the reference ligand information to terminal device 101 via the communication network.
[0047] Please refer to Figure 2, which is a flowchart of a method for training a ligand information generation model provided in an embodiment of the present application. For ease of description, the terminal device 101 or server 102 that executes the method for training a ligand information generation model in the embodiment of the present application is referred to as an electronic device, meaning that the method in the embodiment of the present application can be executed by an electronic device. As shown in Figure 2, the method includes the following steps.
[0048] Step 201 : Obtain sample receptor information and sample ligand information. The binding affinity between the ligand described by the sample ligand information and the receptor described by the sample receptor information is not less than a set affinity.
[0049] In biology, a receptor is a substance that can bind to a ligand, thereby inducing a biological effect. These biological effects include antioxidant activity, blood pressure reduction, and immune cell activation. For example, a receptor can be a protein, nucleic acid sequence, cell membrane, or antigen. Generally, a receptor is also called a target, and a ligand is a substance that can bind to a receptor. For example, a ligand can be a protein, peptide sequence, small molecule, or antibody. A peptide sequence is formed by the dehydration condensation of at least three amino acids. The residue remaining after the dehydration condensation is called an amino acid residue.
[0050] The sample receptor information is used to describe the sample receptor, and specifically can be characteristic data used to describe the sample receptor. The sample receptor can be any type of receptor. For example, if the sample receptor is a protein, the sample receptor information includes the encoding and three-dimensional coordinates of each amino acid residue that constitutes the protein. The sample ligand information is used to describe the sample ligand. Specifically, it can be characteristic data used to describe the sample ligand. The sample ligand can be any type of ligand. For example, if the sample ligand is a polypeptide sequence, the sample ligand information includes the encoding and three-dimensional coordinates of each amino acid residue that constitutes the polypeptide sequence. The method of encoding amino acid residues is not limited here. For example, a one-hot encoding method is used to encode various amino acid residues, so that different amino acid residues correspond to different encodings. That is, the encoding of the amino acid residue is unique and is used to indicate an amino acid residue.
[0051] Ligands can bind to receptors. It is understandable that different ligands bind to different receptors with different strengths. In the field of biology, binding affinity is used to describe the strength of the binding between a ligand and a receptor. The higher the binding affinity, the higher the strength of the binding between the ligand and the receptor. In this example, the binding affinity between the sample ligand and the sample receptor is not less than the set affinity. If the binding affinity between the ligand and the receptor is not less than the set affinity, it means that the binding strength between the ligand and the receptor is high, and if the binding affinity between the ligand and the receptor is less than the set affinity, it means that the binding strength between the ligand and the receptor is low. The embodiments of the present application do not limit the method for determining the set affinity and the numerical value. For example, the set affinity is a numerical value determined based on artificial experience, or the set affinity is a numerical value obtained through experiments. The experimental process will not be repeated here.
[0052] The embodiments of the present application do not limit the method for obtaining sample receptor information and sample ligand information. For example, the electronic device can obtain sample receptor information and sample ligand information input by the user. Alternatively, the electronic device can collect a target polypeptide data set from the Internet, the target polypeptide data set including at least two target polypeptide pairs, one target polypeptide pair including target information and a polypeptide sequence, and the binding affinity between the target protein described by the target information and the polypeptide described by the polypeptide sequence is not less than the set affinity, wherein the target information and the polypeptide sequence both include at least two amino acid texts (such as amino acid words or amino acid characters). The electronic device can train a feature extraction model, and extract the feature data of any target polypeptide pair through the feature extraction model to obtain sample receptor information and sample ligand information. Among them, the feature extraction model is a network for extracting feature data of amino acid texts, and its structure and feature extraction method are not limited here. In this example, the feature data extracted by the feature extraction model is also called embedding (Embedding), based on this, the feature extraction model is called an embedding transformation function. That is, the electronic device maps each amino acid text in the target information into a corresponding text feature through the embedding transformation function to obtain sample receptor information. Similarly, each amino acid text in the peptide sequence is mapped into corresponding text features by embedding a transformation function to obtain the sample ligand information.
[0053] Optionally, for any target peptide pair, the target peptide pair includes target information x and peptide sequence y. Assume that the target information x includes m amino acid texts, and the m amino acid texts are sequentially represented as: x1,…,x m Similarly, the polypeptide sequence y includes n amino acid texts, which are represented in sequence as: y1,…,y n Then: After mapping the target peptide pair by embedding the transformation function EMB(w), we can get the sample receptor information EMB(x1),…,EMB(x m ) and sample ligand information EMB(y1),…,EMB(y n ). The mapping process can be characterized as: in, Characterizes the splicing peptide sequence y after or before the target information x, EMB(x1) characterizes the text feature of the first amino acid included in the target information x, EMB(x m ) represents the text features of the mth amino acid included in the target information x, EMB(y1) represents the text features of the first amino acid included in the polypeptide sequence y, and EMB(y n ) characterizes the text features of the nth amino acid included in the polypeptide sequence y.
[0054] It can be understood that the number of sample data pairs is at least two. Any two sample data pairs include the same or different sample receptor information, and similarly, any two sample data pairs include the same or different sample ligand information.
[0055] Step 202: De-noise the reference noise data based on the sample receptor information using the neural network model to be trained to obtain predicted ligand information.
[0056] In an embodiment of the present application, an electronic device may obtain reference noise data for describing a certain noise. The noise may be any noise, for example, the noise includes at least one of Gaussian noise, Poisson noise, salt and pepper noise, etc. Optionally, the reference noise data is used to describe noise whose probability density function obeys a statistical distribution. The statistical distribution is not limited here. For example, the statistical distribution includes at least one of a normal distribution, a U-type distribution, a J-type distribution, etc. Among them, the probability density function is a function commonly used in statistics and probability theory, which is used to describe the probability of a continuous random variable near a certain value point. The noise whose probability density function obeys a normal distribution is Gaussian noise, and the reference noise data may be data used to describe Gaussian noise. The embodiment of the present application does not limit the method for obtaining the reference noise data. For example, the electronic device may randomly generate the reference noise data, or the electronic device may obtain the reference noise data input by the user, etc.
[0057] Reference noise data can be input into the neural network model, and sample receptor information can also be input into the neural network model. In one possible implementation, the reference noise data is concatenated before or after the sample receptor information to generate concatenated information, which is then input into the neural network model. The neural network model is a network model used to remove noise data. It is used to denoise the reference noise data based on the sample receptor information to obtain predicted ligand information. The predicted ligand information is characteristic data used to characterize the ligand, as predicted by the neural network model.
[0058] The embodiments of this application do not limit the model structure, model parameters, etc. of the neural network model. Exemplarily, the neural network model includes at least one of a linear layer, a nonlinear layer, an activation layer, an attention layer, a convolutional layer, a normalization layer, etc. It is understandable that different neural network model structures will result in different denoising methods.
[0059] In one possible implementation, the neural network model includes at least two denoising networks. Step 202 includes steps 2021 to 2023 (not shown). It is understood that at least two denoising networks are connected in series. A denoising network is a network used to remove noise data, and the structures and functions of each denoising network are similar.
[0060] Step 2021: For the first denoising network, the reference noise data is denoised based on the sample receptor information by the first denoising network to obtain a denoising result of the first denoising network.
[0061] In an embodiment of the present application, the sample receptor information and the reference noise data can be input into the first denoising network, or the sample receptor information can be spliced before or after the reference noise data to obtain spliced information, and the spliced information can be input into the first denoising network. The spliced information is denoised by the first denoising network to obtain the denoising result of the first denoising network.
[0062] The embodiment of the present application does not limit the network structure, model parameters, etc. of the first denoising network. Exemplarily, the first denoising network includes at least one of a linear layer, a nonlinear layer, an activation layer, an attention layer, a convolution layer, a normalization layer, etc. It is understandable that the structure of the first denoising network is different, and the denoising processing method is also different. Optionally, the first denoising network is a BERT (Bidirectional Encoder Representations from Transformers) model or a Transformers network.
[0063] In an exemplary embodiment, the first denoising network includes an attention network and a feedforward network. Step 2021 includes steps A1 to A2 (not shown in the figure). Optionally, the attention network includes at least one of a multi-head attention network, a self-attention network, a masked attention network, etc. The feedforward network includes at least one of a linear layer and a nonlinear layer, etc.
[0064] Step A1: Perform attention processing on the sample receptor information and reference noise data through the attention network to obtain the attention processing result.
[0065] In the embodiments of the present application, sample receptor information and reference noise data are input into an attention network, or the concatenated information obtained by concatenating the sample receptor information and reference noise data is input into the attention network. The attention network then performs attention processing based on the attention mechanism to obtain an attention processing result. It is understood that different attention networks correspond to different attention processing methods. One possible attention processing method is shown below.
[0066] Exemplarily, the sample receptor information includes at least two receptor composition information items, and the reference noise data includes at least one first sub-data item. Step A1 includes: for any first sub-data item, performing attention processing on the first sub-data item and each piece of content information via an attention network to obtain processing results for each piece of content information item; based on the processing results for each piece of content information item, determining the second sub-data item corresponding to each piece of first sub-data item, where each piece of content information item includes at least two pieces of receptor composition information and at least one first sub-data item; and determining the attention processing results based on the second sub-data item corresponding to each piece of first sub-data item.
[0067] In an embodiment of the present application, the sample receptor information used to describe the sample receptor includes at least two pieces of receptor composition information, and any one piece of receptor composition information is used to describe a substance that constitutes the sample receptor. For example, if the sample receptor is a protein, and the protein includes at least one peptide chain, then any one piece of receptor composition information can be used to describe a single peptide chain. For example, the receptor composition information includes the encoding and three-dimensional coordinates of each amino acid residue in the peptide chain. Alternatively, because a peptide chain includes at least two amino acid residues, any one piece of receptor composition information can be used to describe a single amino acid residue. For example, the receptor composition information includes the encoding of the amino acid residue.
[0068] Similarly, the reference noise data used to describe a certain noise includes at least one first sub-data. The number of such noises is at least one, and any first sub-data is used to describe a noise. It is understood that any two first sub-data may describe the same or different noises. For example, two first sub-data may be different data describing Gaussian noise, or one first sub-data may be data describing Gaussian noise and the other first sub-data may be data describing salt and pepper noise.
[0069] In an embodiment of the present application, each content information includes at least two receptor composition information and at least one first sub-data. That is to say, one receptor composition information is one content information, and one first sub-data is also one content information. Therefore, there are at least two content information. For any first sub-data, a linear mapping or nonlinear mapping is performed on the first sub-data through the attention network to obtain a query (Queue, Q) vector corresponding to the first sub-data. For any content information, two different linear mappings or nonlinear mappings are performed on the content information through the attention network to obtain the key (Key, K) vector and value (Value, V) vector corresponding to the content information.
[0070] Next, based on the query vector corresponding to the first sub-data and the key vector corresponding to the content information, an attention weight is determined, and the attention weight is multiplied by the value vector corresponding to the content information to obtain a processing result for the content information. Optionally, the processing result for the content information is determined according to formula (1) shown below.
[0071] Among them, A represents the processing result of content information. Attention represents the attention function used in attention processing, which includes three parameters Q, K, and V. Among them, Q represents the query vector corresponding to the first sub-data, K represents the key vector corresponding to the content information, and V represents the value vector corresponding to the content information. Softmax is a normalized exponential function, and T represents the transposed matrix. d k Characterizes the dimensions of the query vector and key vector. Representing attention weights.
[0072] It is understood that there are at least two pieces of content information, and any one of the first sub-data can be processed with each piece of content information to obtain the processing results of each piece of content information. Subsequently, the processing results of each piece of content information are summed, averaged, weighted, or otherwise calculated to obtain the second sub-data corresponding to the first sub-data.
[0073] Please refer to Figure 3, which is a schematic diagram of an attention processing provided by an embodiment of the present application. As shown in Figure 3, the reference noise data includes the first sub-data X to the first sub-data M, and the sample receptor information includes the receptor composition information P to the receptor composition information F. The first sub-data X is mapped by the attention network to obtain the query vector Q of the first sub-data X. X For each content information i (i takes any value from X to M and P to F) in the first sub-data X to the first sub-data M and the receptor composition information P to the receptor composition information F, the key vector K of the content information i is obtained by mapping the content information i through the attention network. i Sum value vector V i . The query vector Q based on the first sub-data X X and the key vector K of content information i i , determine the attention weight a Xi , and the attention weight a Xi and the value vector V of content information i i The processing results of the content information i are averaged to obtain the second sub-data K corresponding to the first sub-data X.
[0074] By similar processing, the second sub-data corresponding to each first sub-data except the first sub-data X can be obtained, which will not be described in detail here. Afterwards, the second sub-data corresponding to each first sub-data are concatenated to obtain the attention processing result.
[0075] On the one hand, the second sub-data corresponding to any first sub-data is obtained by performing an attention calculation on the first sub-data and the composition information of each receptor. Therefore, the second sub-data is obtained by processing the first sub-data with reference to the composition information of each receptor. This is equivalent to denoising the first sub-data to obtain the second sub-data under the guidance of the sample receptor information. By removing the noise data in the first sub-data under the guidance of the sample receptor information, the ability of the second sub-data to represent the ligand composition information is improved. By splicing the various second sub-data to obtain the attention processing result, the ability of the attention processing result to represent the ligand with the characteristics of binding to the sample receptor is improved.
[0076] On the other hand, the second sub-data corresponding to any first sub-data is obtained by performing attention calculation on the first sub-data and each first sub-data. Therefore, the second sub-data is obtained by processing the first sub-data with reference to each first sub-data, which improves the correlation between the second sub-data and the reference noise data, and makes each second sub-data correlated, thereby improving the accuracy of the attention processing result in characterizing the ligand.
[0077] It is understood that the above-described method for determining the attention processing result is merely exemplary, and other methods may be employed in practical applications. For example, on the one hand, the reference noise data is mapped by the attention network to obtain a query vector. On the other hand, the sample receptor information is mapped twice by the attention network to obtain a key vector and a value vector, respectively. Based on the query vector and the key vector, an attention weight is determined, and the attention weight is multiplied by the value vector to obtain the attention processing result.
[0078] In step A2, the denoising result of the first denoising network is determined based on the attention processing result through the feedforward network.
[0079] In an embodiment of the present application, the attention processing result can be input into a feedforward network, and the attention processing result can be processed by the feedforward network to obtain the processing result of the feedforward network. The processing result of the feedforward network is used as the denoising result of the first denoising network, or mapping, activation, or normalization processing is performed on the processing result of the feedforward network to obtain the denoising result of the first denoising network.
[0080] In an exemplary embodiment, the feedforward network includes at least one of a linear layer and a nonlinear layer, and the feedforward network can perform at least one processing such as linear mapping or nonlinear mapping on the attention processing result to obtain the processing result of the feedforward network.
[0081] Optionally, step A2 includes: mapping the attention processing result through a feedforward network to obtain a mapping result; and mapping the mapping result or setting feature mapping through a feedforward network to obtain a denoising result of a first denoising network.
[0082] In an embodiment of the present application, the feedforward network includes a first mapping layer, a selection layer, and a second mapping layer. The first mapping layer can be a linear layer or a nonlinear layer, and the first mapping layer performs linear mapping or nonlinear mapping on the attention processing result to obtain a mapping result. The selection layer selects the maximum value from the mapping result and the set feature. The maximum value here is the maximum value or the minimum value, which is also the mapping result or the set feature. The second mapping layer can be a linear layer or a nonlinear layer, and the second mapping layer performs linear mapping or nonlinear mapping on the maximum value to obtain the processing result of the feedforward network. Afterwards, the denoising result of the first denoising network is determined based on the processing result of the feedforward network.
[0083] It is understood that the set feature is a pre-set feature, which can be a number or a matrix. For example, the processing result of the feedforward network is determined according to the following formula (2). In formula (2), the set feature is the number 0. FFN(A) = max(0,AW1+b1)W2+b2 Formula (2)
[0084] Where FFN(A) represents the processing result of the feedforward network. max represents the maximum value function used by the selection layer. AW1+b1 represents the mapping result obtained by the first mapping layer, where A represents the attention processing result, W1 represents the weight parameter in the mapping function used in the first mapping layer, and b1 represents the bias parameter in the mapping function used in the first mapping layer. W2 represents the weight parameter in the mapping function used in the second mapping layer, and b2 represents the bias parameter in the mapping function used in the second mapping layer.
[0085] It is understandable that the data involved in the denoising performed by the first denoising network are all feature-level data. In other words, the sample receptor information, reference noise data, and denoising results are all feature-level data.
[0086] Optionally, the neural network model further includes a first embedding layer, which is connected in series before the first denoising network. The electronic device can obtain first media information for describing the sample receptor and second media information for describing the sample ligand. The first media information or the second media information includes at least one of text, audio, image, video, etc. For example, if the sample receptor is a protein, the first media information can be a visual image of the protein. For another example, if the sample ligand is a polypeptide, the second media information can include the type of amino acids included in the polypeptide, as well as the number and three-dimensional coordinates of each amino acid. The first media information and the second media information are both media information, and the media information includes at least two information segments. For example, when the media information is text, the media information includes at least two characters; when the media information is audio, the media information includes at least two audio frames; when the media information is an image, the media information includes at least two pixel values; when the media information is video, the media information includes at least two frame images.
[0087] The first embedding layer can map each information segment included in the first media information to a corresponding embedded feature to obtain sample receptor information. The first embedding layer can also map each information segment included in the second media information to a corresponding embedded feature to obtain sample ligand information. Next, the first denoising network denoises the reference noise data based on the sample receptor information to obtain a denoising result of the first denoising network.
[0088] Step 2022: For the non-first denoising network, the denoising result of the previous denoising network of the non-first denoising network is denoised based on the sample receptor information by the non-first denoising network to obtain the denoising result of the non-first denoising network.
[0089] In an embodiment of the present application, the denoising result of the first denoising network can be input into a second denoising network. The second denoising network denoises the denoising result of the first denoising network based on the sample receptor information according to the denoising principle of the first denoising network, thereby obtaining the denoising result of the second denoising network. The denoising process is not described in detail here. Next, the denoising result of the second denoising network can be input into a third denoising network. The third denoising network denoises the denoising result of the second denoising network based on the sample receptor information according to the denoising principle of the first denoising network, thereby obtaining the denoising result of the third denoising network. The denoising process is not described in detail here. By analogy with the above denoising principle, the denoising result of the last denoising network is obtained.
[0090] From the contents of step 2021 to step 2022, it can be seen that, optionally, by constructing at least two denoising networks connected in series, the reference noise data is gradually denoised until the denoising result of the last denoising network is obtained. This denoising can be expressed as the following formula (3).
[0091] Among them, for a certain denoising network, l t Represents the input of the denoising network, that is, the reference noise data mentioned above or the denoising result of the previous denoising network, and t represents the number of denoising times corresponding to the denoising network. t-1 Characterize the output of the denoising network, that is, the denoising result of the denoising network mentioned above. θ (l t-1 |l t ) represents the denoising network input l t The output after denoising is l t-1 The distribution satisfied by , where θ represents the network parameters of the denoising network. The symbol representing the Gaussian distribution. μ θ (l t ,t) represents the mean of Gaussian distribution, ∑ θ (lt ,t) characterizes the variance of Gaussian distribution.
[0092] Step 2023: Determine the predicted ligand information based on the denoising result of the last denoising network.
[0093] In an embodiment of the present application, the denoising result of the last denoising network can be used as the predicted ligand information, or the predicted ligand information can be determined based on the denoising result of the last denoising network by linear mapping or nonlinear mapping. It can be understood that the predicted ligand information is feature-level data used to characterize the ligand that can bind to the receptor described by the sample receptor information. The ligand here is a ligand of the same type as the sample ligand. For example, if the sample ligand is a polypeptide sequence, the ligand represented by the predicted ligand information is also a polypeptide sequence. Optionally, the predicted ligand information includes the encoding, three-dimensional coordinates, etc. of each amino acid that makes up the polypeptide sequence.
[0094] In an exemplary embodiment, the denoising result of the last denoising network includes at least two prediction sub-features. Step 2023 includes: for each prediction sub-feature, based on the feature similarity between each prediction sub-feature and at least two candidate ligand composition information, screening at least two target ligand composition information from the at least two candidate ligand composition information; sampling from the at least two target ligand composition information to obtain ligand composition information corresponding to each prediction sub-feature; and determining predicted ligand information based on the ligand composition information corresponding to each prediction sub-feature.
[0095] In this embodiment of the present application, the reference noise data includes at least one first sub-data. For any first sub-data, after denoising the first sub-data through each denoising network in sequence according to steps A1 and A2, a prediction sub-feature is obtained as the output of the last denoising network. The prediction sub-feature is feature-level data, which is characteristic data of the substance that constitutes the ligand and is used to describe the substance. For example, if the ligand is a protein, the prediction sub-feature can be a feature used to describe a peptide chain or an amino acid residue.
[0096] The electronic device can obtain a dictionary that includes at least two candidate ligand composition information. Each candidate ligand composition information is characteristic data of a substance that constitutes the ligand and is used to describe the substance. For example, if the ligand is a protein, and the protein is composed of multiple amino acid residues, the candidate ligand composition information can be used to describe amino acid residues such as glycine residues, alanine residues, or tryptophan residues.
[0097] For any predictor feature, the feature similarity can be calculated based on the predictor feature and any candidate ligand composition information according to the similarity calculation formula. The embodiment of the present application does not limit the similarity calculation formula. Exemplarily, the similarity calculation formula includes any one of the calculation formulas for cosine similarity, Euclidean distance, and Manhattan distance. In this way, the feature similarity between any predictor feature and at least two candidate ligand composition information can be calculated.
[0098] Next, target ligand composition information whose feature similarity is not less than the similarity threshold is screened out from at least two candidate ligand composition information. That is, the target ligand composition information is candidate ligand composition information whose feature similarity with the prediction sub-feature is not less than the similarity threshold. The embodiment of the present application does not limit the similarity threshold. Exemplarily, the similarity threshold is a numerical value set according to manual experience. Alternatively, the similarity threshold is any feature similarity among the feature similarities between any prediction sub-feature and at least two candidate ligand composition information. For example, the feature similarities between any prediction sub-feature and at least two candidate ligand composition information are sorted, and the Nth (N is a positive integer) feature similarity after sorting is used as the similarity threshold.
[0099] It is understandable that there are at least two target ligand composition information. Random sampling can be performed from at least two target ligand composition information, or, based on the feature similarity corresponding to each target ligand composition information, sampling can be performed from at least two target ligand composition information to obtain the ligand composition information corresponding to any predictor feature. It should be noted that the feature similarity corresponding to any two target ligand composition information can be the same or different. During the sampling process, the feature similarity corresponding to the target ligand composition information can characterize the probability of sampling the target ligand composition information. For example, if the feature similarity corresponding to the target ligand composition information is 95%, then there is a 95% probability that the target ligand composition information is sampled and a 5% probability that it is not sampled.
[0100] Through the above method, ligand composition information corresponding to each predictor feature can be sampled and obtained. The ligand composition information corresponding to each predictor feature is concatenated to obtain concatenated information, which serves as the predicted ligand information. The predicted ligand information is then used to characterize the ligand. Alternatively, due to the randomness of the sampling of the ligand composition information, the concatenated information may contain noise. Based on this, the concatenated information can be denoised using a target denoising network to obtain the predicted ligand information. The target denoising network is trained based on a large amount of ligand information and has the ability to denoise ligand information. It has learned the sequence characteristics of the ligand information. The network structure and denoising method are not limited here. It should be noted that a large amount of ligand information refers to the amount of ligand information being no less than a threshold. This threshold can be set based on experience, for example, 500. If the concatenated information is noisy, it can be considered noisy ligand information. The target denoising network denoises the noisy ligand information to obtain denoised ligand information, thereby fine-tuning the concatenated information and obtaining more accurate predicted ligand information. Among them, the ligand information after denoising is the predicted ligand information.
[0101] Optionally, the splicing information includes at least two splicing sub-information. For any splicing sub-information, the target denoising network performs attention processing on the splicing sub-information and each splicing sub-information to obtain processing results for each splicing sub-information. Based on the processing results for each splicing sub-information, a denoising result corresponding to the splicing sub-information is determined; and based on the denoising results corresponding to each splicing sub-information, predicted ligand information is determined. This implementation method can be seen in the description of step A1. The implementation principles of the two are similar and will not be repeated here.
[0102] It is understood that in practical applications, the model can output at least one piece of predicted ligand information based on the above-mentioned implementation principles. In the embodiments of the present application, on the one hand, by screening the target ligand composition information from the candidate ligand composition information, the accuracy of the ligand composition information predicted by the model is ensured. On the other hand, by sampling the ligand composition information corresponding to the predicted sub-features from the target ligand composition information, the diversity of the ligand composition information predicted by the model is ensured. Based on this, it is ensured that the model can predict the predicted ligand information with high accuracy and the diversity of the predicted ligand information is ensured.
[0103] Optionally, the neural network model further includes a second embedding layer, which is connected in series after the last denoising network. The predicted ligand information includes at least two features (i.e., ligand composition information). The second embedding layer maps each feature into a corresponding information segment to obtain third media information used to characterize the ligand.
[0104] Step 203: determining a first loss for characterizing the difference between the sample ligand information and the predicted ligand information.
[0105] In an embodiment of the present application, the first loss can be calculated based on the sample ligand information and the predicted ligand information according to the calculation formula of the first loss function. The embodiment of the present application does not limit the first loss function. Exemplarily, the first loss function is any one of the mean square error loss function, the cross entropy loss function, and the relative entropy loss function. Optionally, the first loss is calculated according to the formula (4) shown below.
[0106] in, represents the first loss. n represents the number of sample ligand information and also represents the number of predicted ligand information. y represents the sample ligand information, Characterize predicted ligand information.
[0107] Step 204 : training the neural network model based on the first loss to obtain a ligand information generation model, where the ligand information generation model is used to generate reference ligand information based on the reference receptor information.
[0108] In an embodiment of the present application, the gradient of the first loss relative to the model parameters of the neural network model can be calculated, and the model parameters of the neural network model can be adjusted based on the gradient to perform one training on the neural network model to obtain a trained neural network model.
[0109] If the trained neural network model meets the training termination condition, the trained neural network model is used as the ligand information generation model. If the trained neural network model does not meet the training termination condition, the trained neural network model is used as the neural network model for the next training, and the next training is performed on the neural network model in the manner of steps 201 to 204 to obtain a trained neural network model, and the ligand information generation model is determined based on the trained neural network model.
[0110] The embodiments of the present application do not limit the conditions for the trained neural network model to satisfy the training termination conditions. For example, the conditions for the trained neural network model to satisfy the training termination conditions include: the trained neural network model has been trained a set number of times, or the error of the trained neural network model is within a set error range, etc. The set number of times and the set error range can both be set based on manual experience.
[0111] It can be understood that the first loss is used to characterize the difference between the sample ligand information and the predicted ligand information. By training the neural network model with the first loss, the neural network model can be adjusted in the direction of making the output predicted ligand information close to the sample ligand information, thereby improving the accuracy of the predicted ligand information output by the neural network model. In addition, since the binding affinity between the ligand described by the sample ligand information and the receptor described by the sample receptor information is not less than the set affinity, and in the process of training the ligand information generation model, the predicted ligand information continuously approaches the sample ligand information, the ligand information generation model can generate reference ligand information for describing ligands with high binding affinity to the receptor.
[0112] As mentioned above, the embodiment of the present application does not limit the method for obtaining the reference noise data. A possible implementation is shown below, in which step 205 (not shown in the figure) is included before step 202.
[0113] Step 205 : Add noise to the sample ligand information to obtain reference noise data.
[0114] The embodiment of the present application does not limit the method of adding noise to the sample ligand information. Optionally, based on a noise function of at least one noise such as Gaussian noise or Poisson noise, the sample ligand information is added with noise to obtain reference noise data.
[0115] In an exemplary embodiment, the number of times of adding noise is at least twice. Step 205 includes: Step B1 to Step B2 (not shown in the figure).
[0116] Step B1: For the first noise addition, the first noise addition is performed on the sample ligand information to obtain the noise addition result of the first noise addition.
[0117] In an embodiment of the present application, noise data for the first noisy addition can be randomly generated, and a fusion calculation can be performed on the noise data for the first noisy addition and the sample ligand information to add the noise data for the first noisy addition to the sample ligand information to obtain the noise addition result of the first noisy addition.
[0118] Step B2: for non-first noise addition, perform non-first noise addition on the noise addition result of the last noise addition before the non-first noise addition to obtain the noise addition result of the non-first noise addition, and the noise addition result of the last noise addition is the reference noise data.
[0119] In an embodiment of the present application, noise data for the second noisy addition can be randomly generated, and a fusion calculation is performed on the noise data for the second noisy addition and the noisy result of the first noisy addition to add the noise data for the second noisy addition to the noisy result of the first noisy addition to obtain the noisy result of the second noisy addition.
[0120] Next, noise data for the third noisy addition may be randomly generated, and a fusion calculation may be performed on the noise data for the third noisy addition and the noisy result of the second noisy addition to add the noise data for the third noisy addition to the noisy result of the second noisy addition to obtain the noisy result of the third noisy addition.
[0121] By analogy with the above-mentioned noise addition principle, the sample ligand information is subjected to at least two noise additions in sequence, and the noise addition result of the last noise addition is obtained, which is the reference noise data. Optionally, the noise addition process can be expressed as the following formula (5).
[0122] Among them, l t The data obtained after characterizing the noise addition, that is, the noise addition result of the first noise addition or the noise addition result of the non-first noise addition mentioned above, can be described as the noise addition result of the t-th noise addition. t-1 The data before noise addition, i.e., the sample ligand information mentioned above or the noise addition result of the last noise addition that is not the first noise addition, can be described as the sample ligand information or the noise addition result of the t-1th noise addition. t |l t-1 ) characterizes l t-1 The noise-added l t The distribution that is satisfied. The symbol representing the Gaussian distribution. Characterizes the mean of the Gaussian distribution, β t I represents the variance of the Gaussian distribution. t , I are two hyper parameters.
[0123] In one possible implementation, the number of times of denoising and the number of times of denoising are both T, where T is a positive integer. Step 204 includes: for the i-th denoising, based on the noise data removed during the i-th denoising and the noise data added during the T+1-i-th denoising, determining a second loss corresponding to the i-th denoising, where i is a positive integer between 1 and T; and training a neural network model based on the first loss and the second losses corresponding to each denoising to obtain a ligand information generation model.
[0124] In this embodiment of the present application, the sample ligand information is subjected to T-times of noise addition to obtain reference noise data, and the reference noise data is subjected to T-times of denoising to obtain predicted ligand information. The noise addition and denoising processes are inverse processes to each other. The first noise addition corresponds to the T-th denoising, the second noise addition corresponds to the T-1th denoising, and so on. The T-th noise addition corresponds to the first denoising. Based on this, the i-th denoising corresponds to the T+1-i-th noise addition.
[0125] As mentioned above, the electronic device can randomly generate the noise data for the T+1-ith denoising. In addition, the electronic device can also determine the noise data for the i-th denoising based on the input and output of the i-th denoising network, for example, the reference noise data input to the first denoising network and the denoising result output by the first denoising network, or the denoising result of the previous denoising network input that is not the first denoising network and the denoising result output by the non-first denoising network. The determination method will not be repeated here, for example, the output of the i-th denoising network is subtracted from the input to obtain the noise data for the i-th denoising, or the noise data for the i-th denoising is determined based on the distribution satisfied by the output of the i-th denoising network and the distribution satisfied by the input of the i-th denoising network.
[0126] Next, the second loss corresponding to the i-th denoising can be calculated based on the noise data of the T+1-i-th denoising and the noise data of the i-th denoising according to the calculation formula of the second loss function. The embodiment of the present application does not limit the second loss function. Exemplarily, the second loss function is any one of the mean square error loss function, the cross entropy loss function, and the relative entropy loss function.
[0127] The loss of the neural network model can then be determined based on the first loss and the second losses corresponding to each denoising step. For example, a weighted summation or weighted average operation can be performed on the first loss and the second losses corresponding to each denoising step to obtain the loss of the neural network model. Next, the gradient of the loss of the neural network model relative to the model parameters of the neural network model is calculated, and the model parameters of the neural network model are adjusted based on the gradient to obtain the ligand information generation model. The adjustment process can be seen in the description of step 204 and will not be repeated here.
[0128] By adjusting the model parameters of the neural network model using the second loss corresponding to each denoising step, the neural network model can be optimized to bring the denoised data closer to the noisy data, thereby optimizing the output of the neural network model to bring the predicted ligand information closer to the sample ligand information. This optimization enables the ligand information generation model to relatively accurately restore the sample ligand information through denoising, thereby improving the accuracy of the ligand information generation model.
[0129] In an exemplary embodiment, step 204 includes steps 2041 to 2043 (not shown in the figure).
[0130] Step 2041: Train the neural network model based on the first loss to obtain a reference generation model.
[0131] In an embodiment of the present application, the gradient of the first loss relative to the model parameters of the neural network model can be calculated, and the model parameters of the neural network model can be adjusted based on the gradient to perform one training on the neural network model to obtain a trained neural network model.
[0132] If the trained neural network model meets the training end condition, the trained neural network model is used as the reference generation model. If the trained neural network model does not meet the training end condition, the trained neural network model is used as the neural network model for the next training, and the next training is performed on the neural network model in the manner of steps 201 to 204 to obtain the trained neural network model, and the reference generation model is determined based on the trained neural network model.
[0133] Step 2042: De-noise the reference noise data based on the target receptor information using the reference generation model to obtain the target ligand information.
[0134] In an embodiment of the present application, the electronic device can obtain target receptor information, and the acquisition method is not limited here. For example, the electronic device can obtain the target receptor information input by the user. Alternatively, the electronic device can obtain a receptor dataset, for example, the receptor dataset is a Swiss-Prot dataset or a UniRef30 dataset, both of which are protein-related datasets. Receptor data can be selected from the receptor dataset, and target receptor information can be obtained by mapping the receptor data to corresponding features. For example, protein data with a sequence length of less than 1000 amino acids is extracted from the Swiss-Prot dataset or the UniRef30 dataset, and each amino acid data in the protein data is mapped to a corresponding feature to obtain protein features, thereby obtaining target receptor information. Target receptor information is feature data used to describe the target receptor. The target receptor is any receptor whose receptor composition information is lower than a threshold value, wherein the threshold value can be set based on manual experience. For example, in the aforementioned example, the threshold value is 1000.
[0135] Because the reference generative model is trained based on a neural network model, the implementation principles of the reference generative model and the neural network model are similar. Based on this, the reference noise data and target receptor information can be input into the reference generative model. The reference generative model then denoises the reference noise data based on the target receptor information, according to the implementation described in step 202, to obtain target ligand information. The denoising method is not further described here.
[0136] In one possible implementation, step 2042 includes: denoising the reference noise data based on the target receptor information through a reference generation model to obtain candidate ligand information; if the first quality indicator used to characterize the quality of the candidate ligand information meets the first indicator condition, determining that the candidate ligand information is the target ligand information.
[0137] In the embodiment of the present application, the electronic device can, according to the implementation principles of steps 2021 to 2023, use the reference generation model to denoise the reference noise data based on the target receptor information to obtain candidate ligand information. The denoising process is not further described herein. It is understood that the reference generation model can generate at least two candidate ligand information, and this candidate ligand information is not only highly accurate but also diverse.
[0138] For any candidate ligand information, the electronic device can determine the first quality indicator of the candidate ligand information to describe the quality of the candidate ligand information through the first quality indicator of the candidate ligand information. Generally speaking, the larger the first quality indicator of the candidate ligand information, the better the quality of the candidate ligand information. Among them, the embodiment of the present application does not limit the method for determining the first quality indicator of the candidate ligand information. For example, it can be determined according to any one of the implementation methods A to C as shown below.
[0139] In implementation method A, at least two first conformation information are determined based on the target receptor information and the candidate ligand information, and the first conformation information is used to describe the complex conformation formed after the receptor described by the target receptor information binds to the ligand described by the candidate ligand information; for any first conformation information, a first indicator of any first conformation information is determined, and the first indicator is used to characterize the binding energy of the complex conformation described by any first conformation information; based on the first indicator of each first conformation information, the first quality indicator of the candidate ligand information is determined.
[0140] In an embodiment of the present application, the electronic device can determine at least two binding information of the candidate ligand information based on any search algorithm such as the Monte Carlo simulation algorithm, the genetic algorithm, the particle swarm optimization algorithm, etc. The binding information of the candidate ligand information is used to describe at least one piece of information such as the posture, position, rotation, translation, etc. of the ligand when the ligand described by the candidate ligand information binds to the receptor described by the target receptor information.
[0141] Next, based on the binding information of either the target receptor information or the candidate ligand information, a first conformational information is determined. This transforms the ligand in terms of posture, position, rotation, and translation based on the binding information of the candidate ligand information. The receptor is then bound to the transformed ligand to obtain a complex conformation, which is described by the first conformational information. The complex conformation refers to the structure formed by the binding of the ligand and the receptor.
[0142] Then, for any piece of first conformational information, the electronic device can calculate a first index for the first conformational information using an index function, such as a physical function or a potential energy function. The index function is used to calculate the binding energy of the complex conformation based on van der Waals forces between molecules, chemical bonds between atoms, and the like in the complex conformation. Based on this, the first index for the first conformational information can represent the binding energy of the complex conformation described by the first conformational information.
[0143] In the above manner, the electronic device can determine each piece of first conformational information based on the binding information of the target receptor information and the candidate ligand information, and calculate a first indicator for each piece of first conformational information. Next, based on the first indicator for each piece of first conformational information, the electronic device can determine a median or calculate a mean of the first indicators to obtain a first quality indicator for the candidate ligand information.
[0144] In implementation method B, reference conformation information is selected from a database, and the alignment result between the ligand information included in the reference conformation information and the candidate ligand information meets the alignment condition, and the alignment result between the receptor information included in the reference conformation information and the target receptor information meets the alignment condition; based on the reference conformation information, the target receptor information and the candidate ligand information, the second conformation information is determined, and the second conformation information is used to describe the complex conformation formed after the receptor described by the target receptor information binds to the ligand described by the candidate ligand information; based on the second conformation information, the first quality indicator of the candidate ligand information is determined.
[0145] In an embodiment of the present application, an electronic device can obtain a database that includes at least two candidate conformation information. Any candidate conformation information includes ligand information and receptor information. The candidate conformation information is used to describe the complex conformation formed by the binding of the ligand described by the ligand information and the receptor described by the receptor information. The complex conformation is an actual complex conformation. Among them, different candidate conformation information describes different complex conformations. Since the ligand can bind to the receptor by performing transformations in posture, position, rotation, translation, etc., and the receptor can also correspond to different postures, positions, rotations, translations, etc., the same ligand and receptor can combine to obtain at least one complex conformation. Based on this, any two different complex conformations can correspond to the same or different ligands, the same or different receptors.
[0146] For any candidate conformational information, the electronic device can determine a first alignment result based on the candidate ligand information and the ligand information included in the candidate conformational information. The first alignment result can be used to characterize the degree of alignment between the candidate ligand information and the ligand information included in the candidate conformational information. The ligand described in the candidate ligand information can be aligned with the ligand described in the ligand information included in the candidate conformational information to ensure that the two ligands are arranged as identically as possible. Since a ligand is composed of at least two substances, the two aligned ligands can be considered as two rows of at least two columns of substances. The similarities and differences between the two rows of substances are compared column by column, and the degree of alignment between the candidate ligand information and the ligand information included in the candidate conformational information is determined based on the comparison results for each column. For example, if the ligand is a protein, and a protein comprises at least two amino acid residues, aligning the two proteins means ensuring that the arrangement of the amino acid residues in the two proteins is as identical as possible. If the aligned ligands are considered a row of ligands, then an amino acid residue in a row of ligands can be considered to be located in a column within that row of ligands. In other words, the two aligned rows of ligands can be considered as two rows of multiple columns of amino acid residues. The similarities and differences of the amino acid residues in the two rows are compared column by column, and the degree of alignment of the two proteins is determined based on the comparison results of each column.
[0147] Optionally, the first alignment result is positively correlated with the degree of alignment, i.e., a larger first alignment result indicates a higher degree of alignment. If the first alignment result is not less than an alignment threshold, the first alignment result satisfies the alignment condition; if the first alignment result is less than the alignment threshold, the first alignment result does not satisfy the alignment condition.
[0148] Similarly, for any candidate conformational information, the electronic device can determine a second alignment result based on the target receptor information and the receptor information included in the candidate conformational information. The second alignment result can be used to characterize the degree of alignment between the target receptor information and the receptor information included in the candidate conformational information. Optionally, the second alignment result is positively correlated with the degree of alignment, i.e., a larger second alignment result indicates a higher degree of alignment. If the second alignment result is not less than an alignment threshold, the second alignment result satisfies the alignment condition; if the second alignment result is less than the alignment threshold, the second alignment result does not satisfy the alignment condition.
[0149] The embodiments of the present application do not limit the method for determining the alignment threshold. For example, the alignment threshold is a value set based on manual experience, or the alignment threshold is a value adjusted based on actual conditions. For example, when there are many alignment results higher than 0.5, the alignment threshold can be adjusted accordingly. When there are few alignment results higher than 0.5, the alignment threshold can be adjusted accordingly.
[0150] If the first alignment result of any candidate conformation information satisfies the alignment condition and the second alignment result of the candidate conformation information satisfies the alignment condition, the candidate conformation information is determined as the reference conformation information. In this way, at least one reference conformation information can be determined.
[0151] Next, the electronic device can determine the first distribution information and the second distribution information based on the reference conformation information, the target receptor information and the candidate ligand information. Among them, the electronic device determines the first distribution information and the second distribution information based on this information, the target receptor information and the candidate ligand information by learning some properties of the reference conformation information, such as the posture, rotation, translation, position and other information of the ligand described by the reference conformation information, and the posture, rotation, translation, position and other information of the receptor described by the reference conformation information. The first distribution information is used to describe the distribution satisfied by the distance between adjacent components in the receptor described by the target receptor information and the distribution satisfied by the angle formed by adjacent components (for example, when the receptor is a protein, the components are amino acids or peptide chains, etc.). These distributions reflect the posture, rotation, translation, position and other information of the receptor described by the target receptor information. Similarly, the second distribution information is used to describe the distribution satisfied by the distances between adjacent components (for example, when the ligand is a protein, the components are amino acids or peptide chains, etc.) in the ligand described by the candidate ligand information, as well as the distribution satisfied by the angles formed by adjacent components. These distributions reflect the posture, rotation, translation, position and other information of the ligand described by the candidate ligand information.
[0152] Afterwards, the electronic device determines second conformation information based on the first distribution information and the second distribution information, and uses the second conformation information to describe the conformation of a complex formed after the receptor described by the target receptor information binds to the ligand described by the candidate ligand information.
[0153] It is understood that the target receptor information and the receptor information corresponding to the reference conformational information are highly similar, and the candidate ligand information and the ligand information corresponding to the reference conformational information are highly similar. Generally speaking, the more similar the ligand information, the closer their folding structures are, and the more similar the receptor information, the closer their folding structures are. Based on this, it can be seen that determining the second conformational information based on the reference conformational information can make the complex conformation described by the second conformational information more similar to the actual situation, thereby improving the accuracy of the second conformational information.
[0154] The electronic device can then calculate a second index for the second conformational information using an indicator function such as a physical function or a potential energy function. The second index of the second conformational information characterizes the binding energy of the complex conformation described by the second conformational information. Optionally, the second index of the second conformational information serves as the first quality index for the candidate ligand information. Alternatively, when there are at least two pieces of second conformational information, the first quality index for the candidate ligand information can be obtained by determining the median or calculating the mean of the second indexes based on the second indexes of the respective pieces of second conformational information.
[0155] In implementation method C, the target receptor information and candidate ligand information are input into the binding affinity prediction model to obtain the target binding affinity, which is used to characterize the force when the receptor described by the target receptor information binds to the ligand described by the candidate ligand information; the first quality indicator of the candidate ligand information is determined based on the target binding affinity.
[0156] In the embodiments of the present application, the electronic device can be trained to obtain a binding affinity prediction model, and the training method is not limited here. The embodiments of the present application also do not limit the model structure, model parameters, etc. of the binding affinity prediction model.
[0157] The binding affinity prediction model is a deep learning model that can learn how receptors and ligands bind and predict the binding force. By inputting target receptor information and candidate ligand information into the binding affinity prediction model, the binding affinity prediction model performs feature fusion processing on the target receptor information and candidate ligand information, obtaining target features that characterize the conformation of the complex formed by the binding of the receptor described by the target receptor information and the ligand described by the candidate ligand information. The target binding affinity is then determined based on the target features, and the target binding affinity is used to characterize the binding force between the receptor described by the target receptor information and the ligand described by the candidate ligand information.
[0158] The target binding affinity may be used as the first quality indicator of the candidate ligand information, or the target binding affinity may be mapped as the first quality indicator of the candidate ligand information.
[0159] In practical applications, at least two of Implementations A to C may be combined, that is, the first quality indicator of the candidate ligand information may be determined based on at least two of the first indicator of the first conformational information, the second indicator of the second conformational information, and the target binding affinity. Other methods may also be used to determine the first quality indicator of the candidate ligand information, which will not be described in detail here.
[0160] The first quality index of the candidate ligand information is used to characterize the quality of the candidate ligand information. Optionally, the larger the first quality index of the candidate ligand information, the higher the quality of the candidate ligand information. If the first quality index of the candidate ligand information is not less than the index threshold, the first quality index is determined to meet the first index condition; if the first quality index of the candidate ligand information is less than the index threshold, the first quality index is determined to not meet the first index condition. The index threshold can be a value determined based on manual experience or a value adaptively adjusted based on actual conditions. If the first quality index meets the first index condition, the candidate ligand information corresponding to the first quality index is the target ligand information.
[0161] Step 2043 : Based on the sample receptor information, the sample ligand information, the target receptor information, and the target ligand information, a reference generation model is trained to obtain a ligand information generation model.
[0162] In the embodiment of the present application, the quality of the ligand described by the target ligand information is high, and the binding affinity between the ligand and the receptor described by the target receptor information is high. The target receptor information and the target ligand information can be used as a high-quality data set D artificial , the high-quality dataset is compared with the original dataset D including sample receptor information and sample ligand information native Merge to get the training data set D = D native +D artificial The reference generation model is trained using the training dataset to obtain the ligand information generation model.
[0163] Optionally, the target receptor information and sample receptor information in the training dataset are collectively referred to as training receptor information, and the target ligand information and sample ligand information in the training dataset are collectively referred to as training ligand information. The reference noise data is denoised based on the training receptor information using the reference generative model to obtain first ligand information. A third loss is determined to characterize the difference between the first ligand information and the training ligand information. The reference generative model is trained based on the third loss to obtain a ligand information generation model. This training method can be found in the description of steps 202 to 204; the implementation principles of the two methods are similar and will not be further elaborated here.
[0164] In the embodiments of the present application, the target ligand information is determined based on the target receptor information by using a reference generation model, thereby enriching the model's training data set, making the model's training data diverse and highly accurate. Further training of the model using the training data set can improve the model's accuracy and generalization capabilities. The trained ligand information generation model is used to generate reference ligand information based on the reference receptor information. This content can be seen in the description of Figure 5 below and will not be repeated here.
[0165] As shown in Figure 4, Figure 4 is a schematic diagram of training a ligand information generation model provided in an embodiment of the present application. In this embodiment of the present application, sample information pairs can be obtained, and the sample information pairs include sample receptor information and sample ligand information. First, a neural network model is trained based on the sample information pairs to obtain a reference generation model. Then, receptor database 1 and receptor database 2 are obtained. Receptor database 1 and receptor database 2 include at least two receptor information, and target receptor information is obtained by screening from the at least two receptor information.
[0166] Next, the reference noise data is denoised based on the target receptor information using the reference generative model to obtain candidate ligand information. This process can be performed at least twice to obtain at least two pieces of candidate ligand information. In other words, the reference generative model is used to generate at least two pieces of candidate ligand information based on the target receptor information and the reference noise data.
[0167] For each candidate ligand, a score is assigned based on the target receptor information to obtain the candidate's score. The score corresponds to the first quality indicator of the candidate ligand mentioned above. If the score is greater than a threshold, the candidate is selected as the target ligand. If the score is less than the threshold, the candidate is discarded.
[0168] In this way, the target receptor information and target ligand information can be determined. These target receptor information and target ligand information can be considered a target information pair. Based on the target information pair and the sample information pair, the reference generation model is trained to obtain a ligand information generation model, which is then used to generate reference ligand information.
[0169] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant region. For example, the sample ligand information and sample receptor information involved in this application were all obtained with full authorization.
[0170] The above method first uses a neural network model to denoise reference noise data based on sample receptor information to obtain predicted ligand information. A first loss is then determined based on the sample ligand information and the predicted ligand information. The neural network model is trained based on the first loss, optimizing the neural network model toward bringing the predicted ligand information closer to the sample ligand information, enabling the model to accurately output ligand information. Furthermore, because the binding affinity between the ligand described by the sample ligand information and the receptor described by the sample receptor information is no less than the set affinity, the trained ligand information generation model is capable of generating reference ligand information based on the reference receptor information, and the binding affinity between the ligand described by the reference ligand information and the receptor described by the reference receptor information is high.
[0171] Please refer to Figure 5, which is a flowchart of a method for generating ligand information provided in an embodiment of the present application. For ease of description, the terminal device 101 or server 102 that executes the method for generating ligand information in an embodiment of the present application is referred to as an electronic device. The method can be executed by an electronic device. As shown in Figure 5, the method includes the following steps.
[0172] Step 501: Acquire reference receptor information.
[0173] The embodiments of the present application do not limit the method for obtaining reference receptor information. For example, the electronic device can obtain reference receptor information input by the user. Alternatively, the electronic device can collect a target peptide dataset from the Internet, where the target peptide dataset includes at least two target peptide pairs, and one target peptide pair includes target information. The electronic device maps the target information into text features by embedding a transformation function to obtain reference receptor information. It is understandable that the implementation method of step 501 is similar to the implementation method of step 201. Please refer to the description of step 201 and will not be repeated here.
[0174] Step 502 : De-noise the reference noise data based on the reference receptor information using the ligand information generation model to obtain reference ligand information.
[0175] Among them, the ligand information generation model is trained according to the method related to Figure 2. Reference noise data can be input into the ligand information generation model. In addition, reference receptor information can also be input into the ligand information generation model. In one possible implementation, the reference noise data and the reference receptor information are spliced to obtain spliced information, and the spliced information is input into the ligand information generation model. The ligand information generation model is a denoising network model, which is used to denoise the reference noise data based on the reference receptor information to obtain reference ligand information. The reference ligand information is characteristic data for characterizing the ligand obtained by prediction by the neural network model. It can be understood that the implementation method of step 502 is similar to the implementation method of step 202. Please refer to the description of step 202 and it will not be repeated here.
[0176] In one possible implementation, step 502 includes: denoising the reference noise data based on the reference receptor information through a ligand information generation model to obtain at least two optional ligand information; determining a second quality indicator for characterizing the quality of each optional ligand information; and selecting reference ligand information whose second quality indicator meets the second indicator condition from the at least two optional ligand information.
[0177] In the embodiments of the present application, the content of generating the optional ligand information by the ligand information generation model can be found in the description of the candidate ligand information generation process above and will not be repeated here. Next, the second quality index of the optional ligand information can be determined according to the principle for determining the first quality index of the candidate ligand information. This content can be found in the description of Implementations A to C above and will not be repeated here.
[0178] The second quality index of the optional ligand information is used to characterize the quality of the optional ligand information. Optionally, the larger the second quality index of the optional ligand information, the higher the quality of the optional ligand information. If the second quality index of the optional ligand information is not less than the index threshold, the second quality index is determined to meet the second index condition; if the second quality index of the optional ligand information is less than the index threshold, the second quality index is determined to not meet the second index condition. The index threshold can be a value determined based on manual experience or a value adaptively adjusted based on actual conditions. If the second quality index meets the second index condition, the optional ligand information corresponding to the second quality index is the reference ligand information.
[0179] In practical applications, the ligand information generation model of the embodiment of the present application can be spliced with other models to implement corresponding functions based on these models. For example, a three-dimensional structure prediction model is spliced after the ligand information generation model. After the ligand information generation model generates reference ligand information, the reference ligand information is input into the three-dimensional structure prediction model, and the three-dimensional structure prediction model is used to predict the three-dimensional structure of the ligand after the ligand described by the reference ligand information binds to the receptor. For another example, a complex conformation prediction model is spliced after the ligand information generation model. After the ligand information generation model generates reference ligand information, the reference ligand information is input into the complex conformation prediction model, and the complex conformation prediction model is used to predict the structure formed by the ligand described by the reference ligand information binds to the receptor.
[0180] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, and display, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with relevant laws, regulations, and standards in the relevant region. For example, the reference receptor information involved in this application was obtained with full authorization.
[0181] The ligand information generation model in the above method is trained according to the method related to Figure 2, and can generate reference ligand information based on reference receptor information, and the binding affinity between the ligand described by the reference ligand information and the receptor described by the reference receptor information is high.
[0182] The above describes the method of the embodiment of the present application from the perspective of method steps. The following is a systematic introduction to the method of the embodiment of the present application. The method of the embodiment of the present application can be applied in a variety of scenarios such as drug design and targeted therapy. In these scenarios, the ligand information generation model is trained to generate ligand information based on receptor information through the ligand information generation model, thereby achieving specificity generation tasks. Among them, receptors and ligands include but are not limited to: antigens and antibodies, proteins and proteins, proteins and polypeptide sequences, proteins and small molecules, etc. The following takes the receptor as a protein and the ligand as a polypeptide sequence as an example to illustrate the method of the embodiment of the present application.
[0183] As shown in FIG6 , FIG6 is a schematic diagram of a training method of a polypeptide sequence denoising diffusion model using a sample target protein as a condition provided in an embodiment of the present application. The training process includes a forward diffusion process and a reverse diffusion process.
[0184] The forward diffusion process is the process of gradually adding noise to the embedded features of the peptide sequence text. In the embodiment of the present application, the peptide sequence text includes at least two amino acid words. By mapping each amino acid word to a corresponding vector code, the embedded features of the peptide sequence text are obtained. The embedded features of the peptide sequence text correspond to the sample ligand information mentioned above.
[0185] The embedded features of the peptide sequence text are unnoised feature data, which can be referred to as Noise Result 0. Noise Result 1 is obtained by randomly determining noise data and then adding noise to Noise Result 0 based on the noise data. Noise Result 2 is obtained by randomly determining noise data and then adding noise to Noise Result 1 based on the noise data. Similarly, by performing T noise additions on Noise Result 0, Noise Result T is obtained. Noise Result T is Gaussian noise and corresponds to the reference noise data mentioned above.
[0186] The reverse diffusion process is the process of gradually denoising Gaussian noise. Gaussian noise is the undenoised data, which can be referred to as denoised result 0. Denoised result 0 is denoised based on the sample target protein using the denoising network to obtain denoised result 1. Denoised result 1 is denoised based on the sample target protein using the denoising network to obtain denoised result 2. Similarly, by gradually performing T denoising operations on denoised result 0 based on the sample target protein, denoised result T is obtained. Denoised result T is the embedded features of the peptide sequence text, which can be mapped to the peptide sequence text.
[0187] Based on the embedding features of the peptide sequence text during the denoising process and the embedding features of the peptide sequence text obtained by denoising, the neural network model is trained to obtain a peptide sequence denoising diffusion model with the sample target protein as a condition. This model corresponds to the ligand information generation model mentioned above.
[0188] The polypeptide sequence generation model obtained in the embodiment of the present application can generate a polypeptide sequence based on the text of the target protein. Optionally, the target protein is aspartate 1-decarboxylase alpha chain, and the natural polypeptide and generated polypeptide corresponding to the target protein are shown in Figure 7. The target protein is Bcl-w protein, and the natural polypeptide and generated polypeptide corresponding to the target protein are shown in Figure 8. It can be seen from (1) to (3) in Figure 7 and (1) to (3) in Figure 8 that in most cases, the polypeptide generated by the polypeptide sequence generation model can obtain a lower molecular docking score than the natural polypeptide, indicating that the binding of the generated polypeptide to the target protein is more stable and has a higher binding affinity.
[0189] Peptides (i.e., g1 to g9) were generated based on the glucagon-like peptide 1 receptor (GLP-1R) using a peptide sequence generation model. The generated peptides were compared with the natural peptides to obtain a docking score as shown in (1) in Figure 9 and a distribution comparison diagram of physicochemical similarity as shown in (2) in Figure 9. As shown in Figure 9, the generated peptides can achieve a lower docking score (Vina Score) and a higher physicochemical similarity (PC Score) than the natural peptides, indicating that the generated peptides have a higher binding affinity to the target protein and are closer to the physicochemical properties of existing peptides.
[0190] In addition, on the one hand, a polypeptide is generated based on GLP-1R through a polypeptide sequence generation model, and the ipTM (interface predicted TM-score, predicted interface TM score) confidence of the natural polypeptide and the ipTM confidence of the generated polypeptide are obtained through a tool, and the results shown in (1-1) to (1-3) in Figure 10 are obtained. It can be seen that the generated polypeptide has a higher confidence in binding to the target protein. On the other hand, a polypeptide is generated based on the somatostatin receptor (Somatostatin Receptor 5, SSTR5) through a polypeptide sequence generation model, and the ipTM confidence of the natural polypeptide and the generated polypeptide are obtained through a tool, and the results shown in (2-1) to (2-4) in Figure 10 are obtained. The results show that the embodiment of the present application can quickly and efficiently generate polypeptide sequences with potential high affinity to the target protein, which is of great significance for drug design in personalized medicine and targeted cancer therapy.
[0191] FIG11 is a schematic diagram showing the structure of a training device for a ligand information generation model provided in an embodiment of the present application. As shown in FIG11 , the device includes:
[0192] An acquisition module 1101 is configured to acquire sample receptor information and sample ligand information, wherein the binding affinity between the ligand described by the sample ligand information and the receptor described by the sample receptor information is not less than a set affinity;
[0193] A denoising module 1102 is configured to denoise the reference noise data based on the sample receptor information using a neural network model to be trained to obtain predicted ligand information;
[0194] A determination module 1103 is configured to determine a first loss for characterizing the difference between the sample ligand information and the predicted ligand information;
[0195] The training module 1104 is used to train the neural network model based on the first loss to obtain a ligand information generation model, and the ligand information generation model is used to generate reference ligand information based on reference receptor information.
[0196] In one possible implementation, the neural network model includes at least two denoising networks;
[0197] Denoising module 1102 is used to denoise the reference noise data based on the sample receptor information of the first denoising network to obtain the denoising result of the first denoising network; for non-first denoising networks, denoise the denoising result of the previous denoising network of the non-first denoising network based on the sample receptor information to obtain the denoising result of the non-first denoising network; and determine the predicted ligand information based on the denoising result of the last denoising network.
[0198] In one possible implementation, the first denoising network includes an attention network and a feedforward network;
[0199] The denoising module 1102 is used to perform attention processing on the sample receptor information and the reference noise data through the attention network to obtain an attention processing result; and determine the denoising result of the first denoising network based on the attention processing result through the feedforward network.
[0200] In a possible implementation, the sample receptor information includes at least two receptor composition information, and the reference noise data includes at least one first sub-data;
[0201] The denoising module 1102 is used to perform attention processing on any first sub-data and each content information through an attention network to obtain a processing result of each content information; based on the processing result of each content information, determine the second sub-data corresponding to any first sub-data, where each content information includes at least two receptor composition information and at least one first sub-data; and determine the attention processing result based on the second sub-data corresponding to each first sub-data.
[0202] In one possible implementation, the denoising module 1102 is configured to perform mapping on the attention processing result through a feedforward network to obtain a mapping result; and perform mapping on the mapping result or set features through a feedforward network to obtain a denoising result of the first denoising network.
[0203] In one possible implementation, the denoising result of the last denoising network includes at least two prediction sub-features;
[0204] The denoising module 1102 is used to screen at least two target ligand composition information from at least two candidate ligand composition information for any prediction sub-feature based on the feature similarity between any prediction sub-feature and at least two candidate ligand composition information; sample from at least two target ligand composition information to obtain ligand composition information corresponding to any prediction sub-feature; and determine predicted ligand information based on the ligand composition information corresponding to each prediction sub-feature.
[0205] In a possible implementation, the apparatus further includes:
[0206] The noise adding module is used to add noise to the sample ligand information to obtain reference noise data.
[0207] In a possible implementation, the noise is added multiple times;
[0208] The noise addition module is used to perform the first noise addition on the sample ligand information for the first noise addition to obtain the noise addition result of the first noise addition; for the non-first noise addition, perform the non-first noise addition on the noise addition result of the last noise addition to obtain the noise addition result of the non-first noise addition, and the noise addition result of the last noise addition is the reference noise data.
[0209] In a possible implementation, the number of times of adding noise and the number of times of removing noise are both T, where T is a positive integer;
[0210] The training module 1104 is used to determine the second loss corresponding to the i-th denoising based on the noise data removed during the i-th denoising and the noise data added during the T+1-i-th denoising, where i is a positive integer from 1 to T; and train the neural network model based on the first loss and the second loss corresponding to each denoising to obtain a ligand information generation model.
[0211] In one possible implementation, the training module 1104 is used to train the neural network model based on the first loss to obtain a reference generation model; denoise the reference noise data based on the target receptor information through the reference generation model to obtain the target ligand information; and train the reference generation model based on the sample receptor information, the sample ligand information, the target receptor information and the target ligand information to obtain the ligand information generation model.
[0212] In one possible implementation, the training module 1104 is used to denoise the reference noise data based on the target receptor information through a reference generation model to obtain candidate ligand information; if the first quality indicator used to characterize the quality of the candidate ligand information meets the first indicator condition, the candidate ligand information is determined to be the target ligand information.
[0213] In one possible implementation, the determination module 1103 is also used to determine at least two first conformation information based on the target receptor information and the candidate ligand information, where the first conformation information is used to describe the complex conformation formed after the receptor described by the target receptor information binds to the ligand described by the candidate ligand information; for any first conformation information, determine the first indicator of any first conformation information, where the first indicator is used to characterize the binding energy of the complex conformation described by any first conformation information; and determine the first quality indicator of the candidate ligand information based on the first indicator of each first conformation information.
[0214] In a possible implementation, the apparatus further includes:
[0215] A selection module is used to select reference conformation information from a database, wherein an alignment result between ligand information included in the reference conformation information and candidate ligand information satisfies an alignment condition, and an alignment result between receptor information included in the reference conformation information and target receptor information satisfies an alignment condition;
[0216] Determination module 1103 is also used to determine second conformational information based on reference conformational information, target receptor information and candidate ligand information, where the second conformational information is used to describe the conformation of a complex formed after the receptor described by the target receptor information binds to the ligand described by the candidate ligand information; based on the second conformational information, determine the first quality indicator of the candidate ligand information.
[0217] In one possible implementation, the determination module 1103 is further used to input the target receptor information and the candidate ligand information into a binding affinity prediction model to obtain a target binding affinity, where the target binding affinity is used to characterize the force of binding between the receptor described by the target receptor information and the ligand described by the candidate ligand information; and to determine a first quality indicator of the candidate ligand information based on the target binding affinity.
[0218] The apparatus first uses a neural network model to denoise reference noise data based on sample receptor information to obtain predicted ligand information. A first loss is then determined based on the sample ligand information and the predicted ligand information. The neural network model is trained based on the first loss, optimizing the neural network model toward bringing the predicted ligand information closer to the sample ligand information, enabling the model to accurately output the ligand information. Furthermore, because the binding affinity between the ligand described by the sample ligand information and the receptor described by the sample receptor information is no less than the set affinity, the trained ligand information generation model is capable of generating reference ligand information based on the reference receptor information, and the binding affinity between the ligand described by the reference ligand information and the receptor described by the reference receptor information is high.
[0219] It should be understood that the device provided in FIG. 11 is merely an example of the division of the functional modules described above when implementing its functions. In actual applications, the functions described above can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0220] FIG12 is a schematic diagram showing the structure of a device for generating ligand information provided in an embodiment of the present application. As shown in FIG12 , the device includes:
[0221] An acquisition module 1201 is used to acquire reference receptor information;
[0222] The denoising module 1202 is used to denoise the reference noise data based on the reference receptor information by using the ligand information generation model to obtain reference ligand information. The ligand information generation model is trained according to the method of the first aspect.
[0223] In one possible implementation, the denoising module 1202 is used to denoise the reference noise data based on the reference receptor information through the ligand information generation model to obtain at least two optional ligand information; determine a second quality indicator for characterizing the quality of each optional ligand information; and select reference ligand information whose second quality indicator meets the second indicator condition from the at least two optional ligand information.
[0224] The ligand information generation model in the above device is trained according to the method related to Figure 2, and can generate reference ligand information based on reference receptor information, and the binding affinity between the ligand described by the reference ligand information and the receptor described by the reference receptor information is high.
[0225] It should be understood that the device provided in FIG. 12 is merely an example of the division of the functional modules described above when implementing its functions. In actual applications, the functions described above can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0226] FIG13 shows a block diagram of a terminal device 1300 provided by an exemplary embodiment of the present application. The terminal device 1300 includes a processor 1301 and a memory 1302 .
[0227] The processor 1301 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1301 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1301 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0228] The memory 1302 may include one or more computer-readable storage media, which may be non-transitory. The memory 1302 may also include high-speed random access memory, and non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1302 is used to store at least one computer program, which is used to be executed by the processor 1301 to implement the training method of the ligand information generation model or the generation method of ligand information provided in the method embodiment of the present application.
[0229] In some embodiments, the terminal device 1300 may also optionally include a display screen 1305. The display screen 1305 is used to display a UI (User Interface), which may include graphics, text, icons, videos, and any combination thereof. For example, the UI includes any one or more of the receptor information and ligand information mentioned above.
[0230] Those skilled in the art will understand that the structure shown in FIG13 does not constitute a limitation on the terminal device 1300 , and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0231] FIG14 is a schematic diagram of the structure of the server provided in an embodiment of the present application. The server 1400 may have relatively large differences due to different configurations or performances. It may include one or more processors 1401 and one or more memories 1402, wherein the one or more memories 1402 store at least one computer program, which is loaded and executed by the one or more processors 1401 to implement the training method of the ligand information generation model or the method of generating ligand information provided in the above-mentioned various method embodiments. Exemplarily, the processor 1401 is a CPU. Of course, the server 1400 may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server 1400 may also include other components for implementing device functions, which will not be described in detail here.
[0232] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to enable an electronic device to implement any of the above-mentioned ligand information generation model training methods or ligand information generation methods.
[0233] Optionally, the above-mentioned non-volatile computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0234] In an exemplary embodiment, a computer program is also provided. The computer program is at least one, and the at least one computer program is loaded and executed by a processor to enable an electronic device to implement any of the above-mentioned ligand information generation model training methods or ligand information generation methods.
[0235] In an exemplary embodiment, a computer program product is also provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to enable an electronic device to implement any of the above-mentioned methods for training a ligand information generation model or methods for generating ligand information.
[0236] It should be understood that the term "plurality" used herein refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship. The term "at least one" used herein refers to one or more. Similarly, "at least two" refers to two or more, and "at least three" refers to three or more.
[0237] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0238] The above description is merely an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A training method for a ligand information generation model, wherein: The method is performed by an electronic device, and includes: Acquiring sample receptor information and sample ligand information, wherein the binding affinity between the ligand described by the sample ligand information and the receptor described by the sample receptor information is not less than the set affinity; Denoising the reference noise data based on the sample receptor information using a neural network model to be trained to obtain predicted ligand information; determining a first loss for characterizing a difference between the sample ligand information and the predicted ligand information; The neural network model is trained based on the first loss to obtain a ligand information generation model, and the ligand information generation model is used to generate reference ligand information based on reference receptor information.
2. The method according to claim 1, wherein The neural network model includes at least two denoising networks; the neural network model to be trained denoises the reference noise data based on the sample receptor information to obtain predicted ligand information, including: For a first denoising network, denoising the reference noise data based on the sample receptor information by the first denoising network to obtain a denoising result of the first denoising network; For a non-first denoising network, denoising a denoising result of a previous denoising network of the non-first denoising network based on the sample receptor information by the non-first denoising network to obtain a denoising result of the non-first denoising network; The predicted ligand information is determined based on the denoising result of the last denoising network.
3. The method according to claim 2, wherein: The first denoising network includes an attention network and a feedforward network; Denoising the reference noise data based on the sample receptor information by the first denoising network to obtain a denoising result of the first denoising network includes: performing attention processing on the sample receptor information and the reference noise data through the attention network to obtain an attention processing result; A denoising result of the first denoising network is determined based on the attention processing result through the feedforward network.
4. The method according to claim 3, wherein: The sample receptor information includes at least two receptor composition information, and the reference noise data includes at least one first sub-data; The performing attention processing on the sample receptor information and the reference noise data by the attention network to obtain an attention processing result includes: For any first sub-data, performing attention processing on the first sub-data and each content information through the attention network to obtain processing results of each content information, and determining second sub-data corresponding to the first sub-data based on the processing results of each content information, wherein each content information includes the at least two receptor composition information and the at least one first sub-data; The attention processing result is determined based on the second sub-data corresponding to each of the first sub-data.
5. The method according to claim 3 or 4, wherein: Determining the denoising result of the first denoising network based on the attention processing result through the feedforward network includes: Performing mapping on the attention processing result through the feedforward network to obtain a mapping result; Mapping is performed on the mapping result or the set feature through the feedforward network to obtain a denoising result of the first denoising network.
6. The method according to any one of claims 2 to 5, wherein: The denoising result of the last denoising network includes at least two prediction sub-features; The determining the predicted ligand information based on the denoising result of the last denoising network includes: For any predictor feature, based on the feature similarity between the any predictor feature and at least two candidate ligand composition information, screening at least two target ligand composition information from the at least two candidate ligand composition information; sampling from the at least two target ligand composition information to obtain the ligand composition information corresponding to the any predictor feature; The predicted ligand information is determined based on the ligand composition information corresponding to each prediction sub-feature.
7. The method according to any one of claims 1 to 6, wherein: Before the step of denoising the reference noise data based on the sample receptor information by the neural network model to be trained to obtain the predicted ligand information, the method further includes: Noise is added to the sample ligand information to obtain the reference noise data.
8. The method according to claim 7, wherein: The number of times of adding noise is multiple times; The step of adding noise to the sample ligand information to obtain the reference noise data includes: For the first noise addition, performing the first noise addition on the sample ligand information to obtain a noise addition result of the first noise addition; For non-first noise addition, the non-first noise addition is performed on the noise addition result of the last noise addition before the non-first noise addition to obtain the noise addition result of the non-first noise addition, and the noise addition result of the last noise addition is the reference noise data.
9. The method according to claim 7 or 8, wherein The number of times of adding noise and the number of times of removing noise are both T, where T is a positive integer; and the training of the neural network model based on the first loss to obtain a ligand information generation model includes: For the i-th denoising, determining a second loss corresponding to the i-th denoising based on the noise data removed during the i-th denoising and the noise data added during the T+1-i-th denoising, where i is a positive integer from 1 to T; Based on the first loss and the second loss corresponding to each denoising, the neural network model is trained to obtain a ligand information generation model.
10. The method according to any one of claims 1 to 9, wherein: The step of training the neural network model based on the first loss to obtain a ligand information generation model includes: Training the neural network model based on the first loss to obtain a reference generative model; Denoising the reference noise data based on target receptor information using the reference generation model to obtain target ligand information; The reference generation model is trained based on the sample receptor information, the sample ligand information, the target receptor information, and the target ligand information to obtain the ligand information generation model.
11. The method according to claim 10, wherein: Denoising the reference noise data based on target receptor information using the reference generation model to obtain target ligand information includes: Denoising the reference noise data based on target receptor information using the reference generation model to obtain candidate ligand information; If the first quality indicator used to characterize the quality of the candidate ligand information meets a first indicator condition, the candidate ligand information is determined to be the target ligand information.
12. The method according to claim 11, wherein The method further comprises: Determining at least two first conformational information based on the target receptor information and the candidate ligand information, where the first conformational information is used to describe a conformation of a complex formed after the receptor described by the target receptor information binds to the ligand described by the candidate ligand information; For any piece of first conformational information, determining a first index of the first conformational information, wherein the first index is used to characterize the binding energy of the complex conformation described by the first conformational information; A first quality index of the candidate ligand information is determined based on the first index of each piece of first conformational information.
13. The method according to claim 11, wherein The method further comprises: Selecting reference conformation information from a database, wherein an alignment result between the ligand information included in the reference conformation information and the candidate ligand information satisfies an alignment condition, and an alignment result between the receptor information included in the reference conformation information and the target receptor information satisfies the alignment condition; Determining second conformational information based on the reference conformational information, the target receptor information, and the candidate ligand information, wherein the second conformational information is used to describe a conformation of a complex formed after the receptor described by the target receptor information binds to the ligand described by the candidate ligand information; Based on the second conformational information, a first quality indicator of the candidate ligand information is determined.
14. The method according to claim 11, wherein The method further comprises: Inputting the target receptor information and the candidate ligand information into a binding affinity prediction model to obtain a target binding affinity, wherein the target binding affinity is used to characterize the binding force between the receptor described by the target receptor information and the ligand described by the candidate ligand information; A first quality indicator of the candidate ligand information is determined based on the target binding affinity.
15. A method for generating ligand information, wherein: The method is performed by an electronic device, and includes: Obtain reference receptor information; The reference noise data is denoised based on the reference receptor information by a ligand information generation model to obtain reference ligand information, wherein the ligand information generation model is trained according to any one of the methods of claims 1 to 14.
16. The method according to claim 15, wherein The denoising of the reference noise data based on the reference receptor information by the ligand information generation model to obtain the reference ligand information includes: Denoising the reference noise data based on the reference receptor information using a ligand information generation model to obtain at least two optional ligand information; determining a second quality indicator for characterizing the quality of each optional ligand information; Select reference ligand information whose second quality index satisfies a second index condition from the at least two optional ligand information.
17. A training device for a ligand information generation model, wherein: The device comprises: an acquisition module, configured to acquire sample receptor information and sample ligand information, wherein the binding affinity between the ligand described by the sample ligand information and the receptor described by the sample receptor information is not less than a set affinity; a denoising module, configured to denoise the reference noise data based on the sample receptor information using a neural network model to be trained to obtain predicted ligand information; a determining module, configured to determine a first loss for characterizing a difference between the sample ligand information and the predicted ligand information; A training module is used to train the neural network model based on the first loss to obtain a ligand information generation model, wherein the ligand information generation model is used to generate reference ligand information based on reference receptor information.
18. A device for generating ligand information, wherein: The device comprises: An acquisition module, used for acquiring reference receptor information; A denoising module is used to denoise the reference noise data based on the reference receptor information through a ligand information generation model to obtain reference ligand information, wherein the ligand information generation model is trained according to any one of the methods described in claims 1 to 14.
19. An electronic device, wherein: The electronic device includes a processor and a memory, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor, so that the electronic device implements the method according to any one of claims 1 to 16.
20. A non-volatile computer-readable storage medium, wherein: The non-volatile computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by the processor to enable the electronic device to implement the method according to any one of claims 1 to 16.
21. A computer program, wherein The computer program is at least one, and the at least one computer program is loaded and executed by the processor to enable the electronic device to implement the method according to any one of claims 1 to 16.
22. A computer program product, wherein The computer program product stores at least one computer program, and the at least one computer program is loaded and executed by a processor to enable the electronic device to implement the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Ligand generation method and electronic equipment
CN116364172A
Aptamer generation method based on conditional discrete diffusion model
CN116631499A
PROTAC molecular design method based on graph neural network diffusion model
CN117524352A
Interpretable bio-medical link prediction using deep neural representation
US20190303535A1