Drug candidate library construction method, device, equipment and storage medium
By performing site mutations and graphical characterization on initial proteins, and combining spatial structure loss values with protein property prediction models, a drug candidate library is generated. This solves the problem of predicting the properties of proteins that do not exist in nature, and improves the efficiency and accuracy of new drug discovery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2023-03-22
- Publication Date
- 2026-05-01
AI Technical Summary
Current technologies cannot effectively predict the properties of proteins that do not exist in nature or have no homologous proteins, which slows down the discovery of new drugs.
Novel proteins are generated by site mutation of initial proteins, graphical characterization is performed and spatial structure loss values are calculated. Predicted properties are generated by combining protein property prediction models, target proteins are screened, and amino acid sequences are generated based on the three-dimensional structural data of antigen-antibody complexes to construct a drug candidate library.
It improves the accuracy of predicting the properties of proteins that do not exist in nature, enhances the convenience of new drug discovery and the accuracy of target protein screening, and improves the efficiency and accuracy of drug candidate library generation.
Smart Images

Figure CN116312760B_ABST
Abstract
Description
Methods, apparatus, equipment and storage media for constructing a drug candidate library Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and storage medium for constructing a drug candidate library. Background Technology
[0002] With the development of the internet, artificial intelligence technology has also been applied to the medical field. When constructing a drug candidate library for a certain disease, it is usually necessary to predict the properties of proteins. However, current methods for predicting protein properties can only predict properties based on the comparison of similarity between homologous proteins, and cannot predict the properties of proteins that do not exist in nature or have no homologous proteins. However, for new drug discovery, there is an urgent need for more new candidate protein molecules for high-throughput screening; otherwise, the speed of discovering new drugs from known proteins will inevitably slow down.
[0003] Therefore, how to predict the properties of proteins that do not exist in nature or have no homologous proteins in order to discover new drugs has become an urgent technical problem to be solved. Summary of the Invention
[0004] In view of the above, it is necessary to provide a method, apparatus, equipment and storage medium for constructing a drug candidate library, which can solve the technical problem of how to predict the properties of proteins that do not exist in nature or have no homologous proteins, so as to realize the discovery of new drugs.
[0005] On one hand, this invention proposes a method for constructing a drug candidate library, the method comprising:
[0006] By performing site mutations on the initial protein, a novel protein can be obtained;
[0007] The novel protein is graphically characterized to obtain a novel characterization vector, and the initial protein is graphically characterized to obtain an initial characterization vector.
[0008] Based on the novel representation vector and the initial representation vector, the spatial structure loss value of the novel protein and the initial protein is calculated.
[0009] The novel representation vector and the spatial structure loss value are input into the protein property prediction model corresponding to the disease to be treated to obtain the predicted properties of the novel protein.
[0010] Based on the predicted properties, target proteins are screened from a plurality of the novel proteins;
[0011] Based on the three-dimensional structure data of the antigen-antibody complex and the three-dimensional structure data of the disease antigen of the disease to be treated, an amino acid sequence is generated.
[0012] The target protein and its amino acid sequence are entered into a configuration library to obtain a drug candidate library for the disease to be treated.
[0013] According to a preferred embodiment of the present invention, the step of performing site mutations on the initial protein to obtain a novel protein includes:
[0014] Train a mutant prediction model based on mutation data in a pre-defined mutation library;
[0015] The configuration proteins of the disease to be treated in multiple configuration properties are obtained as the initial proteins;
[0016] Obtain the initial three-dimensional structure data of the initial protein;
[0017] Based on the mutant prediction model, site identification is performed on the initial three-dimensional structural data to obtain mutation sites;
[0018] Based on the mutation site, the initial protein was subjected to site-directed mutagenesis to obtain the novel protein.
[0019] According to a preferred embodiment of the present invention, the protein property prediction model includes the configuration features corresponding to the initial protein and the matching score between the initial protein and each configuration property. The step of inputting the novel representation vector and the spatial structure loss value into the protein property prediction model corresponding to the disease to be treated to obtain the predicted properties of the novel protein includes:
[0020] Based on the convolutional network in the protein property prediction model, the novel representation vector is convolved to obtain the feature information of the novel protein.
[0021] The feature information and the feature difference between the configured features, and the spatial structure loss value are input into the similarity function in the protein property prediction model to obtain the protein similarity.
[0022] Based on the protein similarity and the matching score of the configuration feature on each configuration property, the predicted score of the novel protein on each configuration property is calculated;
[0023] Based on the predicted score, the predicted property is selected from a plurality of the configuration properties.
[0024] According to a preferred embodiment of the present invention, the step of screening for target proteins from a plurality of novel proteins based on the predicted properties is as follows:
[0025] Based on the predicted properties, pharmacologically active proteins are screened from a plurality of the novel proteins;
[0026] Count the number of proteins in the pharmacologically active protein;
[0027] Based on the drug requirement for the disease to be treated and the sequence number of the amino acid sequence, a target quantity is generated;
[0028] If the number of proteins is greater than the target number, the target protein is sequentially selected from the multiple pharmacologically active proteins according to the order of the predicted scores from largest to smallest, and the number of the target protein is equal to the target number.
[0029] According to a preferred embodiment of the present invention, the step of performing graph characterization on the novel protein to obtain a novel characterization vector includes:
[0030] Nodes are constructed based on protein atoms in the novel protein;
[0031] The connection edges of the node are constructed based on the atomic distance between any two atoms in the plurality of protein atoms;
[0032] The node is encoded based on the atom type of the protein atom corresponding to the node, the surface position information of the protein atom, the three-dimensional coordinates of the protein atom, and the number of connecting edges of the node, to obtain the node representation.
[0033] By concatenating multiple node representations, the novel representation vector is obtained.
[0034] According to a preferred embodiment of the present invention, calculating the spatial structure loss value of the novel protein and the initial protein based on the novel representation vector and the initial representation vector includes:
[0035] The novel protein is cleaved to obtain a novel local backbone, and the initial protein is cleaved to obtain an initial local backbone;
[0036] Based on the atom type and the surface position information, a first novel vector is extracted from the novel representation vector, and based on the three-dimensional coordinates and the number of edges, a second novel vector is extracted from the novel representation vector;
[0037] Based on the first novel vector and the first initial vector in the initial representation vector, the atomic weight of each protein atom in the novel local skeleton is calculated, and based on the second novel vector and the second initial vector in the initial representation vector, the atomic distance between each protein atom in the novel local skeleton and the corresponding protein atom in the initial local skeleton is calculated as the target distance.
[0038] The spatial structure loss value is calculated based on the preset skeleton weights, the atomic weights, and the target distance of the novel local skeleton.
[0039] According to a preferred embodiment of the present invention, the amino acid sequence generated based on the three-dimensional structure data of the antigen-antibody complex and the three-dimensional structure data of the disease antigen of the disease to be treated includes:
[0040] Based on the three-dimensional structure data of the complex and the three-dimensional structure data of the antigen, target three-dimensional structure data with a three-dimensional distance less than a preset threshold are identified;
[0041] The compound corresponding to the target three-dimensional structure data is extracted from the antigen-antibody complex as the amino acid sequence.
[0042] On the other hand, the present invention also proposes a drug candidate library construction apparatus, the drug candidate library construction apparatus comprising:
[0043] Mutation units are used to mutate initial proteins at specific sites to obtain novel proteins.
[0044] The characterization unit is used to perform graphical characterization on the novel protein to obtain a novel characterization vector, and to perform graphical characterization on the initial protein to obtain an initial characterization vector.
[0045] The computing unit is used to calculate the spatial structure loss value of the novel protein and the initial protein based on the novel representation vector and the initial representation vector.
[0046] The input unit is used to input the novel representation vector and the spatial structure loss value into the protein property prediction model corresponding to the disease to be treated, so as to obtain the predicted properties of the novel protein.
[0047] A screening unit is configured to screen for a target protein from a plurality of the novel proteins based on the predicted properties;
[0048] The generation unit is used to generate an amino acid sequence based on the three-dimensional structure data of the antigen-antibody complex and the three-dimensional structure data of the disease antigen of the disease to be treated.
[0049] The input unit is used to input the target protein and the amino acid sequence into the configuration library to obtain the drug candidate library for the disease to be treated.
[0050] On the other hand, the present invention also proposes an electronic device, the electronic device comprising:
[0051] Memory, which stores computer-readable instructions; and
[0052] The processor executes computer-readable instructions stored in the memory to implement the drug candidate library construction method.
[0053] On the other hand, the present invention also proposes a computer-readable storage medium storing computer-readable instructions, which are executed by a processor in an electronic device to implement the drug candidate library construction method.
[0054] As can be seen from the above technical solutions, this application characterizes the novel protein and the initial protein in the same dimension, which can improve the accuracy of the calculation of the spatial structure loss value. By generating the predicted properties through the protein property prediction model, the problem of predicting the properties of proteins that do not exist in nature or have no homologous proteins can be solved, which improves the convenience of new drug discovery. Furthermore, by combining the novel characterization vector and the spatial structure loss value to generate the predicted properties, the accuracy of the generated predicted properties can be improved because the changes in spatial structure can be combined with the analysis of the novel protein and the initial protein, thereby improving the accuracy of the screening of the target protein. By combining the relationship between the three-dimensional structure data of the complex and the three-dimensional structure data of the antigen, the amino acid sequence can be directly identified from the antigen-antibody complex, improving the generation efficiency and accuracy of the amino acid sequence. Furthermore, by combining the target protein and the amino acid sequence, a reasonable drug candidate library for the disease to be treated can be generated. Attached Figure Description
[0055] Figure 1 is a flowchart of a preferred embodiment of the drug candidate library construction method of the present invention.
[0056] Figure 2 is a schematic diagram of the generation of novel proteins in the drug candidate library construction method of the present invention.
[0057] Figure 3 is a functional block diagram of a preferred embodiment of the drug candidate library construction device of the present invention.
[0058] Figure 4 is a schematic diagram of the structure of an electronic device that implements the drug candidate library construction method of the present invention. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0060] Figure 1 shows a flowchart of a preferred embodiment of the drug candidate library construction method of the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.
[0061] The method for constructing the drug candidate library can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0062] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0063] The method for constructing the drug candidate library is applied to one or more electronic devices. The electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored computer-readable instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0064] The electronic device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.
[0065] The electronic devices may include network devices and / or user devices. The network devices include, but are not limited to, single network electronic devices, groups of multiple network electronic devices, or cloud computing-based systems consisting of a large number of hosts or network electronic devices.
[0066] The network in which the electronic device is located includes, but is not limited to: the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.
[0067] 101. Site mutations were performed on the initial protein to obtain a novel protein.
[0068] In at least one embodiment of the present invention, the initial protein refers to a configured protein for a disease to be treated in terms of multiple configuration properties. The disease to be treated refers to a disease in the medical field requiring the construction of a drug candidate library; for example, the disease to be treated could be influenza, diabetes, etc. This application does not specifically limit the type of disease to be treated. The multiple configuration properties may include, but are not limited to, affinity, peptide activity, stability, and druggability. The configured protein can be any protein already existing in nature.
[0069] The novel protein refers to a protein obtained by mutating the initial protein, and the novel protein may be a protein that does not exist in nature.
[0070] In at least one embodiment of the present invention, the electronic device performs site mutations on an initial protein to obtain a novel protein, including:
[0071] Train a mutant prediction model based on mutation data in a pre-defined mutation library;
[0072] The configuration proteins of the disease to be treated in multiple configuration properties are obtained as the initial proteins;
[0073] Obtain the initial three-dimensional structure data of the initial protein;
[0074] Based on the mutant prediction model, site identification is performed on the initial three-dimensional structural data to obtain mutation sites;
[0075] Based on the mutation site, the initial protein was subjected to site-directed mutagenesis to obtain the novel protein.
[0076] The preset mutation library stores the mutant protein structures of multiple preset proteins. The mutation data includes the preset structure of each preset protein and the mutant protein structure.
[0077] The mutant prediction model can be trained on the mutation data using a deep learning network. The mutant prediction model is used to predict the specific sites where mutations can occur.
[0078] The initial three-dimensional structure data includes the coordinate information of each protein atom in the initial protein in three-dimensional space.
[0079] The mutation data can be used to train the mutant prediction model, which can then reasonably predict the mutation sites corresponding to the initial protein. Based on these mutation sites, the rationality of the generation of the novel protein can be improved.
[0080] Figure 2 is a schematic diagram illustrating the generation of novel proteins in the drug candidate library construction method of the present invention. In Figure 2, the initial protein includes atoms L1, L2, and L3. The mutation site is atom L1. Atom L1 is mutated to atom R1', and the spatial structure after mutation is recombined to obtain the novel protein, wherein the novel protein includes atoms R1', R2, and R3.
[0081] 102. The novel protein is graphically characterized to obtain a novel characterization vector, and the initial protein is graphically characterized to obtain an initial characterization vector.
[0082] In at least one embodiment of the present invention, the novel representation vector refers to the vector representation of the novel protein on the graph structure, and the initial representation vector refers to the vector representation of the initial protein on the graph structure.
[0083] In at least one embodiment of the present invention, the electronic device performs graph characterization on the novel protein to obtain a novel characterization vector, including:
[0084] Nodes are constructed based on protein atoms in the novel protein;
[0085] The connection edges of the node are constructed based on the atomic distance between any two atoms in the plurality of protein atoms;
[0086] The node is encoded based on the atom type of the protein atom corresponding to the node, the surface position information of the protein atom, the three-dimensional coordinates of the protein atom, and the number of connecting edges of the node, to obtain the node representation.
[0087] By concatenating multiple node representations, the novel representation vector is obtained.
[0088] The protein atoms may include, but are not limited to, carbon atoms, oxygen atoms, nitrogen atoms, and sulfur atoms.
[0089] The connecting edge refers to the edge corresponding to the covalent bond formed by any two atoms whose atomic distance is less than a preset distance.
[0090] The types of atoms include, but are not limited to, carbon, oxygen, nitrogen, and sulfur.
[0091] The surface position information refers to whether the protein atoms are on the surface of the information protein.
[0092] The three-dimensional coordinates refer to the coordinate information of the protein atoms on the spatial coordinate axes.
[0093] By using the atomic distance between any two atoms, the connecting edges can be reasonably constructed, thereby improving the rationality of determining the number of edges. Furthermore, by combining the atom type, the surface position information, the three-dimensional coordinates, and the number of edges, the specific information of the protein atoms in the spatial structure can be accurately characterized, thereby improving the representational ability of the novel representation vector in the spatial structure.
[0094] In other embodiments, the method of generating the initial representation vector is similar to the method of generating the novel representation vector, and will not be described in detail here.
[0095] 103. Based on the novel representation vector and the initial representation vector, calculate the spatial structure loss value of the novel protein and the initial protein.
[0096] In at least one embodiment of the present invention, the spatial structure loss value is used to represent the difference in spatial structure between the novel protein and the initial protein.
[0097] In at least one embodiment of the present invention, the electronic device calculates the spatial structure loss value of the novel protein and the initial protein based on the novel representation vector and the initial representation vector, including:
[0098] The novel protein is cleaved to obtain a novel local backbone, and the initial protein is cleaved to obtain an initial local backbone;
[0099] Based on the atom type and the surface position information, a first novel vector is extracted from the novel representation vector, and based on the three-dimensional coordinates and the number of edges, a second novel vector is extracted from the novel representation vector;
[0100] Based on the first novel vector and the first initial vector in the initial representation vector, the atomic weight of each protein atom in the novel local skeleton is calculated, and based on the second novel vector and the second initial vector in the initial representation vector, the atomic distance between each protein atom in the novel local skeleton and the corresponding protein atom in the initial local skeleton is calculated as the target distance.
[0101] The spatial structure loss value is calculated based on the preset skeleton weights, the atomic weights, and the target distance of the novel local skeleton.
[0102] The novel local framework can be a rigid structure in the novel protein. The initial local framework can be a rigid structure in the initial protein.
[0103] The first novel vector is used to characterize the specific information of protein atoms in the novel protein in terms of the atom type and the surface position information, and the second novel vector is used to characterize the specific information of protein atoms in the novel protein in terms of the three-dimensional coordinates and the number of edges.
[0104] The first initial vector is used to characterize the specific information of the protein atoms in the initial protein in terms of the atom type and the surface position information, and the second initial vector is used to characterize the specific information of the protein atoms in the initial protein in terms of the three-dimensional coordinates and the number of edges.
[0105] The atomic weight can be the vector cosine value between the first novel vector and the first initial vector.
[0106] The target distance can be the distance value between the second novel vector and the second initial vector on distance labels such as Euclidean distance or Mahalanobis distance.
[0107] The preset skeleton weights can be determined based on the stability of the specific rigid structure of the novel local skeleton.
[0108] The formula for calculating the spatial structure loss value is as follows:
[0109]
[0110] Where y represents the spatial structure loss value, N iF N represents the preset skeleton weight corresponding to the i-th novel local skeleton, n represents the number of skeletons in the plurality of novel local skeletons, and N represents the number of skeletons in the plurality of novel local skeletons. jP Let x represent the atomic weight corresponding to the j-th protein atom in the i-th novel local framework, m represent the number of protein atoms in the i-th novel local framework, and x represent the atomic weight corresponding to the j-th protein atom in the i-th novel local framework. j This represents the second novel vector. This represents the second initial vector. This indicates the target distance.
[0111] By segmenting the novel protein, reasonable preset skeleton weights can be set based on the stability of the rigid structure of the novel local skeleton. Then, the atomic weights are determined by the first novel vector and the first initial vector. By combining the preset skeleton weights and the atomic weights to calculate the target distance, the accuracy of determining the spatial structure loss value can be improved.
[0112] 104. Input the novel representation vector and the spatial structure loss value into the protein property prediction model corresponding to the disease to be treated to obtain the predicted properties of the novel protein.
[0113] In at least one embodiment of the present invention, the protein property prediction model includes configuration features corresponding to the initial protein and matching scores between the initial protein and each configuration property.
[0114] The predicted properties may include, but are not limited to: affinity, whether the peptide is active, stability, druggability, etc.
[0115] In at least one embodiment of the present invention, the electronic device inputs the novel representation vector and the spatial structure loss value into a protein property prediction model corresponding to the disease to be treated, and obtains the predicted properties of the novel protein, including:
[0116] Based on the convolutional network in the protein property prediction model, the novel representation vector is convolved to obtain the feature information of the novel protein.
[0117] The feature information and the feature difference between the configured features, and the spatial structure loss value are input into the similarity function in the protein property prediction model to obtain the protein similarity.
[0118] Based on the protein similarity and the matching score of the configuration feature on each configuration property, the predicted score of the novel protein on each configuration property is calculated;
[0119] Based on the predicted score, the predicted property is selected from a plurality of the configuration properties.
[0120] The convolutional network may include convolutional layers in a feedforward neural network and a backward neural network.
[0121] The feature difference can be the feature difference between the feature information and the configuration feature.
[0122] The similarity function can be expressed in the following form:
[0123] d = k1a + k2b;
[0124] Wherein, d represents the protein similarity, a represents the feature difference, b represents the spatial structure loss value, and k1 and k2 represent the hyperparameters of the protein property prediction model, respectively.
[0125] The predictive property refers to the configuration property where the predicted score is greater than the preset score.
[0126] By performing convolution processing on the novel representation vector through the convolutional network, the feature extraction of the novel representation vector can be combined with the contextual semantics, thereby improving the accuracy of the feature information. This avoids the problem of low accuracy of protein similarity caused by inaccurate feature extraction. Furthermore, by combining the matching score of the configuration feature on each configuration property, the prediction score can be reasonably calculated, thereby improving the accuracy of the generation of the prediction property.
[0127] In this embodiment, by generating predictive properties of the novel protein, it is possible to detect whether the novel protein is druggable for the disease to be treated, thereby identifying multiple proteins that can treat the disease and reducing pharmaceutical costs.
[0128] 105. Based on the predicted properties, the target protein is screened from a plurality of the novel proteins.
[0129] In at least one embodiment of the present invention, the number of the target protein is equal to the difference between the amount of drug required for the disease to be treated and the number of amino acid sequences subsequently generated.
[0130] In at least one embodiment of the present invention, the electronic device screens for a target protein from a plurality of novel proteins based on the predicted property:
[0131] Based on the predicted properties, pharmacologically active proteins are screened from a plurality of the novel proteins;
[0132] Count the number of proteins in the pharmacologically active protein;
[0133] Based on the drug demand quantity and amino acid sequence quantity of the disease to be treated, a target quantity is generated;
[0134] If the number of proteins is greater than the target number, the target protein is sequentially selected from the multiple pharmacologically active proteins according to the order of the predicted scores from largest to smallest, and the number of the target protein is equal to the target number.
[0135] The pharmacologically active protein refers to a novel protein whose predicted properties indicate druggability for the disease to be treated.
[0136] The required quantity of the drug can be determined based on the needs of the pharmaceutical user; for example, the required quantity of the drug could be 50.
[0137] The sequence number refers to the total number of amino acid sequences.
[0138] Based on the predictive properties, pharmacodynamic proteins that are druggable for the disease to be treated can be accurately located. Then, by combining the required drug quantity, the number of sequences, and the predictive score, the target protein can be rationally screened from the pharmacodynamic proteins.
[0139] 106. Based on the three-dimensional structure data of the antigen-antibody complex and the three-dimensional structure data of the disease antigen of the disease to be treated, an amino acid sequence is generated.
[0140] In at least one embodiment of the present invention, the three-dimensional structural data of the complex refers to the spatial coordinate information of each complex atom in the antigen-antibody complex.
[0141] The three-dimensional structure data of the antigen refers to the spatial coordinate information of each antigen atom in the disease antigen.
[0142] The amino acid sequence refers to the sequence on the binding surface of the disease antigen and the disease antibody, wherein the disease antibody refers to a protein capable of binding to the disease antigen.
[0143] In at least one embodiment of the present invention, the electronic device generates an amino acid sequence based on the three-dimensional structure data of the antigen-antibody complex and the three-dimensional structure data of the disease antigen of the disease to be treated, including:
[0144] Based on the three-dimensional structure data of the complex and the three-dimensional structure data of the antigen, target three-dimensional structure data with a three-dimensional distance less than a preset threshold are identified;
[0145] The compound corresponding to the target three-dimensional structure data is extracted from the antigen-antibody complex as the amino acid sequence.
[0146] The preset threshold can be set according to actual needs.
[0147] By introducing the preset distance to detect the target three-dimensional structure data, the comprehensiveness of the detection of the target three-dimensional structure data can be improved, thereby improving the accuracy of the determination of the amino acid sequence.
[0148] 107. The target protein and the amino acid sequence are entered into the configuration library to obtain the drug candidate library for the disease to be treated.
[0149] It should be emphasized that, to further ensure the privacy and security of the aforementioned drug candidate database, the database can also be stored in a node of a blockchain.
[0150] In at least one embodiment of the present invention, the configuration library may be a database of any format, and this application does not limit the specific format.
[0151] The drug candidate library refers to the database obtained after the target protein and the amino acid sequence are entered into the configuration library.
[0152] As can be seen from the above technical solutions, this application characterizes the novel protein and the initial protein in the same dimension, which can improve the accuracy of the calculation of the spatial structure loss value. By generating the predicted properties through the protein property prediction model, the problem of predicting the properties of proteins that do not exist in nature or have no homologous proteins can be solved, which improves the convenience of new drug discovery. Furthermore, by combining the novel characterization vector and the spatial structure loss value to generate the predicted properties, the accuracy of the generated predicted properties can be improved because the changes in spatial structure can be combined with the analysis of the novel protein and the initial protein, thereby improving the accuracy of the screening of the target protein. By combining the relationship between the three-dimensional structure data of the complex and the three-dimensional structure data of the antigen, the amino acid sequence can be directly identified from the antigen-antibody complex, improving the generation efficiency and accuracy of the amino acid sequence. Furthermore, by combining the target protein and the amino acid sequence, a reasonable drug candidate library for the disease to be treated can be generated.
[0153] Figure 3 shows a functional block diagram of a preferred embodiment of the drug candidate library construction device of the present invention. The drug candidate library construction device 11 includes a mutation unit 110, a characterization unit 111, a calculation unit 112, an input unit 113, a screening unit 114, a generation unit 115, and a recording unit 116. The module / unit referred to in this invention refers to a series of computer-readable instruction segments that can be acquired by the processor 13 and perform a fixed function, and are stored in the memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0154] Mutation unit 110 is used to perform site mutations on the initial protein to obtain a novel protein;
[0155] The characterization unit 111 is used to perform graph characterization on the novel protein to obtain a novel characterization vector, and to perform graph characterization on the initial protein to obtain an initial characterization vector.
[0156] The calculation unit 112 is used to calculate the spatial structure loss value of the novel protein and the initial protein based on the novel representation vector and the initial representation vector.
[0157] Input unit 113 is used to input the novel representation vector and the spatial structure loss value into the protein property prediction model corresponding to the disease to be treated, so as to obtain the predicted properties of the novel protein.
[0158] Screening unit 114 is used to screen target proteins from a plurality of novel proteins based on the predicted properties;
[0159] The generation unit 115 is used to generate an amino acid sequence based on the three-dimensional structure data of the antigen-antibody complex and the three-dimensional structure data of the disease antigen of the disease to be treated.
[0160] The input unit 116 is used to input the target protein and the amino acid sequence into the configuration library to obtain the drug candidate library for the disease to be treated.
[0161] In at least one embodiment of the present invention, the mutation unit 110 is further used to train a mutant prediction model based on mutation data in a preset mutation library;
[0162] The configuration proteins of the disease to be treated in multiple configuration properties are obtained as the initial proteins;
[0163] Obtain the initial three-dimensional structure data of the initial protein;
[0164] Based on the mutant prediction model, site identification is performed on the initial three-dimensional structural data to obtain mutation sites;
[0165] Based on the mutation site, the initial protein was subjected to site-directed mutagenesis to obtain the novel protein.
[0166] In at least one embodiment of the present invention, the protein property prediction model includes configuration features corresponding to the initial protein and matching scores between the initial protein and each configuration property. The input unit 113 is further configured to perform convolution processing on the novel representation vector based on the convolutional network in the protein property prediction model to obtain the feature information of the novel protein.
[0167] The feature information and the feature difference between the configured features, and the spatial structure loss value are input into the similarity function in the protein property prediction model to obtain the protein similarity.
[0168] Based on the protein similarity and the matching score of the configuration feature on each configuration property, the predicted score of the novel protein on each configuration property is calculated;
[0169] Based on the predicted score, the predicted property is selected from a plurality of the configuration properties.
[0170] In at least one embodiment of the present invention, the screening unit 114 is further configured to screen for pharmacologically active proteins from a plurality of novel proteins based on the predicted properties.
[0171] Count the number of proteins in the pharmacologically active protein;
[0172] Based on the drug requirement for the disease to be treated and the sequence number of the amino acid sequence, a target quantity is generated;
[0173] If the number of proteins is greater than the target number, the target protein is sequentially selected from the multiple pharmacologically active proteins according to the order of the predicted scores from largest to smallest, and the number of the target protein is equal to the target number.
[0174] In at least one embodiment of the present invention, the characterization unit 111 is further configured to construct nodes based on protein atoms in the novel protein;
[0175] The connection edges of the node are constructed based on the atomic distance between any two atoms in the plurality of protein atoms;
[0176] The node is encoded based on the atom type of the protein atom corresponding to the node, the surface position information of the protein atom, the three-dimensional coordinates of the protein atom, and the number of connecting edges of the node, to obtain the node representation.
[0177] By concatenating multiple node representations, the novel representation vector is obtained.
[0178] In at least one embodiment of the present invention, the computing unit 112 is further configured to segment the novel protein to obtain a novel local backbone, and segment the initial protein to obtain an initial local backbone.
[0179] Based on the atom type and the surface position information, a first novel vector is extracted from the novel representation vector, and based on the three-dimensional coordinates and the number of edges, a second novel vector is extracted from the novel representation vector;
[0180] Based on the first novel vector and the first initial vector in the initial representation vector, the atomic weight of each protein atom in the novel local skeleton is calculated, and based on the second novel vector and the second initial vector in the initial representation vector, the atomic distance between each protein atom in the novel local skeleton and the corresponding protein atom in the initial local skeleton is calculated as the target distance.
[0181] The spatial structure loss value is calculated based on the preset skeleton weights, the atomic weights, and the target distance of the novel local skeleton.
[0182] In at least one embodiment of the present invention, the generation unit 115 is further configured to identify target three-dimensional structure data with a three-dimensional distance less than a preset threshold based on the three-dimensional structure data of the complex and the three-dimensional structure data of the antigen.
[0183] The compound corresponding to the target three-dimensional structure data is extracted from the antigen-antibody complex as the amino acid sequence.
[0184] As can be seen from the above technical solutions, this application characterizes the novel protein and the initial protein in the same dimension, which can improve the accuracy of the calculation of the spatial structure loss value. By generating the predicted properties through the protein property prediction model, the problem of predicting the properties of proteins that do not exist in nature or have no homologous proteins can be solved, which improves the convenience of new drug discovery. Furthermore, by combining the novel characterization vector and the spatial structure loss value to generate the predicted properties, the accuracy of the generated predicted properties can be improved because the changes in spatial structure can be combined with the analysis of the novel protein and the initial protein, thereby improving the accuracy of the screening of the target protein. By combining the relationship between the three-dimensional structure data of the complex and the three-dimensional structure data of the antigen, the amino acid sequence can be directly identified from the antigen-antibody complex, improving the generation efficiency and accuracy of the amino acid sequence. Furthermore, by combining the target protein and the amino acid sequence, a reasonable drug candidate library for the disease to be treated can be generated.
[0185] Figure 4 shows a schematic diagram of the electronic device for implementing the drug candidate library construction method of the present invention.
[0186] In one embodiment of the present invention, the electronic device 1 includes, but is not limited to, a memory 12, a processor 13, and computer-readable instructions stored in the memory 12 and executable on the processor 13, such as a drug candidate library construction program.
[0187] Those skilled in the art will understand that the schematic diagram is merely an example of electronic device 1 and does not constitute a limitation on electronic device 1. It may include more or fewer components than shown in the diagram, or combine certain components, or different components. For example, electronic device 1 may also include input / output devices, network access devices, buses, etc.
[0188] The processor 13 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor 13 is the computing core and control center of the electronic device 1, connecting various parts of the electronic device 1 through various interfaces and lines, and executing the operating system of the electronic device 1, as well as various installed application programs and program code.
[0189] For example, the computer-readable instructions can be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units can be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer-readable instructions in the electronic device 1. For example, the computer-readable instructions can be divided into a mutation unit 110, a characterization unit 111, a calculation unit 112, an input unit 113, a filtering unit 114, a generation unit 115, and a recording unit 116.
[0190] The memory 12 can be used to store the computer-readable instructions and / or modules. The processor 13 implements various functions of the electronic device 1 by running or executing the computer-readable instructions and / or modules stored in the memory 12 and calling the data stored in the memory 12. The memory 12 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. The memory 12 may include non-volatile and volatile memory, such as: hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other storage devices.
[0191] The memory 12 can be the external memory and / or internal memory of the electronic device 1. Furthermore, the memory 12 can be a physical memory, such as a memory module, a TF card (Trans-flash Card), etc.
[0192] If the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by instructing related hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when executed by a processor, the computer-readable instructions can implement the steps of the various method embodiments described above.
[0193] The computer-readable instructions include computer-readable instruction code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer-readable instruction code, recording medium, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), and random access memory (RAM).
[0194] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed drug candidate library construction, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0195] Referring to Figure 1, the memory 12 in the electronic device 1 stores computer-readable instructions to implement a method for constructing a drug candidate library, and the processor 13 can execute the computer-readable instructions to achieve the following:
[0196] By performing site mutations on the initial protein, a novel protein can be obtained;
[0197] The novel protein is graphically characterized to obtain a novel characterization vector, and the initial protein is graphically characterized to obtain an initial characterization vector.
[0198] Based on the novel representation vector and the initial representation vector, the spatial structure loss value of the novel protein and the initial protein is calculated.
[0199] The novel representation vector and the spatial structure loss value are input into the protein property prediction model corresponding to the disease to be treated to obtain the predicted properties of the novel protein.
[0200] Based on the predicted properties, target proteins are screened from a plurality of the novel proteins;
[0201] Based on the three-dimensional structure data of the antigen-antibody complex and the three-dimensional structure data of the disease antigen of the disease to be treated, an amino acid sequence is generated.
[0202] The target protein and its amino acid sequence are entered into a configuration library to obtain a drug candidate library for the disease to be treated.
[0203] Specifically, the specific implementation method of the above-mentioned computer-readable instructions by the processor 13 can be referred to the description of the relevant steps in the embodiment corresponding to FIG1, which will not be repeated here.
[0204] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0205] The computer-readable storage medium stores computer-readable instructions, which, when executed by the processor 13, are used to perform the following steps:
[0206] By performing site mutations on the initial protein, a novel protein can be obtained;
[0207] The novel protein is graphically characterized to obtain a novel characterization vector, and the initial protein is graphically characterized to obtain an initial characterization vector.
[0208] Based on the novel representation vector and the initial representation vector, the spatial structure loss value of the novel protein and the initial protein is calculated.
[0209] The novel representation vector and the spatial structure loss value are input into the protein property prediction model corresponding to the disease to be treated to obtain the predicted properties of the novel protein.
[0210] Based on the predicted properties, target proteins are screened from a plurality of the novel proteins;
[0211] Based on the three-dimensional structure data of the antigen-antibody complex and the three-dimensional structure data of the disease antigen of the disease to be treated, an amino acid sequence is generated.
[0212] The target protein and its amino acid sequence are entered into a configuration library to obtain a drug candidate library for the disease to be treated.
[0213] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0214] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0215] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0216] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices described may also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.
[0217] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for constructing a drug candidate library, characterized in that, The method for constructing the drug candidate library includes: performing site mutations on an initial protein to obtain a novel protein; performing graph representation on the novel protein to obtain a novel representation vector, and performing graph representation on the initial protein to obtain an initial representation vector. The process of performing graph representation on the novel protein to obtain the novel representation vector includes: constructing nodes based on protein atoms in the novel protein; constructing connection edges of the nodes based on the atomic distance between any two atoms in the plurality of protein atoms; encoding the node based on the atom type of the protein atom corresponding to the node, the surface position information of the protein atom, the three-dimensional coordinates of the protein atom, and the number of edges of the connection edges of the node to obtain a node representation; concatenating multiple node representations to obtain the novel representation vector; and calculating the spatial relationship between the novel protein and the initial protein based on the novel representation vector and the initial representation vector. The spatial structure loss value includes: segmenting the novel protein to obtain a novel local backbone, and segmenting the initial protein to obtain an initial local backbone; extracting a first novel vector from the novel representation vector based on the atom type and the surface position information, and extracting a second novel vector from the novel representation vector based on the three-dimensional coordinates and the number of edges; calculating the atomic weight of each protein atom in the novel local backbone based on the first novel vector and the first initial vector in the initial representation vector, and calculating the atomic distance between each protein atom in the novel local backbone and the corresponding protein atom in the initial local backbone as the target distance based on the second novel vector and the second initial vector in the initial representation vector; and calculating the spatial structure loss value based on the preset backbone weight of the novel local backbone, the atomic weight, and the target distance. The formula for calculating the spatial structure loss value is as follows: ;in, This represents the spatial structure loss value. Indicates the first The preset skeleton weights corresponding to each novel local skeleton. This indicates the number of skeletons in the plurality of novel local skeletons. Indicates the first The first of the novel local skeletons The atomic weights corresponding to each protein atom Indicates the first The number of protein atoms in a novel local backbone This represents the second novel vector. This represents the second initial vector. The target distance is represented; the novel representation vector and the spatial structure loss value are input into the protein property prediction model corresponding to the disease to be treated to obtain the predicted properties of the novel protein; based on the predicted properties, the target protein is screened from multiple novel proteins; based on the three-dimensional structure data of the antigen-antibody complex and the three-dimensional structure data of the disease antigen of the disease to be treated, an amino acid sequence is generated; the target protein and the amino acid sequence are entered into the configuration library to obtain the drug candidate library for the disease to be treated.
2. The method for constructing a drug candidate library as described in claim 1, characterized in that, The step of performing site mutations on the initial protein to obtain a novel protein includes: training a mutant prediction model based on mutation data in a preset mutation library; obtaining configuration proteins of the disease to be treated in multiple configuration properties as the initial protein; obtaining initial three-dimensional structure data of the initial protein; identifying mutation sites on the initial three-dimensional structure data based on the mutant prediction model to obtain mutation sites; and performing site-directed mutations on the initial protein based on the mutation sites to obtain the novel protein.
3. The method for constructing a drug candidate library as described in claim 2, characterized in that, The protein property prediction model includes the configuration features corresponding to the initial protein and the matching score between the initial protein and each configuration feature. The step of inputting the novel representation vector and the spatial structure loss value into the protein property prediction model corresponding to the disease to be treated to obtain the predicted properties of the novel protein includes: performing convolution processing on the novel representation vector based on the convolutional network in the protein property prediction model to obtain the feature information of the novel protein; inputting the feature information and the feature difference between the configuration features and the spatial structure loss value into the similarity function in the protein property prediction model to obtain the protein similarity; calculating the predicted score of the novel protein on each configuration feature based on the protein similarity and the matching score of the configuration feature on each configuration feature; and selecting the predicted properties from multiple configuration features based on the predicted scores.
4. The method for constructing a drug candidate library as described in claim 3, characterized in that, The step involves screening for target proteins from a plurality of novel proteins based on the predicted properties, and screening for pharmacologically active proteins from a plurality of novel proteins based on the predicted properties. Count the number of proteins in the pharmacologically active protein; Based on the drug requirement for the disease to be treated and the sequence number of the amino acid sequence, a target quantity is generated; If the number of proteins is greater than the target number, the target protein is sequentially selected from the multiple pharmacologically active proteins according to the order of the predicted scores from largest to smallest, and the number of the target protein is equal to the target number.
5. The method for constructing a drug candidate library as described in claim 1, characterized in that, The process of generating an amino acid sequence based on the three-dimensional structure data of the antigen-antibody complex and the three-dimensional structure data of the disease antigen of the disease to be treated includes: identifying target three-dimensional structure data with a three-dimensional distance less than a preset threshold based on the three-dimensional structure data of the complex and the three-dimensional structure data of the antigen; and extracting a compound corresponding to the target three-dimensional structure data from the antigen-antibody complex as the amino acid sequence.
6. A drug candidate library construction apparatus, characterized in that, The drug candidate library construction device includes: a mutation unit for performing site mutations on an initial protein to obtain a novel protein; a characterization unit for performing graph characterization on the novel protein to obtain a novel characterization vector, and performing graph characterization on the initial protein to obtain an initial characterization vector. The graph characterization of the novel protein to obtain the novel characterization vector includes: constructing nodes based on protein atoms in the novel protein; constructing connection edges of the nodes based on the atomic distance between any two atoms in the plurality of protein atoms; encoding the node based on the atom type of the protein atom corresponding to the node, the surface position information of the protein atom, the three-dimensional coordinates of the protein atom, and the number of edges of the connection edges of the node to obtain a node characterization; concatenating multiple node characterizations to obtain the novel characterization vector; and a calculation unit for calculating the relationship between the novel protein and the initial characterization vector. The initial protein spatial structure loss value includes: segmenting the novel protein to obtain a novel local backbone, and segmenting the initial protein to obtain an initial local backbone; extracting a first novel vector from the novel representation vector based on the atom type and the surface position information, and extracting a second novel vector from the novel representation vector based on the three-dimensional coordinates and the number of edges; calculating the atomic weight of each protein atom in the novel local backbone based on the first novel vector and the first initial vector in the initial representation vector, and calculating the atomic distance between each protein atom in the novel local backbone and the corresponding protein atom in the initial local backbone as the target distance based on the second novel vector and the second initial vector in the initial representation vector; calculating the spatial structure loss value based on the preset backbone weight of the novel local backbone, the atomic weight, and the target distance, wherein the formula for calculating the spatial structure loss value is: ;in, This represents the spatial structure loss value. Indicates the first The preset skeleton weights corresponding to each novel local skeleton. This indicates the number of skeletons in the plurality of novel local skeletons. Indicates the first The first of the novel local skeletons The atomic weights corresponding to each protein atom Indicates the first The number of protein atoms in a novel local backbone This represents the second novel vector. This represents the second initial vector. The system comprises: a target distance unit; an input unit for inputting the novel representation vector and the spatial structure loss value into a protein property prediction model corresponding to the disease to be treated, to obtain the predicted properties of the novel protein; a screening unit for screening the target protein from multiple novel proteins based on the predicted properties; a generation unit for generating an amino acid sequence based on the three-dimensional structure data of the antigen-antibody complex and the three-dimensional structure data of the disease antigen of the disease to be treated; and an input unit for inputting the target protein and the amino acid sequence into a configuration library to obtain a drug candidate library for the disease to be treated.
7. An electronic device, characterized in that, The electronic device includes: a memory storing computer-readable instructions; and a processor executing the computer-readable instructions stored in the memory to implement the drug candidate library construction method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions that are executed by a processor in an electronic device to implement the drug candidate library construction method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Method and device for model training, drug screening and affinity prediction
CN114333986A
Intelligent medicine research and development device, storage medium and computer equipment
CN114373520A
Residual artificial neural network to generate protein sequences
WO2023034865A2