Methods and apparatus for polypeptide production
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING NORMAL UNIV AT ZHUHAI
- Filing Date
- 2023-12-01
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]尽管深度学习在生物信息学领域表现出色,但是由于神经网络模型通常需要大量的训练数据,对多肽设计领域而言,获得大量的高质量训练数据是十分困难的,导致利用神经网络模型进行多肽设计受到限制,无法得到更好地多肽
[0064]本申请提供的多肽生成方法及装置,通过对预设多肽信息池中的至少一个已有多肽的氨基酸序列信息进行梯元提取,根据提取到的多个梯元,重构生成至少一个新氨基酸序列信息,根据新氨基酸序列信息与目标蛋白质的结构信息的分子对接结构,从新氨基酸序列信息中确定生成目标新多肽,一方面,对已有多肽的数量要求较低,不需要大量的多肽也可以生成新多肽,避免了对大量训练数据的依赖;另一方面,基于已有多肽的氨基酸序列中的梯元重构新的氨基酸序列信息,可以直观地了解对目标蛋白质更有效的氨基酸片段,使得生成新多肽的过程的解释性更好,提高新多肽在多肽药物设计过程中的实用价值。
Smart Images

Figure CN117612637B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of peptide design technology, and more specifically, to a peptide generation method and apparatus. Background Technology
[0002] Peptides are bioactive molecules composed of short-chain amino acid sequences. Due to their high selectivity and potency, and generally low toxicity, they have received widespread attention in drug development. Designing and screening peptides with drug potential is a complex process that requires consideration of multiple aspects, including the peptide's structure, stability, and interactions with target proteins.
[0003] In recent years, deep learning has achieved remarkable success in many bioinformatics tasks. Deep learning typically uses neural network models to learn the interaction between peptides and target proteins. By training neural network models, new peptides can be predicted.
[0004] Although deep learning has performed well in the field of bioinformatics, neural network models usually require a large amount of training data. For the field of peptide design, it is very difficult to obtain a large amount of high-quality training data, which limits the use of neural network models for peptide design and prevents the generation of better peptides. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of the prior art by providing a method and apparatus for generating peptides, thereby reducing reliance on large amounts of training data and improving the interaction between the generated peptides and target proteins.
[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:
[0007] In a first aspect, embodiments of this application provide a method for generating a polypeptide, the method comprising:
[0008] Obtain a preset peptide information pool and structural information of a target protein, wherein the preset peptide information pool includes: amino acid sequence information of at least one existing peptide;
[0009] The amino acid sequence information of the at least one existing polypeptide is subjected to ladder extraction to obtain multiple ladders, wherein the ladders are repeating amino acid fragments in the amino acid sequence information of the at least one existing polypeptide, or basic amino acids constituting the amino acid sequence information.
[0010] Based on the multiple ladder units, at least one new amino acid sequence information is reconstructed and generated;
[0011] Based on the molecular docking results of the at least one new amino acid sequence information and the structural information of the target protein, a new target polypeptide is determined from the at least one new amino acid sequence information.
[0012] Optionally, the step of performing ladder extraction on the amino acid sequence information of the at least one existing polypeptide to obtain multiple ladders includes:
[0013] Search the amino acid sequence information of the at least one existing polypeptide to identify at least one target existing polypeptide having a first repeating amino acid fragment;
[0014] Based on the first repeating amino acid fragment, the amino acid sequence information of each target existing polypeptide is cut into a set of amino acid fragments;
[0015] Extract the first repeating amino acid fragment from a set of amino acid fragments of a target existing polypeptide to obtain a ladder and a set of remaining fragments corresponding to the target existing polypeptide;
[0016] The amino acid fragment set of other target peptides, the remaining fragment set, and the amino acid sequence information of other existing peptides are searched to determine the second repeating amino acid fragment, until the remaining fragment set of the at least one existing peptide does not contain repeating amino acid fragments, thus obtaining the ladder units corresponding to all repeating amino acid fragments.
[0017] Optionally, the step of reconstructing and generating at least one new amino acid sequence information based on the plurality of ladders includes:
[0018] Initialize a space structure for an amino acid sequence;
[0019] Select multiple target ladder elements from the plurality of ladder elements;
[0020] Based on the positions of the multiple target ladders in the corresponding existing polypeptide amino acid sequence information, the multiple target ladders are filled into the amino acid sequence empty structure until the amino acid sequence empty structure is filled, thereby obtaining the new amino acid sequence information.
[0021] Optionally, selecting multiple target ladder elements from the plurality of ladder elements includes:
[0022] The selection probability of the multiple ladder elements is determined based on their weights.
[0023] Based on the selection probability of the multiple ladder elements, the multiple target ladder elements are selected from the multiple ladder elements.
[0024] Optionally, before determining the selection probability of the plurality of ladder elements based on their weights, the method further includes:
[0025] The weight of the ladder corresponding to each repeating amino acid fragment is determined based on the number of times each repeating amino acid fragment is extracted.
[0026] Optionally, filling the multiple target ladders into the empty amino acid sequence structure according to their positions in the corresponding existing polypeptide amino acid sequence information includes:
[0027] Based on the position of the first target ladder in the amino acid sequence information of the corresponding existing polypeptide, the first target ladder is filled in at the corresponding position of the empty structure of the amino acid sequence;
[0028] If the length of the second target ladder is greater than the length of the first target ladder, and the position of the second target ladder in the corresponding amino acid sequence information partially overlaps with the position of the first target ladder in the corresponding amino acid sequence information, then the first target ladder in the amino acid sequence space structure is replaced with the second target ladder.
[0029] Optionally, the step of determining the generation of a target new polypeptide from the at least one new amino acid sequence information based on the molecular docking results of the at least one new amino acid sequence information and the structural information of the target protein includes:
[0030] The molecular docking result of each new amino acid sequence is determined by scoring the molecular docking between each new amino acid sequence and multiple structural information of the target protein.
[0031] Based on the molecular docking results of the at least one new amino acid sequence information, the target new polypeptide is determined to be generated from the at least one new amino acid sequence information.
[0032] Optionally, the step of determining the generation of a target new polypeptide from the at least one new amino acid sequence information based on the molecular docking results of the at least one new amino acid sequence information and the structural information of the target protein includes:
[0033] Peptide chain structure analysis is performed on the polypeptide corresponding to the at least one new amino acid sequence information to determine the new amino acid sequence information with the target peptide chain structure.
[0034] Based on the molecular docking results of the new amino acid sequence information having the target peptide chain structure and the structural information of the target protein, the target new polypeptide is determined from the new amino acid sequence information having the target peptide chain structure.
[0035] Optionally, the method further includes:
[0036] The new amino acid sequence information of the target new polypeptide is added to the preset polypeptide information pool.
[0037] Secondly, embodiments of this application also provide a polypeptide generation apparatus, the apparatus comprising:
[0038] An information acquisition module is used to acquire a preset peptide information pool and structural information of a target protein. The preset peptide information pool includes: amino acid sequence information of at least one existing peptide.
[0039] A ladder extraction module is used to extract the amino acid sequence information of the at least one existing polypeptide into multiple ladders, wherein the ladders are repeating amino acid fragments in the amino acid sequence information of the at least one existing polypeptide, or basic amino acids constituting the amino acid sequence information.
[0040] The ladder reconstruction module is used to reconstruct and generate at least one new amino acid sequence information based on the multiple ladders;
[0041] The peptide generation module is used to determine the generation of a target new peptide from the at least one new amino acid sequence information based on the molecular docking results of the at least one new amino acid sequence information and the structural information of the target protein.
[0042] Optionally, the ladder element extraction module includes:
[0043] A repeating amino acid fragment search unit is used to search for the amino acid sequence information of the at least one existing polypeptide to determine at least one target existing polypeptide having a first repeating amino acid fragment.
[0044] The amino acid sequence cutting module is used to cut the amino acid sequence information of each target existing polypeptide into a set of amino acid fragments based on the first repeating amino acid fragment;
[0045] A repeating amino acid fragment extraction unit is used to extract the first repeating amino acid fragment from a set of amino acid fragments of a target existing polypeptide, to obtain a ladder and a set of remaining fragments corresponding to the target existing polypeptide.
[0046] The repeating amino acid fragment search unit is also used to search for the amino acid fragment set of other target existing peptides, the remaining fragment set, and the amino acid sequence information of other existing peptides to determine the second repeating amino acid fragment until the remaining fragment set of the at least one existing peptide does not contain a repeating amino acid fragment, thereby obtaining the ladder units corresponding to all repeating amino acid fragments.
[0047] Optionally, the ladder reconfiguration module includes:
[0048] An initialization unit is used to initialize a space structure for an amino acid sequence.
[0049] A ladder element selection unit is used to select multiple target ladder elements from the plurality of ladder elements;
[0050] The ladder filling unit is used to fill the multiple target ladders into the amino acid sequence empty structure according to the positions of the multiple target ladders in the corresponding existing polypeptide amino acid sequence information, until the amino acid sequence empty structure is filled, thereby obtaining the new amino acid sequence information.
[0051] Optionally, the ladder selection unit includes:
[0052] The probability calculation subunit is used to determine the selection probability of the multiple ladder elements based on their weights.
[0053] The ladder element selection subunit is used to select the multiple target ladder elements from the multiple ladder elements according to the selection probability of the multiple ladder elements.
[0054] Optionally, prior to the probability calculation subunit, the device further includes:
[0055] The weight calculation subunit is used to determine the weight of the ladder corresponding to each repeating amino acid fragment based on the number of times each repeating amino acid fragment is extracted.
[0056] Optionally, the ladder filling unit is specifically used to fill the first target ladder at the corresponding position in the amino acid sequence space structure according to the position of the first target ladder in the corresponding existing polypeptide amino acid sequence information; if the length of the second target ladder is greater than the length of the first target ladder, and the position of the second target ladder in the corresponding amino acid sequence information partially overlaps with the position of the first target ladder in the corresponding amino acid sequence information, the first target ladder in the amino acid sequence space structure is replaced with the second target ladder.
[0057] Optionally, the polypeptide generation module is specifically used to score the molecular docking of each new amino acid sequence with multiple structural information of the target protein, determine the molecular docking result of each new amino acid sequence, and determine the generation of the target new polypeptide from the at least one new amino acid sequence based on the molecular docking result of the at least one new amino acid sequence.
[0058] Optionally, the polypeptide generation module is specifically used to perform peptide chain structure analysis on the polypeptide corresponding to the at least one new amino acid sequence information to determine the new amino acid sequence information having the target peptide chain structure; and to determine the generation of the target new polypeptide from the new amino acid sequence information having the target peptide chain structure based on the molecular docking results of the new amino acid sequence information having the target peptide chain structure and the structural information of the target protein.
[0059] Optionally, the device further includes:
[0060] The information pool update module is used to add the new amino acid sequence information of the target new polypeptide to the preset polypeptide information pool.
[0061] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores program instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the program instructions to perform the steps of the polypeptide generation method as described in any of the first aspects.
[0062] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the polypeptide generation method as described in any of the first aspects.
[0063] The beneficial effects of this application are:
[0064] The peptide generation method and apparatus provided in this application extract amino acid sequence information from at least one existing peptide in a preset peptide information pool using a ladder extraction method. Based on the extracted ladders, at least one new amino acid sequence is reconstructed. The target new peptide is then determined from the new amino acid sequence based on the molecular docking structure between the new amino acid sequence and the structural information of the target protein. On the one hand, this method has a lower requirement for the number of existing peptides, allowing for the generation of new peptides without requiring a large number of peptides, thus avoiding reliance on extensive training data. On the other hand, reconstructing new amino acid sequence information based on ladders in the amino acid sequences of existing peptides provides a more intuitive understanding of amino acid fragments more effective for the target protein, making the process of generating new peptides more interpretable and enhancing the practical value of new peptides in peptide drug design. Attached Figure Description
[0065] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0066] Figure 1 A flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 1 ;
[0067] Figure 2 A flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 2 ;
[0068] Figure 3 This is a schematic diagram of a ladder element extraction process provided in an embodiment of this application;
[0069] Figure 4 This is a schematic diagram illustrating another step extraction process provided in an embodiment of this application;
[0070] Figure 5 A flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 3 ;
[0071] Figure 6 A flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 4 ;
[0072] Figure 7 A flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 5 ;
[0073] Figure 8 This is a schematic diagram illustrating the generation of new amino acid sequence information provided in an embodiment of this application;
[0074] Figure 9 A flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 6 ;
[0075] Figure 10 A flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 7 ;
[0076] Figure 11 This is a schematic diagram of the structure of the polypeptide generation device provided in the embodiments of this application;
[0077] Figure 12 A schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0078] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0079] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0080] Furthermore, the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Additionally, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0081] It should be noted that, where there is no conflict, the features in the embodiments of this application can be combined with each other.
[0082] In recent years, deep learning has achieved remarkable success in many bioinformatics tasks. Deep learning typically uses neural network models to learn the interaction between peptides and target proteins. By training neural network models, new peptides can be predicted.
[0083] However, using neural network models for peptide design presents the following technical challenges:
[0084] First, it is highly dependent on data. Neural network models require a large amount of training data to support effective learning. However, in the field of peptide design, the high cost and time consumption of generating experimental data make it very difficult to obtain a large amount of high-quality training data.
[0085] Second, the interpretability is poor. Since the neural network model is a black box, when the model predicts a new polypeptide sequence, researchers find it difficult to understand why the sequence was selected, or what characteristics of the sequence make it have good activity against the target protein.
[0086] To address the problems existing in the prior art, this application proposes a method and apparatus for generating peptides. The method involves extracting the amino acid sequence information of at least one existing peptide from a pre-defined peptide information pool using a ladder extraction process. Based on the extracted ladders, at least one new amino acid sequence is reconstructed. The target new peptide is then determined from the new amino acid sequence based on the molecular docking structure between the new amino acid sequence and the structural information of the target protein. On the one hand, this application has lower requirements for the number of existing peptides, allowing the generation of new peptides without requiring a large number of peptides, thus avoiding reliance on extensive training data. On the other hand, reconstructing new amino acid sequence information based on ladders in the amino acid sequences of existing peptides provides a more intuitive understanding of amino acid fragments more effective for the target protein, making the process of generating new peptides more interpretable and enhancing the practical value of new peptides in peptide drug design.
[0087] The specific implementation of the polypeptide generation method and apparatus provided in this application will be described below with reference to the embodiments.
[0088] Please refer to Figure 1 This is a flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 1 ,like Figure 1 As shown, the method may include:
[0089] S101: Obtain the pre-defined peptide information pool and the structural information of the target protein. The pre-defined peptide information pool includes the amino acid sequence information of at least one existing peptide.
[0090] In this embodiment, peptides are composed of amino acids as basic units, and each peptide is an amino acid sequence composed of multiple amino acids. Existing peptides are peptide structures generated through experiments. Based on the amino acid sequence information of at least one existing peptide, a preset peptide information pool pipPool is constructed.
[0091] The target protein is the protein to which the new polypeptide to be generated is targeted, and the structural information of the target protein is used to indicate the three-dimensional structure of the target protein.
[0092] In some embodiments, the number of amino acid sequence information of existing peptides included in the preset peptide information pool is not limited. That is, if there is only one amino acid sequence information of an existing peptide in the preset peptide information pool, it can be used to construct new amino acid sequence information; if there are multiple amino acid sequence information of existing peptides in the preset peptide information pool, they can also be used to construct new amino acid sequence information.
[0093] S102: Extract ladders from the amino acid sequence information of at least one existing polypeptide to obtain multiple ladders. Each ladder is a repeating amino acid fragment in the amino acid sequence information of at least one existing polypeptide, or a basic amino acid constituting the amino acid sequence information.
[0094] In this embodiment, the amino acid sequence information of at least one existing polypeptide is analyzed to identify repeating amino acid fragments in the amino acid sequence information of at least one existing polypeptide. Each repeating amino acid fragment corresponds to a ladder unit. Then, based on the basic amino acids that constitute the amino acid sequence information, corresponding ladder units are also generated to obtain multiple ladder units.
[0095] In some embodiments, if the preset peptide information pool includes the amino acid sequence information of an existing peptide, then the amino acid sequence information of the existing peptide is searched for repeating amino acid fragments to obtain repeating amino acid fragments in the amino acid sequence information of the existing peptide. Based on the repeating amino acid fragments in the amino acid sequence information of the existing peptide and the basic amino acids constituting the amino acid sequence information of the existing peptide, multiple ladder units are obtained.
[0096] In other embodiments, if the preset peptide information pool includes the amino acid sequence information of multiple existing peptides, then the amino acid sequence information of the multiple existing peptides is searched for repeating amino acid fragments to obtain repeating amino acid fragments in the amino acid sequence information of the multiple existing peptides. Based on the repeating amino acid fragments in the amino acid sequence information of the multiple existing peptides and the basic amino acids constituting the amino acid sequence information of the multiple existing peptides, multiple ladder units are obtained.
[0097] It should be noted that the repeating amino acid fragments in the amino acid sequence information of multiple existing peptides can be repeating amino acid fragments in the amino acid sequence information of one existing peptide, or repeating amino acid fragments in the amino acid sequence information of different existing peptides.
[0098] In some possible implementations, the search for repeating amino acid fragments can begin with the longest repeating amino acid fragment and then gradually search for shorter repeating amino acid fragments until all repeating amino acid fragments consisting of two amino acids have been searched.
[0099] S103: Based on multiple ladder units, reconstruct at least one new amino acid sequence information.
[0100] In this embodiment, based on the number of new amino acid sequence information to be generated and the length of the new amino acid sequence information, ladders are selected from multiple ladders for reconstruction to obtain new amino acid sequence information composed of some ladders from multiple ladders.
[0101] In some embodiments, based on the length of the new amino acid sequence information, a portion of the ladder corresponding to the repeating amino acid fragments can be selected first. The remaining amino acid length can be determined based on the sum of the lengths of the repeating amino acid fragments corresponding to the portion of the ladder and the length of the new amino acid sequence information. Then, based on the remaining amino acid length, a corresponding number of basic amino acids can be selected from the basic amino acids. The new amino acid sequence information is composed of the repeating amino acid fragments corresponding to the portion of the ladder and the corresponding number of basic amino acids.
[0102] S104: Based on the molecular docking results of at least one new amino acid sequence information and the structural information of the target protein, determine the generation of a new target polypeptide from at least one new amino acid sequence information.
[0103] In this embodiment, based on the information of each new amino acid sequence and the structural information of the target protein, molecular docking is performed on each new amino acid sequence and the target protein. The affinity of each new amino acid sequence to bind to the target protein is calculated, and the molecular docking result of each new amino acid sequence to the target protein is obtained. Based on the molecular docking result of at least one new amino acid sequence to the target protein, the new amino acid sequence with the best molecular docking result is selected from at least one new amino acid sequence information. Based on the new amino acid sequence with the best molecular docking result, the target new polypeptide is generated.
[0104] For example, molecular docking software can be used to perform molecular docking between a new amino acid sequence and a target protein, and the molecular docking results between the new amino acid sequence and the target protein can be obtained.
[0105] In some embodiments, the molecular docking results of at least one new amino acid sequence with the target protein can be screened according to a preset molecular docking threshold to determine the new amino acid sequence information that meets the preset molecular docking threshold.
[0106] In other embodiments, at least one new amino acid sequence can be sorted based on the molecular docking results of at least one new amino acid sequence with the target protein, and a preset number of new amino acid sequences with the best molecular docking results can be selected.
[0107] In one possible implementation, the peptide generation method provided in this application embodiment can be defined as a peptide hierarchical reconstructor (PepHiRe) algorithm. This PepHiRe algorithm can be executed on a server. Specifically, through the command-line interface provided by the server on the client, the command-line parameter genPeptides.py is input to reconstruct new amino acid sequence information, and the preset peptide information pool pipPool and the file pdb_files corresponding to the structural information of the target protein are submitted to generate at least one new amino acid sequence information. Then, through the command-line interface, the command-line parameter hdockScore.py for molecular docking is input to perform molecular docking and scoring on at least one new amino acid sequence and the target protein, and the molecular docking result is output.
[0108] The number of target new peptides to be generated can be controlled by the command-line parameter genPeptides.py. Based on the molecular docking results of molecular docking between the new amino acid sequence and the target protein, the corresponding number of new amino acid sequence information with the best docking results can be selected to generate the target new peptide, which makes it more flexible and adaptable when generating target new peptides based on a large peptide information pool.
[0109] The peptide generation method provided in the above embodiments extracts amino acid sequence information from at least one existing peptide in a preset peptide information pool using a ladder extraction method. Based on the extracted ladders, at least one new amino acid sequence is reconstructed. The target new peptide is then determined from the new amino acid sequence based on the molecular docking structure between the new amino acid sequence and the structural information of the target protein. On the one hand, this method has a lower requirement for the number of existing peptides, allowing for the generation of new peptides without requiring a large number of peptides, thus avoiding reliance on extensive training data. On the other hand, reconstructing new amino acid sequence information based on ladders in the amino acid sequences of existing peptides provides a more intuitive understanding of amino acid fragments more effective for the target protein, making the process of generating new peptides more interpretable and improving the practical value of new peptides in peptide drug design.
[0110] In one possible implementation, please refer to Figure 2 This is a flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 2 ,like Figure 2 As shown, the process of extracting the amino acid sequence information of at least one existing polypeptide in step S102 to obtain multiple steps may include:
[0111] S201: Search for the amino acid sequence information of at least one existing polypeptide to identify at least one target existing polypeptide having a first repeating amino acid fragment.
[0112] In this embodiment, the amino acid sequence information of at least one existing polypeptide is searched to identify a first repeating amino acid fragment in the amino acid sequence information of at least one existing polypeptide, and the existing polypeptide having the first repeating amino acid fragment is identified as the target existing polypeptide.
[0113] The first repeating amino acid fragment can be an amino acid fragment that repeats in the amino acid sequence information of an existing polypeptide, or it can be an amino acid fragment that repeats in the amino acid sequence information of multiple existing polypeptides. The first repeating amino acid fragment is the longest repeating amino acid fragment among all repeating amino acid fragments. If there are multiple longest repeating amino acid fragments, one of them can be selected as the first repeating amino acid fragment, and the other longest repeating amino acid fragments will be selected in the next search.
[0114] S202: Based on the first repeating amino acid fragment, cut the amino acid sequence information of each target existing polypeptide into a set of amino acid fragments.
[0115] In this embodiment, based on the position of the first repeating amino acid fragment in the amino acid sequence information of each target existing polypeptide, the amino acid sequence information of each existing polypeptide is segmented into a set of amino acid fragments. Specifically, based on the positions of the first and last amino acids of the first repeating amino acid fragment in the amino acid sequence information of each target existing polypeptide, the amino acid fragments before the first amino acid, the amino acid fragments after the last amino acid, and the first repeating amino acid fragment constitute an amino acid sequence fragment set.
[0116] S203: Extract the first repeating amino acid fragment from a set of amino acid fragments of a target existing polypeptide to obtain a ladder and a set of remaining fragments corresponding to the target existing polypeptide.
[0117] In this embodiment, for a set of amino acid fragments of multiple target existing peptides, a first repeating amino acid fragment is extracted from a set of amino acid fragments of a target existing peptide as a ladder. After extracting the first repeating amino acid fragment from the set of amino acid fragments, the set of amino acid fragments of the target existing peptide becomes the set of remaining fragments. The first repeating amino acid fragments in the amino acid fragment sets of other target existing peptides are still retained in the corresponding amino acid fragment sets.
[0118] S204: Search for the amino acid fragment set, the remaining fragment set, and the amino acid sequence information of other existing peptides to determine the second repeating amino acid fragment, until at least one existing peptide's remaining fragment set does not contain a repeating amino acid fragment, and obtain the ladder corresponding to all repeating amino acid fragments.
[0119] In this embodiment, the search for repeating amino acid fragments is performed again on the remaining fragment set corresponding to a target existing polypeptide, the amino acid fragment set of other target existing polypeptides, and the amino acid sequence information of other existing polypeptides to obtain a second repeating amino acid fragment. The amino acid sequence information containing the second repeating amino acid fragment is also fragmented to obtain the corresponding amino acid fragment set. The second repeating amino acid fragment is extracted from an amino acid fragment set as a ladder unit. Then, the above process is continued until there are no repeating amino acid fragments in the remaining fragment sets of all existing polypeptides. At this point, the search for repeating amino acid fragments is stopped, and the ladder unit corresponding to the repeating amino acid fragment is obtained.
[0120] The length of the second repeating amino acid fragment is equal to or less than the length of the first repeating amino acid fragment.
[0121] It is important to note that after segmenting amino acid sequence information from repeating amino acid fragments, the resulting fragments can only exist independently and cannot be spliced together. That is, the two segments can independently participate in the search for new repeating amino acid fragments, and cannot be spliced together to participate in the search for new repeating amino acid fragments. For example, taking amino acid fragments ADC and BCD as examples, you can only search for whether there are fragments repeating other amino acid fragments in ADC and BCD separately, and you cannot search for whether there are fragments repeating other amino acid fragments in ADCBCD. For instance, if there is an amino acid fragment CB, CB does not have repeating amino acid fragments with either ADC or BCD, not the other way around.
[0122] In some embodiments, when searching for repeating amino acid fragments in amino acid sequence information, in order to ensure the diversity of the generated new amino acid sequence information, it is necessary to limit the maximum length of the repeating amino acid fragments. This is to avoid the reconstructed new amino acid sequence information having low diversity due to the excessive length of the extracted repeating amino acid fragments, resulting in a small number of reconstructable new amino acid sequences and improving the diversity and variability of the reconstructed new amino acid sequence information.
[0123] The following examples illustrate the specific process of extracting ladder units from the amino acid sequence information of existing peptides, using the example of one existing peptide with amino acid sequence information of two existing peptides.
[0124] For example, if the preset peptide information pool contains the amino acid sequence information of an existing peptide, please refer to... Figure 3 This is a schematic diagram of a ladder element extraction process provided in an embodiment of this application, as shown below. Figure 3 As shown, the amino acid sequence information of the existing polypeptide is: CUCGACGACUAUCUCGACAAUGACU.
[0125] By searching for repeating amino acid fragments in the amino acid sequence information, the first repeating amino acid fragment was identified as CUCGAC. The amino acid sequence information was then cut based on the first repeating amino acid fragment CUCGAC, resulting in an amino acid fragment set {CUCGAC, GACUAU, CUCGAC, AAUGACU}. The first repeating amino acid fragment CUCGAC was extracted from this set as a stepper, leaving the remaining fragment set {GACUAU, CUCGAC, AAUGACU}. Continuing the search for repeating amino acid fragments in the remaining fragment set, the second repeating amino acid fragment was identified as GACU. Based on the second repeating amino acid fragment GACU, the remaining fragment set was further cut, resulting in a new amino acid fragment set {GACU, AU, CUCGAC, AAU, GACU}. The second repeating amino acid fragment GACU was extracted from this set as a stepper, leaving the remaining fragment set {AU, CUCGAC, AAU, GACU}. The search for repeating amino acid fragments in the remaining fragment set was then continued. The fragment search identifies the third repeating amino acid fragment as GAC. Based on GAC, the remaining fragment set is further cut to obtain a new amino acid fragment set {AU, CUC, GAC, AAU, GAC, U}. The third repeating amino acid fragment GAC is extracted from this remaining fragment set as a ladder unit, and the remaining fragment set is {AU, CUC, AAU, GAC, U}. The search for repeating amino acid fragments continues in the remaining fragment set, identifying the fourth repeating amino acid fragment as AU. Based on AU, the remaining fragment set is further cut to obtain a new amino acid fragment set {AU, CUC, A, AU, GAC, U}. The fourth repeating amino acid fragment AU is extracted from this remaining fragment set as a ladder unit, and the remaining fragment set is {CUC, A, AU, GAC, U}. There are no repeating amino acid fragments in this remaining fragment set. Therefore, the multiple ladder units corresponding to the amino acid sequence information of the existing polypeptide can be determined as {G, A(3), C(3), U(3) / / AU, GAC / / CUCGAC, GACU}.
[0126] For example, if the preset peptide information pool contains the amino acid sequence information of two existing peptides, please refer to... Figure 4 This is a schematic diagram of another ladder element extraction process provided in an embodiment of this application, as shown below. Figure 4 As shown, the amino acid sequence information of the two existing polypeptides are ABEABDCDCCD and CDABDCDEDCD, respectively.
[0127] By searching for repeating amino acid fragments in these two amino acid sequence information, the first repeating amino acid fragment was identified as ABDCD. Based on the first repeating amino acid fragment ABDCD, the two amino acid sequences were cut to obtain amino acid fragment sets {ABE, ABDCD, CCD} and {CD, ABDCD, EDCD}. The first repeating amino acid fragment ABDCD was extracted from one of these sets as a ladder unit, resulting in the remaining fragment set {ABE, CCD}. The search for repeating amino acid fragments continued in the remaining fragment set {ABE, CCD} and the amino acid fragment set {CD, ABDCD, EDCD} to determine... The second repeating amino acid fragment is DCD. Based on DCD, these two sets are divided to obtain sets {ABE, CCD} and {CD, AB, DCD, E, DCD}. Extracting DCD as a step element, the remaining fragment sets are {ABE, CCD} and {CD, AB, E, DCD}. Continuing the search for repeating amino acid fragments in the remaining fragment sets {ABE, CCD} and {CD, AB, E, DCD}, the third repeating amino acid fragment is determined to be AB. Based on AB, these two sets are divided to obtain sets {AB, E, CCD} and {CD, AB}. Let {E, C, CD} be the set of repeated amino acids. Extract the third repeating amino acid fragment AB as a step element, resulting in the remaining fragment sets {E, CCD} and {CD, AB, E, DCD}. Continue searching for repeating amino acid fragments in the remaining fragment sets {E, CCD} and {CD, AB, E, DCD} to determine the fourth repeating amino acid fragment as CD. Cut these two sets based on the fourth repeating amino acid fragment CD to obtain the sets {E, C, CD} and {CD, AB, E, D, CD}. Extract the fourth repeating amino acid fragment CD as a step element, resulting in the remaining fragment sets {E, C} and {CD, AB, E, D, CD}. Continue searching for repeating amino acid fragments in the remaining fragment sets {E, C, CD} and {CD, AB, E, D, CD}. The search for repeating amino acid fragments in {C} and {CD, AB, E, D, CD} is performed to determine the fifth repeating amino acid fragment as CD. Based on the fifth repeating amino acid fragment CD, these two sets are cut to obtain sets {E, C} and {CD, AB, E, D, CD}. The fifth repeating amino acid fragment CD is extracted as a ladder unit, and the remaining fragment sets are {E, C} and {AB, E, D, CD}. There are no repeating amino acid fragments in the remaining fragment sets. Therefore, it can be determined that the multiple ladder units corresponding to the amino acid sequence information of the existing polypeptide are {A, B, C(2), D(2), E(2) / / AB, CD(2) / / DCD / / ABDCD}.
[0128] The polypeptide generation method provided in the above embodiments searches for the amino acid sequence information of existing polypeptides to identify repeating amino acid fragments, uses these repeating amino acid fragments to cut the amino acid sequence information to obtain a set of amino acid fragments, extracts repeating amino acid fragments from the set of amino acid fragments as ladders, until there are no repeating amino acid fragments in all sets of amino acid fragments. This ensures that even with a small number of polypeptides, a large number of new amino acid sequences can be reconstructed using the ladders corresponding to the repeating amino acid fragments and the basic amino acids. It has low requirements for the number of existing polypeptides and can generate new polypeptides without requiring a large number of polypeptides.
[0129] In one possible implementation, please refer to Figure 5 This is a flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 3 ,like Figure 5 As shown, the process of reconstructing at least one new amino acid sequence information based on multiple ladder units in S103 above may include:
[0130] S301: Initialize an amino acid sequence space structure.
[0131] S302: Select multiple target ladder elements from multiple ladder elements.
[0132] S303: Based on the positions of multiple target ladders in the corresponding existing polypeptide amino acid sequence information, fill the amino acid sequence blank structure with multiple target ladders until the amino acid sequence blank structure is filled to obtain new amino acid sequence information.
[0133] In this embodiment, an amino acid sequence space structure is constructed, wherein the number of spaces in this amino acid sequence space structure is related to the length of the amino acid sequence information of at least one existing peptide in a preset peptide information pool. Specifically, the number of spaces in the amino acid sequence space structure can be determined based on the average length of the amino acid sequence information of at least one existing peptide.
[0134] Select a target ladder from multiple random ladders, fill the empty structure of the amino acid sequence with the target ladder, and then select the next target ladder to continue filling until the empty structure of the amino acid sequence is filled, thus obtaining the new amino acid sequence information.
[0135] The position of the target ladder in the amino acid sequence space structure is consistent with the position of the repeating amino acid fragment corresponding to the target ladder in the amino acid sequence information of the existing peptide. If the repeating amino acid fragment corresponding to the target ladder has multiple positions in the amino acid sequence information of the existing peptide, one of the positions is randomly selected as the position of the target ladder in the amino acid sequence space structure.
[0136] In some embodiments, after each target ladder is filled to its position in the amino acid sequence space structure, a target ladder with a length less than the number of spaces is selected from multiple ladders to fill the remaining spaces based on the number of remaining consecutive spaces in the amino acid sequence space structure. When the number of remaining consecutive spaces is 1, i.e. a single space, the ladder corresponding to the basic amino acid can be selected as the target ladder to fill the single space until the amino acid sequence space structure is filled, thus obtaining new amino acid sequence information.
[0137] In some embodiments, steps S301-S303 described above are performed multiple times according to the preset number of new amino acid sequence information to be generated, thereby obtaining multiple new amino acid sequence information.
[0138] The peptide generation method provided in the above embodiments selects target steps from multiple steps to fill the empty structure of the amino acid sequence until the empty structure of the amino acid sequence is filled, thereby obtaining new amino acid sequence information. This method of generating new amino acid sequence information can intuitively understand the amino acid fragments that make up the peptide, making the generated new peptide more interpretable and improving the practical value of the new peptide in the peptide drug design process.
[0139] In one possible implementation, please refer to Figure 6 This is a flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 4 ,like Figure 6 As shown, the process of selecting multiple target ladder elements from multiple ladder elements in S302 above may include:
[0140] S401: Determine the selection probability of multiple ladder elements based on their weights.
[0141] S402: Select multiple target ladder elements from multiple ladder elements based on the selection probability of multiple ladder elements.
[0142] In this embodiment, when extracting repeating amino acid fragments from the amino acid sequence information of at least one existing polypeptide, the weight of the ladder corresponding to the repeating amino acid fragment is determined. The weight of the ladder is positively correlated with the selection probability of the ladder. That is, the greater the weight of the ladder, the higher the selection probability of the ladder, and the easier it is for the ladder to be selected to fill the empty structure of the amino acid sequence.
[0143] In some embodiments, before determining the selection probability of multiple step elements based on their weights in step element S401, the method may further include:
[0144] The weight of each repeating amino acid fragment is determined based on the number of times it is extracted.
[0145] In this embodiment, the number of times a repeating amino acid fragment is extracted refers to the number of times the repeating amino acid fragment is extracted from the amino acid fragment set. The more times a repeating amino acid fragment is extracted, the greater its corresponding weight. For example, based on the process of extracting ladder elements from the amino acid sequence information ABEABDCDCCD and CDABDCDEDCD described above, it can be seen that the repeating amino acid fragment CD was extracted twice, so the weight of this repeating amino acid fragment is greater than the weight of other repeating amino acid fragments that were extracted once.
[0146] In some embodiments, the weight of the ladder corresponding to the repeating amino acid fragment can be determined based on the number of times the repeating amino acid fragment is extracted and the length of the repeating amino acid fragment.
[0147] Specifically, the weight of the ladder unit corresponding to the repeating amino acid fragment is positively correlated with both the number of times the repeating amino acid fragment is extracted and the length of the repeating amino acid fragment. The longer the repeating amino acid fragment is, the greater its weight. For repeating amino acid fragments of the same length, the more times the repeating amino acid fragment is extracted, the greater its weight. Therefore, the weight of the ladder unit corresponding to the repeating amino acid fragment can be calculated by combining the number of times the repeating amino acid fragment is extracted and the length of the repeating amino acid fragment.
[0148] In some embodiments, the number of extractions of a basic amino acid is based on the number of the same basic amino acid in the remaining fragment set without repeating amino acid fragments, such as... Figure 4 As shown, in the remaining fragment sets {E, C} and {AB, E, D, CD} without repeating amino acid fragments, the extraction times for basic amino acids A and B are 1, and the extraction times for C, D, and E are 2.
[0149] The polypeptide generation method provided in the above embodiments determines the weight of the ladder corresponding to the repeated amino acid fragment based on the number of times the repeated amino acid fragment is extracted, determines the selection probability of the ladder based on the weight of the ladder, and selects the target ladder from multiple ladders based on the selection probability of the ladder. This makes the selection of the target ladder more flexible and improves the diversity of the generated new amino acid sequence information.
[0150] In one possible implementation, please refer to Figure 7 This is a flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 5 ,like Figure 7 As shown, the process of filling multiple target ladders into the empty amino acid sequence structure according to their positions in the corresponding existing polypeptide amino acid sequence information in S303 may include:
[0151] S501: Based on the position of the first target ladder in the amino acid sequence information of the corresponding existing polypeptide, fill the corresponding position of the first target ladder in the empty structure of the amino acid sequence.
[0152] S502: If the length of the second target ladder is greater than the length of the first target ladder, and the position of the second target ladder in the corresponding amino acid sequence information partially overlaps with the position of the first target ladder in the corresponding amino acid sequence information, replace the first target ladder in the amino acid sequence space structure with the second target ladder.
[0153] In this embodiment, a first target ladder is randomly selected from multiple ladders. Based on the position of the repeating amino acid fragment corresponding to the first target ladder in the amino acid sequence information of the existing polypeptide, the first target ladder is filled into the corresponding position in the amino acid sequence space structure. If the repeating amino acid fragment corresponding to the first target ladder has multiple positions in the amino acid sequence information of the existing polypeptide, one of the positions is randomly selected as the position of the first target ladder in the amino acid sequence space structure.
[0154] When the position of the repeating amino acid fragment corresponding to the selected second target ladder in the amino acid sequence information of the existing peptide partially overlaps with the position of the repeating amino acid fragment corresponding to the first target ladder in the amino acid sequence information of the existing peptide, if the length of the second target ladder is greater than the length of the first target ladder, then the first target ladder is deleted from the amino acid sequence blank structure, and the second target ladder is filled into the corresponding position in the amino acid sequence blank structure according to the position of the repeating amino acid fragment corresponding to the second target ladder in the amino acid sequence information of the existing peptide.
[0155] In some embodiments, if the positions of the repeating amino acid fragments corresponding to the second target ladder and the first target ladder partially overlap in the existing amino acid sequence information of the polypeptide, and the length of the second target ladder is greater than that of the first target ladder, it is also necessary to determine that the sum of the number of spaces occupied by the first target ladder in the amino acid sequence space structure and the number of adjacent spaces is greater than or equal to the length of the repeating amino acid fragment corresponding to the second target ladder, so as to avoid the second target ladder covering part of the fragments of other target ladders. That is, it is necessary to determine that the filling position of the second target ladder in the amino acid sequence space structure is not other target ladders besides the first target ladder and spaces.
[0156] For example, please refer to Figure 8 This is a schematic diagram illustrating the generation of new amino acid sequence information provided in an embodiment of this application, such as... Figure 8 As shown, an amino acid sequence space structure with 11 spaces is initialized, starting from... Figure 4Among the multiple ladders shown {A, B, C(2), D(2), E(2) / / AB, CD(2) / / DCD / / ABDCD}, select the first target ladder DCD. Based on the position of DCD in the amino acid sequence information CDABDCDEDCD, fill DCD into spaces 5, 6, and 7 of the amino acid sequence space structure. Select the second target ladder CD. Based on the position of CD in the amino acid sequence information CDABDCDEDCD, fill CD into spaces 10 and 11 of the amino acid sequence space structure. Select the third target ladder CD. Based on the position of CD in the amino acid sequence information CDABDCDEDCD, fill CD into the amino acid sequence space structure. For spaces 1 and 2, select the fourth target ladder ABDCD. Based on the position of ABDCD in the amino acid sequence information ABEABDCDCCD, delete DCD from the amino acid sequence space structure and fill ABDCD into spaces 4, 5, 6, 7, and 8 in the amino acid sequence space structure. Select the fifth target ladder DCD. Based on the position of DCD in the amino acid sequence information CDABDCDEDCD, delete CD from the amino acid sequence space structure and fill DCD into spaces 9, 10, and 11 in the amino acid sequence space structure. Select the sixth target ladder E and fill E into the last space in the amino acid sequence space structure to obtain the new amino acid sequence information.
[0157] The polypeptide generation method provided in the above embodiments can ensure the integrity of the generated new amino acid sequence information by using a shorter target ladder in the space structure of a newly selected target ladder amino acid sequence with a longer length.
[0158] In one possible implementation, please refer to Figure 9 This is a flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 6 ,like Figure 9 As shown, the process of determining the generation of a target new polypeptide from at least one new amino acid sequence information based on the molecular docking results of at least one new amino acid sequence information and the structural information of the target protein in S104 may include:
[0159] S601: Molecular docking is scored based on the relationship between each new amino acid sequence and multiple structural information of the target protein to determine the molecular docking result for each new amino acid sequence.
[0160] S602: Based on the molecular docking results of at least one new amino acid sequence information, determine the generation of a target new polypeptide from at least one new amino acid sequence information.
[0161] In this embodiment, the three-dimensional structure of the target protein can have multiple conformations. For each new amino acid sequence, molecular docking is performed with each of the multiple conformations of the target protein. The affinity of each new amino acid sequence for each of the multiple conformations of the target protein is calculated, and the molecular docking results for each new amino acid sequence with the multiple conformations of the target protein are obtained. Based on the molecular docking results for each new amino acid sequence with the multiple conformations of the target protein, the average score of molecular docking for each new amino acid sequence with the target protein is determined. The corresponding number of new amino acid sequences with the lowest average score are selected from at least one new amino acid sequence. The target new polypeptide is generated based on the new amino acid sequence with the lowest average score. A lower score indicates a better docking effect between the amino acid sequence and the target protein.
[0162] The polypeptide generation method provided in the above embodiments can comprehensively evaluate the binding ability of new amino acid sequence information to target protein by calculating the molecular docking score of each new amino acid sequence information with multiple structural information of the target protein, select the new amino acid sequence information with the best binding effect to the target protein, and thus generate a new polypeptide with the best structural characteristics and the best affinity to the target protein.
[0163] In one possible implementation, please refer to Figure 10 This is a flowchart illustrating the polypeptide generation method provided in the embodiments of this application. Figure 7 ,like Figure 10 As shown, the process of determining the generation of a target new polypeptide from at least one new amino acid sequence information based on the molecular docking results of at least one new amino acid sequence information and the structural information of the target protein in S104 may include:
[0164] S701: Perform peptide chain structure analysis on the polypeptide corresponding to at least one new amino acid sequence information to determine the new amino acid sequence information with the target peptide chain structure.
[0165] S702: Based on the molecular docking results of the new amino acid sequence information with the target peptide chain structure and the structural information of the target protein, the target new polypeptide is generated from the new amino acid sequence information with the target peptide chain structure.
[0166] In this embodiment, peptide chain structure analysis is performed on the polypeptide formed by each new amino acid sequence. A new amino acid sequence with the target peptide chain structure is selected from at least one new amino acid sequence. Based on the new amino acid sequence of the target peptide chain structure and the structural information of the target protein, molecular docking is performed between the new amino acid sequence of the target peptide chain structure and the target protein. The affinity of the new amino acid sequence of the target peptide chain structure and the target protein is calculated, and the molecular docking result is obtained. Based on the molecular docking result of the new amino acid sequence of the target peptide chain structure and the target protein, the new amino acid sequence with the best molecular docking result is selected from the new amino acid sequence of the target peptide chain structure. The target new polypeptide is generated based on the new amino acid sequence with the best molecular docking result.
[0167] For example, the target peptide chain structure can be an α-Helix (α-helix) structure.
[0168] In one possible implementation, the peptide chain structure corresponding to at least one new amino acid sequence can be predicted by inputting the command-line parameter psipredHelix.py through the command-line interface, so that new amino acid sequence information with the target peptide chain structure can be selected.
[0169] In some embodiments, the PSIPRED tool can be used to perform peptide chain structure analysis on the peptides corresponding to the new amino acid sequence information, determine the proportion of the target peptide chain structure in the total peptide structure, select the new amino acid sequence information of the peptides with the target peptide chain structure as the dominant structure, and also select a preset number of peptides with the target peptide chain structure as the dominant structure to add to a special target peptide chain structure peptide pool.
[0170] The polypeptide generation method provided in the above embodiments can determine the new amino acid sequence information with the target peptide chain structure by performing peptide chain structure analysis on the polypeptide corresponding to the new amino acid sequence information, and generate a new target polypeptide with the target peptide chain structure based on the new amino acid sequence information with the target peptide chain structure and the molecular docking structure of the target protein, thereby achieving the generation of high-quality polypeptides with the target peptide chain structure.
[0171] In one possible implementation, the method may further include:
[0172] Add the new amino acid sequence information of the target new peptide to the preset peptide information pool.
[0173] In this embodiment, an iterative method for generating new peptides is provided. By adding the new amino acid sequence information corresponding to the target new peptide to a preset peptide information pool, the steps S101-S104 described above are executed again. Since the new amino acid sequence information corresponding to the target new peptide has a good molecular docking effect with the target protein, adding the new amino acid sequence information corresponding to the target new peptide to the preset peptide information pool again to participate in the generation of the target new peptide can extract repeating amino acid fragments that determine the analytical docking effect with the target protein from the new amino acid sequence information corresponding to the target new peptide as ladders. That is, the new amino acid sequence information corresponding to the target new peptide generated in each round can serve as the guiding information for the next round of newly generated target new peptides, thereby gradually reconstructing and generating the target new peptide that is optimal for the target protein.
[0174] In some embodiments, the number of iterations can be flexibly set to control the computation time for generating the final target new peptide and the result of the output target new peptide.
[0175] The peptide generation method provided in the above embodiments, by adding the new amino acid sequence information of the target new peptide to a preset peptide information pool for iterative generation of new peptides, can emphasize and retain amino acid fragments with high affinity to the target protein in each iteration, thereby providing a better foundation for the generation of new high-quality peptides and improving the molecular docking effect between the generated new peptides and the target protein.
[0176] In one possible implementation, the processes described above—including ladder extraction and reconstruction of existing peptide amino acid sequences to generate new amino acid sequences, molecular docking and scoring of the new amino acid sequences and target proteins, and peptide chain structure analysis of the peptides corresponding to the new amino acid sequences to determine new amino acid sequences with the target peptide chain structure—can be modularized. The modules corresponding to each of these processes can be called using the command-line parameter pepHiRe_main.py. Only a preset peptide information pool and the structural information of the target protein need to be input, along with the desired number of new peptides and the number of iterations. This allows the output of the optimal new peptide after multiple rounds of iterative screening, improving the efficiency and accuracy of new peptide generation and accelerating the drug development process.
[0177] Based on the above method embodiments, this application also provides a polypeptide generation apparatus. Please refer to... Figure 11 This is a schematic diagram of the structure of the polypeptide generation device provided in the embodiments of this application, as shown below. Figure 11 As shown, the device may include:
[0178] The information acquisition module 101 is used to acquire a preset peptide information pool and the structural information of the target protein. The preset peptide information pool includes: the amino acid sequence information of at least one existing peptide.
[0179] The ladder extraction module 102 is used to extract the amino acid sequence information of at least one existing polypeptide into multiple ladders. Each ladder is a repeating amino acid fragment in the amino acid sequence information of at least one existing polypeptide, or a basic amino acid that constitutes the amino acid sequence information.
[0180] The ladder reconstruction module 103 is used to reconstruct and generate at least one new amino acid sequence information based on multiple ladders;
[0181] The peptide generation module 104 is used to determine the generation of a target new peptide from at least one new amino acid sequence information based on the molecular docking results of at least one new amino acid sequence information and the structural information of the target protein.
[0182] Optionally, the ladder element extraction module 102 includes:
[0183] A repeating amino acid fragment search unit is used to search for the amino acid sequence information of at least one existing polypeptide to identify at least one target existing polypeptide having a first repeating amino acid fragment.
[0184] The amino acid sequence cutting module is used to cut the amino acid sequence information of each target existing polypeptide into a set of amino acid fragments based on the first repeating amino acid fragment;
[0185] The repeating amino acid fragment extraction unit is used to extract the first repeating amino acid fragment from a set of amino acid fragments of a target existing polypeptide, resulting in a ladder and a set of remaining fragments corresponding to the target existing polypeptide.
[0186] The repeating amino acid fragment search unit is also used to search for the amino acid fragment set, the remaining fragment set, and the amino acid sequence information of other existing peptides to determine the second repeating amino acid fragment until at least one existing peptide does not have a repeating amino acid fragment in its remaining fragment set, thus obtaining the ladder units corresponding to all repeating amino acid fragments.
[0187] Optionally, the ladder reconfiguration module 103 includes:
[0188] An initialization unit is used to initialize a space structure for an amino acid sequence.
[0189] A ladder element selection unit is used to select multiple target ladder elements from multiple ladder elements;
[0190] The ladder filling unit is used to fill multiple target ladders into the empty amino acid sequence structure according to the positions of multiple target ladders in the corresponding existing polypeptide amino acid sequence information, until the empty amino acid sequence structure is filled, and new amino acid sequence information is obtained.
[0191] Optionally, the ladder selection unit includes:
[0192] The probability calculation subunit is used to determine the selection probability of multiple ladder elements based on their weights.
[0193] The ladder selection subunit is used to select multiple target ladder elements from multiple ladder elements based on the selection probabilities of multiple ladder elements.
[0194] Optionally, prior to the probability calculation subunit, the device may further include:
[0195] The weight calculation subunit is used to determine the weight of the ladder corresponding to each repeating amino acid fragment based on the number of times each repeating amino acid fragment is extracted.
[0196] Optionally, the ladder filling unit is specifically used to fill the first target ladder at the corresponding position in the amino acid sequence space structure according to the position of the first target ladder in the corresponding existing polypeptide amino acid sequence information; if the length of the second target ladder is greater than the length of the first target ladder, and the position of the second target ladder in the corresponding amino acid sequence information partially overlaps with the position of the first target ladder in the corresponding amino acid sequence information, the first target ladder in the amino acid sequence space structure is replaced with the second target ladder.
[0197] Optionally, the peptide generation module 104 is specifically used to score the molecular docking of each new amino acid sequence information with multiple structural information of the target protein, determine the molecular docking result of each new amino acid sequence information, and determine the generation of a target new peptide from at least one new amino acid sequence information based on the molecular docking result of at least one new amino acid sequence information.
[0198] Optionally, the peptide generation module 104 is specifically used to perform peptide chain structure analysis on the peptide corresponding to at least one new amino acid sequence information to determine the new amino acid sequence information with the target peptide chain structure; and to determine the generation of a target new peptide from the new amino acid sequence information with the target peptide chain structure based on the molecular docking results of the new amino acid sequence information with the target peptide chain structure and the structural information of the target protein.
[0199] Optionally, the device may also include:
[0200] The information pool update module is used to add the new amino acid sequence information of the target new peptide to the preset peptide information pool.
[0201] The above-described device is used to execute the method provided in the foregoing embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.
[0202] These modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more microprocessors, or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).
[0203] Please refer to Figure 12 This is a schematic diagram of the electronic device provided in the embodiments of this application, such as... Figure 12 As shown, the electronic device 200 may include a processor 201, a storage medium 202, and a bus. The storage medium 202 stores program instructions executable by the processor 201. When the electronic device 200 is running, the processor 201 communicates with the storage medium 202 via the bus, and the processor 201 executes the program instructions to perform the above-described method embodiment. The specific implementation and technical effects are similar and will not be described in detail here.
[0204] Optionally, the present invention also provides a computer-readable storage medium storing a computer program, which is executed by a processor to perform the above-described method embodiments.
[0205] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0206] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0207] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0208] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0209] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for generating polypeptides, characterized in that, The method includes: Obtain a preset peptide information pool and structural information of a target protein, wherein the preset peptide information pool includes: amino acid sequence information of at least one existing peptide; The amino acid sequence information of the at least one existing polypeptide is subjected to ladder extraction to obtain multiple ladders, wherein the ladders are repeating amino acid fragments in the amino acid sequence information of the at least one existing polypeptide, or basic amino acids constituting the amino acid sequence information. Based on the multiple ladder units, at least one new amino acid sequence information is reconstructed and generated; Based on the molecular docking results of the at least one new amino acid sequence information and the structural information of the target protein, a new target polypeptide is determined from the at least one new amino acid sequence information; The stepwise extraction of the amino acid sequence information of the at least one existing polypeptide yields multiple steps, including: Search the amino acid sequence information of the at least one existing polypeptide to identify at least one target existing polypeptide having a first repeating amino acid fragment; Based on the first repeating amino acid fragment, the amino acid sequence information of each target existing polypeptide is cut into a set of amino acid fragments; Extract the first repeating amino acid fragment from a set of amino acid fragments of a target existing polypeptide to obtain a ladder and a set of remaining fragments corresponding to the target existing polypeptide; Search the amino acid fragment set of other target peptides, the remaining fragment set, and the amino acid sequence information of other existing peptides to determine the second repeating amino acid fragment, until the remaining fragment set of the at least one existing peptide does not contain repeating amino acid fragments, and obtain the ladder units corresponding to all repeating amino acid fragments. The step of reconstructing at least one new amino acid sequence based on the plurality of ladder units includes: Initialize a space structure for an amino acid sequence; Select multiple target ladder elements from the plurality of ladder elements; Based on the positions of the multiple target ladders in the corresponding existing polypeptide amino acid sequence information, the multiple target ladders are filled into the amino acid sequence empty structure until the amino acid sequence empty structure is filled, thereby obtaining the new amino acid sequence information.
2. The method as described in claim 1, characterized in that, The step of selecting multiple target ladder elements from the plurality of ladder elements includes: The selection probability of the multiple ladder elements is determined based on their weights. Based on the selection probability of the multiple ladder elements, the multiple target ladder elements are selected from the multiple ladder elements.
3. The method as described in claim 2, characterized in that, Before determining the selection probability of the plurality of ladder elements based on their weights, the method further includes: The weight of the ladder corresponding to each repeating amino acid fragment is determined based on the number of times each repeating amino acid fragment is extracted.
4. The method as described in claim 1, characterized in that, The step of filling the multiple target ladders into the empty amino acid sequence structure according to their positions in the corresponding existing polypeptide amino acid sequence information includes: Based on the position of the first target ladder in the amino acid sequence information of the corresponding existing polypeptide, the first target ladder is filled in at the corresponding position of the empty structure of the amino acid sequence; If the length of the second target ladder is greater than the length of the first target ladder, and the position of the second target ladder in the corresponding amino acid sequence information partially overlaps with the position of the first target ladder in the corresponding amino acid sequence information, then the first target ladder in the amino acid sequence space structure is replaced with the second target ladder.
5. The method as described in claim 1, characterized in that, The step of determining the generation of a target new polypeptide from the at least one new amino acid sequence information based on the molecular docking results of the at least one new amino acid sequence information and the structural information of the target protein includes: The molecular docking result of each new amino acid sequence is determined by scoring the molecular docking between each new amino acid sequence and multiple structural information of the target protein. Based on the molecular docking results of the at least one new amino acid sequence information, the target new polypeptide is determined to be generated from the at least one new amino acid sequence information.
6. The method as described in claim 1, characterized in that, The step of determining the generation of a target new polypeptide from the at least one new amino acid sequence information based on the molecular docking results of the at least one new amino acid sequence information and the structural information of the target protein includes: Peptide chain structure analysis is performed on the polypeptide corresponding to the at least one new amino acid sequence information to determine the new amino acid sequence information with the target peptide chain structure. Based on the molecular docking results of the new amino acid sequence information having the target peptide chain structure and the structural information of the target protein, the target new polypeptide is determined from the new amino acid sequence information having the target peptide chain structure.
7. The method as described in claim 1, characterized in that, The method further includes: The new amino acid sequence information of the target new polypeptide is added to the preset polypeptide information pool.
8. A polypeptide generation device, characterized in that, The device includes: An information acquisition module is used to acquire a preset peptide information pool and structural information of a target protein. The preset peptide information pool includes: amino acid sequence information of at least one existing peptide. A ladder extraction module is used to extract the amino acid sequence information of the at least one existing polypeptide into multiple ladders, wherein the ladders are repeating amino acid fragments in the amino acid sequence information of the at least one existing polypeptide, or basic amino acids constituting the amino acid sequence information. The ladder reconstruction module is used to reconstruct and generate at least one new amino acid sequence information based on the multiple ladders; A peptide generation module is used to determine the generation of a target new peptide from the at least one new amino acid sequence information based on the molecular docking results of the at least one new amino acid sequence information and the structural information of the target protein; The ladder extraction module is specifically used to: search the amino acid sequence information of the at least one existing polypeptide to determine at least one target existing polypeptide having a first repeating amino acid fragment; Based on the first repeating amino acid fragment, the amino acid sequence information of each target existing polypeptide is cut into a set of amino acid fragments; Extract the first repeating amino acid fragment from a set of amino acid fragments of a target existing polypeptide to obtain a ladder and a set of remaining fragments corresponding to the target existing polypeptide; Search the amino acid fragment set of other target peptides, the remaining fragment set, and the amino acid sequence information of other existing peptides to determine the second repeating amino acid fragment, until the remaining fragment set of the at least one existing peptide does not contain repeating amino acid fragments, and obtain the ladder units corresponding to all repeating amino acid fragments. The ladder reconstruction module is specifically used to: initialize an amino acid sequence space structure; Select multiple target ladder elements from the plurality of ladder elements; Based on the positions of the multiple target ladders in the corresponding existing polypeptide amino acid sequence information, the multiple target ladders are filled into the amino acid sequence empty structure until the amino acid sequence empty structure is filled, thereby obtaining the new amino acid sequence information.