Polypeptide generation method, device, electronic device, and storage medium
By setting up multi-level evaluation sub-modules and iterative optimization in the affinity assessment module, the problem of balancing accuracy and efficiency in peptide drug design is solved, and efficient peptide drug design is achieved.
Patent Information
- Application Number
- CN202310621946.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-30
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2043-05-30
AI Technical Summary
Existing peptide drug design methods cannot balance accuracy and efficiency, resulting in a low cost-effectiveness of peptide drug design processes.
The affinity assessment module is designed with at least two levels of assessment sub-modules. The assessment sub-modules with higher accuracy are assessed later in the order of assessment. The assessment sub-modules at different levels are used for screening with different levels of accuracy and efficiency. The candidate peptides are then optimized through iterative adjustments.
It improves the cost-effectiveness of peptide drug design process, balancing accuracy and efficiency, and enhances the accuracy and efficiency of peptide drug design through module collaboration and iterative optimization.
Smart Images

Figure CN116682489B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical field of biological computing, and more particularly to a polypeptide generation method and device, an electronic device, and a storage medium. BACKGROUND
[0002] Polypeptide drugs generally have low side effects and are one of the ideal choices for disease treatment. The purpose of polypeptide drug design is to develop drug molecules with unique functions, such as developing drugs with high selectivity and affinity for specific protein receptors, for the treatment of diseases or intervention in biological processes. In the related art, polypeptide drug design is usually achieved by experimental determination or calculation methods, but these methods have different precision and efficiency, and cannot balance precision and efficiency, resulting in low "performance-price ratio" of the polypeptide drug design process. SUMMARY
[0003] A polypeptide generation method, device, electronic device, and storage medium are provided.
[0004] According to a first aspect, a polypeptide generation method is provided, including: obtaining a protein target and a reference polypeptide corresponding to the protein target, wherein the protein target is a protein causing a lesion; generating a candidate polypeptide corresponding to the reference polypeptide; and inputting the protein target and the candidate polypeptide into an affinity evaluation module for screening to obtain a target polypeptide, wherein the affinity evaluation module includes at least two levels of evaluation sub-modules arranged in a preset evaluation order, and the later the evaluation sub-module in the evaluation order, the higher the accuracy of the evaluation sub-module.
[0005] According to a second aspect, a polypeptide generation device is provided, including: an obtaining module configured to obtain a protein target and a reference polypeptide corresponding to the protein target, wherein the protein target is a protein causing a lesion; a generating module configured to generate a candidate polypeptide corresponding to the reference polypeptide; and a screening module configured to input the protein target and the candidate polypeptide into an affinity evaluation module for screening to obtain a target polypeptide, wherein the affinity evaluation module includes at least two levels of evaluation sub-modules arranged in a preset evaluation order, and the later the evaluation sub-module in the evaluation order, the higher the accuracy of the evaluation sub-module.
[0006] According to a third aspect, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the polypeptide generation method of the first aspect of the present disclosure.
[0007] According to a fourth aspect, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform a method for generating a polypeptide according to a first aspect of this disclosure.
[0008] According to a fifth aspect, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the method for generating a polypeptide according to a first aspect of this disclosure.
[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0011] Figure 1 This is a schematic flowchart of a method for generating a polypeptide according to the first embodiment of this disclosure;
[0012] Figure 2 This is a schematic flowchart of a method for generating a polypeptide according to a second embodiment of the present disclosure;
[0013] Figure 3 This is a schematic flowchart of a method for generating a polypeptide according to a third embodiment of the present disclosure;
[0014] Figure 4 This is a schematic flowchart of a method for generating a polypeptide according to the fourth embodiment of this disclosure;
[0015] Figure 5 This is a schematic flowchart of a method for generating a polypeptide according to the fifth embodiment of this disclosure;
[0016] Figure 6 This is a schematic diagram of the overall process of a method for generating a polypeptide according to the sixth embodiment of this disclosure;
[0017] Figure 7 This is a block diagram of a polypeptide generation apparatus according to a first embodiment of the present disclosure;
[0018] Figure 8 This is a block diagram of a polypeptide generation apparatus according to a second embodiment of the present disclosure;
[0019] Figure 9 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation
[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0021] Artificial intelligence (AI) is a technical science that studies and develops theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence. Currently, AI technology has the advantages of high automation, high accuracy, and low cost, and has been widely applied.
[0022] Biocomputing refers to a new computing paradigm developed by utilizing the inherent information processing mechanisms of biological systems. Biocomputing research encompasses both devices and systems. Molecular devices are the basic units that utilize ordered systems composed of organic (or biological) materials at the molecular scale to provide information detection, processing, transmission, and storage through physicochemical processes at the molecular level. The structure and computational principles of biocomputing systems differ from traditional computing systems. Their structure is generally parallel and distributed, and information storage often combines short-term and long-term memory, achieved through learning. Biocomputing is a highly integrated discipline; the fusion of biology and computing will bring about significant breakthroughs and progress. Relying on biocomputing engines, massive amounts of biological data can be effectively utilized, transforming drug discovery from a "needle in a haystack" to a "guided search," ultimately benefiting human health.
[0023] The following describes, with reference to the accompanying drawings, a method, apparatus, electronic device, and storage medium for generating polypeptides according to embodiments of the present disclosure.
[0024] Figure 1 This is a schematic flowchart of a method for generating a polypeptide according to the first embodiment of this disclosure.
[0025] like Figure 1 As shown, the method for generating polypeptides according to embodiments of this disclosure may specifically include the following steps:
[0026] S101, obtain the protein target and the corresponding reference peptide.
[0027] Specifically, the execution entity of the polypeptide generation method in this disclosure embodiment can be the polypeptide generation apparatus provided in this disclosure embodiment. The polypeptide generation apparatus can be a hardware device with data processing capabilities and / or the necessary software to drive the hardware device. Optionally, the execution entity may include a workstation, server, computer, user terminal, and other devices. The user terminal includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, and vehicle terminals.
[0028] In this embodiment of the disclosure, the protein target, i.e. the protein that causes the lesion, is any protein that needs to be designed into a peptide drug and is pre-set in the peptide generation device.
[0029] As a feasible implementation, when a user has a need for peptide drug design, a peptide drug design request can be sent to the peptide generation device. The peptide drug design request may include a protein and a corresponding peptide. Correspondingly, the protein in the peptide drug design request can be used as the protein target in this example, and the peptide in the peptide drug design request can be used as the reference peptide corresponding to the protein target.
[0030] The reference peptide is the peptide used as the initial drug in peptide drug design for the protein target. Specifically, the reference peptide can be a pre-defined seed peptide or a randomly initialized peptide. For example, if the protein target is ABCD…EFG, the corresponding reference peptides could be AAAA…AAA, BBBB…BBB, CCCC…CCC, DDDD…DDD, etc. This disclosure uses evolutionary learning to gradually modify these reference peptides to obtain the target peptide. Evolutionary learning can be understood as a process of natural selection and survival of the fittest, allowing the reference peptides to gradually evolve towards a higher affinity, thus obtaining the target peptide. When the reference peptide is a pre-defined seed peptide, the target peptide is obtained by optimizing based on the pre-defined seed peptide. Setting the reference peptide as a randomly initialized peptide allows for entirely new peptide design.
[0031] In some exemplary embodiments, the reference peptide corresponding to the protein target can be obtained based on a pre-stored correspondence between the protein target and the reference peptide.
[0032] S102, Generate the corresponding candidate peptide based on the reference peptide.
[0033] In this embodiment of the disclosure, new peptides can be generated from the reference peptide obtained in step S101 using various methods to serve as candidate peptides. As an example, multiple new candidate peptides can be obtained by randomly selecting unit sites for mutation based on the reference peptide. Specifically, for a single reference peptide, one or more unit sites can be selected for mutation, and the same unit site can be mutated to one or more different values. One reference peptide can correspond to one or more candidate peptides; this embodiment of the disclosure does not impose excessive limitations on this. As another example, multiple new candidate peptides can also be obtained by selecting one or more pre-defined unit sites for mutation based on the reference peptide.
[0034] The number of candidate peptides generated can be a pre-defined third number N0.
[0035] S103, input the protein target and candidate peptide into the affinity assessment module for screening to obtain the target peptide. The affinity assessment module includes at least two levels of assessment sub-modules set in a preset assessment order, and the later the assessment sub-module is in the assessment order, the higher the accuracy.
[0036] In this embodiment of the disclosure, the target peptide is the optimized peptide output by the affinity assessment module. Based on these target peptides, drug experts select peptides with high potential to advance to the next stage of drug development.
[0037] An affinity assessment module can be pre-configured, containing at least two assessment sub-modules with varying levels of precision. Higher precision sub-modules are less efficient. The assessment order of these sub-modules is set according to their precision, with higher precision sub-modules being evaluated later. This approach ensures that the earlier, less precise but more efficient sub-modules screen candidate peptides first, while the later, more precise but less efficient sub-modules perform a more precise second screening. This balances precision and efficiency, improving the cost-effectiveness of the peptide drug design process.
[0038] In summary, the peptide generation method of this disclosure involves inputting a protein target and candidate peptides generated based on a reference peptide corresponding to the protein target into an affinity assessment module for screening to obtain the target peptide. The affinity assessment module includes at least two levels of assessment sub-modules arranged sequentially according to a preset assessment order, with later sub-modules exhibiting higher accuracy. This disclosure, by setting at least two assessment sub-modules with different accuracies within the affinity assessment module, and placing the more accurate sub-module later in the assessment order, allows the earlier, less accurate but more efficient assessment sub-modules to screen candidate peptides first, while the later, more accurate but less efficient sub-modules perform a more precise second screening of the selected candidate peptides. This balances accuracy and efficiency, improving the cost-effectiveness of the peptide drug design process.
[0039] Figure 2 This is a schematic flowchart of a method for generating a polypeptide according to a second embodiment of the present disclosure.
[0040] like Figure 2 As shown, in Figure 1 Based on the illustrated embodiments, the method for generating polypeptides according to the present disclosure may specifically include the following steps:
[0041] S201, obtain the protein target and the corresponding reference peptide.
[0042] S202, Generate corresponding candidate peptides based on reference peptides.
[0043] In this embodiment, steps S201-S202 are the same as steps S101-S102 in the above embodiments, and will not be described again here.
[0044] Step S103 in the above embodiment, "inputting protein targets and candidate peptides into the affinity assessment module for screening to obtain target peptides", may specifically include the following steps S203-S206.
[0045] S203 assesses the affinity between the protein target and the candidate peptide through the first-level evaluation submodule to obtain the first-granularity affinity score for the candidate peptide.
[0046] In this embodiment, the affinity assessment module may specifically include a first-level assessment submodule, a second-level assessment submodule, and a third-level assessment submodule arranged sequentially according to a preset assessment order. The accuracy of the first-level assessment submodule is less than that of the second-level assessment submodule, which is less than that of the third-level assessment submodule. The first-level assessment submodule may assess affinity based on molecular docking methods and / or deep learning methods, etc.
[0047] The protein target and candidate peptides are input into the first-level evaluation submodule. The first-level evaluation submodule evaluates the affinity between the protein target and the candidate peptides and outputs the affinity score corresponding to each candidate peptide, which is recorded as the first-granularity affinity score.
[0048] S204, through the second-level evaluation submodule, evaluates the affinity between the protein target and the first set number of candidate peptides with the highest first-size affinity score, so as to obtain the second-size affinity score corresponding to the first set number of candidate peptides.
[0049] In this embodiment of the disclosure, the second-level evaluation submodule can evaluate affinity based on molecular docking methods and / or deep learning methods, etc.
[0050] Based on the first granularity affinity score, select the first set number of candidate peptides with the highest scores, such as TopN1. Input the protein target and the selected TopN1 candidate peptides into the second-level evaluation submodule. The second-level evaluation submodule evaluates the affinity between the protein target and the TopN1 candidate peptides and outputs the affinity scores corresponding to the TopN1 candidate peptides, which are recorded as the second granularity affinity scores.
[0051] S205, through the third-level evaluation submodule, evaluates the affinity between the protein target and the second set number of candidate peptides with the highest second-size affinity score, so as to obtain the third-size affinity score corresponding to the second set number of candidate peptides.
[0052] In this embodiment of the disclosure, the third-level evaluation submodule can evaluate affinity based on molecular dynamics methods and / or laboratory experimental methods, such as wet laboratory experimental methods.
[0053] Based on the second granularity affinity score, the top N2 candidate peptides with the highest scores are selected. The protein target and the selected Top N2 candidate peptides are then input into the third-level evaluation submodule. The third-level evaluation submodule evaluates the affinity between the protein target and the Top N2 candidate peptides and outputs the affinity scores corresponding to the Top N2 candidate peptides, which are recorded as the third granularity affinity scores.
[0054] S206, the second set number of candidate peptides are identified as the target peptides.
[0055] In this embodiment of the disclosure, the TopN2 candidate peptides output by the third-level evaluation submodule are identified as the target peptides.
[0056] Furthermore, to further improve accuracy, an iterative approach can be adopted: the candidate peptides and affinity scores output by the high-accuracy evaluation submodule are input into the low-accuracy evaluation submodule, allowing the low-accuracy submodule to make targeted adjustments, thereby improving its screening performance. For example... Figure 3 As shown, after step S204 above, the following steps may be included:
[0057] S301, increment the first iteration count by one.
[0058] In this embodiment of the disclosure, the first iteration number is represented by x1. After the second-level evaluation submodule outputs the second granularity affinity score, the first iteration number is incremented by one, i.e., x1 = x1 + 1.
[0059] S302, if the first iteration count reaches the preset first iteration count threshold, then the step of evaluating the affinity between the protein target and the second set number of candidate peptides with the highest second particle size affinity score through the third-level evaluation submodule is executed, and the first iteration count is reset to zero.
[0060] In this embodiment of the disclosure, the number of times the data output by the second-level evaluation submodule needs to be input to the first-level evaluation submodule is preset, and the value of this number plus one is set as the first iteration number threshold. It is determined whether the first iteration number x1 is equal to or greater than the first iteration number threshold. If the first iteration number x1 is equal to or greater than the first iteration number threshold, then step S205 is continued, and the first iteration number is cleared to zero, i.e., x1 = 0.
[0061] S303, if the first iteration count does not reach the first iteration count threshold, then the first-level evaluation submodule is adjusted according to the second granularity affinity score corresponding to the first set number of candidate peptides, and the step of generating the corresponding candidate peptide based on the reference peptide is returned.
[0062] In this embodiment of the disclosure, if the first iteration number x1 is less than the first iteration number threshold, the second granularity affinity score corresponding to the first set number of candidate peptides output by the second-level evaluation submodule is input into the first-level evaluation submodule. The first set number of candidate peptides is used as input, and the second granularity affinity score is used as a label. The first-level evaluation submodule is adjusted accordingly, and the process returns to step S202.
[0063] It should be noted here that before returning to step S202 "generating corresponding candidate peptides based on reference peptides", the following steps may also be included: screening out a portion of peptides from a first set number of candidate peptides based on the second granularity affinity score, and updating the screened portion of peptides as reference peptides.
[0064] In this embodiment of the disclosure, based on the second granularity affinity scores corresponding to a first predetermined number of candidate peptides output by the second-level evaluation submodule, a plurality of peptides with the highest second granularity affinity scores are selected from the first predetermined number of candidate peptides as new reference peptides. Corresponding candidate peptides are then generated based on the new reference peptides.
[0065] like Figure 4 As shown, after step S205 above, the following steps may be included:
[0066] S401, increment the second iteration count by one.
[0067] In this embodiment of the disclosure, the second iteration number is represented by x2. After the third-level evaluation submodule outputs the third-granularity affinity score, the second iteration number is incremented by one, i.e., x2 = x2 + 1.
[0068] S402, if the second iteration count reaches the preset second iteration count threshold, then execute the step of determining the second set number of candidate peptides as the target peptides and reset the second iteration count to zero.
[0069] In this embodiment, the number of times the data output by the third-level evaluation submodule needs to be input to the first-level and second-level evaluation submodules is preset, and the value of this number plus one is set as the second iteration number threshold. It is determined whether the second iteration number x2 is equal to or greater than the second iteration number threshold. If the second iteration number x2 is equal to or greater than the second iteration number threshold, then step S206 is continued, and the second iteration number is cleared to zero, i.e., x2 = 0.
[0070] S403, if the second iteration count does not reach the second iteration count threshold, then adjust the first-level evaluation submodule and the second-level evaluation submodule respectively according to the third granularity affinity score corresponding to the second set number of candidate peptides, and return to the step of generating the corresponding candidate peptide based on the reference peptide.
[0071] In this embodiment of the disclosure, if the second iteration number x2 is less than the second iteration number threshold, the third granularity affinity score corresponding to the second set number of candidate peptides output by the third-level evaluation submodule is input into the first-level evaluation submodule and the second-level evaluation submodule. The second set number of candidate peptides is used as input, and the third granularity affinity score is used as a label. The first-level evaluation submodule and the second-level evaluation submodule are adjusted in a targeted manner, and the process returns to step S202.
[0072] It should be noted here that before returning to step S202 "generating corresponding candidate peptides based on reference peptides", the following steps may also be included: screening out a portion of peptides from a second set number of candidate peptides based on the third granularity affinity score, and updating the screened portion of peptides as reference peptides.
[0073] In this embodiment of the disclosure, based on the third granularity affinity scores corresponding to the second predetermined number of candidate peptides output by the third-level evaluation submodule, several peptides with the highest third granularity affinity scores are selected from the second predetermined number of candidate peptides as new reference peptides. Corresponding candidate peptides are then generated based on these new reference peptides.
[0074] To clearly illustrate the iterative process of the polypeptide generation method in the embodiments of this disclosure, the following is combined with... Figure 5 Provide a detailed description. For example... Figure 5As shown, the protein target (ABCD…EFG) and reference peptides (AAAA…AAA, BBBB…BBB, CCCC…CCC, DDDD…DDD) are input into the generation module. The generation module generates N0 candidate peptides based on the reference peptides. The first-level evaluation submodule assesses the affinity between the protein target and the N0 candidate peptides, outputting a first-granularity affinity score for each of the N0 candidate peptides. The top N1 candidate peptides with the highest first-granularity affinity scores are selected and input into the second-level evaluation submodule. The second-level evaluation submodule assesses the affinity between the protein target and the N1 candidate peptides, outputting a second-granularity affinity score for each of the N1 candidate peptides. The first iteration count is incremented by one. If the first iteration count does not reach the first iteration count threshold, the second granularity affinity scores of the N1 candidate peptides output by the second-level evaluation submodule are input to the generation module and the first-level evaluation submodule. The generation module selects some peptides from the N1 candidate peptides as new reference peptides based on the second granularity affinity scores, and generates N0 new candidate peptides based on the new reference peptides. The first-level evaluation submodule adjusts the scores based on the second granularity affinity scores of the N1 candidate peptides, evaluates the affinity between the protein target and the N0 new candidate peptides, and outputs the first granularity affinity scores of the N0 new candidate peptides. The top N1 candidate peptides with the highest first granularity affinity scores are selected and input to the second-level evaluation submodule. The second-level evaluation submodule evaluates the affinity between the protein target and the N1 candidate peptides and outputs the second granularity affinity scores of the N1 candidate peptides. Increment the first iteration count by one. If the first iteration count does not reach the first iteration count threshold, repeat the above process until the first iteration count reaches the first iteration count threshold. Select the Top N2 candidate peptides with the highest second granularity affinity scores and input them into the third-level evaluation submodule. The third-level evaluation submodule evaluates the affinity between the protein target and the N2 candidate peptides and outputs the third granularity affinity scores of the N2 candidate peptides.The second iteration count is incremented by one. If the second iteration count does not reach the second iteration count threshold, the second granularity affinity scores of the N1 candidate peptides output by the second-level evaluation submodule are input into the generation module, the first-level evaluation submodule, and the third-level evaluation submodule. The generation module selects some peptides from the N2 candidate peptides as new reference peptides based on the third granularity affinity scores, and generates N0 new candidate peptides based on the new reference peptides. The first-level evaluation submodule adjusts the scores based on the third granularity affinity scores of the N2 candidate peptides and evaluates the affinity between the protein target and the N0 new candidate peptides, outputting the first granularity affinity scores of the N0 new candidate peptides. The second-level evaluation submodule then evaluates the affinity between the N2 candidate peptides and the N0 new candidate peptides based on the third granularity affinity scores of the N2 candidate peptides. The third-level affinity score of the candidate peptides is adjusted, and the affinity between the protein target and N1 new candidate peptides is evaluated. The second-level affinity scores of the N1 new candidate peptides are output. Through multiple iterations, when the first iteration count reaches the first iteration count threshold, the top N2 candidate peptides with the highest second-level affinity scores are selected and input into the third-level evaluation submodule. The second iteration count is incremented by one. If the second iteration count does not reach the second iteration count threshold, the above process is repeated until the second iteration count reaches the second iteration count threshold. The N2 candidate peptides output by the third-level evaluation submodule are then identified as target peptides, allowing drug experts to select high-potential peptides to advance to the next stage of drug development.
[0075] In summary, the peptide generation method of this disclosure, by setting at least two evaluation sub-modules with different accuracies in the affinity assessment module, with the higher-accuracy evaluation sub-module being evaluated later, allows the earlier, less accurate but more efficient evaluation sub-module to screen candidate peptides first, while the later, more accurate but less efficient evaluation sub-module performs a more precise screening of the screened candidate peptides. This balances accuracy and efficiency, improving the cost-effectiveness of the peptide drug design process. By inputting the candidate peptides and affinity scores output by the more accurate evaluation sub-module into the less accurate evaluation sub-module, the less accurate evaluation sub-module can make targeted adjustments. Through iterative processing, intermediate data during peptide optimization is utilized. The synergy between modules improves the screening effect of the less accurate evaluation sub-module, thereby further improving the accuracy of the peptide drug design process. The two iterative processes with different iteration ranges result in both the less accurate but more efficient evaluation sub-module and the more accurate but less efficient evaluation sub-module being evaluated more times, further balancing accuracy and efficiency and improving the cost-effectiveness of the peptide drug design process.
[0076] To clearly illustrate the polypeptide generation method of the embodiments of this disclosure, the following is now combined with... Figure 6 Provide a detailed description. Figure 6 This is a schematic diagram of the overall process of a polypeptide generation method according to embodiments of this disclosure.Figure 6 As shown, the method for generating polypeptides according to embodiments of this disclosure includes:
[0077] S601, obtain the protein target and the corresponding reference peptide.
[0078] S602, Generate the corresponding candidate peptide based on the reference peptide.
[0079] S603 assesses the affinity between the protein target and the candidate peptide through the first-level evaluation submodule to obtain the first-granularity affinity score for the candidate peptide.
[0080] S604, through the second-level evaluation submodule, evaluates the affinity between the protein target and the first set number of candidate peptides with the highest first-size affinity score, so as to obtain the second-size affinity score corresponding to the first set number of candidate peptides.
[0081] S605 increments the first iteration count by one.
[0082] S606, determine whether the number of the first iteration has reached the preset threshold for the number of the first iteration. If not, proceed to steps S607-S608. If yes, proceed to step S609.
[0083] S607, the first-level evaluation submodule is adjusted according to the second granularity affinity score corresponding to the first set number of candidate peptides.
[0084] S608, based on the second particle size affinity score, select a subset of peptides from the first predetermined number of candidate peptides, and update the selected subset of peptides as reference peptides. Return to step S202.
[0085] S609, through the third-level evaluation submodule, evaluates the affinity between the protein target and the second set number of candidate peptides with the highest second-level particle size affinity score, so as to obtain the third particle size affinity score corresponding to the second set number of candidate peptides, and resets the first iteration number to zero.
[0086] S610, increment the second iteration count by one.
[0087] S611, determine whether the number of the second iteration has reached the preset threshold for the number of the second iteration. If not, proceed to steps S612-S613. If yes, proceed to step S614.
[0088] S612, the first-level evaluation submodule and the second-level evaluation submodule are adjusted according to the third particle size affinity score corresponding to the second set number of candidate peptides.
[0089] S613, based on the third particle size affinity score, select a subset of peptides from the second predetermined number of candidate peptides, and update the selected subset of peptides as reference peptides. Return to step S202.
[0090] S614, the second set number of candidate peptides are identified as the target peptides.
[0091] Figure 7 This is a block diagram of a polypeptide generation apparatus according to a first embodiment of the present disclosure.
[0092] like Figure 7 As shown, the polypeptide generation apparatus 700 of this embodiment includes: an acquisition module 701, a generation module 702, and a screening module 703. Wherein:
[0093] The acquisition module 701 is used to acquire protein targets and corresponding reference peptides, wherein the protein targets are proteins that cause lesions.
[0094] The generation module 702 is used to generate corresponding candidate peptides based on the reference peptide.
[0095] The screening module 703 is used to input protein targets and candidate peptides into the affinity assessment module for screening to obtain target peptides. The affinity assessment module includes at least two levels of assessment sub-modules set in a preset assessment order, and the accuracy of the assessment sub-modules in the later assessment order is higher.
[0096] It should be noted that the above explanation of the polypeptide generation method embodiments also applies to the polypeptide generation apparatus of the present disclosure embodiments, and the specific process will not be repeated here.
[0097] In summary, the peptide generation apparatus of this disclosure inputs a protein target and candidate peptides generated based on a reference peptide corresponding to the protein target into an affinity assessment module for screening to obtain a target peptide. The affinity assessment module includes at least two levels of assessment sub-modules arranged sequentially according to a preset assessment order, with higher accuracy for sub-modules located later in the assessment order. This disclosure, by setting at least two assessment sub-modules with different accuracies in the affinity assessment module, and placing the more accurate sub-module later in the assessment order, allows the earlier, less accurate but more efficient assessment sub-modules to screen candidate peptides first, while the later, more accurate but less efficient assessment sub-modules perform a more precise second screening of the screened candidate peptides. This balances accuracy and efficiency, improving the cost-effectiveness of the peptide drug design process.
[0098] Figure 8 This is a block diagram of a polypeptide generation apparatus according to a second embodiment of the present disclosure.
[0099] likeFigure 8 As shown, the polypeptide generation apparatus 800 of this embodiment includes: an acquisition module 801, a generation module 802, and a screening module 803.
[0100] The acquisition module 801 has the same structure and function as the acquisition module 701 in the previous embodiment, the generation module 802 has the same structure and function as the generation module 702 in the previous embodiment, and the filtering module 803 has the same structure and function as the filtering module 703 in the previous embodiment.
[0101] Furthermore, the affinity assessment module includes a first-level assessment submodule, a second-level assessment submodule, and a third-level assessment submodule arranged in a preset assessment order. The screening module 803 is further used to: assess the affinity between the protein target and the candidate peptide through the first-level assessment submodule to obtain a first-granularity affinity score corresponding to the candidate peptide; assess the affinity between the protein target and a first predetermined number of candidate peptides with the highest first-granularity affinity score through the second-level assessment submodule to obtain a second-granularity affinity score corresponding to the first predetermined number of candidate peptides; assess the affinity between the protein target and a second predetermined number of candidate peptides with the highest second-granularity affinity score through the third-level assessment submodule to obtain a third-granularity affinity score corresponding to the second predetermined number of candidate peptides; and identify the second predetermined number of candidate peptides as the target peptide.
[0102] Furthermore, the screening module 803 is further configured to: increment the first iteration count by one before evaluating the affinity between the protein target and the second set number of candidate peptides with the highest second particle size affinity score through the third-level evaluation submodule; and when the first iteration count reaches a preset first iteration count threshold, perform the step of evaluating the affinity between the protein target and the second set number of candidate peptides with the highest second particle size affinity score through the third-level evaluation submodule, and reset the first iteration count to zero.
[0103] Furthermore, the screening module 803 is further used to: if the first iteration number does not reach the first iteration number threshold, adjust the first-level evaluation submodule according to the second granularity affinity score corresponding to the first set number of candidate peptides, and trigger the generation module 802 to execute the step of generating the corresponding candidate peptides according to the reference peptides.
[0104] Furthermore, the screening module 803 is further configured to: before the generation module 802 executes the step of generating a corresponding candidate peptide based on the reference peptide, the generation module 802 screens out a portion of the candidate peptides based on the second granularity affinity score, and updates the screened portion of the peptides to the reference peptide.
[0105] Furthermore, the screening module 803 is further configured to: increment the second iteration number by one before determining the second set number of candidate peptides as the target peptide; and when the second iteration number reaches a preset second iteration number threshold, execute the step of determining the second set number of candidate peptides as the target peptide and reset the second iteration number to zero.
[0106] Furthermore, the screening module 803 is further used to: if the second iteration number does not reach the second iteration number threshold, adjust the first-level evaluation submodule and the second-level evaluation submodule respectively according to the third granularity affinity score corresponding to the second set number of candidate peptides, and trigger the generation module 802 to execute the step of generating the corresponding candidate peptide according to the reference peptide.
[0107] Furthermore, the screening module 803 is further configured to: before the generation module 802 executes the step of generating the corresponding candidate peptide based on the reference peptide, the generation module 802 screens out a portion of the peptides from the second set number of candidate peptides based on the third granularity affinity score, and updates the screened portion of the peptides to the reference peptide.
[0108] Furthermore, the first-level evaluation submodule evaluates affinity based on molecular docking methods and / or deep learning methods.
[0109] Furthermore, the second-level evaluation submodule evaluates affinity based on molecular docking methods and / or deep learning methods.
[0110] Furthermore, the third-level evaluation submodule assesses affinity based on molecular dynamics methods and / or laboratory experimental methods.
[0111] Furthermore, the generation module 802 includes: a generation unit 8021, used to randomly select unit sites based on a reference peptide for mutation to obtain candidate peptides.
[0112] Furthermore, the reference peptide is a pre-defined seed peptide or a randomly initialized peptide.
[0113] Furthermore, the generation module 802 is further configured to: generate a third set number of candidate peptides based on the reference peptide.
[0114] It should be noted that the above explanation of the polypeptide generation method embodiments also applies to the polypeptide generation apparatus of the present disclosure embodiments, and the specific process will not be repeated here.
[0115] In summary, the peptide generation apparatus of this disclosure, by setting at least two evaluation sub-modules with different accuracies in the affinity assessment module, with the higher-accuracy evaluation sub-module being evaluated later in the order, allows the earlier, less accurate but more efficient evaluation sub-module to screen candidate peptides first, while the later, more accurate but less efficient evaluation sub-module performs a more precise screening of the screened candidate peptides. This balances accuracy and efficiency, improving the cost-effectiveness of the peptide drug design process. By inputting the candidate peptides and affinity scores output by the more accurate evaluation sub-module into the less accurate evaluation sub-module, the less accurate evaluation sub-module can make targeted adjustments. Through iterative processing, intermediate data in the peptide optimization process is utilized. Through the collaboration between modules, the screening effect of the less accurate evaluation sub-module is improved, thereby further improving the accuracy of the peptide drug design process. By using two iterative processes with different iteration ranges, both the less accurate but more efficient evaluation sub-module and the more accurate but less efficient evaluation sub-module perform more evaluations, further balancing accuracy and efficiency and improving the cost-effectiveness of the peptide drug design process.
[0116] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0117] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0118] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0119] like Figure 9As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. The RAM 903 may also store various programs and data required for the operation of the electronic device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0120] Multiple components in electronic device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of displays, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows electronic device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0121] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as... Figures 1 to 6 The method for generating the polypeptide is illustrated. For example, in some embodiments, the method for generating the polypeptide may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by computing unit 901, one or more steps of the semantic parsing method described above may be performed. Alternatively, in other embodiments, computing unit 901 may be configured to perform the method for generating the polypeptide by any other suitable means (e.g., by means of firmware).
[0122] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0123] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0124] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0125] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0126] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0127] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the management difficulties and weak business scalability inherent in traditional physical hosts and VPS (Virtual Private Server) services. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0128] According to embodiments of this disclosure, this disclosure also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the steps of the polypeptide generation method shown in the above embodiments of this disclosure.
[0129] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0130] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for generating a polypeptide, comprising: obtaining a protein target and a reference polypeptide corresponding to the protein target, wherein the protein target is a protein causing a disease; generating a candidate polypeptide corresponding to the reference polypeptide; and inputting the protein target and the candidate polypeptide into an affinity evaluation module for screening to obtain a target polypeptide, wherein the affinity evaluation module comprises a first evaluation submodule, a second evaluation submodule and a third evaluation submodule arranged in a preset evaluation order, and the later the evaluation submodule in the evaluation order, the higher the accuracy of the evaluation submodule, the affinity between the protein target and the candidate polypeptide is evaluated by the first evaluation submodule to obtain a first granularity affinity score of the candidate polypeptide, the affinity between the protein target and a first set number of candidate polypeptides with the highest first granularity affinity scores is evaluated by the second evaluation submodule to obtain a second granularity affinity score of the first set number of candidate polypeptides, a first iteration number is incremented by one, if the first iteration number reaches a preset first iteration number threshold, the affinity between the protein target and a second set number of candidate polypeptides with the highest second granularity affinity scores is evaluated by the third evaluation submodule to obtain a third granularity affinity score of the second set number of candidate polypeptides, if the first iteration number threshold is not reached, the first evaluation submodule is adjusted according to the second granularity affinity scores of the first set number of candidate polypeptides, the candidate polypeptide corresponding to the reference polypeptide is regenerated, and the first iteration number is cleared, a second iteration number is incremented by one, if the second iteration number reaches a preset second iteration number threshold, the second set number of candidate polypeptides are determined as the target polypeptide, and the second iteration number is cleared, if the second iteration number threshold is not reached, the first evaluation submodule and the second evaluation submodule are adjusted according to the third granularity affinity scores of the second set number of candidate polypeptides, respectively, and the candidate polypeptide corresponding to the reference polypeptide is regenerated. 2.The method of claim 1, further comprising, before the step of regenerating the candidate polypeptide corresponding to the reference polypeptide: selecting part of the candidate polypeptides from the first set number of candidate polypeptides according to the second granularity affinity scores, and updating the selected part of the candidate polypeptides as the reference polypeptide. 3.The method of claim 1, further comprising, before the step of regenerating the candidate polypeptide corresponding to the reference polypeptide: selecting part of the candidate polypeptides from the second set number of candidate polypeptides according to the third granularity affinity scores, and updating the selected part of the candidate polypeptides as the reference polypeptide.
4. The generation method of claim 1, wherein, The first evaluation submodule evaluates the affinity based on a molecular docking method and / or a deep learning method.
5. The generation method of claim 1, wherein, The second evaluation submodule evaluates the affinity based on a molecular docking method and / or a deep learning method.
6. The generation method of claim 1, wherein, The third evaluation submodule evaluates the affinity based on a molecular dynamics method and / or a laboratory experiment method.
7. The generation method of claim 1, wherein, The generating the corresponding candidate polypeptide according to the reference polypeptide comprises: Randomly selecting a unit point for mutation based on the reference polypeptide to obtain the candidate polypeptide.
8. The generation method of claim 1, wherein, The reference polypeptide is a preset seed polypeptide or a randomly initialized polypeptide.
9. The generation method of claim 1, wherein, The generating the corresponding candidate polypeptide according to the reference polypeptide comprises: Generating a third set number of the candidate polypeptides according to the reference polypeptide.
10. A polypeptide generating device, comprising: an acquisition module configured to acquire a protein target and a reference polypeptide corresponding to the protein target, wherein the protein target is a protein causing a pathological change; a generating module configured to generate a corresponding candidate polypeptide according to the reference polypeptide; and a screening module configured to input the protein target and the candidate polypeptide into an affinity evaluation module for screening to obtain a target polypeptide, wherein the affinity evaluation module comprises, in a preset evaluation order, a first-level evaluation submodule, a second-level evaluation submodule and a third-level evaluation submodule, and the later the evaluation submodule in the evaluation order, the higher the accuracy of the evaluation submodule, the first-level evaluation submodule is used to evaluate the affinity between the protein target and the candidate polypeptide to obtain a first-granularity affinity score corresponding to the candidate polypeptide, the second-level evaluation submodule is used to evaluate the affinity between the protein target and a first set number of candidate polypeptides with the highest first-granularity affinity scores to obtain second-granularity affinity scores corresponding to the first set number of candidate polypeptides, a first iteration number is incremented by one, if the first iteration number reaches a preset first iteration number threshold, the third-level evaluation submodule is used to evaluate the affinity between the protein target and a second set number of candidate polypeptides with the highest second-granularity affinity scores to obtain third-granularity affinity scores corresponding to the second set number of candidate polypeptides, if the first iteration number threshold is not reached, the first-level evaluation submodule is adjusted according to the second-granularity affinity scores corresponding to the first set number of candidate polypeptides, the corresponding candidate polypeptides are regenerated according to the reference polypeptide, and the first iteration number is cleared, a second iteration number is incremented by one, if the second iteration number reaches a preset second iteration number threshold, the second set number of candidate polypeptides are determined as the target polypeptides, and the second iteration number is cleared, if the second iteration number threshold is not reached, the first-level evaluation submodule and the second-level evaluation submodule are adjusted according to the third-granularity affinity scores corresponding to the second set number of candidate polypeptides, respectively, and the corresponding candidate polypeptides are regenerated according to the reference polypeptide.
11. The generating device of claim 10, wherein, The screening module is further configured to: before triggering the generating module to execute the regenerating the corresponding candidate polypeptide according to the reference polypeptide, triggering the generating module to screen part of the candidate polypeptides from the first set number of candidate polypeptides according to the second-granularity affinity scores, and updating the screened part of the candidate polypeptides as the reference polypeptide.
12. The generating device of claim 10, wherein, The screening module is further configured to: Before triggering the generating module to perform the regenerating of the corresponding candidate polypeptide according to the reference polypeptide, the generating module is triggered to filter out part of polypeptides from the second set number of candidate polypeptides according to the third granularity affinity score, and update the filtered part of polypeptides as the reference polypeptide.
13. The generating device of claim 10, wherein, The first level evaluation sub-module evaluates the affinity based on a molecular docking method and / or a deep learning method.
14. The generating device of claim 10, wherein, The second level evaluation sub-module evaluates the affinity based on a molecular docking method and / or a deep learning method.
15. The generating device of claim 10, wherein, The third level evaluation sub-module evaluates the affinity based on a molecular dynamics method and / or a laboratory experiment method.
16. The generating device of claim 10, wherein, The generating of the corresponding candidate polypeptide according to the reference polypeptide comprises: Randomly selecting a unit point for mutation based on the reference polypeptide to obtain the candidate polypeptide.
17. The generating device of claim 10, wherein, The reference polypeptide is a preset seed polypeptide or a randomly initialized polypeptide.
18. The generating device of claim 10, wherein, The generating module is further used for: Generating a third set number of candidate polypeptides according to the reference polypeptide.
19. An electronic device comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9.
20. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method according to any one of claims 1-9.
21. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1-9.
Citation Information
Patent Citations
Targeted polypeptide design method and system based on reinforcement learning and molecular simulation
CN115985384A