Methods for the production of unnatural amino acids and their use
By bonding characteristic amino acid groups to target molecular sites to generate non-natural amino acid molecules, the problems existing in the prior art are solved, a method for generating target amino acids is realized, the types of non-natural amino acids are enriched, and the bioactivity, stability and specificity of peptide drugs are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING JINGTAI TECH CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies make it difficult to generate a rich variety of non-natural amino acids, resulting in a lack of novelty and diversity in the design of peptide drugs containing non-natural amino acids, which affects the drug's bioactivity, stability, and specificity.
Candidate amino acid molecules are generated by bonding amino acid characteristic groups to target molecules with a molecular weight less than or equal to a preset molecular weight, and target amino acid molecules are screened out using multiple screening rules, thus enriching the types and diversity of non-natural amino acids.
This approach enriches the non-natural amino acid pool with the target amino acid, enhancing the bioactivity, stability, and specificity of the peptide drug, and has clear drug-like significance.
Smart Images

Figure CN121459997B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of amino acid generation technology, and in particular to a method for generating non-natural amino acids and its application. Background Technology
[0002] Non-canonical amino acids (NCAAs) are playing an increasingly important role in peptide drug design. Compared to natural amino acids, NCAAs offer greater chemical diversity and functionality, significantly enhancing the bioactivity, stability, and specificity of peptide drugs.
[0003] In related technologies, most NCAA-containing peptide drugs are obtained through high-throughput screening techniques or artificial modification of natural amino acid peptide drugs. A few studies have reported methods for designing NCAA-containing peptide drugs, which are based on modifying natural amino acids to obtain NCAAs. This leads to a significant reduction in the diversity of available NCAAs, making it difficult to provide a large number of novel or diverse NCAA molecules to support the screening of lead compounds. Summary of the Invention
[0004] To address or partially address the problems existing in related technologies, this application provides a method for generating non-natural amino acids and its application, which can enrich the novelty and diversity of NCCA types and provide more potential options for the research and development of peptide drugs containing NCCA.
[0005] The first aspect of this application provides a method for generating non-natural amino acids, comprising:
[0006] Step S1: Determine at least one candidate target site in the target molecule, wherein the target molecule is a complete molecule or molecular fragment with a molecular weight less than or equal to a first preset molecular weight;
[0007] Step S2: When there is only one candidate target site, the candidate target site is the target site. An amino acid characteristic group is bonded to the target site to generate the corresponding candidate amino acid molecule. When there are multiple candidate target sites, at least one target site is determined from the multiple candidate target sites, and an amino acid characteristic group is bonded to the at least one target site to generate the corresponding at least one candidate amino acid molecule.
[0008] The amino acid characteristic group is the group shown in Formula I. “ "The bond used to identify the characteristic group of the amino acid and the target site;
[0009] Step S3: Screen the candidate amino acid molecules obtained in step S2 according to the preset amino acid screening rules to obtain the target amino acid molecule, wherein the target amino acid molecule is a non-natural amino acid.
[0010] A second aspect of this application provides an apparatus for generating non-natural amino acids, comprising:
[0011] A site determination module is used to determine at least one candidate target site in a target molecule, wherein the target molecule is a complete molecule or molecular fragment with a molecular weight less than or equal to a first preset molecular weight.
[0012] An amino acid generation module is used to determine at least one target site when there is only one candidate target site, which is the target site, and to bond an amino acid characteristic group to the target site to generate a corresponding candidate amino acid molecule; when there are multiple candidate target sites, at least one target site is determined from the multiple candidate target sites, and an amino acid characteristic group is bonded to the at least one target site to generate at least one corresponding candidate amino acid molecule.
[0013] The amino acid characteristic group is the group shown in Formula I. “ "The bond used to identify the characteristic group of the amino acid and the target site;
[0014] An amino acid screening module is used to screen the candidate amino acid molecules according to a preset amino acid screening rule to obtain a target amino acid molecule, wherein the target amino acid molecule is a non-natural amino acid.
[0015] A third aspect of this application provides an electronic device, comprising:
[0016] Processor; and
[0017] A memory having executable code stored thereon, which, when executed by the processor, causes the processor to perform the method described in the first aspect above.
[0018] A fourth aspect of this application provides a computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the method described in the first aspect above.
[0019] The fifth aspect of this application provides a computer program product comprising computer instructions that, when executed by a processor, implement the method described in the first aspect above.
[0020] The technical solution provided in this application may include the following beneficial results:
[0021] The method for generating non-natural amino acids in this application utilizes a large number of target molecules with molecular weights less than or equal to a preset first preset molecular weight. Amino acid characteristic groups are bonded to the target sites of these target molecules to generate candidate amino acid molecules, greatly enriching the variety of non-natural amino acids. The generated non-natural amino acids exhibit novel structures and high diversity. Furthermore, by screening the candidate amino acid molecules, target amino acids that better meet the required conditions are obtained. This provides a large number of novel or diverse NCCA molecules to support the screening of lead compounds for the development of NCCA-containing peptide drugs, helping to improve the bioactivity, stability, and specificity of peptide drugs, and has clear drug-like significance.
[0022] Furthermore, by screening candidate molecules from a wide range of sources according to a first pre-set screening rule, and by screening molecular fragments of candidate molecules with larger molecular weights, and finally screening the obtained candidate amino acid molecules together, this multiple screening mechanism generates a large number of novel and diverse non-natural amino acids, providing more options for the development of NCCA-containing peptide drugs. At the same time, when the non-natural amino acids obtained through multiple screening are used in drug molecule design, the designed drug molecules have a higher probability of better biocompatibility, a lower probability of toxicity, and a reduced probability of side reactions.
[0023] Furthermore, by obtaining a stable stereoconformation of NCAA monomers, more accurate computational parameters can be extracted. These parameters can then be used for de novo design of NCAA-containing macromolecular drugs (such as peptide drugs), enabling computational analyses of NCAA-containing protein / peptide / antibody docking and mutation. This helps computational software obtain more reliable molecular simulation results, reduces experimental costs, and improves the success rate of drug design. Simultaneously, the NCAA generation method and computational parameter extraction described in this application are universally applicable, capable of generating NCAA monomers from any small molecule structure and further obtaining the corresponding NCAA computational parameters, thus meeting the computational needs of complex peptide / protein / antibody drugs.
[0024] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0025] The above and other objects, features and advantages of this application will become more apparent from the following description of exemplary embodiments of this application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components.
[0026] Figure 1 This is a schematic flowchart illustrating the method for generating non-natural amino acids as shown in the embodiments of this application;
[0027] Figure 2 This is another schematic flowchart illustrating the method for generating non-natural amino acids as shown in the embodiments of this application;
[0028] Figure 3 This is another schematic flowchart illustrating the method for generating non-natural amino acids as shown in the embodiments of this application;
[0029] Figure 4 This is another schematic flowchart illustrating the method for generating non-natural amino acids as shown in the embodiments of this application;
[0030] Figure 5 This is another schematic flowchart illustrating the method for generating non-natural amino acids as shown in the embodiments of this application;
[0031] Figure 6 These are schematic diagrams of the conformations of the target three-dimensional structures in embodiments 1 to 6 of this application;
[0032] Figure 7 This is a schematic diagram of the structure of the apparatus for generating non-natural amino acids shown in the embodiments of this application;
[0033] Figure 8 This is a schematic diagram of the structure of the apparatus for generating non-natural amino acids shown in the embodiments of this application;
[0034] Figure 9 This is a schematic diagram of the structure of an electronic device shown in an embodiment of this application. Detailed Implementation
[0035] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to make this application more thorough and complete, and to fully convey the scope of this application to those skilled in the art.
[0036] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0037] It should be understood that although the terms "first," "second," "third," etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, initial information may also be referred to as second information, and similarly, second information may also be referred to as initial information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0038] In related technologies, NCAA-containing peptide drugs are mostly obtained through high-throughput screening techniques or artificial modification of natural amino acid peptide drugs. When designing NCAA-containing peptide drugs, correctly generating the computational parameters for NCAAs is fundamental to ensuring their performance and efficacy. However, due to the diversity and variability of NCAA structures, macromolecular drug design software (such as Rosetta) typically stores computational parameters for natural amino acids, not NCAAs. Therefore, NCAAs with structures similar to natural amino acids are generally selected for peptide drug design, allowing the computational parameters for natural amino acids to be used instead of NCAAs. However, this approach only addresses the use of NCAAs with similar structures to natural amino acids for drug design, significantly reducing the diversity of available NCAAs and impacting the diversity of NCAA-based drug design.
[0039] To address the aforementioned issues, this application provides a method for generating non-natural amino acids and its applications, which can enrich the novelty and diversity of NCCA types and provide more potential options for the development of NCCA-containing peptide drugs.
[0040] The following is an explanation of the relevant terms used in this application.
[0041] Candidate molecules to be screened: complete molecules or molecular fragments with no limit on molecular weight.
[0042] Candidate molecules: Complete candidate molecules and candidate molecular fragments that meet the first preset screening rules among the candidate molecules to be screened.
[0043] Complete molecule: refers to a molecule that can exist independently and stably, has a complete structure, and retains its chemical properties.
[0044] Molecular fragments: Highly reactive segments produced by the splitting of a complete molecule. Each molecular fragment contains at least one heavy atom carrying at least one unpaired electron and exists in the form of a free radical.
[0045] Target molecule: A complete molecule or molecular fragment with a molecular weight less than or equal to the first preset molecular weight.
[0046] Molecules to be processed: Candidate molecules or candidate molecular fragments with a molecular weight greater than the first preset molecular weight and less than or equal to the second preset molecular weight.
[0047] Chromation site: The designated site, active site, and / or retrosynthetic site on the molecule to be processed for cleavage. In the cleaved molecular fragment, the cleavage site contains unpaired electrons.
[0048] Target site: A heavy atom with a hydrogen atom or a heavy atom with a free radical attached to the target molecule. The target site is used to bond characteristic amino acid groups. The characteristic amino acid groups can bond with the target site through substitution or direct electron pairing.
[0049] The technical solutions of the embodiments of this application are described in detail below with reference to the accompanying drawings.
[0050] See Figure 1 An embodiment of this application illustrates a method for generating non-natural amino acids, comprising:
[0051] S1, identify at least one candidate target site in the target molecule, wherein the target molecule is a complete molecule or molecular fragment with a molecular weight less than or equal to a first preset molecular weight.
[0052] The first preset molecular weight can be selected from 200 Da to 300 Da. For example, the first preset molecular weight can be 200 Da, 250 Da, 300 Da, etc., which are only examples and are not intended to be limiting. For example, the target molecule is a complete molecule or molecular fragment with a molecular weight not exceeding 200 Da. That is to say, the molecular weight of the complete molecule as the target molecule is less than or equal to the first preset molecular weight, and the molecular weight of the molecular fragment as the target molecule is less than or equal to the first preset molecular weight.
[0053] In some embodiments, when the target molecule is a complete molecule, the candidate target site is any heavy atom in the target molecule that is bonded to a hydrogen atom. It is understood that a heavy atom is an atom that is not hydrogen. Preferably, the target molecule can be an organic compound. The complete molecule of the target molecule has one or more heavy atoms bonded to hydrogen atoms; correspondingly, the target molecule may have one or more candidate target sites. In some embodiments, when the target molecule is a molecular fragment, the candidate target site is a heavy atom with a free radical, or the molecular fragment is hydrogenated to form a complete molecule, and then the candidate target site is any heavy atom in the complete molecule that is bonded to a hydrogen atom. It is understood that at least one heavy atom in the molecular fragment has a free radical, and this heavy atom can serve as a corresponding candidate target site. Further, the free radical has unpaired electrons, which can be paired by bonding hydrogen atoms at each free radical, thereby causing the molecular fragment to form a complete molecule. After the complete molecule is formed, if the molecular weight of the complete molecule is less than or equal to a first preset molecular weight, then any heavy atom in the complete molecule bonded to a hydrogen atom is respectively used as a corresponding candidate target site. In other words, the number of candidate target sites for a molecular fragment may change before and after it forms a complete molecule, and the candidate target sites may be all the same or partially the same.
[0054] S2, when there is only one candidate target site, the candidate target site is the target site, and an amino acid characteristic group is bonded to the target site to generate the corresponding candidate amino acid molecule; when there are multiple candidate target sites, at least one target site is determined from the multiple candidate target sites, and an amino acid characteristic group is bonded to the at least one target site to generate the corresponding at least one candidate amino acid molecule.
[0055] It is understood that each target molecule may contain one or more candidate target sites. When there is only one candidate target site, that candidate target site is the target site. When there are multiple candidate target sites, in some embodiments, one or more target sites are determined from among the multiple candidate target sites according to a preset rule, which is to select some or all of the candidate target sites as the target site. For example, among all the candidate target sites in a single target molecule, one, multiple, or all candidate target sites can be randomly selected as the target site. In addition to random selection, one or more specified candidate target sites can also be selected as the target site, without limitation. It is understood that in each target molecule, the number of its target sites is less than or equal to the number of its candidate target sites.
[0056] In this application, the characteristic groups of amino acids include amino (-NH2) and carboxyl (-COOH) groups. The structural formula of the characteristic groups of amino acids is shown in Formula I below, wherein, " "A bond used to identify the binding of a characteristic amino acid group to a target site."
[0057]
[0058] After identifying the target site in the target molecule, in some embodiments, when the target molecule is an intact molecule, the hydrogen atom attached to the target site is replaced by an amino acid characteristic group. That is, by replacing a hydrogen atom attached to a heavy atom, the amino acid characteristic group is bonded to the target molecule, generating the corresponding candidate amino acid molecule.
[0059] In some implementations, when the target molecule is a molecular fragment, an amino acid signature group is attached to the target site. Before the molecular fragment forms a complete molecule, the amino acid signature group is attached to the target site. The free radical at the target site pairs with the unpaired electrons in the free radical at the target site, enabling the characteristic amino acid group to bond with the molecular fragment and generate the corresponding candidate amino acid molecule.
[0060] Without performing deduplication on candidate amino acid molecules, the number of candidate amino acid molecules generated for each target molecule is the same as the number of target sites. That is, each time only one amino acid signature group is bonded to one target site in the target molecule, generating one corresponding candidate amino acid molecule. In other words, when multiple target sites exist on a single target molecule, each target site is traversed, and each time only one amino acid signature group is bonded to one target site, generating one corresponding candidate amino acid molecule.
[0061] S3. Screen each candidate amino acid molecule according to the preset amino acid screening rules to obtain the target amino acid molecule, which is a non-natural amino acid.
[0062] This step is used to screen the candidate amino acid molecules generated in step S2 above. In this application, the preset amino acid screening rules may include one or more rules. By filtering the candidate amino acid molecules, target amino acid molecules that meet the preset screening rules are obtained. In some embodiments, the preset amino acid screening rules include at least one of the following rules:
[0063] (1) The synthesis difficulty of the target amino acid molecule meets the preset conditions.
[0064] (2) The number of dihedral angles of the target amino acid molecule is ≤4.
[0065] (3) The target amino acid molecule is a non-natural amino acid.
[0066] (4) There are no identical molecules among the target amino acid molecules.
[0067] Regarding rule (1) in the preset amino acid screening rules, in some embodiments, the predicted value of the synthesis difficulty corresponding to each candidate amino acid molecule is obtained, and candidate amino acid molecules whose synthesis difficulty meets the preset conditions are retained as target amino acid molecules. For example, the synthesis difficulty corresponding to each candidate amino acid molecule can be predicted according to relevant software such as rdikt software. The synthesis difficulty is quantified by the predicted value. For example, the smaller the predicted value, the lower the synthesis difficulty. Then, molecules that are easier to synthesize are directly selected as target amino acid molecules. In some specific embodiments, candidate amino acid molecules with predicted values lower than the preset difficulty threshold can be retained as target amino acid molecules. In some specific embodiments, after sorting according to their respective predicted values from smallest to largest, multiple candidate amino acid molecules with the highest synthesis difficulty are retained as target amino acid molecules. For example, the predicted value of synthesis difficulty is selected from 0 to 10, where 0 indicates that the molecule has no synthesis difficulty and 10 indicates that the molecule cannot be synthesized. After obtaining the predicted value of the synthesis difficulty corresponding to each candidate amino acid molecule, they are sorted according to the score from smallest to largest. For example, the top 30,000 candidate amino acid molecules can be selected as target amino acid molecules. This design allows for a multi-dimensional combination of screening methods to select and retain an appropriate number of candidate amino acid molecules that are easy to synthesize.
[0068] Regarding rule (2) in the preset amino acid screening rules, in some embodiments, the number of dihedral angles corresponding to each candidate amino acid molecule is obtained, and candidate amino acid molecules with no more than 4 dihedral angles are retained as target amino acid molecules. For example, the number of dihedral angles of each candidate amino acid molecule can be determined using relevant software such as Schrödinger Maestro, GROMACS, Gaussian, etc., and candidate amino acid molecules with no more than 4 dihedral angles are retained as target amino acid molecules, while candidate amino acid molecules with more than 4 dihedral angles are filtered out. By controlling the number of dihedral angles, target amino acid molecules with stable conformation, low synthesis and purification difficulty, and better conformity to drug-likeness rules can be obtained.
[0069] Regarding rule (3) in the preset amino acid screening rules, in some embodiments, when a candidate amino acid molecule is a natural amino acid, the corresponding candidate amino acid molecule is filtered out, and the remaining non-natural amino acid molecule is the target amino acid molecule. It can be understood that the method of this application is used to generate non-natural amino acids, and when any candidate amino acid molecule is a natural amino acid, it can be filtered out accordingly. In other embodiments, after step S1 above, target molecules that can generate natural amino acids after bonding with amino acid characteristic groups can be filtered out in advance. For example, there are currently 20 known natural amino acids. After removing the amino acid characteristic groups of each natural amino acid, the remaining fragment of the natural amino acid is retained as the corresponding reference fragment. If the target molecule is the same as any reference fragment, it means that the target molecule can generate natural amino acids after bonding with amino acid characteristic groups. Accordingly, by filtering out such target molecules first, the subsequent redundant operations of the system can be reduced and the processing efficiency can be improved.
[0070] In some implementations, in response to rule (4) of the preset amino acid screening rules, all generated candidate amino acid molecules are deduplicated to remove redundant candidate amino acid molecules and obtain target amino acid molecules that are different from each other.
[0071] In some optional implementations, when screening candidate amino acid molecules according to multiple rules in the preset amino acid screening rules, the candidate amino acid molecules can be screened sequentially item by item, and the candidate amino acid molecules screened by the previous rule can be screened by the next rule, until the candidate amino acid molecules screened by the last rule are taken as the target amino acid molecules.
[0072] In some optional implementations, when screening candidate amino acid molecules according to multiple rules in the preset amino acid screening rules, the candidate amino acid molecules can be screened independently according to each rule to obtain the preferred amino acid molecules after screening for each rule; after combining and removing duplicates from each group of preferred amino acid molecules, the target amino acid molecule is obtained.
[0073] The above screening methods are only examples. You can also use a custom method to screen out candidate amino acid molecules that meet one or more of the preset amino acid screening rules as target amino acid molecules.
[0074] In this application, each target amino acid molecule is a non-natural amino acid. After obtaining the target amino acid molecule, it can be used to synthesize with macromolecules, such as synthesizing the target amino acid molecule with a polypeptide macromolecule to generate a polypeptide drug.
[0075] As can be seen from this example, the method for generating non-natural amino acids in this application can use a large number of target molecules with a molecular weight less than or equal to a first preset molecular weight, and bond amino acid characteristic groups to the target sites of the target molecules to generate a large number of candidate amino acid molecules, which greatly enriches the types of non-natural amino acids. The generated non-natural amino acids have novel structures and good diversity. Furthermore, after screening the candidate amino acid molecules, target amino acids that better meet the required conditions are obtained. This provides a large number of novel or diverse NCCA molecules to support the screening of lead compounds for the development of NCCA-containing peptide drugs, which helps to improve the bioactivity, stability and specificity of peptide drugs and has clear drug-like significance.
[0076] See Figure 2 An embodiment of this application illustrates a method for generating non-natural amino acids, comprising:
[0077] S210, obtain candidate molecules.
[0078] In this embodiment, S210 is step P1.
[0079] In this application, the candidate molecule is a candidate complete molecule or a candidate molecular fragment. Preferably, the candidate molecule has a molecular weight not exceeding a second preset molecular weight, wherein the second preset molecular weight is selected from 700 Da to 1000 Da (e.g., 700 Da, 750 Da, 800 Da, 850 Da, 900 Da, 950 Da, 1000 Da, etc.). For example, the candidate molecule can be a candidate complete molecule or a candidate molecular fragment with a molecular weight not exceeding 1000 Da.
[0080] S220, determine the molecular weight of the candidate molecule and its size relative to the first preset molecular weight, and obtain the target molecule and / or the molecule to be processed.
[0081] In this embodiment, S220 is step P2.
[0082] In some specific implementations, the molecular weight of the candidate molecule is determined to be greater than or equal to a first preset molecular weight to obtain the target molecule and / or the molecule to be processed. The molecule to be processed is a candidate molecule fragment or a candidate complete molecule with a molecular weight greater than the first preset molecular weight and less than the second preset molecular weight.
[0083] In this step, the molecular weight of each candidate molecule can be obtained and compared with the first preset molecular weight. When the molecular weight of a candidate molecule is less than or equal to the first preset molecular weight, the candidate molecule is used as the target molecule and step S1 is continued.
[0084] In some optional embodiments, this step can also simultaneously obtain the molecule to be processed. When the molecular weight of a candidate molecule is greater than a first preset molecular weight and less than or equal to a second preset molecular weight, the candidate molecule is used as the molecule to be processed and step P3 continues. That is, this step can obtain the target molecule while retaining the molecule to be processed.
[0085] S230, the molecule to be processed is cut at least once to obtain molecular fragments.
[0086] In this embodiment, S230, or step P3, is used to obtain a molecular fragment with a molecular weight smaller than a first preset molecular weight as the target molecule. In some specific embodiments, step P3 specifically includes steps P31 to P33 as described below, and steps P31 to P33 are executed cyclically according to the determination conditions therein. Wherein:
[0087] Step P31 involves cutting the molecule to be processed into multiple candidate molecular fragments.
[0088] It is understandable that when there is only one molecule to be processed, each segmentation will yield two candidate molecular fragments. When there are Q molecules to be processed, each segmentation will yield 2Q candidate molecular fragments, where Q is an integer greater than 1.
[0089] In step P31, in some embodiments, the molecule to be treated is cleaved once at the cleavage site according to a preset cleavage rule to obtain multiple candidate molecular fragments. For example, the chemical bonds in the molecule to be treated can be cleaved using relevant software such as rdkit RECAP, BRICS, or other tools, and each cleavage can obtain two corresponding candidate molecular fragments.
[0090] Furthermore, each molecule to be processed may have one or more slicing sites, and the number of slicing sites in each molecule to be processed is N. When N is an integer greater than 1, in some embodiments, without deduplication, each molecule to be processed is sliced according to one of the following preset slicing rules, which includes, but is not limited to, any one of the following rules:
[0091] (1) Perform a single cut at one of the cut sites of the molecule to be processed, either randomly or by designation, to obtain a total of two candidate molecule fragments.
[0092] (2) Each of the N cleavage sites of the molecule to be treated is cleaved once, and each cleavage yields two corresponding candidate molecular fragments, for a total of 2N candidate molecular fragments.
[0093] (3) Select M cleavage sites on the molecule to be processed randomly or by designation for cleavage. Each cleavage yields two candidate molecular fragments, for a total of 2M candidate molecular fragments, where 1 < M < N.
[0094] In other words, for different molecules to be treated, any of the above-mentioned preset segmentation rules can be selected to segment the molecules to be treated, and each segmentation occurs only at one segmentation site of each molecule to be treated, so that each segmentation can obtain two candidate molecular fragments. When different preset segmentation rules are used, the total number of candidate molecular fragments obtained by segmenting the same molecule to be treated may be the same or different. Different molecules to be treated can be segmented using the same or different preset segmentation rules. For example, without considering deduplication, if a molecule to be treated has 3 active sites, when the above-mentioned preset segmentation rule (1) is selected for segmentation, a total of 2 candidate molecular fragments are obtained. When the above-mentioned preset segmentation rule (2) is selected for segmentation, a total of 6 candidate molecular fragments are obtained. When the above-mentioned preset segmentation rule (3) is selected for segmentation, if only 2 of the segmentation sites are selected for segmentation, a total of 4 candidate molecular fragments are obtained.
[0095] After segmenting all the molecules obtained in step P2 using one or more preset segmentation rules, duplicate fragments can be removed from all candidate molecular fragments to avoid repeated execution of subsequent steps for multiple identical candidate molecular fragments, reducing redundant operations and improving system efficiency. Of course, other custom methods can also be used to segment candidate molecules, which will not be elaborated or limited here.
[0096] In some specific implementations, the cleavage occurs at the cleavage site of the molecule to be treated; wherein, the cleavage site is a designated site, an active site, or a retrosynthetic site;
[0097] In some implementations, the designated site can be any site on the molecule to be treated. That is, any chemical bond connecting any two atoms in the molecule to be treated can be randomly or custom-defined as the designated site for cleavage.
[0098] In some implementations, an active site refers to a chemical bond on an atom or group in the molecule to be treated that can directly interact with a target (such as an enzyme or receptor). For example, the molecules and their derivatives shown in the table below all contain active sites. The active site is defined by the chemical bond it pierces. R1, R2, R3, and R4 are each independently selected from any suitable substituent, including but not limited to H, halogens, alkyl groups, hydroxyl groups, alkoxy groups, aminoalkyl groups, mercaptoalkyl groups, haloalkyl groups, haloalkoxy groups, cycloalkyl groups, heterocyclic groups, aryl groups, and heteroaryl groups. By using the active site as the cleavage site, the pharmacophore can be retained in the cleaved candidate molecular fragments, which helps improve drug design efficiency and is more conducive to generating highly active lead compounds. It should be noted that the molecules and their active sites in the table below are only illustrative examples and are not intended to be limiting.
[0099]
[0100] Similarly, retrosynthetic sites refer to the chemical bonds or linkages in the molecule to be preferentially cleaved during the design of synthetic routes, in order to break down the molecule into simpler precursors. Retrosynthetic sites can be obtained through relevant experience and include, but are not limited to, ester groups, amide groups, hydroxyl groups, double bonds, etc. These are only examples here and will not be elaborated further. By using retrosynthetic sites as cleavage sites, the difficulty of obtaining candidate molecular fragments can be reduced, meaning that candidate molecular fragments are simpler and easier to obtain.
[0101] Step P32: Perform step P31 with the candidate molecular fragment as the candidate molecule to obtain at least one target molecule and / or at least one molecule to be processed.
[0102] In step P32, each candidate molecular fragment is obtained as a candidate molecule and step P2 is executed accordingly. Specifically, the molecular weight of each candidate molecular fragment is compared to a first preset molecular weight. When the molecular weight of a candidate molecular fragment is less than or equal to the first preset molecular weight, it is selected as a molecular fragment in the target molecule. In this case, step P33 is not required, and step S240 is executed directly. When the molecular weight of the candidate molecular fragment is greater than the first preset molecular weight, it is selected as a molecule to be processed, and step P33 is executed. In other words, the molecule to be processed in this embodiment includes not only the aforementioned candidate molecules but also the candidate molecular fragments with molecular weights greater than the first preset molecular weight from step P32.
[0103] Step P33: Perform steps P31 and P32 on the molecules to be processed until a preset termination condition is met; wherein, the preset termination condition is that only the target molecule is obtained after the most recent execution of step P32, but no molecules to be processed, and / or steps P31 and P32 are executed a preset number of times.
[0104] In step P33, the molecules to be processed (i.e., candidate molecular fragments with a molecular weight greater than the first preset molecular weight) obtained in step P32 are repeatedly processed through steps P31 and P32 until a preset termination condition is met, at which point the cycle terminates. For example, if the molecular weight of the candidate molecular fragment obtained in the first round is less than or equal to the first preset molecular weight after the second round of step P31, then the candidate molecular fragment obtained in the second round is used as the target molecule; if the candidate molecular fragment obtained in the second round is greater than the first preset molecular weight, then the candidate molecular fragment obtained in the second round is used as the molecule to be processed and a third round of cycling is performed, and so on.
[0105] In the first optional preset termination condition, steps P31 and P32 are executed for each molecule to be processed obtained in step P2. The loop terminates when the cumulative number of iterations of steps P31 and P32 executed on a single molecule to be processed obtained in step P2 reaches a preset number. For example, the preset number of iterations can be 2 to 5 times, which is only illustrative and not limited. By controlling the cumulative number of iterations, system resources are saved and system processing efficiency is improved. It can be understood that if the candidate molecule fragment obtained in the last round of segmentation at the time of loop termination is still larger than the first preset molecular weight, the candidate molecule fragment can be discarded.
[0106] In the second optional preset termination condition, for each molecule to be processed obtained in step P2, the number of cycles in steps P31 and P32 can be unlimited until all candidate molecule fragments obtained most recently are less than or equal to the first preset molecular weight, and there are no molecule to be processed that are greater than the first preset molecular weight. At this time, it means that the preset termination condition has been met and the cycle is terminated.
[0107] Optionally, the first and second preset termination conditions can be used simultaneously. When the system is detected to meet either preset termination condition first, the loop terminates.
[0108] In this step, each molecule to be processed obtained in step P2 is cut at least once in the manner described above, thereby obtaining molecular fragments with a molecular weight less than or equal to the first preset molecular weight.
[0109] S240, identify at least one candidate target site in the target molecule, wherein the target molecule is a complete molecule or molecular fragment with a molecular weight less than or equal to a first preset molecular weight.
[0110] Step S240 in this embodiment is the same as step S1, and the relevant description of step S1 above can be referred to in the same way.
[0111] Furthermore, target molecules are pre-classified according to whether they contain cyclic groups. In some embodiments, when the target molecule is a complete molecule, it can be a complete chain molecule or a complete cyclic molecule. Chain molecules do not contain cyclic groups, while cyclic molecules contain at least one cyclic group. The positions of the corresponding candidate target sites differ depending on the structure of the target molecule.
[0112] In some embodiments, when the target molecule is a complete chain molecule, the candidate target sites are selected from heavy atoms whose chain ends are connected to hydrogen atoms. The chain molecule can be a straight-chain structure or a branched structure. Specifically, when the target molecule is a straight-chain structure, the candidate target sites are located at the two chain ends of the straight chain and are heavy atoms connected to at least one hydrogen atom; correspondingly, the number of candidate target sites in this target molecule is at most two. Similarly, when the target molecule is a branched structure, the candidate target sites are located at the chain ends of the main chain and each branch and are heavy atoms connected to hydrogen atoms; the number of candidate target sites in this target molecule is at most (2+Z), where Z is the number of branches.
[0113] In some implementations, the heavy atom connecting the hydrogen atom is a carbon atom. That is, when the target molecule is a complete chain molecule, if any chain end is a carbon atom connected to at least one hydrogen atom, then that carbon atom at the chain end is a candidate target site. Conversely, if the heavy atom connecting the hydrogen atom at the chain end is not a carbon atom, then that heavy atom may not be considered a candidate target site.
[0114] In some implementations, when the target molecule is a complete cyclic molecule, if the complete cyclic molecule does not contain branches, the candidate target site is selected from heavy atoms with hydrogen atoms on the ring. That is, for a target molecule that has only a ring structure, the candidate target site is a heavy atom located on the ring, and the heavy atom is connected to at least one substituted hydrogen atom for substitution by amino acid characteristic groups.
[0115] In some embodiments, when the target molecule is a complete cyclic molecule, if the complete cyclic molecule contains branches and the ring group is located at the chain end, then the candidate target sites are selected from heavy atoms with hydrogen atoms attached to the ring group and / or heavy atoms with hydrogen atoms attached to each chain end. Similarly, the ends of the branches of the cyclic molecule are also chain ends, and when the chain end is a heavy atom with at least one hydrogen atom attached, the heavy atom in the chain end can be used as a candidate target site.
[0116] In some implementations, when the target molecule is a complete cyclic molecule, if the complete cyclic molecule contains branches and the cyclic group is located in the middle of the chain, the candidate target sites are selected from heavy atoms with hydrogen atoms at the ends of each chain.
[0117] In other words, for complete cyclic molecules, we can divide them into the three cases mentioned above, and determine the corresponding candidate target sites for each case. Preferably, in the above-mentioned complete cyclic molecules, when the heavy atom connecting each hydrogen atom is a carbon atom, the heavy atom can be considered as a candidate target site. Conversely, if the heavy atom connecting the hydrogen atom is not a carbon atom, then the heavy atom can be disregarded as a candidate target site.
[0118] It is understandable that when a target molecule possesses both a chain structure and a ring group, the target molecule contains one or more ring groups, which may be located in the middle of the chain or at the end of the chain. Correspondingly, ring groups located in the middle of the chain do not have candidate target sites. Candidate target sites are carbon atoms with hydrogen atoms attached at the ends of each chain, and carbon atoms with hydrogen atoms attached to ring groups located at the ends of the chain.
[0119] It should be noted that if the target molecule does not have any candidate target sites, the target molecule can be filtered out without further steps.
[0120] In some embodiments, when the target molecule is a molecular fragment, the candidate target site is a cleavage site. It is understood that the cleavage site has unpaired electrons, and the chemical bond at the cleavage site can directly bond to the characteristic groups of the amino acid. In some embodiments, when the target molecule is a molecular fragment, after the molecular fragment is formed into a complete molecule, it is used as the aforementioned complete chain molecule or complete cyclic molecule, and at least one corresponding candidate target site is determined. Specifically, a hydrogen atom can be attached to the cleavage site of the molecular fragment, thereby pairing with the unpaired electrons of the cleavage site, and the molecular fragment forms a complete molecule. If the complete molecule is a chain molecule, the corresponding candidate target site is determined according to the aforementioned method for determining candidate target sites for chain molecules. If the complete molecule is a cyclic molecule, the corresponding candidate target site is determined according to the aforementioned method for determining candidate target sites for cyclic molecules, which will not be repeated here.
[0121] S250, when there is only one candidate target site, the candidate target site is the target site, and an amino acid characteristic group is bonded to the target site to generate the corresponding candidate amino acid molecule; when there are multiple candidate target sites, at least one target site is determined from the multiple candidate target sites, and an amino acid characteristic group is bonded to the at least one target site to generate the corresponding at least one candidate amino acid molecule.
[0122] In this embodiment, S250 is step S2, and the description of step S2 in the above embodiments can be referred to similarly. It can be understood that when the target molecule is a complete molecule, in some embodiments, the hydrogen atom attached to the target site is replaced by an amino acid characteristic group, thereby achieving the bonding of the amino acid characteristic group at the target site. It can be understood that the number K of target sites in a single target molecule is greater than or equal to 1, and only one amino acid characteristic group is bonded to one target site at a time, generating a corresponding candidate amino acid molecule.
[0123] Specifically, when the number of target sites K in the target molecule is equal to 1, an amino acid characteristic group is attached to the heavy atom of the target site by substitution, generating a corresponding candidate amino acid molecule. More specifically, when the heavy atom is attached to multiple hydrogen atoms, any one of the hydrogen atoms can be substituted to bond the amino acid characteristic group, generating the corresponding candidate amino acid molecule.
[0124] When the number of target sites K in the target molecule is greater than 1, optionally, by traversing all target sites and reacting with only one hydrogen atom of one target site at a time, the characteristic amino acid group is attached to the carbon atom of that target site, thereby generating K corresponding candidate amino acid molecules. Optionally, among the K target sites, only one target site can be specified or randomly selected, and one amino acid characteristic group can be used to replace the hydrogen atom attached to that target site. Optionally, multiple target sites can be specified or randomly selected among the K target sites to undergo substitution reactions, generating multiple corresponding candidate amino acid molecules, i.e., the number of candidate amino acid molecules generated can be less than K.
[0125] It should be noted that the target molecule itself may or may not have amino acid characteristic groups; after an amino acid characteristic group is bonded to any target site of the target molecule, each candidate amino acid molecule generated will have at least one amino acid characteristic group.
[0126] It is understood that among the multiple candidate amino acid molecules generated by the scheme claimed in this application, there is a possibility that some candidate amino acid molecules are the same. In order to reduce redundant operations and reduce the system data processing load, alternatively, deduplication can be performed to retain only one molecule among multiple identical candidate amino acid molecules before proceeding to the subsequent step S260. Alternatively, a deduplication operation can be performed once after step S260 to delete redundant target amino acid molecules, ensuring that each target amino acid molecule is different from the others.
[0127] S260: Screen each candidate amino acid molecule according to the preset amino acid screening rules to obtain the target amino acid molecule, which is a non-natural amino acid.
[0128] In this embodiment, S260 is step S3. Similarly, refer to the above description of step S3, and will not be repeated here.
[0129] As can be seen from this example, the method for generating non-natural amino acids in this application can use candidate molecules from a wide range of sources. The molecular weight is used to determine whether to use the candidate molecules directly as target molecules or to use the cleaved molecular fragments as target molecules. Target sites that can bond to characteristic amino acid groups are determined according to a preset method, generating a large number of non-natural amino acids. This greatly enriches the types of non-natural amino acids. The generated non-natural amino acids have novel structures and good diversity. Furthermore, after screening according to preset amino acid screening rules, target amino acids that better meet the required conditions are obtained. This provides a large number of novel or diverse NCCA molecules to support the screening of lead compounds for the development of NCCA-containing peptide drugs, which helps to improve the bioactivity, stability, and specificity of peptide drugs and has clear drug-like significance.
[0130] See Figure 3 and Figure 4 An embodiment of this application illustrates a method for generating non-natural amino acids, comprising:
[0131] S310: Obtain multiple candidate molecules to be screened.
[0132] In this embodiment, S310 is step P0.
[0133] In this embodiment, the candidate molecules to be screened are organic molecules with complete structures and no limit on molecular weight. In some embodiments, a large number of organic molecules can be collected as candidate molecules from various publicly available databases or other sources. For example, candidate molecules can be collected from databases such as the ChEMBL Database, ZINC Database, and FDA Approved Drug Products. Alternatively, AIGC technology can be used to automatically generate virtual candidate molecules to obtain millions of candidate molecules.
[0134] S320: According to the first preset screening rule, multiple candidate molecules to be screened are screened to obtain one or more candidate molecules.
[0135] In this embodiment, S320 is the specific step P1.
[0136] To improve data processing efficiency and conserve system computing resources, this step filters multiple candidate molecules according to a first preset screening rule. The selected candidate molecules conform to the first preset screening rule. The first preset screening rule includes at least one of the following rules:
[0137] (1) The molecules have a standardized format. For example, the standardized format can be the SMILES formula, two-dimensional structural formula, three-dimensional structural formula, etc., so as to clearly and uniformly represent each molecule and read it by computer, which is conducive to the efficient execution of the technical solution of this application.
[0138] (2) The molecules have complete structures rather than fragmented structures. That is to say, each complete molecule obtained by screening has an independent and complete structure, rather than a segment cut or extracted from a complete molecule.
[0139] (3) The atoms in the molecule do not contain atoms other than H, C, N, O, F, Cl, S, P, Br and I. That is to say, the elements in each molecule are selected from H, C, N, O, F, Cl, S, P, Br and I, so that when the non-natural amino acids obtained in this application are used in drug molecule design, there is a greater probability of obtaining drug molecules with better biocompatibility, a lower probability of drug molecule toxicity, and a reduced probability of side reactions.
[0140] (4) The molecular weight of the molecule is ≤800 Da. By controlling the molecular weight of the molecule, it is more conducive to ensuring the purification, transport, absorption and metabolism properties of the drug molecule designed based on the non-natural amino acids obtained in this application.
[0141] (5) The molecule does not contain radioactive isotope atoms. It is understood that by not using radioactive isotope atoms, the safety of the resulting drug molecule is better guaranteed when the non-natural amino acids obtained in this application are used in drug molecule design.
[0142] (6) The molecule meets at least one of the following drug class rules: Rule 5, Rule 4, or Rule 3. Specifically, drug class rules are empirical criteria used to predict whether a compound has drug-like properties. Among them, Rule 5 includes: (1) molecular weight ≤ 500 Da; (2) number of hydrogen bond donors ≤ 5; (3) number of hydrogen bond acceptors ≤ 10; (4) lipid-water partition coefficient ≤ 5; (5) number of rotatable bonds ≤ 10. Rule 4 includes: (1) molecular weight ≤ 250 Da; (2) number of hydrogen bond donors ≤ 3; (3) number of hydrogen bond acceptors ≤ 3; (4) lipid-water partition coefficient ≤ 3. Rule 3 includes: (1) molecular weight ≤ 300 Da; (2) number of rotatable bonds ≤ 3; (3) lipid-water partition coefficient ≤ 3. Depending on different drug requirements, different drug class rules can be selected to filter and screen candidate molecules.
[0143] Based on one or more of the first preset screening rules mentioned above, multiple candidate molecules can be screened and filtered before proceeding to the next step S330 for the candidate molecules that meet the first preset screening rules. Preferably, candidate molecules that meet all the rules in the first preset screening rules can be screened, so that the target amino acid molecules generated subsequently can help improve the efficacy, safety and feasibility of drug molecule development.
[0144] It should be noted that step S320 can be executed either before or after S330; that is, the execution order of S320 can be adjusted. It can be understood that if S320 is executed after S330, then the input data for S320 will be the target molecule and the molecule to be processed output by S330.
[0145] S330, determine the molecular weight of the candidate molecule and its relationship with the first preset molecular weight to obtain the target molecule and / or the molecule to be processed.
[0146] In this embodiment, S330 is step P2. Similarly, the relevant description of step S220 is not repeated here.
[0147] It should be noted that if step S320 has been executed before step S330, the input data for this step is the candidate molecules that meet the first preset screening rules. If step S320 has not been executed before step S330, the input data for this step is the candidate molecules to be screened in S310.
[0148] S340, the molecule to be processed is cut at least once to obtain molecular fragments.
[0149] In this embodiment, S340 refers to steps P31 to P33. Similarly, the relevant description of step S230 is not repeated here.
[0150] S350, the molecular fragments are screened according to the second preset screening rules to obtain molecular fragments that meet the second preset screening rules.
[0151] In this embodiment, S350 is step P4, which can be either executed or not.
[0152] After obtaining multiple molecular fragments, this application can further screen the molecular fragments to perform subsequent step S360 on molecular fragments that meet the second preset screening rules. The second preset screening rules include at least one of the following rules:
[0153] (1) The number of heavy atoms in the molecular fragment is 5 to 50 (the number of heavy atoms can be, for example, 5-10, 5-20, 5-30, 5-40, 10-20, 10-30, 10-40, 15-25, 15-35, or 15-45, etc.). After obtaining all molecular fragments corresponding to the same molecule to be processed, by calculating the number of heavy atoms in each molecular fragment, molecular fragments with inconsistent heavy atom numbers can be filtered out, thereby controlling the molecular weight of the molecular fragments and selecting a better molecular fragment as the starting point for drug molecule design.
[0154] (2) The molecular fragment conforms to at least one of the following rules: Drug Class 5, Drug Class 4, or Drug Class 3. The Drug Class rules are the same as those described above and will not be repeated here.
[0155] (3) The molecular fragment is not any of the following fragments:
[0156] CH3-, -CH2CH(CH3)2, -CH2CH2SCH3, -CH2-phenyl, , , -CH2OH, -CHOH(CH3), -CH2CONH2, -CH2CH2CONH2, -CH2SH, -CH2COOH, -CH2CH2COOH, -CH2CH2CH2CH2NH2, -CH2CH2CH2NHC(=NH)NH2.
[0157] (4) The molecular fragments that form a complete molecule are not any of the following:
[0158] CH4, CH3CH2SCH3, toluene, p-hydroxytoluene , CH3OH, CH3CONH2, CH3CH2CONH2, CH3SH, CH3COOH, CH3CH2COOH, CH3CH2CH2CH2NH2, CH3CH2CH2NHC(=NH)NH2.
[0159] (5) When the complete molecule formed by the molecular fragment is CH3CH2OH, CH3CH2CH3, CH3CH(CH3)2 or CH2(CH3)CH2CH3, the candidate target site of the complete molecule is the carbon atom located at the end of the chain.
[0160] It is understandable that among the 20 known natural amino acid molecules, after removing the characteristic amino acid groups from the natural amino acid molecules, the remaining molecular fragments are the fragments listed in item (3) above. By pairing the unpaired electrons of the fragments with hydrogen atoms, the molecules listed in item (4) above can be obtained. The method of this application aims to generate non-standard amino acid molecules by pre-removing molecular fragments that can generate natural amino acid molecules, thereby reducing redundant operations in subsequent steps and improving processing efficiency. Among them, CH3CH2OH and CH3CH2CH 3、 In both CH3CH(CH3)2 and CH2(CH3)CH2CH3, the chain terminus is -CH3, and the candidate target site at the chain terminus is the carbon atom in the methyl group. When a characteristic amino acid group replaces any hydrogen atom attached to the carbon atom at the chain terminus, a native amino acid will not be formed.
[0161] In some implementations, when a distinct "-" marker is detected within a molecular fragment, the "-" is identified as the target site.
[0162] S360, identify at least one candidate target site in the target molecule, wherein the target molecule is a complete molecule or molecular fragment with a molecular weight less than or equal to a first preset molecular weight.
[0163] Step S360 in this embodiment is the same as step S1. Similarly, the relevant descriptions of steps S1 and S240 can be referred to, and will not be repeated here.
[0164] If the target molecule is a molecular fragment, in some optional embodiments, this embodiment may further determine at least one corresponding candidate target site after some or all of the molecular fragments are formed into a complete molecule and the molecular weight of the formed complete molecule is less than or equal to a first preset molecular weight.
[0165] In some alternative implementations, after the molecular fragments are formed into complete molecules, the original molecular fragments can be retained and the corresponding candidate target sites can be identified for subsequent steps.
[0166] S370, when there is only one candidate target site for the target molecule, the candidate target site is the target site, and an amino acid characteristic group is bonded to the target site to generate the corresponding candidate amino acid molecule; when there are multiple candidate target sites for the target molecule, at least one target site is determined from the multiple candidate target sites, and an amino acid characteristic group is bonded to the at least one target site to generate the corresponding at least one candidate amino acid molecule.
[0167] In this embodiment, S370 is step S2. Similarly, the relevant descriptions of steps S2 and S250 above can be referred to, and will not be repeated here.
[0168] S380: Screen each candidate amino acid molecule according to the preset amino acid screening rules to obtain the target amino acid molecule, which is a non-natural amino acid.
[0169] In this embodiment, S380 is step S3. Similarly, the description of S3 will not be repeated here.
[0170] As illustrated in this example, the method for generating non-natural amino acids in this application involves screening a wide range of candidate molecules according to a first preset screening rule. Candidate molecules with molecular weights less than or equal to the first preset molecular weight are then selected as target molecules. Furthermore, molecular fragments obtained by slicing candidate molecules with molecular weights greater than the first preset molecular weight are also selected as target molecules. These fragments are then screened according to a second preset screening rule. Finally, candidate amino acid molecules obtained by bonding characteristic amino acid groups to the target molecules are also screened. This multiple screening mechanism generates a large number of novel and diverse non-natural amino acids, providing more options for the development of NCCA-containing peptide drugs. Simultaneously, when non-natural amino acids obtained through multiple screening are used in drug molecule design, the probability of the designed drug molecules being more biocompatible is higher, the probability of toxicity is lower, and the probability of side reactions is reduced.
[0171] After obtaining a large number of non-natural amino acids through the above embodiments, the method of this application further includes generating computational parameters for non-natural amino acids for reading by macromolecular drug design software, which facilitates computer design of drugs.
[0172] See Figure 5 Based on the target amino acid molecules generated in the above embodiments, a method for generating non-natural amino acids according to an embodiment of this application further includes:
[0173] S4 bonds an acetyl group (-C(O)CH3) to the N-terminus of the target amino acid molecule and an N-methyl group (-NCH3) to the C-terminus to generate the corresponding dipeptide compound.
[0174] In this step, by extending the target amino acid molecule to a dipeptide, stereoconformity constraints can be introduced, which helps to simulate the environment of the monomer after it is applied to peptide drugs. This allows for a more accurate determination of the monomer's stereoconformity preference, thereby screening out more suitable monomer conformations.
[0175] To obtain a dipeptide compound of the target amino acid molecule, in some specific embodiments, an acetyl group (-C(O)CH3) is bonded to the N-terminus of the target amino acid molecule, and an N-methyl group (-NCH3) is bonded to the C-terminus. Exemplarily, tools such as pymol, rdkit, Schrödinger Maestro, and ChemDraw can be used to perform end-capping operations at the N- and C-termini of the target amino acid molecule; that is, acetyl substitution (-C=OCH3) is performed at the N-terminus (-NH2 terminus), and N-methyl substitution (-NHCH3) is performed at the C-terminus (-COOH terminus), generating the dipeptide structure of the target amino acid molecule.
[0176] S5, obtain the target stereoconformation corresponding to the dipeptide compound.
[0177] In some specific implementations, step S5 includes:
[0178] Step S51: Obtain at least one initial stereoconformation of the dipeptide compound;
[0179] Step S52: Select the target stereo configuration from each initial stereo configuration according to the preset configuration screening rules.
[0180] It is understandable that the number of chiral centers contained in each dipeptide compound may differ. Accordingly, after the initial stereoconformation is formed, the dipeptide compound may contain L-type or D-type amino acid structures with only a single stereoconformation in its side chain, or it may contain L-type or D-type amino acid structures with multiple stereoconformations in its side chain. For example, when the side chain of a dipeptide compound has one chiral center, its initial stereoconformation will correspondingly have both L-type and D-type conformations. When the side chain of a dipeptide compound has multiple chiral centers, various initial stereoconformations will be generated through permutation and combination.
[0181] Furthermore, tools such as Schrödinger Maestro, Discovery Studio, Gaussian, and GROMACS can be used to optimize the energy minimization of each initial stereoconformation of the dipeptide compound. After energy minimization, more stable target stereoconformations can be screened out.
[0182] To obtain more stable conformations and reduce the processing of redundant data, for multiple initial stereoformations of the same dipeptide compound, some embodiments include preset conformation screening rules including at least one of the following rules:
[0183] (1) The initial stereo conformation with a single conformation of the side chain is taken as the target stereo conformation.
[0184] (2) Among the various stereo configurations of the side chain, the initial stereo configuration with the lowest energy is selected as the target stereo configuration.
[0185] (3) The initial stereoconformation of the L-type amino acid structure is taken as the target stereoconformation.
[0186] Preferably, a more stable target stereo configuration can be selected from multiple initial stereo configurations by simultaneously integrating multiple rules.
[0187] S6: Extract and store the computational parameters corresponding to the target amino acid in the target stereoconformation of the dipeptide compound.
[0188] The computational parameters corresponding to the monomer of the target amino acid molecule in the target stereoconformation are extracted and stored. When the target amino acid molecule is needed for the design of peptide macromolecules, the corresponding computational parameters can be quickly invoked to participate in the drug molecule design. In this way, a large number of rich and novel NCCAs can be involved in drug molecule design, expanding the selection range of NCCA-containing peptide drugs without being limited to NCCAs with structures similar to natural amino acids, and providing more binding possibilities for peptide drugs.
[0189] Furthermore, computational parameters are essential for macromolecular drug design software such as Rosetta to perform molecular structure prediction and molecular docking. The computational parameters for a monomer include at least one of the following parameters, for example:
[0190] (1) Names of non-natural amino acids (can be customized or named according to chemical nomenclature);
[0191] (2) All atom names in non-natural amino acids, and corresponding atom type information and atom charge information; in order to facilitate the calculation software to call parameter information, for example, the atom name and atom type can be defined according to the standard classification and naming rules of protein main chain atoms in software such as GROMOS or OPLS-AA.
[0192] (3) Information on the interatomic bonds in non-natural amino acids;
[0193] (4) Structural property information of non-natural amino acids, such as conformational information (L-type or D-type) and structural information (α-amino acid or β-amino acid);
[0194] (5) Information about neighboring atoms of non-natural amino acids, such as: neighboring atom names, distance from the neighboring atom to the farthest heavy atom;
[0195] (6) Information on the first atom name of the side chain of non-natural amino acids;
[0196] (7) Atomic information and hybridization information of movable dihedrals in non-natural amino acids, such as SP3 or SP2;
[0197] (8) Coordinate information of the dihedral angle in non-natural amino acids.
[0198] More preferably, the calculation parameters for non-natural amino acids also include the different rotation angles of different rotational isomers of the non-natural amino acid and the total energy of the non-natural amino acid molecule corresponding to each rotational isomer. The total energy is the rotational energy of the single bonds in the rotational isomer, which is the torsional strain energy generated by the resistance to rotation of single bonds in the molecule. When a molecule rotates around a single bond, it needs to overcome the repulsive forces between groups; this energy barrier determines whether rotational isomers can interconvert. The lower the total energy, the more stable the molecule.
[0199] Optionally, this application can extract the corresponding calculation parameters for each target amino acid molecule according to the above steps, and then collect and store them to form a corresponding NCCA calculation parameter database. Furthermore, the database can be updated in real time or periodically to obtain a more complete and accurate NCCA calculation parameter database for future use.
[0200] As this example illustrates, the method for generating non-natural amino acids in this application obtains a stable stereoconformation of the dipeptide protecting NCCA, thereby extracting more accurate computational parameters for NCAA. These parameters can then be used to achieve de novo design of large NCAA-containing drugs (e.g., peptide drugs), enabling computational analysis such as protein / peptide / antibody docking and mutation. This helps computational software obtain more reliable molecular simulation results, reduces experimental costs, and increases the success rate of drug design. Furthermore, the NCAA generation method and parameter extraction described in this application are universally applicable, capable of generating NCAA monomers from any small molecule structure and further obtaining the corresponding NCAA computational parameters, thus meeting the computational needs of complex peptide / protein / antibody drugs.
[0201] The application of the method for generating non-natural amino acids in this application will be described below with reference to more specific embodiments.
[0202] Example 1
[0203] In this Example 1, we introduce the generation of the corresponding non-natural amino acid based on molecule 1, and the corresponding calculation parameters.
[0204] 1. Obtain molecule 1, whose structural formula is shown in Formula 1.1 below.
[0205] 2. The molecular weight of molecule 1 is less than the preset threshold of 200 Da, and molecule 1 is a cyclic molecule with a branched chain. The target site is determined to be the methyl group at the end of the branched chain.
[0206] 3. At the target site of molecule 1, i.e., the methyl group at the end of the branched chain, an amino acid characteristic group is added in order to replace the hydrogen atom, generating the corresponding non-natural amino acid molecule, the structural formula of which is shown in Formula 1.2 below.
[0207] 4. The non-natural amino acid molecule 1.2 is end-capped by adding ACE (-C=OCH3) to the N-terminus and NME (-NHCH3) to the C-terminus to expand it into the corresponding dipeptide compound, the structural formula of which is shown in Formula 1.3 below.
[0208] 5. An initial stereoconformation of dipeptide compound 1.3 was generated, and after energy minimization calculations, the L-shaped conformation was selected as the target stereoconformation, as shown below. Figure 6 Conformation A is shown in .
[0209] 6. Calculation parameters for extracting non-natural amino acid molecules 1.2 from target stereoform A. The calculation parameters include those shown in Tables 1 to 10.
[0210] (1) The names of non-natural amino acid molecules 1.2 are shown in Table 2.
[0211] Table 2
[0212]
[0213] (2) The atomic names, atomic types and atomic charge information contained in non-natural amino acid molecules 1.2 are shown in Table 3.
[0214] Table 3
[0215]
[0216] (3) Information on the bonding bonds of monomer atoms in non-natural amino acid molecules 1.2 is shown in Table 4.
[0217] Table 4
[0218]
[0219] (4) Structural properties of non-natural amino acid molecules 1.2 are shown in Table 5.
[0220] Table 5
[0221]
[0222] In this context, PROTEIN indicates that the molecule is an amino acid of a protein, ALPHA_AA indicates that the amino acid is an alpha-type amino acid, and L_AA indicates that the amino acid is an L-type amino acid.
[0223] (5) Information on neighboring atoms of non-natural amino acid molecules 1.2 is shown in Table 6.
[0224] Table 6
[0225]
[0226] NBR_ATOM represents "nearest atom," which is the atom closest to the geometric center of the non-natural amino acid molecule; in Example 1, this atom is C-beta, denoted by CB. NBR_RADIUS represents the radius of gyration of the non-natural amino acid molecule, which is used to define the overall size of the non-natural amino acid molecule, and is measured in angstroms.
[0227] (6) Information on the first atom name of the side chain of non-natural amino acid molecule 1.2 is shown in Table 7.
[0228] Table 7
[0229]
[0230] (7) The atomic information and hybridization information of the movable dihedral angle of the non-natural amino acid molecule 1.2 are shown in Table 8.
[0231] Table 8
[0232]
[0233] The first line records information about the first dihedral angle. "CHI" indicates a dihedral angle, "1" indicates this line records information about the first dihedral angle, and "N, CA, CB, CG" indicates the first dihedral angle is composed of the amino N atom of a non-natural amino acid and the C-alpha, C-beta, and C-gamma atoms of the non-natural amino acid. The second line records information about the second dihedral angle. "CHI" indicates a dihedral angle, "2" indicates this line records information about the second dihedral angle, and "CA, CB, CG, CD2" indicates the second dihedral angle is composed of the C-alpha, C-beta, C-gamma, and second C-delta atoms of the non-natural amino acid. The third line records summary information about the dihedral angles, where "2, 3, 2" indicates the non-natural amino acid contains a total of two dihedral angles, with the first dihedral angle being sp3 hybridized and the second dihedral angle being sp2 hybridized.
[0234] (8) The dihedral coordinates (ICOOR_INTERNA) of non-natural amino acid molecules 1.2 are shown in Table 9.
[0235] Table 9
[0236]
[0237] Here, `parent atom` specifies the starting atom of the dihedral angle, `child atom` is the atom closest to and bonded to `parent atom`, `angle atom` is the atom located in the middle of the dihedral angle, and `torsion atom` is the ending atom of the dihedral angle. `Phi Angle` represents the dihedral angle of the `child`, `parent`, `angle`, and `torsion` atoms, `Theta` represents the improper angle (non-normal dihedral angle), and `Distance` represents the distance between the `child` and `parent` atoms.
[0238] (9) Non-natural amino acid molecules 1.2 Rotation angle and energy information of different rotational isomers, as shown in Table 10, taking the skeleton dihedral angles PHI (between amino N and C alpha) and PSI (between C alpha and carbonyl C) as examples with values of -170° and -170°.
[0239] Table 10
[0240]
[0241] In this diagram, the first column (UNK) represents non-natural amino acid molecules; the second column represents the values of the backbone dihedral angles PHI and PSI; the third column represents the sampling angle markers for the SP3 hybridized and SP2 hybridized side chain dihedral angles; and the fourth column represents the probability of the side chain appearing at the next angle within the sampling swing angle range. Each sampling angle marker has a corresponding preset sampling angle. For SP3 hybridized side chain dihedral angles, the preset sampling angles include -60°, 60°, and 180° (3 options); for SP2 hybridized side chain dihedral angles, the preset sampling angles include -180° and 180° (2 options).
[0242] Taking the first row of Table 10 as an example, it records that when the dihedral angles PHI and PSI of the non-natural amino acid skeleton are -170° and -170° respectively, the dihedral angle of the SP3 hybridized side chain is selected at the third sampling angle (180°), and the dihedral angle of the SP2 hybridized side chain is selected at the second sampling angle (180°). The probability of the side chain appearing at this angle is 0.301889 (calculated by Rosetta software). The probability is used to represent the score of total energy. A low probability means that the conformation is unlikely to exist.
[0243] Tables 2 to 9 above record the calculated parameters for non-natural amino acid molecules 1.2. These parameters can be applied to drug design calculations in computer design software such as Rosetta. The following examples can extract the parameters listed in Tables 2 to 9. The values of each parameter are determined based on actual conditions and will not be elaborated upon here.
[0244] Example 2
[0245] In this Example 2, the process of generating the corresponding non-natural amino acid based on molecule 2 and obtaining the corresponding calculation parameters is introduced.
[0246] 1. Obtain molecule 2, whose structural formula is shown in Formula 2.1 below.
[0247] 2. Since the molecular weight of molecule 2 is less than the preset threshold of 200 Da, and molecule 2 is a straight-chain molecule, its target site is determined to be the methyl groups at the two ends of the straight chain. Based on the identical structure of the two chain ends, duplicates can be removed, and only one chain end can be selected as the target site for subsequent steps.
[0248] 3. Add an amino acid characteristic group to the methyl group at the end of molecule 2 in order to replace the hydrogen atom, and generate the corresponding non-natural amino acid molecule, the structural formula of which is shown in Formula 2.2.
[0249] 4. The non-natural amino acid molecule 2.2 was end-capped by adding ACE (-C=OCH3) to the N-terminus and NME (-NHCH3) to the C-terminus to expand it into the corresponding dipeptide compound, the structural formula of which is shown in Formula 2.3 below.
[0250] 5. The initial stereoconformation of dipeptide compound 2.3 was generated. The L-shaped conformation was selected as the target stereoconformation, and after energy minimization calculations, the target stereoconformation was obtained as follows: Figure 6 The conformation B is shown in the figure.
[0251] 6. Calculation parameters for extracting non-natural amino acid molecules 2.2 from target stereoform B. The calculation parameters include the corresponding parameters of the nine types in Example 1 above, which will not be repeated here.
[0252]
[0253] Example 3
[0254] In this Example 3, the process of generating the corresponding non-natural amino acid based on molecule 3 and obtaining the corresponding calculation parameters is introduced.
[0255] 1. Obtain molecule 3, whose structural formula is shown in Formula 3.1 below.
[0256] 2. The molecular weight of molecule 3 is less than the preset threshold of 200 Da, and molecule 3 is a branched molecule, so its target site is determined to be located at the end of the chain. After deduplication, molecule 3 has two different chain ends, which are used as target sites.
[0257] 3. At the methyl end of one chain of molecule 3, an amino acid characteristic group is added in a manner that replaces a hydrogen atom to generate the corresponding non-natural amino acid molecule, the structural formula of which is shown in Formula 3.2 below; at the methyl end of the other chain of molecule 3, an amino acid characteristic group is added in a manner that replaces a hydrogen atom to generate the corresponding non-natural amino acid molecule, the structural formula of which is shown in Formula 3.3 below.
[0258] 4. The two non-natural amino acid molecules, formulas 3.2 and 3.3, from step 3 above are capped, i.e., ACE (-C=OCH3) is added to the N-terminus and NME (-NHCH3) is added to the C-terminus to expand them into corresponding dipeptide compounds. The structural formulas of the two dipeptide compounds are shown in formulas 3.4 and 3.5 below, respectively.
[0259] 5. Generate the initial stereoconformation corresponding to dipeptide compound 3.4, select the L-shaped conformation as the target stereoconformation, and perform energy minimization calculations to obtain the target stereoconformation (e.g., Figure 6 (As shown in conformation C). Since the side chain of dipeptide compound 3.5 is chiral, it can generate two L-type dipeptide conformations. After minimizing the energy of the two L-type dipeptide conformations, the dipeptide conformation with the lowest energy is selected as the corresponding target stereoconformation (e.g., ...). Figure 6 (As shown in conformation D).
[0260] 6. Calculation parameters for non-natural amino acid 3.2 are extracted from target stereoform C, and calculation parameters for non-natural amino acid 3.3 are extracted from target stereoform D. The calculation parameters include the corresponding parameters of the nine types described in Example 1 above, which will not be elaborated upon here.
[0261]
[0262] Example 4
[0263] In this Example 4, the process of generating the corresponding non-natural amino acid based on molecule 4 and obtaining the corresponding calculation parameters is introduced.
[0264] 1. Obtain molecule 4, whose structural formula is shown in Formula 4.1 below.
[0265] 2. The molecular weight of molecule 4 is less than the preset threshold of 200 Da, and molecule 4 is an unbranched cyclic molecule. Its target sites are determined to be located at each methylene group on the ring. After deduplication, molecule 4 has two distinct target sites on the ring.
[0266] 3. At the methylene position on one of the rings of molecule 4, an amino acid characteristic group is added in a manner that replaces a hydrogen atom to generate the corresponding non-natural amino acid molecule, the structural formula of which is shown in Formula 4.2 below; at the other methylene position on the ring of molecule 4, an amino acid characteristic group is added in a manner that replaces a hydrogen atom to generate the corresponding non-natural amino acid molecule, the structural formula of which is shown in Formula 4.3 below.
[0267] 4. Non-natural amino acids 4.2 and 4.3 were capped by adding ACE (-C=OCH3) to their N-terminus and NME (-NHCH3) to their C-terminus to expand them into corresponding dipeptide compounds. The structural formula of the dipeptide compound corresponding to non-natural amino acid 4.2 is shown in Formula 4.4, and the structural formula of the dipeptide compound corresponding to non-natural amino acid 4.3 is shown in Formula 4.5.
[0268] 5. Initial stereoconformities of dipeptide compound 4.4 are generated. Since the carbon atom at the target site on the ring is chiral, two L-shaped initial stereoconformities are generated. After minimizing the energy of each of the two L-shaped initial stereoconformities, the L-shaped conformation with the lowest energy is selected as the target stereoconformation for dipeptide compound 4.4, as follows: Figure 6 The conformation E is shown in the diagram. Similarly, since the target carbon atom on the ring of dipeptide compound 4.5 is chiral, two initial L-shaped stereoconformations are generated. After minimizing the energy of each of the two initial L-shaped stereoconformations, the L-shaped conformation with the lowest energy is selected as the target stereoconformation for dipeptide compound 4.5, as shown in the diagram. Figure 6 The conformation F in the figure is shown.
[0269] 6. Calculation parameters for non-natural amino acid 4.2 are extracted from target stereoform E, and calculation parameters for non-natural amino acid 4.3 are extracted from target stereoform F. The calculation parameters include the corresponding parameters of the nine types described in Example 1 above, which will not be elaborated upon here.
[0270]
[0271] Example 5
[0272] In this Example 5, the process of generating the corresponding non-natural amino acid based on molecule 5 and obtaining the corresponding calculation parameters is introduced.
[0273] 1. Obtain molecule 5, whose structural formula is shown in Formula 5.0 below.
[0274] 2. The molecular weight of molecule 5 is greater than the preset threshold of 200 Da. The active site of molecule 5 is successively cleaved until the molecular weight of each molecular fragment is less than 200 Da, resulting in a total of 4 molecular fragments. Their structural formulas are shown in Equations 5.1 to 5.4, respectively. The cleavage site of each molecular fragment is the target site for bonding characteristic amino acid groups.
[0275] 3. At each cleavage site of molecular fragments 5.1 to 5.4, amino acid characteristic groups are added in a manner that replaces hydrogen atoms to generate corresponding non-natural amino acid molecules 5.5 to 5.8, respectively, with the structural formulas shown in Formulas 5.5 to 5.8.
[0276] 4. Non-natural amino acid molecules 5.5 to 5.8 were capped by adding ACE (-C=OCH3) to their N-terminus and NME (-NHCH3) to their C-terminus, thus expanding them into corresponding dipeptide compounds with structural formulas as shown in formulas 5.9 to 5.12.
[0277] 5. Initial stereoforms for dipeptide compounds 5.9 to 5.12 were generated. Since the side chains of dipeptide compounds 5.9 to 5.12 are all chiral, each dipeptide compound can generate two initial stereoforms: L-type and D-type. The corresponding L-type conformation was selected, and energy minimization calculations were performed to obtain the four corresponding target stereo structures shown below. Figure 6 The conformations G, H, I, and J are shown in the figure.
[0278] 6. Calculation parameters of 5.5 to 5.8 corresponding to the non-natural amino acid molecules were extracted from the target stereoforms G to J, respectively. The calculation parameters include the corresponding parameters of the nine types in Example 1 above, which will not be described in detail here.
[0279]
[0280] Example 6
[0281] In this Example 6, the process of generating the corresponding non-natural amino acid based on molecule 6 and obtaining the corresponding calculation parameters is introduced.
[0282] 1. Obtain molecule 6, whose structural formula is shown in Formula 6.1 below.
[0283] 2. Since the molecular weight of molecule 6 is greater than the preset threshold of 200 Da, one active site is selected as the cleavage site for each cleavage. The active site of molecule 6 is cleaved successively until the molecular weight of each fragment is less than 200 Da, resulting in a total of four molecular fragments. The structural formulas of each fragment are shown in Equations 6.2 to 6.5 below. The cleavage site of each molecular fragment is the target site for connecting the characteristic amino acid groups. Since the number of heavy atoms in molecular fragment 6.5 is less than 5, it does not possess drug-like properties and can be directly filtered and deleted without proceeding to the next step to generate non-natural amino acids.
[0284] 3. At each target site in molecular fragments 6.2 to 6.4, amino acid characteristic groups are added by replacing hydrogen atoms to generate corresponding non-natural amino acid molecules, with structural formulas shown in formulas 6.6 to 6.8. Since the number of dihedral angles in non-natural amino acid molecule 6.6 is greater than 4, it is filtered and deleted according to the preset screening rules, and no further step is required.
[0285] 4. Non-natural amino acid molecules 6.7 and 6.8 were respectively capped by adding ACE (-C=OCH3) to their N-terminus and NME (-NHCH3) to their C-terminus to expand them into corresponding dipeptide compounds, the structural formulas of which are shown in Equations 6.9 and 6.10 below.
[0286] 5. Initial stereostructures of dipeptide compounds 6.9 and 6.10 were generated. Since the side chains of dipeptide compounds 6.9 and 6.10 are chiral, each dipeptide compound can generate two initial stereostructures: L-type and D-type. The corresponding L-type conformation was selected, and energy minimization calculations were performed to obtain the two corresponding target stereostructures shown below. Their structures are as follows: Figure 6 The conformations K and L are shown in the figure.
[0287] 6. Calculation parameters for non-natural amino acid molecules 6.7 are extracted from the target stereoform K, and calculation parameters for non-natural amino acid molecules 6.8 are extracted from the target stereoform L. The calculation parameters include the corresponding parameters of the nine types described in Example 1 above, which will not be elaborated upon here.
[0288]
[0289] An embodiment of this application also provides a polypeptide comprising non-natural amino acids generated by the generation methods of any of the above embodiments.
[0290] The non-natural amino acids generated in this application, or the polypeptides containing non-natural amino acids, can be used in the preparation of pharmaceuticals.
[0291] Corresponding to the aforementioned application function implementation method embodiments, this application also provides a non-natural amino acid generation device, electronic device, and corresponding embodiments.
[0292] Figure 7 This is a schematic diagram of the structure of the apparatus for generating non-natural amino acids as shown in the embodiments of this application.
[0293] See Figure 7 An embodiment of this application illustrates an apparatus for generating non-natural amino acids, which includes a site determination module 1, an amino acid generation module 2, and an amino acid screening module 3. Wherein:
[0294] The site determination module 1 is used to determine at least one candidate target site in the target molecule, wherein the target molecule is a complete molecule or molecular fragment with a molecular weight less than or equal to a first preset molecular weight.
[0295] The amino acid generation module 2 is used to determine at least one target site when there is only one candidate target site, and to bond amino acid characteristic groups at the target site to generate the corresponding candidate amino acid molecule; when there are multiple candidate target sites, at least one target site is determined among the multiple candidate target sites, and amino acid characteristic groups are bonded at at least one target site to generate at least one corresponding candidate amino acid molecule.
[0296] The amino acid screening module 3 is used to screen the candidate amino acid molecules obtained in step S2 according to the preset amino acid screening rules to obtain the target amino acid molecules, which are non-natural amino acids.
[0297] In some specific implementations, the site determination module 1 is used to determine one or more target sites from multiple candidate target sites according to preset rules, wherein the preset rules are to select some or all of the candidate target sites as target sites.
[0298] In some specific implementations, the site determination module 1 is used to select candidate target sites as any heavy atom in the target molecule that is connected to a hydrogen atom when the target molecule is a complete molecule; and to select candidate target sites as heavy atoms with free radicals in the target molecule when the target molecule is a molecular fragment, or to form a complete molecule by adding hydrogen to the molecular fragment, and then select candidate target sites as any heavy atom in the complete molecule that is connected to a hydrogen atom.
[0299] In some specific implementations, the site determination module 1 is used to select the target site from heavy atoms with hydrogen atoms attached to the ends of each chain when the target molecule is a complete chain molecule.
[0300] In some specific embodiments, the site determination module 1 is used to determine the following when the target molecule is a complete cyclic molecule: if the complete cyclic molecule does not contain branches, the candidate target site is selected from the heavy atoms of the ring group of the target molecule that are connected to hydrogen atoms; if the complete cyclic molecule contains branches and the ring group is located at the end of the chain, the candidate target site is selected from the heavy atoms of the ring group of the target molecule that are connected to hydrogen atoms and / or the heavy atoms of each chain end of the target molecule that are connected to hydrogen atoms; if the complete cyclic molecule contains branches and the ring group is located in the middle of the chain, the candidate target site is selected from the heavy atoms of each chain end of the target molecule that are connected to hydrogen atoms.
[0301] See Figure 8 In some embodiments, the generating apparatus further includes a molecule acquisition module 4 for acquiring candidate molecules. The molecule acquisition module is also used to acquire multiple candidate molecules to be screened.
[0302] In some embodiments, the generating apparatus further includes a molecular weight determination module 5, used to determine the size of the molecular weight of the candidate molecule and the first preset molecular weight, to obtain the target molecule and / or the molecule to be processed, wherein the molecule to be processed is a candidate molecule fragment or a candidate complete molecule with a molecular weight greater than the first preset molecular weight and less than the second preset molecular weight.
[0303] In some embodiments, the generating apparatus further includes a molecular segmentation module 6, used to segment the molecule to be processed at least once to obtain molecular fragments. In some specific embodiments, the molecular segmentation module is used to segment the molecule to be processed at least once according to steps P31 to P33 to obtain molecular fragments.
[0304] The generating device also includes a first screening module 7, which is used to screen multiple candidate molecules to be screened according to a first preset screening rule, and obtain one or more candidate molecules that meet the first preset screening rule.
[0305] The generating device also includes a second screening module 8, which is used to screen molecular fragments according to a second preset screening rule to obtain molecular fragments that conform to the second preset screening rule.
[0306] In some embodiments, the generating apparatus of this application further includes a parameter generating module 9, used to bond an acetyl group (-C(O)CH3) to the N-terminus of the target amino acid molecule and an N-methyl group (-NCH3) to the C-terminus to generate the corresponding dipeptide compound; obtain the target stereoconformation corresponding to the dipeptide compound; extract and store the calculated parameters of the target amino acid molecule from the target stereoconformation.
[0307] This example demonstrates that the non-natural amino acid generation device of this application, by bonding amino acid characteristic groups to target sites of different types of target molecules and filtering them through various screening rules, yields a large number of novel and high-quality non-natural amino acids. This facilitates the selection of NCCAs in peptide drugs and provides more potential options for the design of NCCA-containing peptide drugs. Furthermore, by pre-extracting and saving the calculation parameters of various NCCAs, accurate NCCA structural parameters are provided, meeting the computational needs of drug design in computer design software such as Rosetta. This helps obtain more reliable molecular simulation results, reduces wet experiment costs, and improves the success rate of drug design.
[0308] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated further here.
[0309] Figure 9 This is a schematic diagram of the structure of an electronic device shown in an embodiment of this application.
[0310] See Figure 9 The electronic device 1000 includes a memory 1010 and a processor 1020.
[0311] The processor 1020 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0312] Memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. ROM may store static data or instructions required by processor 1020 or other modules of the computer. Permanent storage devices may be read-write storage devices. Permanent storage devices may be non-volatile storage devices that retain stored instructions and data even when the computer is powered off. In some embodiments, permanent storage devices use mass storage devices (e.g., magnetic or optical disks, flash memory) as permanent storage devices. In other embodiments, permanent storage devices may be removable storage devices (e.g., floppy disks, optical drives). System memory may be a read-write storage device or a volatile read-write storage device, such as dynamic random access memory. System memory may store some or all of the instructions and data required by the processor during operation. Furthermore, memory 1010 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and disks and / or optical disks may also be used. In some embodiments, memory 1010 may include a removable storage device that is readable and / or writable, such as a laser disc (CD), a read-only digital multifunction optical disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, an ultra-high density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not contain carrier waves or transient electronic signals transmitted wirelessly or via wired connections.
[0313] The memory 1010 stores executable code, which, when processed by the processor 1020, can cause the processor 1020 to execute part or all of the methods described above.
[0314] Furthermore, the method according to this application can also be implemented as a computer program or computer program product, which includes computer program code instructions for performing some or all of the steps in the method described above.
[0315] Alternatively, this application may be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium) storing executable code (or computer program or computer instruction code) that, when executed by a processor of an electronic device (or server, etc.), causes the processor to perform part or all of the steps of the methods described above according to this application.
[0316] This application also provides a computer program product, which includes computer instructions that, when executed by a processor, implement the method described above.
[0317] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for generating non-natural amino acids, characterized in that, include: Step S1: Determine at least one candidate target site in the target molecule, wherein the target molecule is a complete molecule or molecular fragment with a molecular weight less than or equal to a first preset molecular weight; Step S2: When there is only one candidate target site, the candidate target site is the target site. The amino acid characteristic group is bonded to the target site to generate the corresponding candidate amino acid molecule. When there are multiple candidate target sites, at least one target site is determined from the multiple candidate target sites, and amino acid characteristic groups are bonded to the at least one target site to generate at least one corresponding candidate amino acid molecule; The amino acid characteristic group is the group shown in Formula I. " "The bond used to identify the characteristic group of the amino acid and the target site; Step S3: Screen the candidate amino acid molecules obtained in step S2 according to preset amino acid screening rules to obtain target amino acid molecules, wherein the target amino acid molecules are non-natural amino acids; wherein the preset amino acid screening rules include at least one of the following rules: (1) The synthesis difficulty of the target amino acid molecule meets the preset conditions; (2) The number of dihedral angles in the target amino acid molecule is ≤4; (3) The target amino acid molecule is a non-natural amino acid; (4) Any two target amino acid molecules selected are not the same; Before step S1, the generation method further includes steps P1 and P2; after step P2 and / or before step S1, the generation method further includes step P3; wherein: Step P1 is: obtaining candidate molecules; Step P2 is: determining the molecular weight of the candidate molecule and the size of the first preset molecular weight to obtain the target molecule and / or the molecule to be processed; wherein, the molecule to be processed is a candidate molecule fragment or a candidate complete molecule with a molecular weight greater than the first preset molecular weight and less than or equal to the second preset molecular weight. Step P3 is: to perform at least one segmentation on the molecule to be processed to obtain molecular fragments; wherein, step P3 specifically includes: step P31, performing one segmentation on the molecule to be processed to obtain multiple candidate molecular fragments; step P32, using the multiple candidate molecular fragments as multiple candidate molecules to perform step P2 to obtain at least one target molecule and / or molecule to be processed; step P33, performing steps P31 and P32 on at least one molecule to be processed in step P32 respectively, until a preset termination condition is reached; wherein, the preset termination condition is that only the target molecule is obtained after the most recent execution of step P32, and no molecule to be processed is obtained, and / or steps P31 and P32 are executed a preset number of times.
2. The generation method according to claim 1, characterized in that: When the target molecule is a complete molecule, the candidate target site is any heavy atom in the target molecule that is connected to a hydrogen atom; When the target molecule is a molecular fragment, the candidate target site is a heavy atom with a free radical in the target molecule, or the molecular fragment is hydrogenated to form a complete molecule, and then the candidate target site is any heavy atom in the complete molecule that is connected to a hydrogen atom.
3. The generation method according to claim 2, characterized in that, The complete molecule is a complete chain molecule or a complete ring-containing molecule; When the target molecule is a complete chain molecule, the target site is selected from heavy atoms with hydrogen atoms attached to the ends of each chain; When the target molecule is a complete cyclic molecule If the complete cyclic molecule does not contain branches, the candidate target site is selected from the heavy atom on the cyclic group of the target molecule that is attached to a hydrogen atom; If the complete cyclic molecule contains branches and the cyclic group is located at the end of the chain, then the candidate target site is selected from the heavy atoms of the cyclic group of the target molecule that are connected to hydrogen atoms and / or the heavy atoms of each chain end of the target molecule that are connected to hydrogen atoms. If the complete cyclic molecule contains branched chains and the cyclic group is located in the middle of the chain, then the candidate target sites are selected from the heavy atoms at the ends of each chain of the target molecule that are connected to hydrogen atoms.
4. The generation method according to claim 3, characterized in that, The heavy atom connecting the hydrogen atom is a carbon atom; or Determining at least one target site among the plurality of candidate target sites includes: One or more target sites are determined from the plurality of candidate target sites according to a preset rule, wherein the preset rule is to select some or all of the candidate target sites as target sites.
5. The generation method according to claim 1, characterized in that, The cleavage described in step P31 occurs at the cleavage site of the molecule to be treated; the cleavage site is a designated site, an active site, and / or a retrosynthetic site; and / or Step P31 involves performing a single segmentation at the segmentation site of the molecule to be processed according to a preset segmentation rule to obtain multiple candidate molecular fragments.
6. The generation method according to claim 5, characterized in that, The number of cleavage sites for each molecule to be processed is N. When N is an integer greater than 1, for each molecule to be processed, cleavage is performed using one of the following preset cleavage rules, which includes: (1) Perform a single cut at one of the cut sites of the molecule to be processed, either randomly or by designation, to obtain a total of two candidate molecule fragments; (2) The N cleavage sites of the molecule to be treated are cleaved once, and each cleavage yields two corresponding candidate molecular fragments, for a total of 2N candidate molecular fragments; (3) Select M cleavage sites on the molecule to be processed randomly or by designation for cleavage. Each cleavage yields two candidate molecular fragments, for a total of 2M candidate molecular fragments, where 1 < M < N.
7. The generation method according to claim 1, characterized in that, Prior to step P1, the method further includes: Step P0: Obtain multiple candidate molecules to be screened; Step P1 specifically includes: screening the plurality of candidate molecules to be screened according to the first preset screening rule to obtain one or more candidate molecules.
8. The generation method according to claim 7, characterized in that, The first preset filtering rule includes at least one of the following rules: (1) The molecule has a standardized format; (2) The atoms in the molecule do not contain any atoms other than H, C, N, O, F, Cl, S, P, Br and I; (3) The molecular weight of the molecule is ≤800 Da; (4) The molecule does not contain radioactive isotope atoms; (5) The molecule is a non-ionized molecule; (6) The molecule conforms to at least one of the following rules: Drug Class 5, Drug Class 4, or Drug Class 3; (7) The molecule has a complete structure.
9. The generation method according to claim 1, characterized in that, Following step P3, the method further includes: Step P4: Screen the molecular fragments according to the second preset screening rule to obtain molecular fragments that meet the second preset screening rule, and use them as the target molecules in step S1; wherein, the second preset screening rule includes at least one of the following rules: (1) The number of heavy atoms in the molecular fragment is 5 to 50; (2) The molecular fragment conforms to at least one of the following rules: Drug Class 5, Drug Class 4, or Drug Class 3; (3) The molecular fragment is not any of the following fragments: -CH3, -CH(CH3)2, -CH2CH(CH3)2, -CH(CH3)(CH2CH3), -CH2CH2SCH3, -CH2-O-O-O, 、 、-CH2OH、-CHOH(CH3)、-CH2CONH2、-CH2CH2CONH2、-CH2SH、-CH2COOH、-CH2CH2COOH、-CH2CH2CH2CH2NH2、 -CH2CH2CH2NHC(=NH)NH2; (4) The molecular fragments that form a complete molecule are not any of the following: CH4, CH3CH2SCH3, toluene, p-hydroxytoluene , CH3OH, CH3CONH2, CH3CH2CONH2, CH3SH, CH3COOH, CH3CH2COOH, CH3CH2CH2CH2NH2, CH3CH2CH2NHC(=NH)NH2; (5) When the complete molecule formed by the molecular fragment is CH3CH2OH, CH3CH2CH3, CH3CH(CH3)2 or CH2(CH3)CH2CH3, the candidate target site of the complete molecule is the carbon atom located at the end of the chain.
10. The generation method according to any one of claims 1 to 9, characterized in that, After step S3, the method further includes: Step S4: An acetyl group (-C(O)CH3) is bonded to the N-terminus of the target amino acid molecule, and an N-methyl group (-NCH3) is bonded to the C-terminus to generate the corresponding dipeptide compound; Step S5: Obtain the target stereoconformation corresponding to the dipeptide compound; Step S6: Extract and store the calculation parameters of the target amino acid molecule from the target stereoconformation.
11. The generation method according to claim 10, characterized in that, Step S5 includes: Step S51: Obtain at least one initial stereoconformation of the dipeptide compound; Step S52: Select the target stereo configuration from each of the initial stereo configurations according to the preset configuration selection rules; the preset configuration selection rules include at least one of the following rules: (1) The side chains of the initial stereoformation have a single conformation; (2) It has the lowest energy among multiple initial stereoformations; (3) The initial stereotype is an L-shaped conformation.
12. An apparatus for generating non-natural amino acids, characterized in that, include: A site determination module is used to determine at least one candidate target site in a target molecule, wherein the target molecule is a complete molecule or molecular fragment with a molecular weight less than or equal to a first preset molecular weight. The amino acid generation module is used to bond amino acid characteristic groups to the target site when there is only one candidate target site, which is the target site, to generate the corresponding candidate amino acid molecule. When there are multiple candidate target sites, at least one target site is determined from the multiple candidate target sites, and amino acid characteristic groups are bonded to the at least one target site to generate at least one corresponding candidate amino acid molecule; The amino acid characteristic group is the group shown in Formula I. " "The bond used to identify the characteristic group of the amino acid and the target site; An amino acid screening module is used to screen candidate amino acid molecules according to preset amino acid screening rules to obtain target amino acid molecules, wherein the target amino acid molecules are non-natural amino acids; wherein the preset amino acid screening rules include at least one of the following rules: (1) The synthesis difficulty of the target amino acid molecule meets the preset conditions; (2) The number of dihedral angles in the target amino acid molecule is ≤4; (3) The target amino acid molecule is a non-natural amino acid; (4) Any two target amino acid molecules selected are not the same; The generating apparatus further includes a molecule acquisition module for acquiring candidate molecules; and / or The generating device further includes a molecular weight determination module, used to determine the molecular weight of the candidate molecule and its relationship to a first preset molecular weight, thereby obtaining the target molecule and / or the molecule to be processed. The molecule to be processed is a candidate molecule fragment or a candidate complete molecule with a molecular weight greater than the first preset molecular weight and less than a second preset molecular weight; and / or The generating device further includes a molecular segmentation module, used to segment the molecule to be processed at least once according to steps P31 to P33 to obtain molecular fragments; wherein, step P31, the molecule to be processed is segmented once to obtain multiple candidate molecular fragments; step P32, the multiple candidate molecular fragments are respectively used as multiple candidate molecules and judged by the molecular weight determination module to obtain at least one target molecule and / or molecule to be processed; step P33, steps P31 and P32 are executed on at least one molecule to be processed in step P32 respectively until a preset termination condition is reached; wherein, the preset termination condition is that only the target molecule is obtained after the most recent execution of step P32, but no molecule to be processed, and / or steps P31 and P32 are executed a preset number of times.
13. The generating apparatus according to claim 12, characterized in that: The generating device further includes a first screening module, used to screen multiple candidate molecules to be screened according to a first preset screening rule, to obtain one or more candidate molecules that meet the first preset screening rule; and / or The generating device further includes a second screening module, used to screen the molecular fragments according to a second preset screening rule to obtain molecular fragments that conform to the second preset screening rule.
14. An electronic device, characterized in that, include: processor; as well as A memory having executable code stored thereon, which, when executed by the processor, causes the processor to perform the generation method as described in any one of claims 1-11.
15. A computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the generation method as described in any one of claims 1-11.
16. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a processor, implement the generation method according to any one of claims 1-11.
Citation Information
Patent Citations
Molecular generation method and device, equipment and storage medium
CN114171134A
Cyclic peptide design method, compound structure generation method, device and electronic equipment
CN114333985A