De novo designed macrocyclic oligoamides
Patent Information
- Application Number
- JP2024541614
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-26
- Filing Date
- 2023-01-12
- Publication Date
- 2026-01-19
AI Technical Summary
The prior art is difficult to quickly and comprehensively search and discover compounds with novel biological activities, especially in cases where there are many different skeletons, and it is difficult to randomly explore the skeleton torsion angles and design arrays.
Using the universal cyclic oligoamide designed by Denovo, low-energy sorted cyclic oligoamides were generated by computational methods, using the computational potential energy of AC-BPRO-NHME and AC-Ben2-Nme2, combined with the six-dimensional conversion of terminal amides, the potential universal cyclic compounds were identified, and their chemical diversity was verified by X-ray and NMR structures.
The rapid discovery and design of universal ring compounds is achieved, the efficiency and accuracy of the discovery of compound biological activity is improved, and the search process of large-scale chemical space is simplified.
Smart Images

Figure 00000174_0000 
Figure 00000174_0001 
Figure 00000174_0002
Abstract
Description
[Technical field]
[0001] cross reference This application claims priority to U.S. Provisional Application No. 63 / 299,633, filed January 14, 2022, and U.S. Provisional Application No. 63 / 369,441, filed July 26, 2022, both of which are incorporated by reference in their entireties. [Background technology]
[0002] A computational method that allows rapid and exhaustive searching of the space of possible macrocycles would greatly facilitate the discovery of novel biologically active compounds; however, no such method currently exists. When many different backbones are involved, the random search for combinations of backbone torsion angles and subsequent sequence design is difficult. [Brief description of the drawings]
[0003] [Figure 1]Overview of the macrocycle discovery approach. (A) Chemical structures of the monomeric building blocks used to construct small drug-like macrocycles: amino acids (a, b, g, d, e), aminobenzoates (c, f, j), aminomethylbenzoates (h, k, n), aminophenylacetic acids (i, l, o), aminomethylphenylacetic acids (m, p, q), oxazolidines / thiazolidines (r), oxazoles / thiazoles (t), thioethers (s), aminomethylpicolinates (u), and triazoles (v). Backbone atoms are colored to highlight differences between the building blocks. (B) Starting from the calculated potential energy surfaces of Ac-Bpro-NHMe and Ac-Ben2-NMe2, construction and identification of low-energy conformers of cyclo-(Bpro-Ben2-Bpro-Ben2). Combinations of these low-energy conformers provide access to thousands of Ac-Bpro-Ben2-NMe2 dimers. For each dimer, we calculate the 6D transformation between the planes defined by the terminal amides, then compartmentalize. We identify all pairs of dimers in close proximity to each other by quick lookup, and minimize the resulting ensemble of macrocycle conformers. [Diagram 2] Chemical and structural diversity of small de novo macrocycles. (A) Distribution of unique sequences sampled across all species of four-residue macrocycles. The 15 most abundant species are labeled. (B) Torsional diversity of sampled four-residue macrocycle species. Heatmap pixels represent four-residue macrocycle species generated from two-residue fragments on the x and y axes. The inset at the top right shows the full map across all 224 combinations of monomeric species; a 10x10 subset of this is expanded in the main panel for clarity. Two representative examples of torsional compartments are shown in boxes for two unique species. (C) Principal moments of inertia distribution (second row) as well as representative conformers (bottom) sampled by macrocycles built from different species (top row). [Diagram 3]The X-ray and NMR structures of the locally encoded macrocycles are very close to the designed models. Rows A-E, crystal structures. Column I, chemical structure; column II, designed model; column III: surface representation of the designed model; column IV: superposition of the designed model and experimentally determined coordinates; column V: chemical properties and apparent permeability; the number of hydrogen bond donors (HBDonor) and hydrogen bond acceptors (HBAcceptor) were determined using rdkit. ahah was modeled with phenylalanine and chemically synthesized with pyridylalanine (row E). Rows F-H, NMR structures. Column II shows the ROE used for distance-constrained ensemble generation instead of the designed model. Experimental structures to design model RMSD are reported for global backbone atoms for the X-ray structures and for ensemble average backbone coordinates for the NMR structures. [Figure 4] The X-ray crystal structures and NMR ensemble structures of the hydrogen-bond-containing macrocycles are very close to their design models. Columns I–V are the same as in Fig. 3. Column II, rows E–G: ROEs used to generate the distance-constrained ensemble. Arrows are the same as in Fig. 3. [Diagram 5] Potential energy surfaces have been calculated for 130 chemicals. The chemical species compartment for each chemical is listed next to its abbreviation. [Figure 6] Species distribution of unique sequences sampled across all four-residue macrocycles discovered in our expanded search (left), all three- and four-residue macrocycles in the Cambridge structural database (center), and all three- and four-residue macrocycles in the PubChem database (right). [Figure 7] Within a species, many conformers of many sequences have different backbone torsion compartments sampled non-uniformly. (A) Schematic of the torsion binning scheme. Each conformer of each sequence belonging to a given species (aaa in this example) is constructed and its backbone torsion angles are measured and partitioned into 60 degree intervals. The number of conformers in the 16 most densely sampled compartments is shown in the bar chart. (B) The most sampled 16 compartments for several 3- and 4-residue macrocycles. Non-uniform sampling is evident. [Figure 8] Approximately half of the unique macrocycles listed herein follow the rule of five. (A) Distribution of unique sequences that follow Lipinski's rule of five (Ro5) and do not (bRo5). (B) Distribution of macrocycle molecular weights and calculated atomic log(P). (C) Histogram showing the distribution of hydrogen bond acceptor atoms in each macrocycle. (D) Histogram showing the distribution of hydrogen bond donor atoms in each macrocycle. [Figure 9] Effect of side chain substitution on molecular shape. (A) Principal moments in all-atom inertia analysis of the minimum energy conformers for a small sequence of three species. Different preferences for PMI ratios are evident within each species. (B) Superposition of the minimum energy conformers of the macrocycles aaap-PRO-PHE_pp-mTIC-AMACBEN3 and aaap-PRO-mPHE_pp-mTIC-AMACBEN3. These diastereomers differ in the α-stereocenters in their phenylalanine monomers, yet their respective phenylalanine and Tic monomers are oriented in similar configurations. Epimerization occurs, resulting in macrocycles with very similar shapes, despite their different torsional sections. [Figure 10]The low-energy conformers identified from the hashed ensembles are very similar to the X-ray crystal structure of the macrocycle. The backbone atoms are highlighted in Figure 1. In each energy landscape, the RMSD is calculated for the lowest energy conformer sampled from the ensemble. Each dot represents the endpoint of the AIMNet minimization of a single conformer generated by the hash-based assembly. The predicted low-energy conformers for each macrocycle are represented as sticks. The crystallographic coordinates of each macrocycle are represented as sticks. Cyclotricosine contains two different conformers of the macrocycle in the asymmetric unit in its X-ray crystal structure. The conformers are mirror images of each other. The two conformers were sampled in the hashed ensemble and are nearly identical in energy. The two predicted conformers are superimposed on the asymmetric unit of cyclotricosine. [Figure 11] In non-minimized ensembles, the energy landscapes of macrocycles that sample only a few torsion compartments often contain deep energy minima that sample a single state with Pnear>0.8. Ensembles of three-residue macrocycles were minimized (1735 sequences from macrocycles sampling fewer than 5 compartments and 889 sequences from macrocycles sampling more than 100 compartments) and the quality of the resulting landscapes was assessed with Pnear. The energy landscapes of macrocycles in non-minimized ensembles that sample many torsion compartments often contain shallow energy minima that sample multiple states. [Figure 12]The NMR ensembles or X-ray crystal structures of the three designed macrocycles are not close to their predicted low-energy conformations. The designed models are shown as sticks and spheres. The 20 lowest energy conformers from unconstrained MD simulations are shown as lines. The X-ray crystal structures are shown as sticks. The distance-constrained ensemble of aaag contains three different conformers that are all different from the predicted low-energy conformer of the molecule. The distance-constrained ensemble of aabs contains a single conformer at the level of the backbone torsion angles, but this conformer is stabilized by a hydrogen bond between the carbonyl of the aminobutyric acid monomer and the NH of the cysteine-containing monomer. [Figure 13]Cis-amides are a common motif in four-residue macrocycles built primarily with α-amino acids. (A) Proportion of conformers containing only trans-amides or at least one cis-amide in each sequence within a given species. Cis-amides are present in many of the conformers of macrocycles sampled from species composed of three α-amino acids and one unnatural amino acid, particularly those built from backbones containing many fixed torsion angles (e.g., 3-aminobenzoic acid / f-containing macrocycles, or 4-aminomethylphenylacetic acid / q-containing macrocycles). Long unconstrained backbones (e.g., ε-amino acids (e)) and backbones with obvious turn-inducing motifs (e.g., 2-aminobenzoic acid (c) and 1,2- or 1,3-aminomethylphenylacetic acid (m,p, respectively)) relieve the strain associated with cis-amides, primarily with tertiary amides. (B) Ramachandran plot of all α-amino acids with cis-amides for the species in panel A. The dots represent N-alkylated amino acids (e.g., N-methylalanine); the dots represent proline-like amino acids; and the dots represent N-protoamino acids (e.g., alanine). These tertiary amides, whether proline-like residues, N-methylated residues, or peptoids, reside in four distinct regions of Ramachandran space; the same region of Ramachandran space is also sampled by N-protoamino acids, including cis-amides. Representative D-proline conformers in two of the four discrete compartments are represented by sticks. These conformers facilitate abrupt turns in the backbone of these macrocycles. These conformers are not exclusive to proline-like amino acids, but are sampled by most α-amino acids in these macrocycles. With cis amides, certain combinations of these Φ and Ψ torsion angles result in α-amino acid conformers that facilitate sharp turns of the backbone within a single residue; regions containing trans amides result in more extended α-amino acid conformers in the same region of Ramachandran space. [Figure 14] The aaam mimics the spacing and orientation of residues i, i+4 of an α-helix. X-ray crystal structure of an aaam (left) superimposed on an idealized α-helix (right). [Figure 15] PAMPA results for all macrocycles prepared. The dashed line indicates the position where Papp=1x10-6 cm / s. [Figure 16] Effect of hash table resolution on the quality of macrocycle closure in two-residue macrocycles. Hash tables were constructed for each monomer at high (0.5 Å, 10 degrees), medium (0.75 Å, 12.5 degrees), low (1.0 Å, 15 degrees), and very low (2.0 Å, 30 degrees) Cartesian (distance) and origin (rotation angle) resolutions. The resulting NCO alignment error for each conformer was calculated and plotted in a KDE-normalized histogram. The higher the hash table resolution, the smaller the RMSD error in the fragment alignment. [Figure 17] Identification of candidate low energy sequences for each torsion section sampled across basic residues in a species. (A) Schematic of the algorithm for identifying potential low energy sequences. (B) Example of the process used to identify the lowest energy rotamer for each monomer in a representative single torsion section of the macrocycle aaas-SAR_pp-SAR-AGLY-SUGA. First, the backbone torsion angles for the representative section are calculated. Then, for each monomer, backbone torsion angles are assumed for the representatives in the torsion section, and the lowest energy rotamer is identified for each possible substitution of that monomer. Finally, candidate sequences are constructed by identifying the substitutions that result in the lowest energy in each monomer. [Figure 18] Residuals from fitting linear regression models to predict AIMNet energies from sequence composition (top panel). Zscore analysis of all 8,280 macrocycles used to build the linear regression models (bottom panel). Macrocycles with Pnear>0.8 and Zscore<-2.0 were manually screened, and some were chemically synthesized. [Figure 19] Analytical ultra-performance liquid chromatography (UPLC) and liquid chromatography / mass spectrometry (LCMS) mass spectra of exemplary macrocyclic aas oligoamides. [Figure 20] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic akak oligoamides. [Figure 21] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic ahah oligoamides. [Figure 22] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aabc(1) oligoamides. [Figure 23] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aaam(1) oligoamides. [Figure 24] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aaam oligoamides. [Diagram 25] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aaap oligoamides. [Figure 26] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aaaq oligoamides. [Figure 27] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aabc(2) oligoamides. [Figure 28] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aalm oligoamides. [Figure 29] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aaam(4) oligoamides. [Diagram 30] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aaam(5) oligoamides. [Diagram 31] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aaam(6) oligoamides. [Diagram 32] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aaam(8) oligoamides. [Diagram 33] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aabs(1) oligoamides. [Diagram 34] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aaap(1) oligoamides. [Diagram 35] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aaap(2) oligoamides. [Diagram 36]Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aaap(3) oligoamides. [Figure 37] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic asbc(3) oligoamides. [Figure 38] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aabb(1) oligoamides. [Figure 39] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aaag(1) oligoamides. [Diagram 40] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aaag(4) oligoamides. [Diagram 41] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aabg(1) oligoamides. [Diagram 42] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aabi(2) oligoamides. [Diagram 43] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aabi(3) oligoamides. [Diagram 44] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aabi(4) oligoamides. [Diagram 45] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aabi(5) oligoamides. [Diagram 46] Analytical UPLC and LCMS mass spectra of exemplary macrocycle bbbb(3) oligoamides. [Figure 47] Analytical UPLC and LCMS mass spectra of exemplary macrocyclic aagb(1) oligoamides. [Figure 48] An example of a macrocycle calculation system is shown below. [Figure 49] 1 is a flow diagram of an exemplary process for generating a set of monomer conformations. [Figure 50] 1 is a flow diagram of an exemplary process for adaptively sampling the structural parameter space of a monomer. [Figure 51]An example of adaptively sampling the structural parameter space of a monomer to determine a conformational set of the monomer is shown. [Figure 52] 1 is a flow diagram of an exemplary process for generating data defining a set of molecular fragment conformations based on possible conformations of monomers in a monomer sequence of the molecular fragment. [Figure 53] An example of a macrocycle identification system is shown. [Figure 54] 1 is a flow diagram of an exemplary process for identifying macrocycles using discretized start-to-end transformations and discretized end-to-start transformations. [Figure 55] 1 is a flow diagram of an exemplary process for predicting whether a molecule capable of adopting a set of macrocycle conformations will adopt a rigid macrocycle conformation. [Figure 56] A first molecular fragment and a second molecular fragment are shown having a conformation that cooperates to form a closed circular loop. [Figure 57] A backbone superposition of a set of molecular fragment conformations is shown. [Figure 58] 1 shows a translation vector that defines the translation from the beginning of the molecular fragment to the end of the molecular fragment. [Figure 59] 1 shows a start-end conversion dictionary and a end-end-start conversion dictionary.Like reference numbers and names in the various drawings indicate like elements. Summary of the Invention
[0004] In a first aspect, the present disclosure provides 3-4 residue non-natural macrocyclic oligoamides comprising a 3- or 4-monomer residue species selected from the group consisting of monomers a, b, c, d, e, f, g, h, I, j, k, l, m, n, o, p, q, r, s, t, u, and v as defined in Table 1, or a salt thereof.
[0005] In one embodiment, (a) The monomer names a, b, g, d, and e are amino acids; where (i) Monomer designation a includes or consists of all L- and D-proteinogenic amino acids and all L- and D-nonstandard alpha amino acids; peptoids (N-alkylated glycines), and their optical isomers; (ii) monomer designation b includes or consists of all B3 amino acids containing L- or D-proteinogenic side chains, and optical isomers thereof; (iii) Monomer designation d comprises or consists of δ-amino acids and their optical isomers; (iv) Monomer designation g includes or consists of gamma amino acids and their optical isomers; and (v) Monomer designation e includes or consists of ε amino acid and its optical isomers; (b) Monomer designations c, f, and j include or consist of aminobenzoic acids and their optical isomers; (c) Monomer designations h, k, and n include or consist of aminomethylbenzoic acids and their optical isomers; (d) Monomer designations i, l, o include or consist of aminophenylacetic acids and their optical isomers; (e) The monomer designations m, p, and q include or consist of aminomethylphenylacetic acids and their optical isomers; (f) Monomer designation r includes or consists of oxazolidines and thiazolidines and their optical isomers, but includes all L- and D-proteinogenic side chains substituted at the second atom in the backbone for both oxazolidinones and thiazolidinones; (g) Monomer designation t includes or consists of oxazoles and thiazoles and their optical isomers, but includes all L- and D-proteinogenic side chains substituted at the second atom in the backbone for both oxazolidines and thiazolidinines; (h) Monomer designation s includes or consists of thioethers and their optical isomers; (i) The monomer name includes or consists of aminomethylpicolinic acids and their optical isomers; and (j) Monomer designation v comprises or consists of triazoles and their optical isomers.
[0006] In another embodiment, the macrocyclic oligoamide comprises a species selected from the group of species listed in Table 2, or circularly permuted forms thereof. In a further embodiment, the macrocyclic oligoamide comprises at least one monomer not included in the monomer designation a; or at least one residue not included in the monomer designations a, b, g, d, e. In a further embodiment, the macrocyclic oligoamide is membrane permeable. In various embodiments, the macrocyclic oligoamide comprises the structure of any macrocyclic oligoamide disclosed in any figure herein, the circularly permuted structure, or a salt thereof; or wherein the macrocyclic oligoamide comprises or consists of the structure of any macrocyclic oligoamide shown in any of Figures 3-4, 12, and 19-47, or a salt thereof; or comprises or consists of the structure of any one of compounds 1-218 in Table 3, or a salt thereof. In other embodiments, the macrocyclic oligoamide comprises replacement of one monomeric subunit, two monomeric subunits, three monomeric subunits, or a total of four monomeric subunits.
[0007] In one embodiment, the disclosure provides a library comprising 10, 50, 100, 500, 1000, 5000, 10000, 25000, 35000 or more macrocyclic oligoamides of any embodiment or combination of embodiments described herein. In another embodiment, the disclosure provides a method of using the macrocyclic oligoamides and / or libraries of the embodiments for any suitable purpose, including but not limited to panning the library to identify one or more macrocyclic oligoamides that bind to a compound of interest, for therapeutic treatment, diagnostics, and / or for any use.
[0008] In another aspect, the disclosure provides one or more computer-implemented methods for identifying a macrocycle conformation of a molecule comprising a first molecular fragment and a second molecular fragment, the method comprising the steps of: obtaining a conformation set for the first molecular fragment; determining, for each conformation of the first molecular fragment, values of a set of parameters for a start-to-end transformation that defines a translation and rotation of a start end of the first molecular fragment relative to a end end of the first molecular fragment in that conformation; obtaining a conformational set of a second molecular fragment; determining, for each conformation of the second molecular fragment, values of a set of parameters for a start-to-end transformation that defines a translation and rotation of a terminal end of the second molecular fragment relative to a starting end of the second molecular fragment in that conformation; and processing the parameter values for each of the start-to-end and end-to-start transformations to identify a set of one or more macrocycle conformations of the molecule; Here, each macrocyclic conformation of the molecule comprises respective conformations of the first molecular fragment and the second molecular fragment which cooperate to form a closed circular loop.
[0009] In one embodiment, processing the parameter values for each of the start-to-end and end-to-start transformations to identify a set of one or more macrocycle conformations of the molecule comprises, for each macrocycle conformation, the steps of: Determining that a parameter value of an start-to-end transformation for the conformation of the first molecular fragment and a parameter value of an end-to-start transformation for the conformation of the second molecular fragment satisfy a loop closure criterion.
[0010] In another embodiment, determining that the parameter values of the start-to-end transformation for the conformation of the first molecular fragment and the parameter values of the end-to-start transformation for the conformation of the second molecular fragment satisfy a loop closure criterion comprises the steps of: Determining that a discretization of the start-end transformation parameter values for the conformation of the first molecular fragment is equal to a discretization of the end-end transformation parameter values for the conformation of the second molecular fragment.
[0011] In another embodiment, the method further comprises the steps of: determining, for each macrocycle conformation, a respective energy value associated with that macrocycle conformation; identifying a macrocycle conformation from the plurality of macrocycle conformations that has a minimum energy value; For each macrocycle conformation, (i) the macrocycle conformation, and (ii) the macrocycle conformation with the smallest energy value; determining respective similarity measures between and predicting whether the molecule comprising the first molecular fragment and the second molecular fragment will adopt a rigid macrocycle conformation based on the energy value and the similarity measure of the macrocycle conformation.
[0012] In a further embodiment, the method further comprises physically synthesizing the molecule.
[0013] In one embodiment, the method further comprises the steps of: determining one or more conformations of said physically synthesized molecule; and Comparing the conformation of the physically synthesized molecule with the macrocycle conformation identified for that molecule.
[0014] The present disclosure also provides a system that includes: One or more computers; and one or more storage devices communicatively connected to the one or more computers; The one or more storage devices store instructions for execution by the one or more computers when the one or more computers execute the operations of each method in any of the embodiments disclosed herein.
[0015] The present disclosure further provides one or more non-transitory computer storage media that store instructions for execution by one or more computers when the one or more computers execute the operations of each of the embodiments disclosed herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] All references cited herein are incorporated herein by reference in their entirety.
[0017] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. As used herein, "and" is used as a synonym for "or" unless expressly stated otherwise.
[0018] As used herein, amino acid residues are represented by the following abbreviations: Alanine (Ala; A), asparagine (Asn; N), aspartic acid (Asp; D), arginine (Arg; R), cysteine (Cys; C), glutamic acid (Glu; E), glutamine (Gln; Q), glycine (Gly; G), histidine (His; H), isoleucine (Ile; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), valine (Val; V), and alpha-aminoisobutyric acid (AIB, B). D-form amino acid residues are represented by the abbreviation "D" before the amino acid residue abbreviation. L-form amino acid residues are represented by the abbreviation only. Note that glycine and alpha-aminoisobutyric acid are non-chiral.
[0019] As used herein, "salt" refers to both acid and base addition salts.
[0020] In any aspect of the present invention, any of the embodiments can be used in combination unless the context clearly dictates otherwise.
[0021] In a first aspect, the present disclosure provides unnatural 3-4 residue macrocyclic oligoamides comprising species of 3 or 4 monomeric residues selected from the group consisting of monomers a, b, c, d, e, f, g, h, I, j, k, l, m, n, o, p, q, r, s, t, u, and v, as defined herein. As non-limiting examples, the macrocyclic oligoamides of the present disclosure may be used as membrane-permeable "shuttles" to translocate "cargo" attached thereto across cell membranes and as catalysts in chemical reactions.
[0022] The macrocycles herein are cyclic polymers. The macrocycles disclosed herein are oligoamides, in which each of the three or four monomer residues in the molecule contains an amide group as well as a carboxylic acid. As can be seen from the examples, the inventors have shown that the chemical species of the macrocycle, regardless of its side chain substitution, defines the spherical shape of the macrocycle. As described herein, the macrocycles are classified into different chemical species according to similar substructures within their complex monomers.
[0023] Each monomer is assigned a single property that defines the number, hybridization, and elemental identity of the bonded atoms along the backbone starting from the N-terminal amide nitrogen and ending at the C-terminal amide carbon. If multiple paths exist between these starting and terminal atoms, the atom path is defined as the shortest path along the bonded atoms between the amide nitrogen and the carbonyl carbon.
[0024] Table 1 shows the required atoms and their hybridization along the backbone that a monomer must contain in order to be assigned to a particular chemical species. Substitution of any of these required atoms is permitted as long as the number of atoms and the hybridization of the backbone are not changed by the substitution. Non-limiting examples include: L-alanine, D-alanine, L-proline, and D-proline are all substitutions for the name "a". On the other hand, azaglycine (see pubs.acs.org / doi / 10.1021 / acs.joc.9b02539 for its structure) does not belong to "a" because the number of atoms along the backbone changes to [7, 7, 6].
[0025] [Table 1]
[0026] Figure 1A shows the required backbone structures for each monomer name. For each monomer name, any compound that contains the highlighted backbone structure shown in Figure 1 belongs to that name. The non-highlighted portions of the monomers shown in Figure 1A are not requirements for a monomer to belong to that monomer name.
[0027] In some embodiments, the monomer designation a is as follows: ● Amino acids (monomer names a, b, g, d, and e): a includes all L- and D-proteinogenic amino acids and all L- and D-nonstandard alpha-amino acids; peptoids (N-alkylated glycines); and optical isomers thereof. ○ b includes all B3-amino acids containing L- or D-proteinogenic side chains and their optical isomers. B3 amino acids are β-amino acids whose side chains are located at the β-carbon in the backbone. The β-carbon is the carbon adjacent to the amide nitrogen. o d is the name for delta amino acids and their optical isomers; o g is the name for gamma amino acid and its optical isomer; and o e is the name for ε amino acid and its optical isomer; ● Aminobenzoic acids (monomer names: c, f, j); ● Aminomethylbenzoic acids and their optical isomers (monomer names: h, k, n); ● Aminophenylacetic acids and their optical isomers (monomer names: i, l, o); ● Aminomethylphenylacetic acids and their optical isomers (monomer names: m, p, q); ● Oxazolidines / thiazolidines (monomer designation: r) include all L- and D-proteinogenic side chains substituted at the second atom in the backbone, both for oxazolidinones and thiazolidinones, and their optical isomers; ● Oxazoles / thiazoles (monomer designation: t) include all L- and D-proteinogenic side chains substituted at the second atom in the backbone, both for oxazolidines and thiazolidinines, and their optical isomers; ● Thioethers (monomer name: s), Aminomethylpicolinic acids and their optical isomers (monomer name: u), and ● Triazoles and their optical isomers (monomer name: v).
[0028] Further exemplary monomer embodiments include, but are not limited to, the residues shown in Figure 5. The monomer name (also referred to herein as "section") is provided next to each compound. Each monomer below uses a unique 3-8 letter code to refer to the monomer as well as its respective chemical section, according to the rules in Table 1.
[0029] In another embodiment, the monomer is selected from the group consisting of the following as defined by the chemical structures in FIG. A monomer designation "a" selected from the group consisting of: AGLY (glycine), ALA (alanine), ABU (aminobutyric acid), SER (serine), ASN (asparagine), CLA (chloroalanine), VAL (valine), TBG (tert-butylglycine), NVL (norvaline), LEU (leucine), TBA (tert-butylalanine), PHG (phenylglycine), PHE (phenylalanine), FPA (perfluorophenylalanine), NAP1, NAP2, ANT9, PYR3, DHA (dehydroalanine), AIB (aminoisobutyric acid), ACPC (1-aminocyclopropanecarboxylic acid), ACBC (1-aminocyclobutanecarboxylic acid), ACPenC, ACPhenC, ACHC (1-aminocyclohexanecarboxylic acid), SAR (sarcosine), PTABU, PTIPA, P TTBA, PTANI, PTAMBA, NMALA (N-methylalanine), NMABU (n-methylaminobutyric acid), NMVAL (n-methylvaline), NMPHG (n-methylphenylglycine), NMLEU (n-methylleucine), NMPHE (n-methylphenylalanine), NMSER (N-methylserine), NMASN (n-methylasparagine), AZE (azetidine), PRO (proline), PIP (pipecolic acid), TIQ, TIC, NHM, THP, DHP (dehydroproline), FPS, FPR, HPS (gamma-S-hydroxyproline), HPR (gamma-S-hydroxyproline), OXO, PPS (gamma-S-phenylproline), PPR (gamma-R-phenylproline), AMP (alpha-methylproline), and IDC, and their optical isomers; A monomer designation "b" selected from the group consisting of: BGLY (β-alanine), B3ALA (β-3-homoalanine), B3VAL (β-3-homovaline), B3PHG (β-3-phenylalanine), B3PHE (β-3-homophenylalanine), B2ALA (β-2-alanine), BAZE (3-azetidinecarboxylic acid), BPRO (β-3-homoproline), BPIP (β-3-homopipecolic acid), PRR, NIP (nipecotic acid), LARD, ACPC12C, ACPC12T, ACBC12C, ACBC12T, ACPenC12C, ACPenC12T, ACHC12C, ACHC12T, and ABOC, and their optical isomers; Monomer designation "c" selected from the group consisting of BEN2 (2-aminobenzoic acid) and NMBEN2 (2-N-methylaminobenzoic acid); Monomer designation "d" selected from the group consisting of D4ALA, DGLY (5-aminopentanoic acid), ACHC14C, and ACHC14T; Monomer designation "e" selected from the group consisting of EGLY, ACHA14C, ACHA14T, AMC14C, and AMC14T; A monomer designation "f" selected from the group consisting of BEN3 and NMBEN3; a monomer designation "g" selected from the group consisting of GGLY, G4ALA, GPN, INIP, LARE, ACBC13C, ACBC13T, ACPenC13C, ACPeneC13C, ACPenC13t, GAPC, GAPT, ACHC13C, and ACHC13T; The monomer designation "h" is AMBEN2 (2-aminomethylbenzoic acid); Monomer designation "i" is ACBEN2 (2-aminophenylacetic acid); Monomer designation "j" selected from the group consisting of BEN4 (4-aminobenzoic acid) and NMBEN4 (4-n-methyl-aminobenzoic acid); The monomer name "k" is AMBEN3 (3-aminomethylbenzoic acid); The monomer name "l" is ACBEN3 (3-aminophenylacetic acid); The monomer designation "m" is AMACBEN2 (2-aminomethylphenylacetic acid); The monomer name "n" is AMBEN4 (4-aminomethylbenzoic acid); The monomer designation "o" is ACBEN4 (3-aminophenylacetic acid); The monomer designation "p" is AMACBEN3 (3-aminomethylphenylacetic acid); The monomer designation "q" is AMACBEN4 (4-aminomethylphenylacetic acid); a monomer designation "r" selected from the group consisting of AGLS, AGLR, OZS, OZR, TZS, TZR, and VIC, and optical isomers thereof; The monomer name "s" is SUGA; a monomer designation "t" selected from the group consisting of OZL, TZL, and HUCP, and optical isomers thereof; The monomer name "u" is HUCQ; and / or Monomer designation "v" selected from the group consisting of CLICK and CLICKS.
[0030] For purposes of attributing species of the macrocyclic oligoamides of the present disclosure, first, the names of the composite monomers are determined using Table 1. The resulting monomer name strings are then circularly permuted into the lowest priority string to determine an alphabetical representation. For example, the macrocycle cyclo-(glycine-glycine-glycine-β_glycine) (SEQ ID NO:1) has a macrocyclic species of aaab. The species aaab can be circularly permuted into four identical species: aaba, abaa, baaa, aaab. Alphabetizing this list produces the following list: aaab, aaba, abaa, baaa. Thus, each listed species below includes its circularly permuted form. The first section of the alphabetical representation list is the macrocyclic species of the chemical.
[0031] The macrocycles of the present disclosure may comprise any non-natural species of 3 or 4 monomeric residues selected from the group consisting of: a, b, c, d, e, f, g, h, I, j, k, l, m, n, o, p, q, r, s, t, u, and v. In one embodiment, the macrocyclic oligoamides comprise a species selected from the group listed in Table 2, or a circularly permuted version thereof.
[0032] [Table 2] TIFF2025503658000003.tif242154TIFF2025503658000004.tif242153TIFF2025503658000005.tif242154TIFF2025503658000006.tif242154TIFF2025503658000007.tif242154TIFF2025503658000008.tif242154TIFF2025503658000009.tif242153TIFF2025503658000010.tif242153TIFF2025503658000011.tif242153TIFF2025503658000012.tif242154TIFF2025503658000013.tif242153TIFF2025503658000014.tif242154TIFF2025503658000015.tif242153TIFF2025503658000016.tif242153TIFF2025503658000017.tif242154TIFF2025503658000018.tif242156TIFF2025503658000019.tif242154TIFF2025503658000020.tif242154TIFF2025503658000021.tif242153TIFF2025503658000022.tif242155TIFF2025503658000023.tif242153TIFF2025503658000024.tif242154TIFF2025503658000025.tif242154TIFF2025503658000026.tif242154TIFF2025503658000027.tif242153TIFF2025503658000028.tif242153TIFF2025503658000029.tif242154TIFF2025503658000030.tif242154TIFF2025503658000031.tif242153TIFF2025503658000032.tif242153TIFF2025503658000033.tif242153TIFF2025503658000034.tif242154TIFF2025503658000035.tif242152TIFF2025503658000036.tif242153TIFF2025503658000037.tif242153TIFF2025503658000038.tif242153TIFF2025503658000039.tif242153TIFF2025503658000040.tif242154TIFF2025503658000041.tif242153TIFF2025503658000042.tif242154TIFF2025503658000043.tif242153TIFF2025503658000044.tif242154TIFF2025503658000045.tif242154TIFF2025503658000046.tif242154TIFF2025503658000047.tif242153TIFF2025503658000048.tif242154TIFF2025503658000049.tif242154TIFF2025503658000050.tif242154TIFF2025503658000051.tif242153TIFF2025503658000052.tif242153TIFF2025503658000053.tif242153TIFF2025503658000054.tif242153TIFF2025503658000055.tif242153TIFF2025503658000056.tif242154TIFF2025503658000057.tif242153TIFF2025503658000058.tif242153TIFF2025503658000059.tif242152TIFF2025503658000060.tif242153TIFF2025503658000061.tif242154TIFF2025503658000062.tif242154TIFF2025503658000063.tif242154TIFF2025503658000064.tif242153TIFF2025503658000065.tif242152TIFF2025503658000066.tif242153TIFF2025503658000067.tif111159.
[0033] As will be appreciated by those of skill in the art, the monomers may be substituted in any manner appropriate for the intended application, so long as the resulting monomer contains the requisite backbone structure shown in FIG. 1A.
[0034] In one embodiment, the macrocyclic oligoamide comprises at least one residue not included in the monomer designation a. In another embodiment, the macrocyclic oligoamide comprises at least one non-amino acid residue (i.e.: at least one residue not included in the monomer designations a, b, g, d, e). In one embodiment, the macrocyclic oligoamide has 3 residues. In another embodiment, the macrocyclic oligoamide has 4 residues.
[0035] In a further embodiment, the macrocyclic oligoamide has a species selected from the group consisting of aaar, aarb, abar, aavb, aabr, and aatb, or circularly permuted forms thereof, which represent some of the most abundant species designed using the methods described in the Examples (see FIG. 2).
[0036] In one embodiment, the macrocycle oligoamide has a species selected from the group consisting of akak, aaaq, and aabc, or a circularly permuted form thereof. These embodiments show species for which X-ray and NMR structures are provided for exemplary members of the species as described in the Examples (see FIG. 3). In one non-limiting embodiment, the macrocycles aaam and aaaq provide two different exemplary means of mimicking the extended conformation of poly-alpha amino acid peptides, such as many protease and kinase substrates, and useful for targeting these enzymes.
[0037] In another embodiment, the macrocyclic oligoamide has the species aas or a circularly permuted version thereof, where the macrocyclic oligoamide comprises two hydrogen bonds: one between backbone amides and the other with the primary amide of the "s" monomer. In one embodiment, the one or more hydrogen bonds comprise hydrogen bonds between backbone amides containing non-α-amino acids that help stabilize the macrocycle. In other embodiments, the macrocycle is constructed from eight different monomer species (a, b, g, h, i, l, m, p) arranged into six unique macrocyclic species (aaam, aaap, aabi, aagb, aalm, ahah) and stabilized by transannular hydrogen bonds between backbone amides containing non-α amino acids. As described in the Examples, the aalm-based macrocycles contain hydrogen-bonding fragments constructed from monomers with predominantly sp2 hybridized atoms in the backbone (lm); the aagb-based macrocycles contain hydrogen-bonding fragments of a backbone that contains many more sp3 hybridized atoms than are present in the α-amino acid backbone (gb); and the aabi-based macrocycles contain hydrogen-bonding fragments that combine these two features (bi). Six of the macrocycles contain consecutive fragments of α-amino acids that mimic β-turns common at protein / protein interfaces; the aagb-, aabi-, and aaam-based macrocycles contain turns similar to type I β-turns; and the aalm-based macrocycles contain turns similar to type II β-turns. N-methylated amino acid residues present in both aaap-based macrocycles adopt cis amides, resulting in type VI-like turns. Despite the presence of these distinct β-turn-like features in the aaam macrocycle, the spacing and orientation of the two phenylalanine side chains mimics that of the side chains located at i and i+4 of an α-helix (Figure S10; this spacing is absent in either of the two macrocycles based on aaap, which is an isomer of aaam at the backbone atom level.) This mimicry of β-turn and helical configurations allows the macrocycle of this embodiment to be used for targeting proteins that recognize these structural elements.
[0038] In a further embodiment, the macrocyclic oligoamide has the species aaam or akak, or circularly permuted versions thereof, where either nitrogen in the backbone is a tertiary amide.
[0039] In one embodiment, the macrocyclic oligoamide has the species aaaq or a circularly permuted version thereof, where one of the monomers "a" contains a pentafluorophenylalanine residue or an optical isomer thereof and residue "q" contains a 4-aminomethylphenylacetic acid residue or an optical isomer thereof.
[0040] In a further embodiment, the macrocyclic oligoamide has a species selected from the group consisting of ahah, aabi, and aalm, or circularly permuted forms thereof, for which X-ray and NMR structures are provided for exemplary members of the species as described in the Examples (see FIG. 4).
[0041] In one embodiment of any of the embodiments or combinations thereof described herein, the macrocyclic oligoamides are membrane permeable. As described in the Examples, a large percentage of the macrocycles tested were highly membrane permeable. In one embodiment, the macrocyclic oligoamides have a log(PAMPA) of 0.01 to 0.01, as determined using the Parallel Artificial Membrane Permeability Assay (PAMPA) described in the Examples. app ) is greater than -6. In one embodiment, the macrocyclic oligoamide does not contain exposed polar groups (such as side chain hydroxyl groups and / or primary amides). Such embodiments tend to have lower permeabilities.
[0042] In various further embodiments, the macrocyclic oligoamide comprises or consists of any macrocyclic oligoamide structure shown in any of the Figures, its optical isomers, its circularly permuted forms, or a salt thereof. In one embodiment, the macrocyclic oligoamide comprises or consists of any macrocyclic oligoamide structure shown in any of Figures 3-4, 12, and 19-47, or a salt thereof. As described herein, each monomer within these macrocyclic oligoamides may be independently substituted in any manner appropriate for the intended application, so long as the resulting monomer contains the required backbone structure shown in Table 1.
[0043] In one embodiment of any of the embodiments described herein, the macrocyclic oligoamide comprises a substitution of one monomeric subunit, two monomeric subunits, three monomeric subunits, or a total of four monomeric subunits. Any substitution may be made as appropriate for the intended purpose. As described in the examples below, a majority (%) of the macrocyclic oligoamides of the present disclosure are membrane permeable, and therefore can be complexed with any moiety of interest for which such membrane permeability is useful. In non-limiting embodiments, the moiety may include a therapeutic agent, a diagnostic agent, a marker, a linker, a dye, a purification tag, a peptide, a small molecule, a nucleic acid, etc.
[0044] In a further embodiment, the macrocyclic oligoamide comprises or consists of the structure, or a salt thereof, of any one of compounds 1-218 in Table 3. As described in the Examples, the compounds in Table 3 are designed for binding to specific targets.
[0045] [Table 3] JPEG2025503658000069.jpg242155JPEG2025503658000070.jpg232159JPEG2025503658000071.jpg220159JPEG2025503658000072.jpg228159JPEG2025503658000073.jpg228159JPEG2025503658000074.jpg232159JPEG2025503658000075.jpg220159JPEG2025503658000076.jpg242155JPEG2025503658000077.jpg233159JPEG2025503658000078.jpg242155JPEG2025503658000079.jpg242158JPEG2025503658000080.jpg242155JPEG2025503658000081.jpg242158JPEG2025503658000082.jpg232159JPEG2025503658000083.jpg214159JPEG2025503658000084.jpg218159JPEG2025503658000085.jpg241159JPEG2025503658000086.jpg242159JPEG2025503658000087.jpg242155JPEG2025503658000088.jpg238159JPEG2025503658000089.jpg214159JPEG2025503658000090.jpg242154JPEG2025503658000091.jpg213159JPEG2025503658000092.jpg222159JPEG2025503658000093.jpg233159JPEG2025503658000094.jpg238159JPEG2025503658000095.jpg242159JPEG2025503658000096.jpg242156JPEG2025503658000097.jpg223159JPEG2025503658000098.jpg242156JPEG2025503658000099.jpg242155JPEG2025503658000100.jpg233159JPEG2025503658000101.jpg209159JPEG2025503658000102.jpg222159JPEG2025503658000103.jpg239159JPEG2025503658000104.jpg228159JPEG2025503658000105.jpg70159.
[0046] In a further embodiment, the present disclosure provides a library of macrocyclic oligoamides comprising two or more cyclic peptides and / or complexes according to any embodiment or combination of embodiments disclosed herein. In various embodiments, the library may comprise at least 5, 10, 25, 50, 75, 100, 250, 500, 1000, 5000, 10000, 25000, 35000, or more macrocyclic oligoamides according to any embodiment or combination of embodiments disclosed herein. The library may be used in any suitable application, including but not limited to panning the library to identify macrocyclic oligoamides that bind to a compound of interest.
[0047] In another embodiment, the disclosure provides uses and methods of using the macrocyclic oligoamides and / or libraries of any of the embodiments or combinations of embodiments disclosed herein. In various non-limiting embodiments, the macrocyclic oligoamides and / or libraries may be used to pan the libraries to identify one or more macrocyclic oligoamides that bind to a compound of interest, to add reactive moieties for therapeutic treatment, diagnostics, and / or any particular use. In one embodiment, the macrocyclic oligoamides may be used, for example, to deliver a substituted moiety across a cell membrane. For example, the method may include administering a substituted macrocyclic oligoamide or a macrocyclic oligoamide linked to a small molecule therapeutic agent to enable delivery of the therapeutic agent to a cell to exert a desired therapeutic effect.
[0048] Macrocycle Design Embodiments In another aspect, this document generally describes a system, implemented as a computer program on one or more computers at one or more locations, for identifying macrocycles.
[0049] Throughout this specification, a "molecular fragment" may refer to a sequence of one or more monomers, where each monomer is selected from a monomer set. A monomer refers to a molecule that can react with other monomers to form a sequence (chain) of monomers. The monomer set may include, for example, alpha amino acids, beta amino acids, aminobenzoic acids, oxazoles, thiazoles, or any other suitable monomer.
[0050] A molecular fragment (or monomer) may include a first portion that is designated as the "beginning" of a molecular fragment (or monomer) and a second portion that is designated as the end of the molecular fragment (or monomer). A portion of a molecular fragment (or monomer) may be designated as the beginning or end based on any suitable criteria. In particular, the beginning and end of a molecular fragment (or monomer) are generally considered to be such that the end of one fragment (or monomer) may be chemically bonded to, reacted with, or fused to the beginning of another fragment (or monomer), and vice versa. For example, for a molecular fragment that includes a sequence of amino acid monomers, the N-terminus of the first amino acid of the sequence may be designed as the beginning of the molecular fragment, and the C-terminus of the final amino acid of the sequence may be designed as the end of the molecular fragment. As another example, for an amino acid monomer, the N-terminus of the amino acid may be designed as the beginning of the monomer, and the C-terminus of the amino acid may be designed as the end of the monomer.
[0051] A molecule comprising a first molecular fragment and a second molecular fragment can be described as being a macrocycle; that is, having a macrocycle conformation when the conformations of the first molecular fragment and the second molecular fragment cooperate to form a closed circular loop. For example, the conformations of the first molecular fragment and the second molecular fragment cooperate to form a closed circular loop when the end of the first fragment binds to (or reacts with or fuses to) the beginning of the second fragment and the beginning of the first fragment binds to (or reacts with or fuses to) the end of the second fragment.
[0052] The conformation of a molecule (or a molecular fragment) refers to the spatial arrangement (e.g., three-dimensional (3D) spatial arrangement) of the atoms in the molecule (or molecular fragment). The 3D spatial positions of the atoms can be expressed in any suitable coordinate system, for example, a Cartesian coordinate system.
[0053] A first set of values (e.g., transformation parameter values defining a transformation) may be said to be "equal to" a second set of values if each first value in the first set of values is equal to a corresponding second value in the second set of values.
[0054] Particular implementations of the subject matter described herein can be implemented so as to realize one or more of the following advantages.
[0055] Conventional methods for identifying macrocycles generate the first cyclic conformer of a polyglycine backbone, and then perform sequence optimization to systematically enumerate large macrocycles. These methods require random sampling of backbone torsion angles to facilitate the identification of closed ring conformations of the macrocycle, making it difficult to extend the generation of macrocycles formed from monomers with a wide range of backbone chemistries. This random exploration of backbone torsion angle combinations and subsequent sequence design can become unwieldy, for example, when many different types of monomer building blocks are included in the backbone. For example, from 22 building block monomers, nearly 60,000 two-dimensional chemical structures can be constructed for two-, three-, and four-residue macrocycles, each of which can be further diversified by potentially millions of different side chain combinations. This chemical space is so large that it is not possible to rely on random sampling to identify just a few linear sequences that can close to form a ring.
[0056] The system described herein performs computationally tractable operations in sampling the vast chemical space of possible macrocycles. More specifically, to identify possible macrocycle conformations of a molecule composed of a first molecular fragment and a second molecular fragment, the system can characterize each conformation of the first and second molecular fragments by a respective rigid body transformation. A rigid body transformation of a molecular fragment can be defined as a translation and rotation of one end (end) of the molecular fragment relative to the other end (end) of the molecular fragment. The system can process the rigid body transformations to identify complementary conformations of the first and second molecular fragments that cooperate to form a closed circular loop (and thus also a macrocycle).
[0057] The system can efficiently identify the complementary conformations of the first and second molecular fragments by generating a "dictionary" that represents a set of conformations of the first and second molecular fragments. In particular, to generate a dictionary of molecular fragments, the system can discretize the rigid transformations that correspond to the conformations of the molecular fragments, and map each discretized rigid transformation to a respective "key" (e.g., represented by an integer value). The system can then create the dictionary by adding each key to the dictionary and associating each key with a set of "values," where each value represents a conformation that has a rigid transformation that corresponds to the key.
[0058] The system can generate a first dictionary of start-to-end rigid transformations of the conformations of the first molecular fragment and a second dictionary of end-to-start rigid transformations of the second molecular fragment. The system can then identify keys common to the two dictionaries. For each common key, the system can identify a conformation of the first molecular fragment associated with the key in the first dictionary and a conformation of the second molecular fragment associated with the key in the second dictionary as jointly defining a macrocycle. By discretizing the rigid transformations and representing them in a dictionary, the computational complexity of identifying complementary conformations of the first and second molecular fragments can be significantly reduced. In particular, the system achieves lower computational complexity by exploiting the observation that rigid transformations corresponding to conformations are not uniformly distributed (i.e., are not uniformly distributed in the space of possible rigid transformations) and are often clustered into a small number of groups. By discretizing the space of possible rigid transformations, the number of unique rigid transformations corresponding to conformations of molecular fragments can be significantly reduced, and dictionaries (such as those described above) allow the subsequent identification of complementary conformations by efficient operations (e.g., intersection operations to identify common keys).
[0059] One or more embodiments of the subject matter herein are set forth in detail in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
[0060] Figure 48 illustrates an example of a macrocycle computing system 100. Macrocycle computing system 100 is an example of a system implemented as a computer program on one or more computers located at one or more locations;
[0061] Macrocycle computational system 100 is configured to process data defining monomer set 102 for the purpose of identifying macrocycle molecular set 114 ("macrocycles").
[0062] The monomer set 102 may include, for example, alpha amino acids, beta amino acids, aminobenzoic acids, oxazoles, thiazoles, or any other suitable monomers. For example, the monomer set 102 may be provided to the macrocycle computing system 100 through an application program interface (API) or user interface made available by the macrocycle computing system 100.
[0063] Each macrocycle 114 is a molecule formed from a pair of molecular fragments having respective conformations that cooperate to form a closed circular loop. Each of the molecular fragments includes a respective sequence of one or more monomers from the monomer set 102. Figure 56 (discussed in further detail below) provides an illustrative example of a macrocycle formed from a pair of molecular fragments.
[0064] The macrocycle computational system 100 includes a monomer conformation system 104, a fragment generation system 108, and a macrocycle identification system 600; each of which is described below.
[0065] The monomer conformation system 104 is configured to determine a set of one or more conformations for each monomer 102 from the set of monomers 102. A conformation of a monomer refers to the 3D spatial arrangement of atoms in the monomer. A conformation of a monomer may be represented in any suitable manner, for example, by a list of 3D spatial coordinates of the atoms in the monomer. The monomer conformation system 104 may generate any suitable number of conformations for each monomer (e.g., 5 conformations, 10 conformations, 100 conformations, or 1000 conformations). The monomer conformation system 104 may generate different numbers of conformations for different monomers. An exemplary process for generating monomer conformations is described with reference to Figures 49-50.
[0066] In some embodiments, the macrocycle computational system 100 can exclude the monomer conformational system 104; i.e., the monomer conformational system 104 can be excluded from being included in the macrocycle computational system 100. For example, rather than generating the monomer conformations 106, the macrocycle computational system 100 can receive as input a library in which the monomer conformations 106 are predefined.
[0067] The fragment generation system 108 includes: (i) molecular fragment set 110; and (ii) a respective set of conformations 112 for each molecular fragment 110; The method is configured to generate data that specifies:
[0068] Each molecular fragment 110 includes a sequence of one or more monomers 102 from the set of monomers 102. While a particular molecular fragment 110 may include only a single monomer, another molecular fragment 110 may include multiple monomers (e.g., 2 monomers, 3 monomers, 4 monomers, 10 monomers, 20 monomers, or any other suitable number of monomers). The fragment generation system 108 may generate any suitable number of molecular fragments (e.g., 1000 molecular fragments, 100000 molecular fragments, 1000000 molecular fragments, etc.).
[0069] To generate molecular fragments 110, the fragment generation system 108 can select the length of the molecular fragment, i.e., the number of monomers in the monomer sequence of the molecular fragment. The fragment generation system 108 can then select a monomer for each position in the monomer sequence of the molecular fragment, for example, by sampling one monomer 102 from the set of monomers 102.
[0070] The fragment generation system 108 can generate, for each position in a monomer sequence of a molecular fragment 110, data defining a conformational set 112 for the molecular fragment 110 based on the monomer conformational set 106 for the monomer at that position. An exemplary process for generating data defining a conformational set for a molecular fragment based on possible conformations of monomers in the monomer sequence of the molecular fragment is described with reference to Figure 52.
[0071] A macrocycle identification system 600 is configured to process the set of molecular fragments 110 and their corresponding conformations 112 to generate data defining a set of macrocycles 114. Each of the macrocycles 114 is a molecule formed from a pair of molecular fragments 110 having fragment conformations 112 that cooperate to form a closed circular loop. An example macrocycle identification system 600 is described in further detail with reference to Figures 53-55.
[0072] The macrocycles 114 identified by the macrocycle computation system 100 can be utilized in any of a variety of possible downstream processes. A few examples of downstream processes that may utilize the macrocycles 114 identified by the macrocycle computation system 100 are described below.
[0073] In some embodiments, a macrocycle 114 identified by the macrocycle computational system 100 can be physically synthesized (e.g., in a laboratory). The physically synthesized macrocycle can be used in a variety of ways. For example, the properties of the physically synthesized macrocycle, such as chemical stability, toxicity, solubility, reactivity, etc., can be determined by appropriate experimental techniques. As another example, the macrocycle can be physically synthesized for therapeutic use (e.g., as a drug). The therapeutic agent can be administered to a subject, such as a mouse, cat, dog, pig, or human, for example, to exert a therapeutic effect in the subject. As another example, the conformation of the physically synthesized macrocycle can be determined by experimental techniques (e.g., by x-ray crystallography) and compared to the conformation predicted by the macrocycle computational system 100 (e.g., as part of the validation of the macrocycle computational system 100). As another example, the physically synthesized macrocycle can be made accessible to a binding site; for example, the surface of a protein, for purposes of determining the binding affinity of the macrocycle with respect to the binding site.
[0074] In some embodiments, the macrocycles 114 identified by the macrocycle computation system 100 can be provided for use in downstream computational simulation and computational analysis processes. For example, high-fidelity computational simulations (e.g., density functional theory (DFT) based simulations) can be performed to determine the energy of the macrocycle. As another example, a structural representation of the macrocycle can be processed by a computational analysis system to predict the affinity of the macrocycle for a binding site (e.g., the surface of a protein).
[0075] Figure 49 is a flow diagram of an exemplary process 200 for generating a conformational set of monomers. For convenience, process 200 is described as being performed by one or more computer systems located at one or more locations. For example, a macrocycle computational system suitably programmed in accordance with the present disclosure (e.g., macrocycle computational system 100 of Figure 48) can perform process 200.
[0076] The system initializes a set of points in a structural parameter space, where each point in the structural parameter space defines a respective candidate conformation of the monomer (202). More specifically, each point in the structural parameter space can specify a value for each structural parameter included in a set of structural parameters that together define a candidate conformation of the monomer. A structural parameter set can include, for example, a set of torsion (dihedral) angles that define bond angles between atoms in the monomer. A structural parameter set can include any suitable number of structural parameters (e.g., 3 structural parameters, 10 structural parameters, or 30 structural parameters).
[0077] The system can initialize the set of points in the structural parameter space in any suitable manner. For example, the system can initialize the set of points in the structural parameter space by randomly sampling a predefined number of points in the structural parameter space. As another example, the system can initialize the set of points in the structural parameter space by selecting points that form a uniform grid over the structural parameter space. The system can initialize the set of points in the structural parameter space to include any suitable number of points (e.g., 100 points, 1000 points, or 100000 points).
[0078] In steps 204-208, the system iteratively augments the set of points in the structural parameter space to include additional points. Iteratively augmenting the set of points has the effect of adaptively sampling points from the structural parameter space with the goal of efficiently identifying candidate conformations having low energy. For convenience, steps 204-208 are described with reference to the "current" set of points in the structural parameter space at the "current" iteration (i.e., the "current" iteration in the series of iterations being performed by the system).
[0079] For each point in the current set of points, the system determines (204) the energy of the candidate conformation identified by the point. If the current iteration is the first iteration in the series of iterations, the system determines the energy for each of the candidate conformations corresponding to each point in the initial set of points. For iterations after the first iteration, the system may only be required to determine the energy of the candidate conformations corresponding to points added to the current set of points in the previous iteration. More specifically, after determining the energy of the candidate conformations corresponding to a point, the system may store the energy value, but is not required to regenerate the energy value of the point in each iteration.
[0080] The system can determine the energy of the candidate conformation by any suitable means. For example, the system can use an energy prediction machine learning model to process a model input based on a set of structural parameters that define the candidate conformation to generate a model output that defines a predicted energy of the candidate conformation for the purpose of determining the energy of the candidate conformation. The energy prediction machine learning model can have any suitable machine learning model architecture that enables the model to perform its described function. For example, the energy prediction machine learning model can be a neural network model, a random forest model, a support vector machine model, etc. In an embodiment in which the energy prediction machine learning model is a neural network model, the neural network model can include any suitable number (e.g., 5 layers, 10 layers, or 50 layers) of neural network layers (e.g., fully connected layers, convolutional layers, attention layers, etc.) of any suitable type of neural network layers connected in any suitable configuration (e.g., as a linear array of layers).
[0081] Referring to R. Zubatyuk, J.S. Smith, J. Leszczynski, O. Isayev, “Accurate and transferable multitask prediction of chemical properties with an atoms-in-molecules neural network,” Science Advances, 9 August 2019, Volume 5, Issue 8, we describe an example of an energy prediction machine learning model that can be used to determine the energies of candidate conformations.
[0082] The energy prediction machine learning model can be trained with a set of training examples, where each training example corresponds to a training conformation and includes: (i) training inputs for the energy prediction machine learning model; and (ii) the target output of the energy prediction machine learning model. The training input of the training example can be based on the structural parameters that define the corresponding training conformation. The target output of the training example can define the energy of the corresponding training conformation. The energy of the training conformation can be determined, for example, using high-fidelity computer simulation (e.g., simulation based on density functional theory (DFT)).
[0083] The system selects one or more new points for inclusion in the current point set based on the energies of conformations corresponding to points in the current point set (206). The system may select new points in a manner that encourages dense sampling of regions of structural parameter space determined to correspond to low energy conformations, while also encouraging exploration of undersampled portions of structural parameter space. One exemplary process for selecting new points for inclusion in the current point set is described with reference to Figure 50.
[0084] The system determines (208) whether a termination criterion is met, which may be, for example, that the system has performed steps 204-208 a predefined number of iterations or that the current point set contains at least a threshold number of points.
[0085] In response to determining that the termination criteria are not met, the system returns to step 204 and performs yet another iteration that augments the current set of points in the structural parameter space.
[0086] In response to determining that the termination criteria have been met, the system determines a conformation set for the monomer based on the current set of points in the structural parameter space (i.e., as the final iteration of steps 204-208) (210). For example, the system can identify each point in the current set of points that corresponds to a candidate conformation having an energy that meets (e.g., is below) a threshold for being a respective conformation of the monomer.
[0087] Figure 50 is a flow diagram of an exemplary process 300 for adaptively sampling a structural parameter space, where each point in the structural parameter space defines a respective candidate conformation of a monomer. For convenience, the process 300 is described as being performed by one or more computer systems located at one or more locations. For example, a macrocycle computing system, such as the macrocycle computing system 100 of Figure 48, suitably programmed in accordance with the present disclosure, can perform the process 300.
[0088] The system obtains (302) a current set of points in the structural parameter space (i.e., the current set of points in the structural parameter space that define the current sampling).
[0089] The system determines 304 a triangulation, e.g., a Delaunay triangulation of the current point set. A triangulation of a point set refers to a simplicial complex whose vertices are determined by the point set and that contains the convex hull of the point set. A triangulation can divide the region of the construction parameter space contained in the convex hull of the point set into a set of subregions; these subregions are conveniently called "simplices" of the triangulation.
[0090] The system determines a respective score for each simplex of the triangulation 306. For example, the system may determine the scores for the simplexes of the triangulation based on (i) and (ii) below: (i) the volume of the simplex of the triangulation; and (ii) the energy of the points at the vertices of the triangulation simplices. In one particular example, the system may determine a score for a triangulation simplices based on the product of (i) and (ii): (i) the volume of the simplex of the triangulation; and (ii) the negative exponential of the product of the energies of the points at the vertices of the simplex of the triangulation.
[0091] The system selects one or more simplexes of the triangulation based on the scores for the simplexes of the triangulation 308. For example, the system can select a predetermined number of the triangulation simplexes with the highest scores, or the system can select each simplex of the triangulation that has a score above a threshold.
[0092] The system may generate one or more new points to add to the point set based on the selected simplex of the triangulation 310. For example, for each selected simplex of the triangulation, the system may determine a point that is at the center of the triangulation simplex and then add the center point of the triangulation simplex to the point set.
[0093] Augmenting the point set in this manner facilitates exploration of the structural parameter space, for example by scoring the triangulation simplices based on their volume, which increases the likelihood that new points will be selected from triangulation simplices that have large volumes. Additionally, the system allows for dense sampling of regions of structural parameter space that are determined to contain low energy points by scoring the triangulation simplices based on, for example, the energy of the points at the vertices of the triangulation simplices.
[0094] Figure 51 shows an example of an adaptively sampled structural parameter space for a monomer, such as that sampled to determine a set of conformations for the monomer, e.g., as described with reference to Figures 49-50. In this example, the structural parameter space has one dimension corresponding to the Φ torsion angle 404 of the monomer and another dimension corresponding to the Ψ torsion angle 402 of the monomer. The sampling is performed in a manner that efficiently identifies a low energy conformation of the monomer, as described above.
[0095] Figure 52 is a flow diagram of an exemplary process 500 for generating data defining a set of conformations of a molecular fragment based on possible conformations of monomers in the monomer sequence of the molecular fragment. For convenience, process 500 is described as being performed by one or more computer systems located at one or more locations. For example, a macrocycle computing system, such as macrocycle computing system 100 of Figure 48, suitably programmed in accordance with the present disclosure, can perform process 500.
[0096] The system obtains (502) a set of conformations for each monomer in a monomer sequence of the molecular fragment.
[0097] The system determines (504) a set of fragment conformations for the molecular fragment. Each fragment conformation corresponds to selecting a particular monomer conformation for each monomer in the monomer sequence of the molecular fragment. The system can determine the number of fragment conformations given by Π_(i=1)^N n_i, where N represents the number of positions in the monomer sequence of the molecular fragment; i represents the index of a position in the monomer sequence of the molecular fragment; and n_i represents the number of monomer conformations of the monomer at position i in the monomer sequence of the molecular fragment.
[0098] For each identified fragment conformation of the molecular fragment, the system determines (506) the 3D spatial positions of the atoms in the molecular fragment when the molecular fragment adopts the fragment conformation. More specifically, for a given fragment conformation, the system can sequentially determine the 3D spatial positions of the atoms in each monomer at each position of the monomer sequence of the molecular fragment, starting from a first position. For a first position in the monomer sequence, the spatial position of each atom of the monomer at that position can be directly defined by the conformation of the monomer at that position. For each position after the first position in the monomer sequence, the spatial position of each atom of the monomer can be determined by applying rotation and translation operations to the spatial positions of the atoms defined by the conformation of the monomer at that position. The rotation and translation operations can be selected to align the beginning of the monomer at the current position with respect to the end of the monomer at the previous position.
[0099] Figure 53 illustrates an example of a macrocycle identification system 600. Macrocycle identification system 600 is an example of a system implemented as a computer program on one or more computers at one or more locations, including the systems, components, and techniques described below.
[0100] The macrocycle identification system 600 is configured to process a set of molecular fragments 110 and their corresponding conformations 112 to generate data that identify a set of macrocycles 114. Each of the macrocycles 114 is a molecule formed from a pair of molecular fragments 110 having respective fragment conformations 112 that cooperate to form a closed circular loop.
[0101] Macrocycle identification system 600 includes a transformation engine 602 and a loop closure engine 608, which are described in further detail below.
[0102] The transformation engine 602 is configured to generate data that defines a “start-end” transformation 604 and an “end-start” transformation 606 for each fragment conformation 112 .
[0103] The start-end transformation 604 for the fragment conformation 112 of the molecular fragment 110 defines a translation and rotation of the start of the molecular fragment 110 relative to the end of the molecular fragment 110. That is, the start-end transformation defines translation and rotation operations that, when applied to the coordinates of atoms at the start of the molecular fragment, produce transformed coordinates that align with the coordinates of atoms at the end of the molecular fragment.
[0104] The start-to-end transformation is parameterized by a set of transformation parameters in a transformation parameter space. For example, a transformation parameter set that parameterizes the start-to-end transformation may specify a rotation matrix that specifies a rotation operation in 3D space (e.g., a 3×3 rotation matrix) and a translation vector that specifies a translation operation in 3D space (e.g., a 3×1 translation vector).
[0105] The end-to-start transformation 606 for the fragment conformation 112 of the molecular fragment 110 defines a translation and rotation of the end of the molecular fragment 110 relative to the start of the molecular fragment 110. That is, the end-to-start transformation defines translation and rotation operations that, when applied to the coordinates of atoms at the end of the molecular fragment, produce transformed coordinates that align with the coordinates of atoms at the start of the molecular fragment.
[0106] A terminal-start transformation is parameterized by a transformation parameter set in a transformation parameter space. For example, a transformation parameter set that parameterizes a terminal-start transformation may specify a rotation matrix that specifies a rotation operation in 3D space (e.g., a 3×3 rotation matrix) and a translation vector that specifies a translation operation in 3D space (e.g., a 3×1 translation vector).
[0107] A loop closure engine 608 is configured to process the start-end transformations 604 and the end-end transformations 606 with the goal of identifying a set of macrocycles 114. To accomplish this, the loop closure engine 608 identifies pairs of fragment conformations 112 (i.e., including a first fragment conformation and a second fragment conformation) where the start-end transformation of the first fragment conformation, together with the end-end transformation of the second fragment conformation, satisfy the loop closure criteria.
[0108] In general, a first fragment conformation and a second fragment conformation cooperate to form a (at least approximately) closed circular loop if they satisfy loop closure criteria. Some examples of loop closure criteria are described below.
[0109] In some embodiments, a loop closure criterion is satisfied for a start-end transformation and a end-end transformation if the transformation parameter values defining the start-end transformation are equal to the transformation parameter values defining the end-end transformation. In these embodiments, a pair of fragment conformations cooperate to form a closed circular loop if a start-end transformation for a first fragment conformation and a end-end transformation for a second fragment conformation satisfy the loop closure criterion.
[0110] In some embodiments, a loop closure criterion is satisfied for the start-end transformation and the end-end transformation if the discretization of the transformation parameter values defining the start-end transformation is equal to the discretization of the transformation parameter values defining the end-end transformation. In these embodiments, a pair of fragment conformations cooperate to form a loop that is at least approximately circularly closed if the start-end transformation for a first fragment conformation and the end-end transformation for a second fragment conformation satisfy the loop closure criterion. (The closeness of approximation is determined by the fit of the discretization of the transformation parameter values defining the start-end transformation and the end-end transformation.) As described below with reference to FIG. 54, discretizing the transformation parameter values defining the start-end transformation and the end-end transformation may enable the loop closure criterion to be efficiently evaluated across a large number of fragment conformations.
[0111] Discretizing values of a first set of possible values may mean mapping the values to discrete values taken from a second set of possible values, where the second set of possible values includes a smaller number of values than the first set of possible values. The parameter values that define the start-end and end-end transformations are taken from a continuous space of possible values, e.g., a continuous Euclidean space. Discretizing parameter values that define a transformation refers to mapping continuous transformation parameter values that define the transformation from a discrete (i.e., non-continuous) set of transformation parameter values to corresponding discrete transformation parameter values. (The "goodness" of the discretization characterizes how densely the discrete set of transformation parameter values contains the continuous space of transformation parameter values.)
[0112] The conformation of a molecule formed from a pair of molecular fragments (i.e., including a first molecular fragment and a second molecular fragment) is defined by the conformation of the first molecular fragment and the conformation of the second molecular fragment. The first and second molecular fragments can each adopt a conformation included in a conformation set, so that the molecule can assume multiple possible conformations. If the conformation of the first molecular fragment and the conformation of the second molecular fragment satisfy the loop closure criterion, then the conformation of a molecule formed from a pair of molecular fragments 110 (including a first molecular fragment and a second molecular fragment) is said to satisfy the loop closure criterion.
[0113] Based on whether the possible conformations of the molecule satisfy the loop closure criteria, the macrocycle identification system 600 may determine that a molecule formed from a first molecular fragment and a second molecular fragment is a macrocycle 114. If at least one conformation of a molecule satisfies the loop closure criteria, the macrocycle identification system 600 may determine that the molecule is a macrocycle 114. If at least a predefined number of conformations of a molecule satisfy the loop closure criteria, the macrocycle identification system 600 may determine that the molecule is a macrocycle. If at least a predefined percentage of conformations of a molecule satisfy the loop closure criteria, the macrocycle identification system 600 may determine that the molecule is a macrocycle 114.
[0114] In some embodiments, the macrocycle identification system 600 predicts whether a molecule that can adopt a set of macrocycle conformations will form a rigid macrocycle based on the energy distribution and conformation of the molecule that meets the loop closure criteria. An exemplary process for predicting whether a molecule will form a rigid macrocycle is described with reference to FIG.
[0115] The set of fragment conformations 112 may include a large number of fragment conformations, e.g., billions of different fragment conformations. Thus, it may be computationally infeasible to directly evaluate the loop closure criteria for each pair of fragment conformations. An example of an efficient and computationally tractable process for identifying macrocycles by evaluating the loop closure criteria for discretized transformation parameter values is described with reference to FIG.
[0116] Figure 54 is a flow diagram of an exemplary process 700 for identifying macrocycles using discretized start-to-end transformations and discretized end-to-start transformations. For convenience, the process 700 is described as being performed by one or more computer systems located at one or more locations. For example, a macrocycle identification system, such as the macrocycle identification system 600 of Figure 53, suitably programmed in accordance with the present disclosure, can perform the process 700.
[0117] The system obtains a collection of fragment conformations including a respective set of fragment conformations for each of a plurality of molecular fragments (702). The collection of fragment conformations can include any suitable number of fragment conformations, for example, one million different fragment conformations, one billion different fragment conformations, or one trillion different fragment conformations. The molecular fragments and fragment conformations can be generated, for example, by a fragment generation system described with reference to FIG.
[0118] The system determines 704 a start-to-end transformation and an end-to-start transformation for each fragment conformation in the collection of fragment conformations. The start-to-end transformation of a fragment conformation defines a translation and rotation of the start of the fragment conformation relative to the end of the fragment conformation. The end-to-start transformation of a fragment conformation defines a translation and rotation of the end of the fragment conformation relative to the start of the fragment conformation.
[0119] The system discretizes the start-end transformation and the end-end transformation (706). More specifically, for each transformation (i.e., each start-end transformation and each end-end transformation), the system discretizes the transformation parameter values that define the transformation. The system may discretize the set of transformation parameter values that define the transformation using any suitable discretization technique. For example, the transformation parameter space may be a Euclidean space, in which case the system may divide the Euclidean space into hypercubes (e.g., having predefined edge lengths and volumes). The system may discretely represent the set of transformation parameter values (i.e., represented in a continuous Euclidean space) by an index of a hypercube that contains the set of transformation parameter values.
[0120] The system generates (708) a dictionary corresponding to the begin-end transformation (an "begin-end" dictionary) and a dictionary corresponding to the end-end transformation (an "end-end" dictionary).
[0121] The start-end dictionary includes (i) a set of keys and (ii) one or more values associated with each key, where each key indicates a respective discretized start-end transformation, and each value associated with a key identifies a fragment conformation with respect to the discretized start-end transformation indicated by the key.
[0122] To generate the start-end dictionary, the system may instantiate the discretized start-end transformation set and then deduplicate the discretized start-end transformation set; that is, deduplicate by removing duplicate copies of the discretized start-end transformation. The system may then process each discretized start-end transformation from the deduplicated set to generate a corresponding key for the start-end dictionary. The system may generate keys from the start-end transformations in any of a variety of ways. For example, the system may process a set of discretized parameter values representing the start-end transformations using a hash function to generate the corresponding key. The hash function may be, for example, a multiplicative hash function, an algebraic coding hash function, a Fibonacci hash function, or the like. Each key may be represented by an appropriate numeric data, for example, an integer value. In addition to generating the keys for the start-end dictionary, the system can associate each key with a set of values that identify each fragment conformation using the discretized start-end transformation indicated by the key.
[0123] To generate the terminal start dictionary, the system may instantiate the set of discretized terminal start transformations and then deduplicate the set of discretized terminal start transformations; that is, deduplicate by removing duplicate copies of the discretized terminal start transformations. The system may then process each discretized terminal start transformation from the de-duplication set to generate a corresponding key for the terminal start dictionary. The system may generate keys from terminal start transformations in any of a variety of ways. For example, the system may process a set of discretized parameter values representing a terminal start transformation using a hash function to generate a corresponding key. The hash function may be, for example, a multiplicative hash function, an algebraic coding hash function, a Fibonacci hash function, or the like. Each key may be represented by an appropriate numeric data, for example, an integer value. In addition to generating the keys for the end-start dictionary, the system can associate each key with a set of values that identify each fragment conformation using the discretized end-start transformation indicated by the key.
[0124] The number of keys in the start-end dictionary may be significantly reduced (e.g., by an order of magnitude or more) from the original number of start-end transformations. In particular, the set of parameter values representing each start-end transformation is generally not uniformly distributed throughout the space of transformation parameter values, but rather clusters into multiple sets. Thus, discretization of start-end transformations allows a large number of continuous-value start-end transformations to be mapped onto a much smaller set of discretized start-end transformations. Therefore, each key in the start-end dictionary may be associated with multiple discretized start-end transformations (e.g., 10 transformations, 100 transformations, or 1000 transformations). Figure 57, described in further detail below, illustrates the clustering of start-end transformations and end-end transformations in the transformation space.
[0125] Similarly, the number of keys in the terminal start dictionary may be significantly reduced (e.g., by an order of magnitude or more) from the original number of terminal start transformations. Each key in the terminal start dictionary may be associated with multiple discretized terminal start transformations (e.g., 10 transformations, 100 transformations, or 1000 transformations).
[0126] The system utilizes the start-end dictionary and the end-end dictionary to identify pairs of fragment conformations that satisfy the loop closure criteria (710). In particular, the system can identify a set of one or more "joint" keys that are common to both the start-end dictionary and the end-end dictionary. For example, the system can identify joint keys by performing an intersection operation on (i) a set of keys in the start-end dictionary and (ii) a set of keys in the end-end dictionary. For each joint key, the system can identify each pair of fragment conformations that includes (i) a first fragment conformation associated with the joint key in the start-end dictionary and (ii) a second fragment conformation associated with the joint key in the end-end dictionary, such that the pair satisfies the loop closure criteria.
[0127] Discretizing the transforms and creating dictionaries, as described above, can reduce the computational complexity of identifying pairs of transforms that satisfy the loop closure criterion by an order of magnitude or more, as compared to, for example, directly evaluating the loop closure criterion for all pairs of continuous-valued transforms. In particular, discretizing the transforms can collapse the number of unique transforms into a significantly smaller number of unique discrete transforms, and intersection operations can be utilized to efficiently compare dictionaries in order to identify pairs of transforms that satisfy a common key and the loop closure criterion.
[0128] The system identifies one or more macrocycles based on the identified fragment conformational pairs that satisfy the loop closure criteria (712). For example, the system can determine that a molecule formed from a molecular fragment pair is a macrocycle if the molecule has a set of conformations that satisfy the loop closure criteria.
[0129] Figure 55 is a flow diagram of an exemplary process 800 for predicting whether a molecule capable of adopting a group of macrocycle conformations will adopt a rigid macrocycle conformation. For convenience, process 800 will be described as being performed by one or more computer systems located at one or more locations. For example, a macrocycle identification system, such as macrocycle identification system 600 of Figure 53, suitably programmed in accordance with the present disclosure, can perform process 800.
[0130] The system obtains a set of macrocycle conformations for the molecule 802. The system can obtain the set of macrocycle conformations for the molecule, for example, by the process described with reference to Figure 54.
[0131] For each macrocycle conformation, the system determines a respective energy value associated with the macrocycle conformation (804). The system can determine the energy values for the macrocycle conformations using, for example, an energy prediction machine learning model, such as that described with reference to Figure 49.
[0132] The system identifies (806) a macrocycle conformation from among the plurality of macrocycle conformations that has a minimum energy value.
[0133] For each macrocycle conformation, the system determines a respective similarity measure between (i) the macrocycle conformation and (ii) the macrocycle conformation having the smallest energy value 808. The system can assess the similarity between pairs of macrocycle conformations based on, for example, the root mean square deviation (RMSD) of the atomic positions between the pairs of macrocycle conformations.
[0134] The system determines whether the molecule is predicted to adopt a rigid macrocycle conformation based on the portion of the macrocycle conformation of the molecule that has (i) an energy that meets an energy threshold and (ii) a similarity to a minimum energy macrocycle conformation that meets a similarity threshold (810). The system can determine that the energy of the macrocycle conformation meets the energy threshold, for example, if the energy of the macrocycle conformation is lower than the energy threshold. The system can determine that the similarity of the macrocycle conformation to a minimum energy macrocycle conformation meets the similarity threshold, for example, if the similarity is higher than the similarity threshold. The system can determine that the molecule is predicted to adopt a rigid macrocycle conformation, for example, if the portion of the macrocycle conformation that simultaneously meets the energy threshold and the similarity threshold has at least the threshold.
[0135] FIG. 56 shows a first molecular fragment and a second molecular fragment that have conformations that cooperate to form a closed circular loop.
[0136] Figure 57 shows a backbone superposition of a set of conformational structures of molecular fragments. The start of the fragment conformations are aligned and the end of the fragment conformations are clustered into minority groups. As can be seen from the above discussion with reference to Figure 54, the system described herein can exploit the cluster distribution of fragment conformations to efficiently identify macrocycle conformations.
[0137] Figure 58 shows a translation vector that defines the translation from the beginning of a molecular fragment to the end of the molecular fragment, which, as explained above with reference to Figure 54, can be part of a beginning-end transformation that the systems described herein can use to identify macrocycle conformations.
[0138] Figure 59 illustrates a start-end transformation dictionary 1202 and a end-end transformation dictionary 1204. Each key in the start-end dictionary 1202 represents a discretized start-end transformation and is associated with a set of values that identify a fragment conformation having the discretized start-end transformation represented by the key. Each key in the end-end dictionary 1204 represents a discretized end-end transformation and is associated with a set of values that identify a fragment conformation having the discretized end-end transformation represented by the key. As discussed above with reference to Figure 54, the system described herein can efficiently identify a macrocycle conformation 1208 by a value associated with a key common to both dictionaries, e.g., key 1206.
[0139] The term "configured" is used herein with respect to systems and computer program components. With reference to a system of one or more computers configured to perform a particular operation or action, we mean a system that has installed thereon software, firmware, hardware, or a combination thereof that causes the system to perform the operation or action. With reference to one or more computer programs configured to perform a particular operation or action, we mean that the one or more programs contain instructions that, when executed by a data processing device, cause the device to perform the operation or action.
[0140] The subject matter and functional operations described herein may be implemented in digital electronic circuitry, tangibly embodied computer software or firmware, computer hardware (including the structures disclosed herein and their structural equivalents, or a combination of one or more of them). The subject matter described herein may be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a non-transitory tangible storage medium for execution by or control of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access storage device, or a combination of one or more of them. Alternatively, or in addition, the program instructions may be encoded in an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a receiving device suitable for execution by a data processing device.
[0141] The term "data processing apparatus" refers to data processing hardware and includes all kinds of apparatus, devices, and equipment for processing data, including programmable processors, computers, or multiple processors or multiple computers. The apparatus may also be or include special purpose logic circuits, such as FPGAs (field programmable gate arrays) or ASICs (application specific integrated circuits). In addition to hardware, the apparatus may optionally include code that creates an environment for running computer programs, such as code that constitutes processor firmware, protocol stacks, database management systems, operating systems, or one or more combinations thereof.
[0142] A computer program, which may also be referred to or described as a program, software, software application, app, module, software module, script, or code, may be written in any form of programming language, including compiled, interpreted, declarative, or procedural languages; and may be deployed in any form, whether as a stand-alone program or as a module, component, subroutine, or other suitable unit for use in a computing environment. A program may correspond to a file in a file system, but is not necessarily a file. A program may be stored as part of a file that holds other programs or data (e.g., a document in a markup language, a single file specific to the program, or multiple linked files, e.g., one or more scripts stored in a file that contains one or more modules, subprograms, or code portions). A computer program may be deployed to run on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected via a data communications network.
[0143] As used herein, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specialized functions. Typically, an engine is implemented as one or more software modules or components installed on one or more computers located at one or more locations. In some cases, one or more computers are specialized for a particular engine; in other cases, multiple engines may be installed and executed on the same computer or computers.
[0144] The processes and logic flows described herein can be performed by one or more programmable computers executing one or more computer programs to perform functions that operate on input data and generate output. The processes and logic flows can also be performed by special purpose logic circuitry (e.g., FPGAs or ASICs), or by a combination of special purpose logic circuitry and one or more programmed computers.
[0145] A computer suitable for executing a computer program may be based on a general-purpose microprocessor or a special-purpose microprocessor, or on any other type of central processing unit. In general, the central processing unit receives instructions and data from a read-only memory or a random access memory, or on both. The essential elements of a computer are a central processing unit that performs or executes instructions, and one or more memory devices that store instructions and data. The central processing unit and memory can be supplemented with or incorporate special-purpose logic circuits. In general, a computer also includes or is controllably connected to one or more mass storage devices (e.g., magnetic disk, magneto-optical disk, or optical disk) that store data, to receive data from or transfer data to the mass storage device, or both. However, a computer does not have to have such devices. In addition, a computer can be incorporated into other devices, examples of which include a mobile phone, a personal digital assistant (PDA), a portable audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive.
[0146] Computer-readable media suitable for storing computer program instructions and data include all kinds of non-volatile memory media and memory devices, examples of such computer-readable media include semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0147] To provide for user interaction, embodiments of the subject matter described herein can be implemented in a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) that displays information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) that allows the user to input to the computer. Other types of devices can be used to provide for user interaction as well; for example, the feedback provided to the user can be any type of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and the input from the user can be any type, including auditory, speech, or tactile input. Furthermore, a computer can interact with a user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser on a user's device in response to a request received from the web browser. A computer can also interact with a user by sending a text message or other form of message to a personal device (e.g., a smartphone running a messaging application) and receiving a response message from the user in return.
[0148] A data processing device for implementing a machine learning model may also include, for example, special purpose hardware accelerator units for handling the machine learning training or generation, i.e., inference, normal and numerical portions of the workload.
[0149] The machine learning model can be implemented and deployed using a machine learning framework, for example, the TensorFlow framework or the Jax framework.
[0150] An embodiment of the subject matter described herein can be implemented in a computing system that includes a back-end component (e.g., a back-end component as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer having a graphical user interface, a web browser, or an app) with which a user can interact in implementing the subject matter described herein, or includes any combination of one or more of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include local area networks (LANs) and wide area networks (WANs), e.g., the Internet.
[0151] The computing system may include a client and a server. Clients and servers are generally remote from each other and typically communicate through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., HTML pages) to a user device, e.g., a user device intended to display data, and receives user input from a user interacting with the device acting as a client. Data generated at the user's device (e.g., data generated as a result of user interaction) can be received from the device at the server.
[0152] Although this specification contains details of many specific implementations, these should not be considered as limiting the scope of the invention or the claims; rather, they should be considered as descriptions of features that may be specific to certain embodiments of a particular invention. Certain features described herein for separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described for a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, although features may be described above as functioning in a particular combination, or even claimed as such, one or more features in a claimed combination may, in some cases, be deleted from the combination, and the claimed combination may be subject to subcombinations or variations of the subcombinations.
[0153] Similarly, although operations may be shown in the figures or recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order or sequence shown, or that all of the operations shown be performed, to achieve desirable results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the program components and systems described may generally be incorporated together in a single software product or packaged in multiple software products.
[0154] Although certain embodiments of the subject matter have been described above, other embodiments are within the scope of the following claims. For example, the items recited in the claims can be performed in a different order and still achieve desirable results. By way of example, the processes illustrated in the accompanying figures need not necessarily be in the particular order shown, nor necessarily sequentially numbered, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
[0155] Working Example Small macrocycles consisting of only three or four standard and nonstandard amino acids are among the most pharmacologically potent natural products isolated to date, but their abundance in nature is so low that there is currently no method to systematically explore their structural space. A general computational method is reported here to identify a large sampling of closed macrocycles formed by combinations of α-, β-, γ-, δ-, and ε-type aminobenzoates, aminophenylacetic acids, aminomethylbenzoates, and aminomethylpicolinates, along with oxazoles, thiazoles, oxazolidines, thiazolidines, triazoles, and thioethers. To perform a computationally tractable search, we rapidly compare rigid body transformations between the terminal amides of one and two monomer units by hashing to identify combinations of conformers that close to form macrocycles. The above method was used to identify 14.9 million macrocycles constructed from three or four monomers containing over 42,000 unique combinations of standard and non-standard backbones and up to 32 membered rings; of these, approximately 7.4 million meet the rule of five criteria for drug-like compounds. We chemically synthesized 30 macrocycles predicted to adopt a single low-energy state and determined X-ray or NMR structures for 18; 15 of these closely matched the corresponding computational models. Our results greatly increase the number of known small macrocycles and should facilitate structure-based drug design and other applications that rely on small molecule libraries.
[0156] Macrocycles composed of up to four α- or β-amino acids, aminobenzoates, oxazoles, or thiazoles possess biological activities ranging from antifungal or antibiotic properties to cancer cytotoxicity and pain relief. Despite the diversity of biological activities observed in this class of natural products, chemists and biologists have mainly limited their exploration of this space to closely related homologs of those already sampled; in fact, the majority of currently approved macrocyclic drugs for human use are primarily derived from natural products. Computational methods that allow rapid and exhaustive search of the space of possible macrocycles could greatly facilitate the discovery of natural product-like compounds with novel biological activity, but such methods do not currently exist. Methods have been developed to systematically enumerate large macrocycles formed from all α-amino acids by first generating cyclic conformers of the polyglycine backbone and then performing sequence optimization, but these methods require random sampling of backbone torsion angles to facilitate the identification of closed cyclic conformers of the macrocycle, and therefore cannot be easily extended to generate macrocycles formed from monomers with a wide range of backbone chemistries. The random search of backbone torsion angle combinations and subsequent sequence design is easy when the macrocycle is constructed from a single backbone building block (e.g., only α-amino acids), but becomes difficult to handle when many different backbones are involved. There are approximately 60,000 different 2D chemical structures for the 2-, 3-, and 4-residue macrocycles that can be constructed from the 22 building blocks shown in Figure 1A, but each of the 60,000 can be further diversified because of the millions of potential combinations of different side chains. This chemical space is too large to rely on random sampling to identify only a few linear sequences that can close to form a ring.
[0157] We set out to develop a computationally tractable method to sample the very large chemical space of possible compact macrocycles (Figure 1). We investigated a three-step approach to perform this search on an input set of monomers with highly diverse chemical structures. In the first step, low-energy conformers are generated for each of the monomers. In the second step, the rigid body transformations associated with these monomer conformations and all combinations that pair the monomer conformations into dipeptide-like structures are calculated and stored in a hash table. In the third step, macrocycles are rapidly and systematically generated by identifying pairs of entries in the hash table whose sum of rigid body transformations is close to zero (so the fragments form a ring together).
[0158] To identify low-energy conformers for a given set of monomers, we developed an adaptive grid search of the backbone and side-chain torsion angles that form a potential energy surface for each of the monomers. We used an atom-in-molecules deep neural network (AIMNet) modified with London dispersion force correction for the energy evaluation; it requires only atomic coordinates and atom numbers as input, and therefore does not require complex atom typologies such as many molecular mechanics-based energy functions, and is as accurate as DFT-based methods, which are several orders of magnitude more expensive to evaluate. From these energy surfaces, we extract low-energy conformers to use as building blocks for macrocycle construction. To enable rapid identification of one- or two-monomer conformer pairs that will generate macrocycles by close combination, six-dimensional rigid-body transformations between the terminal amide and dipeptide conformers of each low-energy monomer are computed in both the N-to-C and C-to-N directions and stored in a hash table. By simply taking the intersection of the hash values present in the N-to-C hash table of one component and the C-to-N hash table of another component, pairs of these components that share the complementary terminal orientations required to form a closed macrocycle are almost instantly identified (e.g., the green hash entry pairs on the left and right in Figure 1b). This hashing approach requires much less computational time than would be required to generate explicit backbone coordinates for each of the different possible monomer combinations and evaluate the ring closure of a huge number of possible conformers; it also does not require information about the number, identity, and connectivity of atoms between the terminal amides, since only their relative orientations in space are considered for ring closure determination. The generality of this approach allows for the construction of macrocycles from all possible unique combinations of monomers shown in Figures 1A and 5.Although Figure 1B outlines the above process of constructing a four-residue macrocycle from building blocks of dipeptide conformers, this approach can be generalized to any size; for example, a three-residue macrocycle is discovered by searching for dimers that close with a monomer (or vice versa), while a two-residue macrocycle is identified by searching for monomers that close with other monomers.
[0159] Enumeration of closed macrocyclic skeletons We used our approach to explore a very large space of three- and four-residue macrocycles and their appropriate optical isomers, as well as variants of C-terminal dimethylamides, constructed from 130 unique monomers (Figure 5). These monomers are divided into 22 classes (hereafter referred to as species); these species are differentiated by the atomic number and atomic hybridization of the atoms in the backbone of each monomer (Figure 1A); we represent each species with a letter (e.g., a for alpha amino acids; b for beta amino acids; d for delta amino acids, etc.; see legend to Figure 1A). From the potential energy surface of each monomer, we created hash tables for both monomers and dimers, and then performed a systematic search through all combinations of monomer / dimer, dimer / monomer, and dimer / dimer hash tables generating ensembles for each chemical for which a matching hash value could be found. These hash tables contain nearly 9 billion unique conformations for the monomeric and dimeric subunits; exploring ring closures with explicit conformer generation across all combinations takes approximately 10 19 would require macrocycle constructions, which is far more than is feasible; our hashing approach reduces the complexity required to evaluate all of these combinations from O(n2) to O(n).
[0160] A systematic search of the hash table yielded 14.9 million closed macrocycles, including 9-32 membered rings, belonging to 3494 three-residue species and 38544 four-residue species (comprising all unique circular permutations of cyclic species). The chemical space of three- and four-residue species (Figure 2A) is greatly expanded over previous searches (Figure 7); 206 unique chemicals belonging to 23 species have high-resolution structures published in the Cambridge Structural Database (CSD), and 13,932 chemicals belonging to 397 species are represented in the pubchem database. Previous computationally designed macrocycles were constructed from only one chemical species (alpha amino acids). Our 14.9 million closed macrocycles are distributed across a range of molecular weights and calculated octanol / water partition coefficients, but roughly half meet the rule of five (Ro5) criteria for drug-like compounds (Figure 8) (13).
[0161] Each species samples a different 3D shape (Fig. 2B and C). To characterize the structural diversity of each species, we compartmentalized the backbone torsion angles into 60° compartments and represented each macrocycle conformer as a string of these compartments. Most of the thousands of conformers generated for each species had only a few compartment strings (Fig. 9), likely reflecting the torsional preferences of the constituent amino acids. Principal moment of inertia (PMI) analysis of conformers across all the compartment strings sampled for each sequence revealed that different macrocycle species sample distinct regions of shape space (14). In some cases (e.g., in the case of aaaa and aaac), the specific 3D shapes sampled were highly restricted by ring-closure constraints, regardless of sequence and torsional diversity (Fig. 9).
[0162] For each monomer sequence, the hash-based ring closure generates an ensemble of macrocycle structural models built from low-energy monomer conformers. Given sufficient computational power, for each of these ensembles, a minimization in AIMNet can be performed (to incorporate monomer / monomer interactions, strain due to ring closure, etc.) and the energy landscape mapped by the set of all low-energy conformers of the sequence (e.g., the Boltzmann weight P near Using an approximation; (15)), it is possible to assess the degree to which a monomer sequence uniquely encodes a single low-energy conformer.
[0163] To evaluate this hash-based conformational sampling and subsequent AIMNet minimization approach, we used it to map the energy landscape of eight macrocycles with known structures composed of three or four α-amino, β-amino, and aminobenzoic acid monomers. The lowest energy conformers in the hashed ensembles after AIMNet minimization had backbone RMSDs of 0.3 Å or better compared to the corresponding X-ray crystal structures in all eight cases (FIG. 10). Thus, the hashed sampling method coupled with AIMNet calculations is capable of predicting the structures of small macrocycles from their monomer sequences. However, performing a complete energy calculation for all 14 million generated ensembles is not computationally tractable; minimization of the approximately 23 billion conformers sampled across the 14 million ensembles required roughly 10 8 Instead, we focused on two subclasses of closed macrocycles that are hypothesized to be particularly likely to have a single unique ground state: those with local interactions that strongly favor one or a few closed ring states, and those that contain nonlocal hydrogen bonds between backbone amides.
[0164] To identify the first local code subclass, we generated conformational ensembles from a hash table containing only monomers whose minimum energy states are within 1 kcal / mol; such macrocycles were optimized in torsion from the depth of the complex building blocks. Because the resulting macrocycle space is still very large, for computational tractability, we probed for the use of the number of different torsion compartment strings sampled in these ensembles as an indication of the degree to which the torsion biases together with the ring-closure constraints are sufficient for the identification of a single low-energy minimum. After AIMNet minimization of a subset of the ensembles, we found that when the number of torsion compartment strings sampled is 5 or less, there is typically a deep energy minimum in the AIMNet landscape near a single state with Pnear>0.9 (Figure 11). We identified ∼380,000 three- and four-residue sequences whose hashed ensembles contained more than 50 closed-ring conformers across ∼5 torsion sections, and performed AIMNet minimization on ∼4800 of them; roughly 85% of these had a single low-energy structure with Pnear > 0.9.
[0165] We have synthesized 13 such macrocycles and have been able to determine the structures of 10 of them using X-ray crystallography and NMR spectroscopy (Fig. 3, 12; 19-47). Six of the macrocycles produced crystals of sufficient quality in vapor diffusion experiments for structure determination using direct methods; seven of the macrocycles produced NMR spectra in DMSO-d6 with sufficient dispersion to allow unambiguous assignment of the major ROEs in the ROESY spectra; for these we created 3D models by relaxing conformers generated from distance restraints based on the built-in ROEs in unconstrained molecular dynamics (MD) simulations and selecting the minimum energy conformer (see Methods). Eight of the eleven experimentally determined structures had RMSDs of 0.8 Å or less in superposition to their respective designed models, with small differences in ring puckering of five- and six-membered rings in single side-chain and backbone torsion angle pairs (larger differences were observed for the remaining three macrocycles (Figure 12)). For two of the eleven macrocycles, it was possible to determine both X-ray crystallographic and NMR structures. In these two cases (aabc and aaam), both the NMR and X-ray crystallographic structures were very similar to the respective designed models and to each other (Figure 13).
[0166] Macrocycles with solved structures (Figure 3) are constructed from a total of eight different monomer species (a, b, c, h, k, m, q, s) arranged into six unique macrocycle species (aas, aaam, aaaq, aabc, ahah, akak). The designed aas contain two types of hydrogen-bonding interactions: one between the backbone amides and the other involving the primary amide of the monomer. Two macrocycle aaams and one macrocycle akak both contain proline-like monomers that sample cis amides. In the NMR structures of the two macrocycle aaams, these motifs result in strong ROE between the Cα protons of the proline-like monomers and the Cα protons of their immediate N-terminal neighbors; these macrocycles contain α amino acids in a highly extended β-sheet-like conformation, and the cis amides rapidly rotate the backbone to allow ring closure (Figure 14). The macrocycle aaaq contains an extended pentafluorophenylalanine residue that embeds its amide NH relative to 4-aminomethylphenylacetic acid, shielding it from solvent and causing a downfield shift of the amide NH with a temperature shift coefficient (Tcoeff) of 0.9 ppb / K (the backbone amides in the remaining monomers of this design are upfield shifted, all with Tcoeffs less than -5.0 ppb / K; see Supplementary Data). The macrocycles aaam and aaaq provide two different means to mimic the extended conformation of poly-α-amino acid peptides, similar to the conformations adopted by many protease and kinase substrates (16, 17), and are therefore useful when targeting these enzymes.
[0167] To identify the second subclass of macrocycles (those containing transannular hydrogen bonds), we created a dimer hash table with the goal of increasing the sampling of these rare interactions (those formed by the terminal amides of conformers); for inclusion, we created a hash table that includes conformers whose terminal amides form hydrogen bonds that increase the monomer energy cutoff to within 4.0 kcal / mol of the global minimum. With the goal of reducing local strain, we selected monomer species and low energy side chain rotamers (given backbone torsion angles) after identifying macrocycle species that contain long-range hydrogen bonding interactions. After ensemble enumeration and AIMNet minimization based on the hash table, we selected designs that were experimentally characterized to have much lower energies than predicted by simple model-based monomer compositions (Methods).
[0168] Seventeen macrocycles designed to contain long-range hydrogen-bonding interactions were prepared, and high-resolution structures could be determined for seven of them (Figures 4 and 19-47). Four macrocycles formed crystals of sufficient quality in vapor diffusion experiments for structure determination using direct methods; three macrocycles showed NMR spectra with sufficient dispersion to allow unambiguous assignment of the major ROEs in the ROESY spectra. Eight of the 17 macrocycles showed 1D-NMR in DMSO-d6, suggesting a mixture of multiple states; six of these collapsed to a single major species in CDCl3. Only three macrocycles showed sufficient dispersion for unambiguous assignment, while ten of the 17 had sufficient dispersion to allow measurement of amide Tcoeffs, consistent with the presence of designed transannular hydrogen bonds in all ten cases. All seven examples of experimentally determined structures superimpose onto their respective designed models with RMSDs of less than 0.8 Å. As with the macrocycle shown in Figure 3, minor differences between the designed models of these hydrogen-bond-containing macrocycles and the experimental structures include differences in ring puckering and a 180 degree rotation of the backbone amide.
[0169] The macrocycles shown in Figure 4 are constructed from eight different monomer species (a, b, g, h, i, l, m, p) arranged into six unique macrocycle species (aaam, aaap, aabi, aagb, aalm, aahh) and stabilized by transannular hydrogen bonds between backbone amides containing non-α amino acids. The aalm-based macrocycles contain hydrogen-bonding fragments constructed from monomers with predominantly sp2 hybridized atoms in the backbone (lm); the aagb-based macrocycles contain backbone hydrogen-bonding fragments that contain many more sp3 hybridized atoms than are present in the α amino acid backbone (gb); and the aabi-based macrocycles contain hydrogen-bonding fragments that combine these two features (bi). Six of the macrocycles contain a continuous segment of α-amino acids that mimic a β-turn common to protein / protein interfaces; the aagb-, aabi-, and aaam-based macrocycles contain a turn similar to a type I β-turn; and the aalm-based macrocycle contains a turn similar to a type II β-turn. The N-methylated amino acid residues present in both aaap-based macrocycles adopt a cis-amide, resulting in a type VI-like turn. Despite the presence of these distinct β-turn-like features in the aaam macrocycle, the spacing and orientation of the two phenylalanine side chains mimics the spacing and orientation of side chains located at i and i+4 of an α-helix (Figure 15: this spacing is absent in either of the two aaap-based macrocycles, which are aaam isomers at the backbone atom level). This mimicry of β-turn and helix configurations is useful for targeting proteins that recognize these structural elements.
[0170] Passive membrane diffusion measurements were performed on all 29 macrocycles designed by the two approaches using the parallel artificial membrane permeability assay (PAMPA) (Figures 3, 4, and 16). Nearly all of the macrocycles were highly membrane permeable. Sixteen of the 29 macrocycles showed log(P app) was greater than -6; and only three were not detected in the acceptor wells. Under the same conditions, the small molecule drug propranolol exhibited a log(P app )=-5.36 (Figure 16). As expected, designs with exposed polar groups had lower permeability, especially those with hydroxyls and primary amides in the side chains. Overall, the hydrogen bond-containing designs were more permeable than designs based on torsional optimization; this is likely due to fewer exposed backbone polar groups.
[0171] Compound synthesis All the macrocycles shown in Table 3 were synthesized. Below is the detailed protocol for the synthesis of PJS-MPRO-1-j (compound 218 in Table 3). A similar protocol was used for the preparation of other macrocycles. Dichloromethane was added to a vessel containing CTC resin (0.1 mmol, 0.65 mmol / g, 0.15 g) and 2-(2-((((9H-fluoren-9-yl)methoxy)carbonyl)amino)phenyl)acetic acid (0.38 mg, 0.10 mmol, 1.00 equiv.) and nitrogen was bubbled through. Diisopropylethylamine (4 equiv.) (DIEA) was added dropwise and stirred for 2 h; at this point methanol (0.15 mL) was added. The mixture was bubbled through for an additional 30 min. The reaction vessel was removed from the solvent and washed five times with dimethylformamide (DMF). After washing, a solution of 20% piperidine in DMF was added to the reaction vessel and then bubbled for 30 minutes; at this point, the vessel was stripped of solvent and washed five times with DMF. The reaction vessel was added with the Fmoc-protected amino acid solution followed by the activating reagent solution. The reaction was allowed to continue for 1 hour with stirring by bubbling nitrogen, at which point the reaction vessel was stripped of solvent and washed five times with DMF. This cycle of deprotection, washing, coupling, and washing was repeated until the full-length linear peptide was completed, at which point the N-terminal Fmoc protecting group was removed. After completion of the linear peptide, the resin was washed three times with methanol and then dried under vacuum. The reagents used in the coupling steps for the synthesis of PJS-MPRO-1-j are shown in Table 4.
[0172] [Table 4]
[0173] The linear side-chain protected peptide was cleaved from the resin by adding cleavage solution (1% trifluoroacetic acid in dichloromethane) to the resin at room temperature. This cleavage was performed twice (3 min each) with continuous nitrogen bubbling. The filtrates of the cleavage reaction were collected, pooled, and diluted with 100 mL of dichloromethane. TBTU (64.2 mg, 2 eq.) and HOBt (27 mg, 2 eq.) were added to dilute the linear protected peptide. The pH of the solution was adjusted to about 8.0 with DIEA. The mixture was stirred at 25° C. for 30 min. The reaction mixture was then washed with 1 M HCl, and the organic layer was concentrated to dryness under reduced pressure. The resulting residue was treated with a global deprotection solution (92 / 5% trifluoroacetic acid, 2.5% 3-mercaptopropionic acid, 2.5% TIS, 2.5% water) and stirred for 30 min. The mixture was precipitated with cold isopropyl ether and centrifuged at 3000 rpm for 3 min. The resulting insoluble material was then precipitated two more times with additional isopropyl ether, removing the solvent each time, and the resulting solid, containing the crude cyclic peptide, was dried under vacuum for 2 hours to remove residual isopropyl ether.
[0174] The peptide was then purified from this material using preparative reverse-phase HPLC (A: water with 0.075% TFA, B: ACN) to give PJS-MPRO-1-j (3.1 mg, 94.6% purity, 5.14% yield) as a white solid.
[0175] Activity Test Compounds were tested for their ability to alter the function of one of the following protein targets: Mycobacterium tuberculosis ClpP1P2 protease, MCL-1, SARS-CoV-2 major protease, FKBP51, Mycobacterium tuberculosis DnaN, tubulin, insulin-degrading enzyme, thrombin, and MDM2.
[0176] Assay The main protease of SARS-CoV-2 Compounds were tested in 10 dose IC50 singlets with 3-fold serial dilutions starting at 200 μM. The protease activity was monitored by measuring the increase in fluorescent signal over time.
[0177] CLPP1P2 Compounds were tested in 10 dose IC50 singlets with 3-fold serial dilutions starting at 200 μM. The protease activity was monitored by measuring the increase in fluorescent signal over time.
[0178] MCL-1 Compounds were tested at 10 doses IC50 singlet with 3-fold serial dilutions starting at 200 μM. Compounds were added to the enzyme solution by using Acoustic Technology. Compounds were incubated with the enzyme for 10 minutes at room temperature. Then, each compound was added and the mixture was incubated for another 10 minutes. Finally, anti-GST-Tb was added. After 60 minutes of incubation, HTRF signals were measured.
[0179] Tubulin The assay was performed using a commercially available Cytoskeleton tubulin polymerization assay kit (SKU: BK011) according to the manufacturer's instructions.
[0180] Thrombin The assay was performed using a commercially available Abcam Thrombin Activity Assay Kit (SKU: ab197007) according to the manufacturer's instructions.
[0181] MDM2 The assay was performed using a commercially available PerkinElmer alphalisa assay kit (part number: AL3168C) according to the manufacturer's instructions.
[0182] DanN The assay was carried out by a competitive assay using biolayer interferometry.
[0183] FKBP51 Compounds were tested in 10 dose IC50 singlets with 3-fold serial dilutions starting at 200 μM.
[0184] Activity Compound 11 (4.58 μM), compound 16 (13.5 μM), compound 18 (28.2 μM), compound 19 (40.6 μM), compound 29 (22.8 μM), compound 119 (18.9 μM), compound 121 (50.1 μM), and compound 122 (12.3 μM) were the most active in the MCL-1 assay.
[0185] Compound 32 (100 μM) and compound 35 (20 μM) were the most active in the MDM2 assay.
[0186] Compound 131 (4.07 μM), compound 132 (1.25 μM), compound 133 (0.844 μM), compound 134 (1.22 μM), compound 135 (1.47 μM), compound 136 (1.47 μM), compound 137 (1.27 μM), compound 138 (2.14 μM), compound 139 (2.37 μM), compound 140 (1.53 μM), compound 141 (1.84 μM), compound 142 (1.89 μM), compound 143 (2.65 μM), compound 144 (1.26 μM), and compound 218 (17.8 μM) were the most active in the assay for the SARS-CoV-2 major protease.
[0187] conclusion Small macrocycles commonly found in nature exist in solution as a mixture of conformers. Our approach allows for the rapid exploration of a chemically diverse, large macrocycle space with the aim of identifying macrocycles that exist primarily as a single conformer. Both of the two strategies described herein (one using conformationally constrained building blocks that favor specific structures, and one incorporating longer-range hydrogen-bonding interactions) yielded macrocycles with well-defined structures. Our finding that the macrocycle species defines its globular shape independent of its side-chain substitutions suggests that diversity-oriented synthesis across a wide range of shapes should focus on increasing the number of species that can be combinatorially synthesized, rather than the diversity of side chains placed on a single species or a small number of species. Encouraging for future therapeutic applications, nearly half of the macrocycles we characterized are membrane permeable (log(Papp)>-6.0 in PAMPA), and this percentage is likely to be much higher with explicit design with respect to permeability (e.g., disfavoring compounds with exposed NH groups). The chemical space can be easily expanded by including more diverse side chains or additional non-standard scaffolds in the hashing step, and almost all compounds in this class can be easily synthesized using standard manual Fmoc-based solid-phase peptide synthesis followed by solution-phase cyclization (see Methods).
[0188] The very large and diverse collection of compounds described herein offers a great new tool in drug discovery. The rigidity of the molecule should lead to specific target affinity due to lower entropy loss upon target binding, and off-target binding should be reduced due to a smaller number of different states; the rule-of-five compatibility in the generation should provide membrane permeability and other desirable pharmacological properties. Libraries of rule-of-five compatible macrocycles existing in a single state can be screened in silico and / or experimentally to identify new lead compounds that bind to the target of interest. For targets to which small molecule fragments are already known to bind, custom rigid macrocycle libraries that introduce this specific functionality can be easily created by searching a hash table for the fully closed macrocycle that introduces the fragment. References 1. V. Sarojini, AJ Cameron, KG Varnava, WA Denny, G. Sanjayan, Cyclic Tetrapeptides from Nature and Design: A Review of Synthetic Methodologies, Structure, and Function. Chem. Rev. 119, 10318-10359 (2019). 2. JT Mhlongo, E. Brasil, BG de la Torre, F. Albericio, Naturally Occurring Oxazole-Containing Peptides. Mar. Drugs. 18 (2020), doi:10.3390 / md18040203. 3. T. Degenkolb, W. Gams, H. Brueckner, Natural cyclopeptaibiotics and related cyclic tetrapeptides: structural diversity and future prospects. Chem. Biodivers. 5, 693-706 (2008). 4. M. J. Ferracane, A. C. Brice-Tutt, J. S. Coleman, G. G. Simpson, L. L. Wilson, S. O. Eans, H. M. Stacy, T. F. Murray, J. P. McLaughlin, J. V. Aldrich, Design, Synthesis, and Characterization of the Macrocyclic Tetrapeptide cyclo[Pro-Sar-Phe-d-Phe]: A Mixed Opioid Receptor Agonist-Antagonist Following Oral Administration. ACS Chem. Neurosci. 11, 1324-1336 (2020). 5. E. M. Driggers, S. P. Hale, J. Lee, N. K. Terrett, The exploration of macrocycles for drug discovery--an underexploited structural class. Nat. Rev. Drug Discov. 7, 608-624 (2008). 6. M. Muttenthaler, G. F. King, D. J. Adams, P. F. Alewood, Trends in peptide drug discovery. Nat. Rev. Drug Discov. 20, 309-325 (2021). 7. K. T. Mortensen, T. J. Osberger, T. A. King, H. F. Sore, D. R. Spring, Strategies for the Diversity-Oriented Synthesis of Macrocycles. Chem. Rev. 119, 10288-10317 (2019). 8. G. Sangouard, A. Zorzi, Y. Wu, E. Ehret, M. Schuettel, S. Kale, C. Diaz-Perlas, J. Vesin, J. Bortoli Chapalay, G. Turcatti, C. Heinis, Picomole-Scale Synthesis and Screening of Macrocyclic Compound Libraries by Acoustic Liquid Transfer. Angew Chem Int Ed Engl. 60, 21702-21707 (2021). 9. P. Hosseinzadeh, G. Bhardwaj, V. K. Mulligan, M. D. Shortridge, T. W. Craven, F. Pardo-Avila, S. A. Rettie, D. E. Kim, D.-A. Silva, Y. M. Ibrahim, I. K. Webb, J. R. Cort, J. N. Adkins, G. Varani, D. Baker, Comprehensive computational design of ordered peptide macrocycles. Science. 358, 1461-1466 (2017). 10. E. A. Coutsias, K. W. Lexa, M. J. Wester, S. N. Pollock, M. P. Jacobson, Exhaustive conformational sampling of complex fused ring macrocycles using inverse kinematics. J. Chem. Theory Comput. 12, 4674-4687 (2016). 11. R. Zubatyuk, J. S. Smith, J. Leszczynski, O. Isayev, Accurate and transferable multitask prediction of chemical properties with an atoms-in-molecules neural network. Sci. Adv. 5, eaav6490 (2019). 12. E. Caldeweyher, S. Ehlert, A. Hansen, H. Neugebauer, S. Spicher, C. Bannwarth, S. Grimme, A generally applicable atomic-charge dependent London dispersion correction. J. Chem. Phys. 150, 154122 (2019). 13. C. A. Lipinski, F. Lombardo, B. W. Dominy, P. J. Feeney, Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings. Adv. Drug Deliv. Rev. 46, 3-26 (2001). 14. W. H. B. Sauer, M. K. Schwarz, Molecular shape diversity of combinatorial libraries: a prerequisite for broad bioactivity. J. Chem. Inf. Comput. Sci. 43, 987-1003 (2003). 15. G. Bhardwaj, V. K. Mulligan, C. D. Bahl, J. M. Gilmore, P. J. Harvey, O. Cheneval, G. W. Buchko, S. V. S. R. K. Pulavarti, Q. Kaas, A. Eletsky, P.-S. Huang, W. A. Johnsen, P. J. Greisen, G. J. Rocklin, Y. Song, T. W. Linsky, A. Watkins, S. A. Rettie, X. Xu, L. P. Carter, D. Baker, Accurate de novo design of hyperstable constrained peptides. Nature. 538, 329-335 (2016). 16. P. K. Madala, J. D. A. Tyndall, T. Nall, D. P. Fairlie, Update 1 of: Proteases universally recognize beta strands in their active sites. Chem. Rev. 110, PR1-31 (2010). 17. C. J. Miller, B. E. Turk, Homing in: mechanisms of substrate targeting by protein kinases. Trends Biochem. Sci. 43, 380-394 (2018). 18. J. Gavenonis, B. A. Sheneman, T. R. Siegert, M. R. Eshelman, J. A. Kritzer, Comprehensive analysis of loops at protein-protein interfaces for macrocycle design. Nat. Chem. Biol. 10, 716-722 (2014). 19. M. Guharoy, P. Chakrabarti, Secondary structure based analysis and classification of biological interfaces: identification of binding motifs in protein-protein interactions. Bioinformatics. 23, 1909-1918 (2007). 20. B. L. Sibanda, T. L. Blundell, J. M. Thornton, Conformation of beta-hairpins in protein structures. A systematic classification with applications to modelling by homology, electron density fitting and protein engineering. J. Mol. Biol. 206, 759-777 (1989). 21. P. N. Lewis, F. A. Momany, H. A. Scheraga, Chain reversals in proteins. Biochim. Biophys. Acta. 303, 211-229 (1973). 22. DB Diaz, SD Appavoo, AF Bogdanchikova, Y. Lebedev, TJ McTiernan, G. Dos Passos Gomes, AK Yudin, Illuminating the dark conformational space of macrocycles using dominant rotors. Nat. Chem. 13, 218-225 (2021).
[0189] calculation method The entire computational pipeline is implemented using several python packages. The following sections provide implementation details and elaborate on the computational details.
[0190] Residue parameters To facilitate the construction and manipulation of coordinates for both monomers and macrocycles, we created a python dictionary that contains various parameters for each residue. These parameters store multiple Boolean values for the residue's chemical identity, a python list of various atom indices, and the residue's SMILES string. We use these residue parameters at various points throughout the scripts and workflows below. Below is an abbreviated version of the residue_params dictionary that contains only glycine: residue_params = { 'AGLY' : { 'smiles' : 'CC(=O)NCC(=O)NC', 'dofs' : [[1, 3, 4, 5], [3, 4, 5, 7]], 'proline_like' : False, 'preproline' : False, 'chiral' : False, 'lower_smiles' : 'CC(=O)', ’residue_smiles’ : ’NCC(=O)’, ’upper_smiles’ : ’NC’, ’lower_indices’ : [0, 1, 2, 9, 10, 11], ’residue_indices’ : [3, 4, 5, 6, 12, 13, 14], ’upper_indices’ : [7, 8, 15, 16, 17, 18], ’lower_connect_indices’ : [1, 2, 3], ’upper_connect_indices’ : [5, 6, 7], ’backbone_indices’ : [0, 1, 3, 4, 5, 7, 8], ’atoms’ : [6, 6, 8, 7, 6, 6, 8, 7, 6, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1], ‘atom_types’ : [’SP3’, ’SP2’, ’SP2’, ’SP2’, ’SP3’, ’SP2’, ’SP2’, ’SP2’, ’SP3’, ’S’, ’S’, ’S’, ’S’, ’S’, ’S’, ’S’, ’S’, ’S’, ’S’], ‘bonds’ : [[0, 1], [1, 2], [1, 3], [3, 4], [4, 5], [5, 6], [5, 7], [7, 8], [0, 9], [0, 10], [0, 11], [3, 12], [4, 13], [4, 14], [7, 15], [8, 16], [8, 17], [8, 18]], ‘bond_types’ : [’SINGLE’, ’DOUBLE’, ’SINGLE’, ’SINGLE’, ’SINGLE’, ’DOUBLE’, ’SINGLE’, ’SINGLE’, ’SINGLE’, ’SINGLE’, ’SINGLE’, ’SINGLE’, ’SINGLE’, ’SINGLE’, ’SINGLE’, ’SINGLE’, ’SINGLE’, ’SINGLE’], 'rotamers': 2, 'charges' : [-0.18, 0.25, -0.38, -0.5, 0.03, 0.27, -0.39, -0.47, -0.08, 0.12, 0.1, 0.11, 0.23, 0.15, 0.11, 0.25, 0.14, 0.13, 0.11], }, }
[0191] Boltzmann Adaptive Sampling We developed a novel sampling protocol that generates the full potential energy surface of monomeric residues; we name this protocol Boltzmann adaptive sampling (BAS).
[0192] The iterative process of Delaunay triangulation construction, point selection, and molecular energy evaluation is repeated until the potential energy surface of the molecule converges. The BAS scheme utilizes the Delaunay triangulation to maintain the relationship of all sampled points to each other. In this protocol, every point of the Delaunay triangulation is a dihedral angle value of a sampled conformation. The molecular conformation and energy associated with the point are preserved. Bond torsions with many possible values, such as Φ and Ψ of α-amino acids, were sampled in BAS; bond torsions with a very small number of possible angles, such as amide bonds, are considered "rotamers" that are not adaptively sampled, but instead are fixed in the cis or trans configuration and evaluated in two separate runs. Similarly, dihedral angles that always vary in concert with each other (i.e., the ring flips) were treated as separate "rotamers" and evaluated in two separate runs.
[0193] The sampling protocol is initialized by sampling a sparse regular lattice with 7 points per dimension (i.e., rotatable bonds). This turns out to sample the most common dihedral angles (-180°, -120°, -60°, 0°, 60°, 120°, 180°). Note that the initial and final dihedral angles are identical, but periodicity is not considered in the implementation of the Delaunay triangulation that works in higher dimensions. After evaluation of the initial lattice conformation, a Delaunay triangulation is generated to guide further sampling. Each simplex in the Delaunay triangulation is evaluated using a loss function derived from the Boltzmann probability (see below). The simplex with the maximum loss is then selected. After selecting a simplex, the center of the simplex is chosen with a small random perturbation to avoid deterministic sampling. That point corresponds to the next twist set to be evaluated. After evaluating the point by minimization of the conformation with the new dihedral angles, the point is returned to the Delaunay triangulation. As a result, a simplex is selected that splits into multiple new simplexes. This process of simplex identification, simplex center evaluation, and Delaunay triangulation update can be performed individually or in batches. Although performing this process on only one simplex at a time is optimal for sampling the simplex with the maximum loss, we perform the process in small batches (<32) because updating the Delaunay triangulation takes a significant amount of time. To allow 99% of the run time to be spent evaluating the molecular energy, we developed a dynamic controller that adjusts the batch size. This controller was especially necessary in high dimensions because computing the Delaunay triangulation scales poorly with the number of dimensions and samples. Without a controller to adjust the batch size, the protocol would spend most of its time selecting sampling points instead of evaluating their energies. The loss function that determines the next simplex to sample is the most important part of this protocol.We devise what we call the volume-adjusted Boltzmann probability, which guides the adaptive lattice to sporadically sample both small simplexes with points of very low relative energy or large simplexes that were not explored, thereby avoiding the adaptive lattice from further subdivision of the lowest energy simplex, which would result in narrow sampling.
[0194] Once a new point is selected for evaluation, a monomer is constructed and subjected to AIMNet minimization. We found that without minimization, the resulting potential energy surface for the molecule has sharp discontinuities. We found that minimizing the conformation under bond torsion angle constraints produces a smooth potential energy surface that qualitatively reproduces known potential energy surfaces, such as that of alanine. It was essential to perform a full minimization of the molecule; this is because small changes in bond lengths and bond angles (which are not constrained in the minimization) can have a large effect on the energy of a conformation. We also found that using the nearest conformation as a starting point and imposing constraints before imposing larger degrees of freedom improved run time and convergence of the minimization trajectory. This is likely due to the nearest neighbors having similar perturbations to the smaller degrees of freedom.
[0195] An important metric for the performance of a novel sampling protocol is the rate of convergence of the potential energy surface. The potential energy surface can be generated by linearly interpolating the Delaunay triangulation. The more points sampled in the Delaunay triangulation, the more accurate the linear interpolation of the points will be in predicting the novel interpolated points. Once a sufficient number of points have been sampled, the potential energy surface for a given molecule will converge, and adding additional points will not affect the potential energy surface. Once the potential energy surface has converged, the energy of any conformation can be accurately predicted.
[0196] We found that compared to a naive grid search, BAS identifies low energy conformers of monomers, especially for monomers that are torsionally constrained or contain many rotatable bonds, such as aminobutyric acid. For the monomers used in this study, BAS performed well to generate potential energy surfaces, but it proved difficult to sample the potential energy surfaces of molecules with more than six degrees of freedom. This barrier is due to the time to compute the Delaunay triangulation, which is exacerbated by the fact that the higher the dimension, the more sampling points are generally required. For example, for a million points and six dimensions, it takes the same amount of time to update the Delaunay triangulation for 100 additional points as it takes to energy minimize these 100 new points. This discrepancy between the triangulation update and the energy evaluation of the new points on the surface becomes more dramatic as the dimension increases (i.e., more bond torsions are sampled). For these reasons, BAS is ideal for sampling the potential energy surfaces of monomers rather than macrocycles.
[0197] We used BAS to generate potential energy surfaces for each monomer shown in Figure 5. In these calculations, surfaces were generated for each monomer containing both a C-terminal methylamide and a C-terminal dimethylamide modification. The latter is preproline, denoted by adding _pp after the residue name. The potential energy surfaces are saved to disk as HDF5 files containing the AIMNet-minimized conformers of the monomer and their associated AIMNet energies (expressed in kcal / mol) for use in subsequent hashing and macrocycle construction (see below). In the following methods, these potential energy surfaces are referred to as the "conformer database". The following commands show how to generate a potential energy surface for glycine: OMP_NUM_THREADS=1 python run_residue.py --residue AGLY --rotamer 0 --number_of_threads 4
[0198] Sequence compatibility check In the following method, frequent checks are made for monomer compatibility when forming a dimer from two monomers before generating a hash table or ensemble of macrocycles and continuing the run. Two monomers are compatible if they obey the following rules: In a dipeptide, if monomer 2 is proline_like (e.g., proline, N-methylated, proline-like, etc.), then the conformer of monomer 1 must come from the preproline conformer database of monomer 1; if monomer 2 is not proline_like, then the conformer of monomer 1 must not come from the preproline conformer database of monomer 1. To facilitate this check, in the residue_params dictionary, we attribute to each monomer two Boolean properties; one that stores whether the monomer is proline_like or not; the other that stores whether the monomer is preproline or not.
[0199] Creating a Hash Table We generate a monomer hash table in a three-step process. In the first step, conformers of the identified monomer are loaded into memory from the conformer database created above. Conformers of the monomer can be filtered for inclusion in the hash table by specifying an energy threshold (in kcal / mol). If the energy of a conformer is greater than this threshold compared to the global minimum energy of the monomer, the conformer is not included in the hashing. In the second step, we build the coordinate frame of the terminal amide using the amide nitrogen, oxygen, and carbon atoms, calculate the relative transformation between these coordinate frames, and partition using a hash function available in the python package xbin. We do the above for the N-terminal coordinate frame relative to the C-terminal coordinate frame and the C-terminal coordinate frame relative to the N-terminal coordinate frame. In the third step, we sort the conformers that belong in the partition and finally use them to build a getpy dictionary object and its associated numpy array, which we call the key-value array. The dictionaries contain as keys the unique set of compartmentalization transformations computed across all conformers incorporated in the hashing in step 1. The values accessed with these keys are start and stop index tuples used to select a sub-slice of the key-value array. Indexing the key-value array with this tuple produces a numpy array of conformer indices that are used to access conformers of monomers in the conformer database that belong to a particular compartmentalization transformation (described in more detail below). These dictionaries and their respective key-value arrays are saved to disk as binary files for later use.
[0200] In this process, the size of the partitions used for hashing is determined by two parameters: Cartesian (distance) resolution and origin (rotation angle) resolution. Increasing one discretizes the 6D space into a few larger partitions, or decreasing the other discretizes the 6D space into many smaller partitions. This can dramatically affect the quality and number of ring closures observed in subsequent steps. A large partition in 6D space results in a hash table with a smaller number of partitioning transformations, each with many conformers. A small partition in 6D space results in a hash table with many partitioning transformations, each with few conformers. To adjust the size of these partitions, we created hash tables of various resolutions for each monomer and then evaluated the quality of the two-residue macrocycles obtained by searching these hash tables for ring closure. For each conformer obtained by this search, we created a macrocycle by superimposing the N-terminal amide pair of one monomer onto the C-terminal amide of the second monomer. The RMSD between these two pairs of amides in the resulting conformers was then calculated (FIG. 10). High quality ring closures are macrocycles with a small RMSD. This occurs when the two conformers used to build the macrocycle share a very similar relative orientation. Low quality ring closures are macrocycles with a large RMSD. This occurs when the two conformers used to build the macrocycle have different relative orientations that are placed in the same compartment by a large discretization of 6D space in hashing. We found that as the size of the compartment increases, the resulting macrocycle ensemble contains more conformers, but the quality of the ring closure is often poor. As the size of the compartment decreases, the resulting macrocycle ensemble contains fewer conformers, but the quality of the ring closure increases. Although it is possible to identify poor ring closures in the AIMNet minimization, we focused on identifying high quality ring closures. Therefore, in this study, either an origin (rotation angle) resolution of 15.0 degrees and a Cartesian coordinate (distance) resolution of 1.0 Å, or an origin (rotation angle) resolution of 10.0 degrees and a Cartesian coordinate (distance) resolution of 0.5 Å, were used.
[0201] The command below indicates that a hash table for each monomer in HDF5_DIR will be created containing conformers that are higher than the global minimum and less than or equal to 6 kcal / mol, using 1.0 Angstrom Cartesian (distance) resolution and 15.0 degree origin (rotation angle) compartmentalization resolution; this hash table will be saved in HASH_DIR. python make_monomer_NC-CN_hash. py -kT 6.0 -car_resl 1.0 -ori_resl 15.0 -hash_dir<full path to HASH_DIR> -hdf5_dir<full path to HDF5_DIR>
[0202] The hash table for a dimer is created in an almost identical process. First, the conformers of the two monomers (monomer 1 and monomer 2) are evaluated for their sequence compatibility and then loaded into memory. Second, each pair of conformers of monomer 1 and monomer 2 are aligned to build a dimer. In this alignment, the C-terminal amide of monomer 1 is aligned to the N-terminal amide of monomer 2. The relative transformation between the coordinate frames of the terminal amides (N-terminus of monomer 1 and C-terminus of monomer 2) is calculated and partitioned. The resulting getpy dictionary and associated key-value array are saved to disk for future use. Now, instead of one index, the key-value array contains two indexes: the first is used to identify the conformer of monomer 1 in the dimer; the second is used to identify the conformer of monomer 2 of the dimer that belongs to the partitioning transformation in question. The following command shows the creation of a hash table for proline / alanine dipeptides: python make_dimer_NC-CN_hash.py -m1 PRO -m2 ALA -kT 1.0 -cart_resl 0.5 -ori_resl 10.0 -hash_dir<full path to HASH_DIR> -hdf5_dir<full path to HDF5_dir>
[0203] A hash table containing dimers with hydrogen bond interactions is created by first identifying the conformation of the dimer that contains the hydrogen bond and then using only that conformer in the hashing. Hydrogen bonds are identified by a fast structure score term. This score term ranges between 0 and 1, with higher values producing more linear hydrogen bonds between amides. In this work, we created a hash table containing hydrogen bonds that included conformers with the score criteria above 0.75, above the global minimum, and below 4 kcal / mol. The following command demonstrates the above for the same proline alanine dipeptide, ensuring that there are no conformers containing cis amides (unless the monomer is proline_like), and that the quality of the hydrogen bond interactions is at least higher than 0.75: python make_dimer_NC-CN_hash.py -m1 PRO -m2 ALA -kT 4.0 -cart_resl 0.5 -ori_resl 10.0 --filter_out_cis --hbond_threshold 0.5 -hash_dir<full path to HASH_DIR> -hdf5_dir<full path to HDF5_dir>
[0204] Alphabetical notation of chemical species and sequences Each monomer is assigned a single letter that defines the number and hybridization of atoms along the backbone starting from the N-terminal amide nitrogen and ending at the C-terminal amide carbon. These single letters or species represent the first layer of the alphabetic notation. Additionally, each monomer is assigned a priority based on its arbitrary position in the list of all monomers. This priority represents the second layer of the alphabetic notation. Combined, the first and second layers of the alphabetic notation generate an offset that is used to replace an input sequence with the lowest priority alphabetic notation species. This ensures that a single sequence of monomers will generate a single output macrocycle during ensemble generation, regardless of their circular permutation. A method to implement this alphabetic notation is included in coarse_cluster.py.
[0205] The first layer of the alphabetic representation is used to group circularly permuted scaffolds into the same chemical species. For example, the chemical species aaab, aaba, abaa, and baaa are all exactly the same circularly permuted chemical species. Given an input sequence of monomers, we assign the sequence to a chemical species and then replace the sequence with the lowest priority circular permutation in the alphabetic representation. In the above example, aaab, aaba, abaa, and baaa would all be replaced with aaab.
[0206] A second layer of alphabetical notation is used to ensure that the input sequence of monomers produces only a single macrocyclic sequence, regardless of circular permutation. For chemical species that contain circular symmetry elements (e.g., aaaa, abab, etc.), degenerate circular permutations exist. For example, acac can be substituted into four chemical species: acac, caca, acac, and caca. In the alphabetical notation of the chemical species, the first and third substitutions are identical in this case. In such cases, the monomer priority order is used to select between these degenerate chemical species.
[0207] Ensemble Generation We generate an ensemble in a four-step procedure. First, load the required hash tables into memory (as getpy dictionaries). Second, identify common keys between both hash tables and extract the conformer indices that belong to these common partitions from their respective key value arrays. Third, combine every combination of conformer indices in the common partition from N to C with every combination of conformer indices in the common partition from C to N. This procedure results in a two-dimensional numpy array that we call IJKL. The IJKL array is (M_conformers x i_residues), where M_conformers is the number of macrocycle conformers in the ensemble and i_residues is the number of monomers in the macrocycle. For example, a 100-membered ring ensemble of the three-residue ring aaa-AGLY-AGLY-AGLY will have an IJKL array of (100, 3), where IJKL[0,:] is the index of the monomer conformers used to build the 0th conformer of the macrocycle. IJKL[0,0] is the first conformer index of the first glycine monomer of the macrocycle; IJKL[0,1] is the first conformer index of the second glycine monomer of the macrocycle; and so on. Third, the second dimension of the IJKL array is replaced with the lowest priority sequence according to the alphabetization process detailed above. Finally, the replaced IJKL array is placed into an HDF5 file named according to the lowest priority sequence of the macrocycle and saved to disk.
[0208] The following commands show how to generate 3- and 4-residue macrocycles. Note that in all of these commands, the Cartesian (distance) and origin (rotation angle) resolutions are specified and must match those used to create the hash tables. It is also important to separate the dimer and monomer hash tables into separate HASH_DIRs. These commands collectively attempt to generate ensembles and output them to CSV_DIR / SUB_DIR. Saving the ensembles in a small number of subdirectories helps ensure that file quota limits are not reached for large-scale sampling. Possible write collisions that can occur due to circular permutation of arrays are further avoided by appending a unique string to the output ensemble name.
[0209] The following command:<HASH_DIR> This shows an ensemble search of all three-residue macrocycles containing a 1.0 kcal / mol proline / alanine dipeptide that can be constructed from the monomer hash table: python build_3mers_ensemble.py -m1 PRO -m2 ALA -cart_resl 0.5 -ori_resl 10.0 -kT 1.0 -hash_dir<full path to monomer HASH_DIR> -csv_dir<full path to where output ensembles will be saved> -sub_dir<subdirectory within CSV_DIR> -subfile<unique string> -hdf5_dir<full path to HDF5_dir containing conformer databases>
[0210] The following command closes the proline-alanine dipeptide:<HASH_DIR> This shows the ensemble generation of all possible four-residue macrocycles that can be constructed from the dimer hash table: python build_4mers_ensemble.py -m1 PRO -m2 ALA -cart_resl 0.5 -ori_resl 10.0 -hash_dir<full path to HASH_DIR> -csv_dir<full path to where output ensembles will be saved> -sub_dir<subdirectory within CSV_DIR> -subfile<unique string>
[0211] The following command shows an attempt to generate an ensemble of only the macrocycle aaaa-ALA-ALA-ALA_pp-PRO. python build_4mers_ensemble.py -m1 PRO -m2 ALA -m3 ALA -m4 ALA_pp -cart_resl 0.5 -ori_resl 10.0 -hash_dir<full path to HASH_DIR> -csv_dir<full path to where output ensembles will be saved> -sub_dir<subdirectory within CSV_DIR> -subfile<unique string>
[0212] After generating these ensembles in subdirectories with unique names associated with them, the following command compiles the unique conformers that exist separately across these different ensembles into a single HDF5 file for each species by combining the unique conformers of all sequences within a species: python combine_seperated_ensembles.py <chemotype><full path to WORK_DIR> Note that WORK_DIR must contain a directory named CSV_DIR as well as HDF5. The combined ensemble will be saved in WORK_DIR / HDF5 / CHEMOTYPE / .
[0213] Generation of macrocycle 3D coordinates in ensembles stored in HDF5 Before performing AIMNet minimization or any other analysis, the IJKL array that stores the ensemble, in addition to the (N_atoms, ) array of atom numbers, is converted to a (M_conformers x N_atoms x 3) array of atomic coordinates by a process we call construction. Construction is performed from the in-memory ensemble IJKL array just before any analysis / minimization which reduces the amount of space required to store the ensemble on disk. The paragraphs below provide a detailed description of the construction procedure for the 4-residue macrocycle.
[0214] First, we load all conformers for the monomers in a given sequence from their respective conformer databases as (X_conformers, Y_atoms, 3) numpy arrays of atomic coordinates. We then use the IJKL array to index these atomic conformations on the X_conformers axis for each monomer. For a four-residue macrocycle, this generates four masked arrays, one for each monomer in the sequence. These arrays are (M_conformers, Y_atoms, 3) in shape, where M represents the number of conformers in the macrocycle ensemble and Y represents the number of atoms in that particular monomer. Second, we subject these masked arrays to an initial building cycle, where the N-terminal amide of monomer i+1 is aligned to the C-terminal amide of monomer i. For a four-residue macrocycle, we run this loop from i=1 to i=3. Once this three residue fragment is constructed, the N-terminal amide of the final monomer is aligned to the C-terminal amide of the third monomer, while at the same time the C-terminal amide of the final monomer is aligned to the N-terminal amide of the first monomer.
[0215] Third, the C-terminal amide of monomer i is then realigned to align with the N-terminal amide of monomer i+1, while the N-terminal amide of monomer i is realigned with the C-terminal amide of monomer i-1 to obtain a macrocycle from each monomer. This process is repeated while tracking the average RMSD between the amides that are linked together. The cycle is completed when the average RMSD no longer decreases or when a maximum number of cycles is reached. Typically, we limited this construction loop to no more than 200 cycles.
[0216] At the end of this cycle, the four aligned (M_conformers, Y_atoms, 3) XYZ arrays are joined together to build a final (M_conformers, N_atoms, 3) XYZ array of the macrocycle, which we call the macrocycle_xyz array, in addition to an (N_atoms, ) array of atomic coordinates, which we call the macrocycle_atoms array. These arrays can be used in subsequent calculations. The methods for performing these steps are implemented in utils_build.py.
[0217] Twisted Plot Clustering Torsion compartment clustering is performed by the following steps: First, construct the 3D coordinates for all conformers in the ensemble in memory as described above. The resulting macrocycle_xyz array is masked to contain only the XYZ coordinates of the backbone atoms. These XYZ coordinates are converted to internal coordinates creating what we call the degrees of freedom array (dofs array), a (M_confomers, N_backbone_atoms, 3) array of bond lengths, bond angles, and bond torsions. The dihedral angles are extracted from the dofs array and converted from angles to integers using the following equation: parcels = ((np.around(dofs[:,:,-1] / torsion_resolution) + 3) % 6 - 3). astype(np.int8) Now set the torsion_resolution to 60 degrees. Then save these partitions to a new HDF5 for further analysis. Use the following command to perform this torsion partitioning: python analyze_macrocycle_torsions.py -i<full path to ensemble> --write_cluster_hdf5 -out_dir<path to save output HDF5> -pdb_dir<path to save PDB files of cluster representatives> -hdf5_dir<path to conformer databases>
[0218] Principal moments of inertia analysis Principal moments of inertia ratios for torsion clusters of multiple species are calculated as described. We have implemented this calculation in the analyze_macrocycle_torsions.py script. The following command adds two PMI ratios from the torsion clustering dataset labeled I1_I3 and I2_I3 to the output HDF5: python analyze_macrocycle_torsions.py --moi -i<full path to ensemble> --write_cluster_hdf5 -out_dir<path to save output HDF5> -pdb_dir<path to save PDB files of cluster representatives> -hdf5_dir<path to conformer databases> In our analysis, only backbone heavy atoms are used in determining the two PMI ratios.
[0219] AIMNet minimization and Pnear calculation Minimization of the macrocycle ensemble was performed in the AIMNet score function modified with dftd4 correction using the ase python package using the BFGS optimizer. FixInternals constraints were used to constrain each macrocycle conformer to ensure that no chemical reactions occurred during minimization. Geometry optimizations were performed for no more than 100 steps or until the maximum force did not exceed 0.005 eV / A. In these optimizations, maxstep was set to 0.5 Å. This minimization scheme typically takes approximately 1-3 minutes of CPU time per macrocycle conformer to complete depending on the number of atoms in the macrocycle.
[0220] We implemented this minimization in the script minimze_macrocyle_ensemble.py, which first builds the 3D coordinates of the macrocycle from the input ensemble as described above, and then performs the minimization on the conformers. The script uses the flag -j to enable parallel processing of the minimization on multiple CPU cores. In our tests, each parallel job requires approximately 3 GB of RAM. After minimization of all conformers, the lowest energy conformer obtained from the ensemble is identified. Alignment is performed with respect to the backbone atoms of all other conformers and the backbone atoms of the low energy conformers. RMSD is calculated for these alignments. The RMSD and AIMNet energy are used to create the energy landscape shown in Figure 1B and Figure 10. The Pnear value of the landscape is calculated using the following formula with kT=0.62 and λ=0.3 Å:
[0221] The following command shows the minimization of 1000 random conformers from one ensemble. This command parallelizes the minimization to 20 CPU cores: python minimize_macrocycle_ensemble.py -i<full path to ensemble> --endpoint_only --limit 1000 -j 20 --save_hdf5 -hdf5_dir<path for conformer databases> -pdb_dir<path to PDB_DIR> -png_dir<path to PNG_DIR> This generates an SVG file of the energy landscape in PNG_DIR, a PDB file containing up to the 100 lowest energy conformers of the macrocycle in PDB_DIR, and an HDF5 file containing the XYZ of the minimized ensemble, their RMSD relative to the lowest energy conformer in the ensemble, and their AIMNet-minimized energies is saved in PDB_DIR. We performed a manual check to see if the Pnear of each of the energy landscapes of the lowest energy conformer macrocycles was greater than 0.75.
[0222] Design of hydrogen-bonded macrocycle sequences We used an enhanced pipeline to identify specific sequences that can form long-range hydrogen bonds from the above. first identified a set of basic residues in each chemical species. Then, for all possible four-residue macrocycles that contain combinations of only these basic residues, an ensemble was generated by searching a hash table containing hydrogen bonds for ring closures. The ensembles obtained above were then clustered using the analyze_macrocycle_torsions.py script, which generates a pdb file containing a single representative of each sampled torsion section in the ensemble. The analyze_compatible_sequences.py script was then used to identify possible low-energy sequences for each of the representative torsion sections. An ensemble was then generated for each of the candidate low-energy sequences and subjected to minimization as described above.
[0223] In the analyze_compatible_sequences.py script, candidate low energy sequences are identified by a one-body side-chain packing algorithm. First, we identify possible side-chain substitutions at each position or in the sequence of the base residue macrocycle (e.g., mutating glycine to alanine, phenylalanine, etc.). We then rank-order the substitutions based on their suitability for the backbone torsion angles of the representative torsion section being processed. We query the potential energy surface of each possible substitution at each position in the macrocycle with the goal of identifying the relative energy of the lowest energy rotamer possible for each substitution given the backbone torsion angles of the base residues in the macrocycle. This process allows the identification of monomers that are most likely to have the torsion angles present in the representative torsion section. This is done for each position in the macrocycle being processed. The result of this process is a list of sequences consisting of ideal substitutions of the base residues in each of the representative torsion sections. The inventors have shown that this analysis produces sequences that, when minimized, reduce the occurrence of non-ideal torsion angles, such as L-amino acids placed at positions that have positive Φ torsion angles.
[0224] To generate a list of candidate sequences for ensemble generation, use the following command: python analyze_compatible_sequences.py -i<full path to pdb from torsion binning> -hdf5_dir<path for conformer databases> -out_dir<path to save output lists> We also used this process to identify the three lowest energy monomers at each position and generate candidate sequences for each combination at each position by adding the -max_monomers 3 flag.
[0225] We minimized approximately 8280 hydrogen-bond-containing macrocycles resulting from the sequence design process described above. To prioritize this subset of macrocycles for manual inspection and synthesis, we built a simple linear regression model to predict the AIMNet energy of a macrocycle based on its sequence composition. We created a matrix in which each column is associated with a specific monomer and each row is associated with a macrocycle sequence. In this matrix, for a given macrocycle, we count the number of specific monomers in each sequence. We then created a vector in which each row contains the AIMNet energy of the low-energy conformer of the macrocycle. We then utilized the python package sklearn to train the weights for each monomer in the matrix with a linear regression model. We found that this model predicts the AIMNet energy of the low-energy conformer for a given macrocycle to within 5 energy units. We then prioritized macrocycles by checking whether their ZScore (calculated from minimized AIMNet energy + predicted AIMNet energy) was less than -2. In this process, for a given macrocycle sequence, macrocycles whose AIMNet energy was much more negative than the predicted energy (i.e., favored) were identified.
[0226] Materials and Methods General matters Peptide synthesis The preparation of the macrocycles described herein was carried out by WuXi apptec. Typically, WuXi created peptides on CTC resin using Fmoc-based solid-phase synthesis utilizing repeated couplings with either HBTU / DIEA, HATU / DIEA, or DIC / HOBt. The linear peptide was cleaved from the resin and then cyclized in solution. Typically, peptide cyclization was performed in either DCM or DMF using either TBTU / DIEA, HATU / DIEA, or EDCl / HOBt. The crude cyclization reaction was then purified using reversed-phase HPLC. In WuXi's technology, these methods typically yielded 0.5-20 mg of purified macrocycles as white or off-white powders after lyophilization. Below is a more detailed protocol for the synthesis of aaap:
[0227] CTC resin (0.1 g, 0.1 mmol, 0.99 mmol / g) was swollen in DCM with shaking under nitrogen atmosphere. A solution of Fmoc-3-aminomethyl-phenylacetic acid (38.74 mg, 0.1 mmol, 1.0 equiv.) in DCM was added to the resin and stirred. DIEA (4.0 equiv.) was added dropwise to the resin. The resulting mixture was mixed under shaking for 2 h. The solution was then discarded and the resin was capped with methanol (0.1 mL) for 30 min. The loaded capped resin was then washed with DMF. Linear peptides were synthesized by performing repeated Fmoc-deprotection reactions with 20% piperidine in DMF, followed by coupling reactions with 6 equiv. of Fmoc-protected amino acids, 5.7 equiv. of HATU, and 12 equiv. of DIEA added sequentially. The resin was washed between each deprotection and each coupling reaction with DMF. The progress of these reactions was monitored using the ninhydrin test. After the final Fmoc deprotection reaction was completed, the linear peptide was cleaved from the resin using a solution of 1% (v / v) trifluoroacetic acid in DCM. This resin cleavage treatment was performed twice, each treatment for 3 min. The combined filtrates (from the two runs) containing the cleaved linear peptide were then diluted with 100 mL of DCM and HATU (76 mg, 2 equiv.) was added. The solution was adjusted to pH 8.0 with DIEA. The resulting mixture was stirred at room temperature for 0.5 h, after which 1 M HCl was added. The organic phase was collected and concentrated under vacuum to obtain crude macrocyclic aaap. The crude reaction mixture was dissolved in water:CAN and purified by preparative reversed-phase HPLC (25-55% ACN in 55 min). The pure fractions were combined and lyophilized to obtain 14.1 mg of aaap with a purity of about 99.4%.
[0228] NMR spectroscopy 1D 1 H-NMR and 2D 1 H-NMR spectra were recorded on a Bruker NOE 600 MHz spectrometer equipped with a QCI-F cryoprobe. TOCSY spectra were recorded with a spin-lock mixing time of 80 ms. ROESY spectra were recorded with a spin-lock mixing time of 200 ms. Data were processed with TopSpin.
[0229] VT-NMR Variable temperature NMR experiments were carried out to determine the amide hydrogen temperature shift coefficients in the range 298–318 K either in DMSO-d6 or CDCl3. 1 H-NMR spectra were recorded: temperature shift coefficients less than 4 ppb / K were assumed to represent hydrogens shielded from the solvent, whereas coefficients greater than 4 ppb / K were assumed to indicate hydrogens exposed to the solvent.
[0230] Structure determination by NMR NMR-based distance restraints were used in combination with molecular dynamics simulations to determine the 3D structures of peptide macrocycles. ROE-based distance restraints were determined from the average integral series of the diagonal ROE cross peaks. These integrals were converted to atomic distances using the inverse sixth power relationship. The conversion of integrated ROE to atomic distances was calibrated using reference distances of 1.78 Å for the diastereotopic methylene protons or 2.49 Å for the 1,2-substituted phenyl protons. The calculated distances were then adjusted by 10% to obtain the range of allowed distances for each distance restraint. The structural hashing scheme detailed above was used to obtain an ensemble of macrocycle conformers that obeyed each distance restraint. First, we created hash tables containing monomers with energies up to 6 kcal / mol higher than the global minimum. We then hashed searched these tables for ring closures. This search yielded 10 6 ~10 8 We then screened the conformers based on whether they satisfied all of the experimentally obtained distance constraints, resulting in an ensemble of approximately 400 conformers for the peptide. For each of the 400 conformers, we performed 50 picosecond unconstrained molecular dynamics simulations with xTB at 298K with either implicit DMSO or implicit CHCl3 (Bannwarth et al., 2019). Overall, this resulted in approximately 20 nanoseconds of simulation time for each peptide. All 20 lowest energy structures over the approximately 20 nanosecond simulation were taken as the ensemble structure for the peptide.
[0231] Single crystal X-ray diffraction of peptides Crystal diffraction data were collected from a single crystal at a synchrotron (APS 24ID-C) at 100 K. Unit cell refinement and data reduction were performed using the XDS and CCP4 suites (Kabsch 2010; Winn et al. 2011). The structure was identified by direct methods, and then refined using SHELXL-2018 / 3, where F 2 The structure was refined by full matrix least squares using the method of 1.2×U for the linking atom; anisotropic displacement parameters were used for non-H atoms (Sheldrick 2015b; Sheldrick 2015a). Coot / Shelxle was used to assist in the structure analysis (Emsley and Cowtan 2004; Huebschle et al. 2011). Hydrogen atoms adjacent to heavy atoms were aligned with a 1.2×U for the linking atom. eq The calculations were performed at an ideal position with the isotropic displacement parameters set as follows:
[0232] PAMPA Stock solutions of each peptide were prepared gravimetrically by dissolving approximately 1-3 mg of peptide in either 100% MeOH or 100% DMSO; thus resulting in approximately 1-10 mM peptide solutions. These stock solutions were stored at 4°C when not in use. Samples and calibration curves for PAMPA experiments were prepared by diluting these stock solutions to obtain 1.2 mL solutions containing 20 μM peptide in PBS supplemented with either 5% DMSO or 10% MeOH. Calibration curves for each peptide were generated by serially diluting the 20 μM peptide two-fold to obtain six samples with the following concentrations: 20, 10, 5, 2.5, 1.25, and 0.625 μM. These standard curves were used to determine the concentration of each peptide in both donor and acceptor wells of the PAMPA plate. All PAMPA experiments were performed using the commercially available Corning Gentest Precoated PAMPA Plate System (Part Number: 353015) according to the manufacturer's instructions. Briefly, 300 μL of the 1.2 mL solution was dispensed into donor wells; 200 μL of either 5% DMSO in PBS or 10% MeOH in PBS was taken and added to acceptor wells. Each peptide was prepared in triplicate. After approximately 19 hours, the concentration of each peptide in the donor and acceptor wells was quantified by LCMS. These concentrations were used to calculate Papp according to the manufacturer's instructions.
[0233] [Table 5] TIFF2025503658000108.tif242150TIFF2025503658000109.tif242152TIFF20255036580 00110.tif242151TIFF2025503658000111.tif242150TIFF2025503658000112.tif242150< / chemotype>
Claims
1. Non-natural macrocyclic oligoamides of 3-4 residues comprising a three- or four-monomer residue species selected from the group consisting of monomers a, b, c, d, e, f, g, h, I, j, k, l, m, n, o, p, q, r, s, t, u, and v as defined in Table 1, or salts thereof.
2. 2. The macrocyclic oligoamide of claim 1 ; where (a) The monomer names a, b, g, d, and e are amino acids; where (i) The monomer designation a includes or consists of all L- and D-proteinogenic amino acids and all L- and D-non-standard α-amino acids; peptoids (N-alkylated glycines), and optical isomers thereof; (ii) the monomer designation b comprises or consists of all RNase B3 amino acids containing L- or D-proteinogenic side chains, and optical isomers thereof; (iii) the monomer designation d comprises or consists of δ amino acids and their optical isomers; (iv) the monomer designation g comprises or consists of γ-amino acids and their optical isomers; and (v) The monomer designation e comprises or consists of ε amino acid and its optical isomer; (b) Monomer names c, f, and j include or consist of aminobenzoic acids and their optical isomers; (c) Monomer names h, k, and n include or consist of aminomethylbenzoic acids and their optical isomers; (d) Monomer designations i, l, and o include or consist of aminophenylacetic acids and their optical isomers; (e) The monomer names m, p, and q include or consist of aminomethylphenylacetic acids and their optical isomers; (f) Monomer designation r includes or consists of oxazolidines and thiazolidines and their optical isomers, but includes all L- and D-proteinogenic side chains substituted at the second atom in the backbone for both oxazolidinones and thiazolidinones; (g) Monomer designation t includes or consists of oxazoles and thiazoles and their optical isomers, but includes all L- and D-proteinogenic side chains substituted at the second atom in the backbone for both oxazolidines and thiazolidinines; (h) The monomer designation s includes or consists of thioethers and their optical isomers; (i) The monomer name includes or consists of aminomethylpicolinic acids and their optical isomers; and (j) The monomer designation v comprises or consists of triazoles and their optical isomers; Macrocyclic oligoamides.
3. 2. The macrocyclic oligoamide of claim 1, wherein the monomer is as defined by the chemical structure of Figure 5 and Table 5: AGLY, ALA (alanine), ABU, SER (serine), ASN (asparagine), CLA, VAL (valine), TBG, NVL (norvaline), LEU (leucine), TBA, PHG (phenylglycine), PHE (phenylalanine), FPA, NAP1, NAP2, ANT9, PYR3, DHA, AIB (aminoisobutyric acid), ACPC, ACBC, ACPenC, ACPhenC, ACHC, SAR (sarcosine), PTABU, PTIPA, PTTBA, PTANI, PTAMBA, NMALA (n-methylalanine), NMABU, NMVAL (n-methylvaline), NMPHG, NMLEU (n-methylleucine), NMPHE (n-methylphenylalanine), NMSER (n-methylserine), NMASN (n-methylasparagine), AZE (azetidine), PRO (proline), PIP, TIQ, TIC, NHM, THP, DHP, FPS, FPR, HPS, HPR, OXO, PPS, PPR, AMP, and IDC, and optical isomers thereof; a monomer designation "b" selected from the group consisting of BGLY (β-alanine), B3ALA (β-3-homoalanine), B3VAL (β-3-leucine), B3PHG (β-3-phenylalanine), B3PHE (β-3-homophenylalanine), B2ALA, BAZE, BPRO, BPIP, PRR, NIP, LARD, ACPC12C, ACPC12T, ACBC12C, ACBC12T, ACPenC12C, ACPenC12T, ACHC12C, ACHC12T, and ABOC, and optical isomers thereof; a monomer designated "c" selected from the group consisting of BEN2 and NMBEN2; Monomer designation "d" selected from the group consisting of D4ALA, DGLY, ACHC14C, and ACHC14T; a monomer designation "e" selected from the group consisting of EGLY, ACHA14C, ACHA14T, AMC14C, and AMC14T; a monomer designated "f" selected from the group consisting of BEN3 and NMBEN3; a monomer designation "g" selected from the group consisting of GGLY, G4ALA, GPN, INIP, LARE, ACBC13C, ACBC13T, ACPenC13C, ACPeneC13C, ACPenC13t, GAPC, GAPT, ACHC13C, and ACHC13T; Monomer name "h" which is AMBEN2; Monomer name "i" which is ACBEN2; a monomer designation "j" selected from the group consisting of BEN4 and NMBEN4; Monomer name "k" which is AMBEN3; Monomer name "l" which is ACBEN3; Monomer name “m” which is AMACBEN2; Monomer name "n" which is AMBEN4; Monomer name "o" which is ACBEN4; Monomer name “p” which is AMACBEN3; Monomer name “q” which is AMACBEN4; a monomer designation "r" selected from the group consisting of AGLS, AGLR, OZS, OZR, TZS, TZR, and VIC, and optical isomers thereof; Monomer designation "s" which is SUGA; a monomer designated "t" selected from the group consisting of OZL, TZL, and HUCP, and optical isomers thereof; Monomer name "u" which is HUCQ; and / or Click and Clicks selected from the group consisting of monomer designation "v"; A macrocyclic oligoamide selected from the group consisting of:
4. 10. The macrocyclic oligoamide of claim 1, comprising a species selected from the group of species listed in Table 2, or circularly permuted versions thereof.
5. 2. The macrocyclic oligoamide of claim 1, wherein the macrocyclic oligoamide comprises at least one monomer not included in the monomer designation a.
6. 2. The macrocyclic oligoamide of claim 1, wherein the macrocyclic oligoamide comprises at least one residue not included in the monomer designations a, b, g, d, and e.
7. 10. The macrocyclic oligoamide of claim 1, wherein the macrocyclic oligoamide is a four-residue macrocycle.
8. 10. The macrocyclic oligoamide of claim 1, having a species selected from the group consisting of aaar, aarb, abar, aavb, aabr, and aatb, or circularly permuted forms thereof.
9. 10. The macrocyclic oligoamide of claim 1, having a species selected from the group consisting of akak, aaaq, and aabc, or circularly permuted forms thereof.
10. 10. The macrocyclic oligoamide of claim 1, having the species aas or a circularly permuted form thereof, wherein the macrocyclic oligoamide comprises two hydrogen bonds, one between the backbone amides and the other comprising the primary amide of the "s" monomer.
11. 10. The macrocyclic oligoamide of claim 1, having the species aaam or akak, or circularly permuted forms thereof, wherein any nitrogen in the backbone is a tertiary amide.
12. 10. The macrocyclic oligoamide of claim 1, having the chemical species aaaq or a circularly permuted form thereof, wherein one of the "a" monomers comprises a pentafluorophenylalanine residue or an optical isomer thereof, and the "q" residue comprises a 4-aminomethylphenylacetic acid residue or an optical isomer thereof.
13. 10. The macrocyclic oligoamide of claim 1, having a species selected from the group consisting of ahah, aaam, aabi, aagb, aaap, and aalm, or circularly permuted forms thereof.
14. 10. The macrocyclic oligoamide of claim 1, wherein the macrocyclic oligoamide comprises one or more hydrogen bonds.
15. 15. The macrocyclic oligoamide of claim 14, wherein the one or more hydrogen bonds comprise hydrogen bonds between backbone amides comprising non-alpha amino acids.
16. 10. The macrocyclic oligoamide of claim 1, wherein the macrocyclic oligoamide is membrane-permeable.
17. 10. The macrocyclic oligoamide of claim 1, wherein the macrocyclic oligoamide does not contain any exposed polar groups, such as side chain hydroxyl groups and / or primary amides.
18. 10. The macrocyclic oligoamide of claim 1, comprising the structure of any macrocyclic oligoamide disclosed in any drawing herein, a circularly permuted structure thereof, or a salt thereof; or wherein said macrocyclic oligoamide comprises or consists of the structure of any macrocyclic oligoamide shown in any of Figures 3-4, 12, and 19-47, or a salt thereof; The compound of claim 1, wherein the compound is a compound selected from the group consisting of: Macrocyclic oligoamides.
19. 10. The macrocyclic oligoamide of claim 1, comprising substitution of one, two, three, or all four types of monomeric subunits.
20. 20. The macrocyclic oligoamide of claim 19, wherein the substitution comprises a functional group.
21. 21. The macrocyclic oligoamide of claim 20, wherein the functional group comprises a therapeutic moiety, a diagnostic moiety, and / or a reactive moiety.
22. A library comprising 10, 50, 100, 500, 1,000, 5,000, 10,000, 25,000, 35,000 or more types of macrocyclic oligoamides according to claim 1.
23. A method of using the macrocyclic oligoamides of claim 1 and / or the library of claim 22 for any suitable purpose, including but not limited to panning the library to identify one or more macrocyclic oligoamides that bind to a target compound, therapeutic treatment, diagnostics, and / or adding reactive moieties for any use.
24. 1. A method implemented by one or more computers for identifying a macrocycle conformation of a molecule comprising a first molecular fragment and a second molecular fragment, the method comprising: obtaining a conformation set for the first molecular fragment; determining, for each conformation of the first molecular fragment, values of a set of parameters for a start-to-end transformation that defines the translation and rotation of a start end of the first molecular fragment relative to a end end of the first molecular fragment in that conformation; obtaining a conformational set of a second molecular fragment; determining, for each conformation of the second molecular fragment, values of a set of parameters for a start-to-end transformation that defines the translation and rotation of the end of the second molecular fragment relative to the start of the second molecular fragment in that conformation; and processing the parameter values for each of the start-to-end and end-to-start transformations to identify a set of one or more macrocycle conformations of the molecule; wherein each macrocyclic conformation of the molecule comprises respective conformations of the first molecular fragment and the second molecular fragment that cooperate to form a closed circular loop; A method comprising: