Systems and methods for identifying multi-targeting molecules

In silico methods and systems address the challenges of screening multi-targeting molecules by computationally evaluating candidate molecules' interactions, providing efficient and scalable identification of molecules with desired specificities.

WO2025231366A1PCT designated stage Publication Date: 2025-11-06FLAGSHIP PIONEERING INNOVATIONS VII LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/027503
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-03
Filing Date
2025-05-02
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Existing wet lab techniques for screening candidate molecules face challenges in cost and scaling, especially for multi-targeting molecule profiles, necessitating the development of in silico methods for identifying multi-targeting molecules.

Method used

In silico methods and systems are developed to identify multi-targeting molecules through computational design and analysis, utilizing a set of reference structures and a computer-simulated library of candidate molecules, evaluating their interactions with multiple partner elements, and employing scoring functions to identify molecules with specific interactions.

Benefits of technology

These methods efficiently identify multi-targeting molecules, such as polypeptides, with desired interactions, enabling cost-effective and scalable screening beyond traditional wet lab limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000026_0001
    Figure IMGF000026_0001
  • Figure IMGF000027_0001
    Figure IMGF000027_0001
  • Figure IMGF000028_0001
    Figure IMGF000028_0001
Patent Text Reader

Abstract

Methods for identifying multi-targeting molecules by analyzing their interactions with multiple partner elements within given reference structures are provide. In certain embodiments, the methodology includes providing reference structures and an in silico library of candidate molecules, and assessing the interactions between candidate and partner molecules. Molecules that bind with multiple partners are identified as multi-targeting molecules. The library can include different types of molecules, such as polypeptide molecules. The methods, in some embodiments, further involves predicting the structure of candidate molecules and evaluating their structure in the context of the reference structure. The invention also provides systems configured to execute this methodology and the multi -targeting molecule identified through this process.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR IDENTIFYING MULTI-TARGETING MOLECULESCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This Application claims priority to United States Provisional Patent Application No. 62 / 642,460, entitled “SYSTEMS AND METHODS,” filed May 3, 2024, which is hereby incorporated by reference.BACKGROUND

[0002] Potential molecule space, including protein-based molecules, is massive. Molecules, including proteins, can cover a vast array of industrial applications, including therapeutics, such as human therapeutics or veterinary applications, as well as agricultural, sustainability, and industrial applications. Molecules with multi-targeting specificity are highly desirable but present design and screening challenges. Wet lab techniques for screening candidate molecules face challenges in cost and scaling, especially for more complicated multi-targeting molecule profiles. Accordingly, a need exists for in silico methods — and associated systems for performing them— that overcome these and other challenges, e.g., to identify multi-targeting molecules.SUMMARY

[0003] The invention provides, inter alia, in silico methods — and associated systems for performing them— that overcome these and other challenges, e.g., to identify multi-targeting molecules. In one aspect, the invention provides methods for identifying a multi-targeting molecules through computational design and analysis. These methods, in certain embodiments, involve providing a set of reference structures and an in silico (computer-simulated) library of candidate molecules. These candidate molecules are then evaluated for their potential to interact with two or more of the partner elements within the reference structures. Molecules showing an ability to interact with multiple partner elements are identified as multi-targeting molecules. The application also provides systems involving a microprocessor and computer-readable instructions for performing this methodology. In other aspects, the invention provides the multi -targeting molecule identified by this method, such as polypeptides, and, in such embodiments, envisionsnucleic acids encoding them, host cells that can express them, and method of identifying and using them that will be readily apparent to the skilled artisan upon reviewing the application.

[0004] One aspect of the present disclosure provides a method of identifying a multi-targeting molecule. The method comprises providing a plurality of three-dimensional reference structures, each member of the plurality of three-dimensional reference structures comprising an initial molecule element in complex with a different interacting partner element in a plurality of interacting partner elements.

[0005] In some such embodiments, each three-dimensional reference structure in the plurality of three-dimensional reference structures comprises a different interacting partner element in the plurality of interacting partner elements determined by one or more of: NMR, cryo-EM, or X- ray diffraction.

[0006] In some such embodiments, each three-dimensional reference structure in the plurality of three-dimensional reference structures comprises a different interacting partner element in the plurality of interacting partner elements determined by one or more of: NMR, cryo-EM, or X- ray diffraction.

[0007] In some such embodiments, each three-dimensional reference structure in the plurality of three-dimensional reference structures comprises a different interacting partner element in the plurality of interacting partner elements determined by NMR with an RMSD resolution of less than about: 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 Angstrom, or less.

[0008] In some such embodiments, each three-dimensional reference structure in the plurality of three-dimensional reference structures comprises a different interacting partner element in the plurality of interacting partner elements determined by X-ray diffraction with a crystallographic resolution of less than about: 3.5, 3.4, 3.3, 3.2, 3.1, 3.0, 2.9, 2.8, 2.7, 2.6, 2.5, 2.4, 2.3, 2.2, 2.1, 2.0, 1.9, 1.8, 1.7, 1.6, Angstroms or less.

[0009] In some such embodiments, a three-dimensional reference structure in the plurality of three-dimensional reference structures is formed by computationally docking the initial molecule element to an interacting partner element in the plurality of interacting partner elements.

[0010] In some embodiments, the plurality of three-dimensional reference structures comprises at least about 2, 3, 4, 5, 6, 7, 8, 9, 10 structures, or more.

[0011] In some embodiments, the plurality of three-dimensional reference structures comprises at least 20, 30, 40, 50, 100, 200, 400, 500, 1000, 2000, 4000, 5000, or 10000 structures.

[0012] In some embodiments, a plurality of candidate molecules is provided in silico. Each candidate molecule in the plurality of candidate molecules is within a defined evolutionary chemical space relative to the initial molecule element.

[0013] In some embodiments, the plurality of candidate molecules is a plurality of polypeptides.

[0014] In some embodiments, the initial molecule element is a polymer, the plurality of candidate molecules is a plurality of polypeptides comprising one or more of: a) a first collection of candidate molecules that collectively represent at least about 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of all single point mutations of a portion (c. ., about: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100%) of a sequence of the initial molecule element, including, in some embodiments; b) a second collection of candidate molecules that collectively represent at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, or 40% of all possible double point mutations of the sequence of the initial molecule element; c) a third collection of candidate molecules that collectively represent at least 1%, 2%, 3%, 4%, 5% of all possible 3-order mutations through M-order mutations of the initial molecule element, in which M is a length of the sequence of the initial molecule element; or d) any combination of any of the foregoing a), b), and c).

[0015]

[0016] In some embodiments, a plurality of three-dimensional testing complexes is formed. Each respective testing complex comprises a candidate molecule three-dimensionally threaded into a place of the initial molecule element in a corresponding three-dimensional reference structure in the plurality of three-dimensional reference structures thereby replacing the initial molecule element in the corresponding three-dimensional reference structure to form the respective three-dimensional testing complex.

[0017] In some embodiments, the plurality of candidate molecules is evaluated for interaction with the plurality of interacting partner elements using a scoring function that jointly scores an interaction between a candidate molecule of the plurality of candidate molecules and each respective three-dimensional testing complex in the plurality of testing complexes that includes the candidate molecule to form a single score. Satisfaction of a threshold by the single score satisfies identifies the candidate molecule as a multi-targeting molecule.

[0018] In some embodiments, any one or any combination of a) the initial molecule element being a first polypeptide, b) interacting partner element in the plurality of interacting partnerelements being a polypeptide; c) the plurality of candidate molecules being a plurality of polypeptide molecules; and d) the evaluating of the plurality of candidate molecules for interaction with the plurality of interacting partner elements of the plurality of three-dimensional testing complexes comprises evaluating the plurality of candidate molecules for binding against each of the interacting partner elements in the plurality of interacting partner elements.

[0019] In some embodiments, the first polypeptide is a polypeptide ligand.

[0020] In some embodiments, an interacting partner element in the plurality of interacting partner elements is a receptor polypeptide.

[0021] In some embodiments, the multi-targeting molecule comprises one or more of: a) agonist activity for at least one interacting partner in the plurality of interacting partner elements; b) antagonist activity for at least one interacting partner in the plurality of interacting partner elements; c) neutral binding activity for at least one interacting partner in the plurality of interacting partner elements; d) non-binding activity for at least one interacting partner in the plurality of interacting partner elements; e) any combination of a), b), c), or d) including partial agonist activity, partial antagonist activity, as well as condition-specific activity of any of the foregoing, such as pH dependent activity for an interacting partner element in the plurality of interacting partner elements.

[0022] In some embodiments, an ensemble of sequence variants of the multi -targeting molecule is identified that preserve the structure of the multi-targeting molecule to an RMSD resolution of less than about: 5, 4, 3, 2, 1, 0.5, or 0.3 Angstroms, or less. In some such embodiments, the identification of the ensemble of sequence variants is by AlphaFold, OmegaFold, ESMFold, molecular docking, molecular dynamics, RosettaFold, FoldX, RFDiffusion, ProteinMPNN, or dTERMen.

[0023] In some embodiments, test, in vitro, in vivo, or ex vivo, the interactions of a multitargeting molecule identified by the method.

[0024] In some embodiments, the multi-targeting molecule is a polypeptide, the plurality of interacting partner elements are polypeptides and the interaction comprises binding.

[0025] In some embodiments, a first interacting partner element in the plurality of interacting partner elements is a receptor (e.g., a G protein-coupled receptor).

[0026] In some embodiments, the initial molecule element is a first small molecule, the first candidate molecule of the plurality of candidate molecules is a second small molecule, and thescoring function normalizes a score of the interaction between the first candidate molecule and a three-dimensional testing complex by a graph edit distance between the first candidate molecule and the initial molecule element.

[0027] In some embodiments, the initial molecule element is a first polypeptide, the first candidate molecule of the plurality of candidate molecules is a second polypeptide, and the scoring function normalizes a score of the interaction between the first candidate molecule and a three-dimensional testing complex by a mutation distance between the first candidate molecule and the initial molecule element.

[0028] In some embodiments, the initial molecule element is a first polypeptide, each candidate molecule of the plurality of candidate molecules is different polypeptide, and a candidate molecule in the plurality of candidate molecules is within the defined evolutionary chemical space when a sequence of the candidate molecule and a sequence of the initial molecule element share at least seventy percent sequence identity with the proviso that the sequence of the candidate molecule includes at least two or more sequence substitutions relative of the sequence of the initial molecule element.

[0029] In some embodiments, the initial molecule element is a first small molecule, each candidate molecule of the plurality of candidate molecules is a different small molecule, and a candidate molecule in the plurality of candidate molecules is within the defined evolutionary chemical space when a graph edit distance of the candidate molecule and the initial molecule element is between 2 and 5.

[0030] In some embodiments, the initial molecule element is a first small molecule, each candidate molecule of the plurality of candidate molecules is a different small molecule, and a candidate molecule in the plurality of candidate molecules is within the defined evolutionary chemical space when a graph edit distance of the candidate molecule and the initial molecule element is between 4 and 10.

[0031] In some embodiments, the initial molecule element is a first small molecule, each candidate molecule of the plurality of candidate molecules is a different small molecule, and a candidate molecule in the plurality of candidate molecules is within the defined evolutionary chemical space when a graph edit distance of the candidate molecule and the initial molecule element is between 6 and 14.

[0032] In some embodiments, the method further comprises synthesizing the multi-targeting molecule.

[0033] In some such embodiments, a multi -targeting molecule identified by any method of the present disclosure is provided.

[0034] Exemplary software available to the skilled artisan that can be used in various embodiments of the invention are described herein, but the skilled artisan will readily appreciate that other software can be deployed without departing from the principles first described herein.

[0035] The figures below serve to help illustrate the principles and utilities of the methods and systems provided by the invention.BRIEF DESCRIPTION OF THE DRAWINGS

[0036] FIGS. 1A, IB, and 1C summarize in silica DMS profiles for targets 1, 2, and 3 against peptide X. Heatmaps show binding energy for variants predicted to bind at the receptor binding site (RMSD < 5 A) or having improved binding energy (AAG < 0).

[0037] FIG. 2 illustrates a heatmap summarizing integrated in silico DMS profiles for multibinders’ identification. The heatmap shows variants predicted to bind at the receptor binding site (RMSD < 5 A) or having improved binding energy (AAG < 0) for at least one of the target receptors.

[0038] FIG. 3 illustrates a bar graph summarizing experimentally validated multi-binder peptides (33) ranked by Equation 1. Using such a strategy identified 5 out of the 6 full multibinder peptides in the top 15.

[0039] FIG. 4 illustrates a bar graph summarizing 5,000 randomly generated and scored sequences with same mutational load to the experimentally synthesized peptides (10-13 mutations). Only 5% of the random sequences scored above the threshold identified after prioritizing experimentally validated peptides with our scoring function.

[0040] FIG. 5 illustrates how a first set of nonbinders on the left side of plot did not satisfy a threshold for binders while all 16 triple binders in a second set of nonbinders satisfied a threshold for binders in accordance with an embodiment of the present disclosure.

[0041] FIG. 6 illustrates a schematic overview of a process for identifying binders to targets in accordance with an embodiment of the present disclosure.

[0042] FIG. 7 illustrates a computer system in accordance with some embodiments of the present disclosure.

[0043] FIGS. 8A, 8B, 8C, 8D, and 8E illustrate methods for identifying a multi-targeting molecule in accordance with some embodiments of the present disclosure in which optional elements are indicated by dashed line boxes.

[0044] FIG. 9 illustrates the difference between chemical and exact graph edit distance in accordance with the prior art.

[0045] FIG. 10 illustrates pseudocode in accordance with one embodiment of the present disclosure.DETAILED DESCRIPTION

[0046] The invention provides, inter alia, methods, non-transitory computer-readable instructions (e.g., on a non-transient storage medium), and systems, for identifying multitargeting molecules, such as multi-targeting polypeptides.

[0047] A “multi-targeting molecule” is a molecule with specific interactions with an ensemble of interaction partners. For example, a multi-targeting molecule that is a polypeptide ligand may, in some embodiments, specifically bind all “n” of a set of n interacting partners that are receptors, such as G-protein coupled receptors. In other illustrative embodiments, a multitargeting molecule (e.g., polypeptide ligand) might specifically bind n-1 interacting partners (e.g., G-protein coupled receptors) in an ensemble but not specifically bind another interacting partner (e.g., non-binding activity, for example, de-tuning a binding activity that was present in a reference molecule or starting molecule, e.g., the base molecule of a library). In other illustrative embodiments, the specific interaction may be selected for desired species specificity — for example, to human, cyno, and murine homologues of an interacting molecule (or a plurality of interacting molecule); or, in other embodiments, to bind, e.g., human and cyno, but not murine. The skilled artisan can readily envision how the methods, systems, and molecules provided by the invention can facilitate a wide array of design parameters for desired specificity. In other embodiments, multi-targeting can involve design to specifically interact (for example to specifically bind, e.g., to inhibit) groups of pathogenic molecules (e.g., from bacteria, viruses, parasites, etc.) but not a human analog, such a human molecule with some shared structural features, for example, that may give rise to disease pathologies, such as the rheumatic symptomsassociated with Lyme disease. In other embodiments, specificity can be engineered according to disease state, e.g., for targeting a disease state molecule, such as an activating mutation, such as an EGFR mutation. In still other embodiments, which can be combined with any of the foregoing, molecules can be modified for biophysical properties, for example, pH, isoelectric point, hydrophobicity, etc., so any of the forgoing properties can be condition depended. For example, to take pH into account, molecules’ protonation state at the pH can be determined before determining specific interaction, such as when calculating binding energy.

[0048] The interaction of a multi-targeting molecule and an ensemble of interaction partners may entail any interaction, including, for example, neutral binding, agonism, antagonism, biased signaling (partial agonism or antagonism, such as exhibiting only a portion of downstream effects of a reference molecule), non-binding, etc., and in any combination amongst members of the ensemble of interacting partners. A multi-targeting molecule can be any molecular entity but in certain embodiments it is a polypeptide. In certain particular embodiments, it is a short peptide (e.g., comprising an amino acid sequence comprising (or consisting or consisting essentially of) less than about: 100, 90, 80, 70, 60, 50, 40, 30, or 20 amino acids, or less, e.g., 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, or 5 amino acids). An interacting partner element can be any molecular entity but in certain embodiments it is a polypeptide. In certain particular embodiments, it is a polypeptide (e.g., comprising an amino acid sequence comprising (or consisting or consisting essentially of) at least about: 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000 amino acids, or more). For polypeptide embodiments of a subject molecule, interacting molecule or candidate molecule can comprise any amino acid, including both the 20 canonical amino acids as well as non-canonical amino acids — which can include one or more D- isomers of the 20 canonical amino acids or L or D isomers of non-canonical amino acids. In some embodiments, a variant of a polypeptide, e.g., library elements (candidate molecules), relative to a reference sequence, comprise substitution(s) with non-canonical amino acids — which can include one or more D-isomers of the 20 canonical amino acids or L or D isomers of non-canonical amino acids. In some embodiments, relative to a reference molecule a candidate molecule comprises conservative substitutions or highly conservative substitutions, relative to the reference sequence comprises conservative substitutions or highly conservative substitutions, relative to the reference sequence. “Conservative substitutions” relative to a reference sequence means a givenamino acid substitution has a value of 0 or greater in BLOSUM62. “Highly conservative substitutions” relative to a reference sequence means a given amino acid substitution has a value of 1 or greater (e.g., in some embodiments, 2, or more) in BLOSUM62.

[0049] Interactions between interacting partners — e.g., a subject molecule element, candidate molecule, or multi-targeting molecule with an interacting partner element — can be any type of molecule interaction, e.g., binding (whether direct or indirect) or catalytic interaction. In certain embodiments, the interaction is by binding, such as direct non-covalent binding. Binding can be evaluated by any computation method known in the art, such as heuristic scoring functions based on thresholding evaluation methods, Bayesian fitting of evaluation method results to fit experimentally validated results (either by classification or regression), or training a machine learning model to predict the experimentally validated scores (either by classification or regression). Exemplary tools for performing any of these methods include AlphaFold, OmegaFold, ESMFold, molecular docking, molecular dynamics, RosettaFold, FoldX, or dTERMen, including adaptations and modifications thereof.

[0050] Interaction, in some embodiments, may entail activity in addition to binding (e.g., direct non-covalent binding), such as downstream molecular signaling, such as agonism, antagonism, or biased signaling — e.g., where a multi -targeting molecule is a signaling molecule that, relative to a given activity of a reference molecule exhibits about: 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 95, 96, 97, 98, 99%, or more, of the activity , e.g., at least about: 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000% activity — thus in some embodiments, exhibits enhanced activity (whether activating, inhibiting, or biased). In other certain embodiments a multi-targeting molecule is a signaling molecule that, relative to a given activity of a reference molecule exhibits less than about: 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1%, or less of the activity, e.g., less than about: 0.5, 0.4, 0.3, 0.2, 0.1, 0.075, 0.05, 0.025, 0.01% of the activity — thus a reduced (or substantially eliminated) activity (whether activating, inhibiting, or biased). In certain embodiments, activities are substantially preserved (e.g., within about: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 1,3 14, 15, 16, 17, 18, 19, or 20% of reference activity, such as within about: 10% or 5%, or less), while introducing structural diversity.

[0051] Figure 7 illustrates a computer system 100 for identifying a multi -targeting molecule.

[0052] Referring to Figure 7, in typical embodiments, computer system 100 comprises one or more computers. For purposes of illustration in Figure 7, the computer system 100 isrepresented as a single computer that includes all of the functionality of the disclosed computer system 100. However, the present disclosure is not so limited. The functionality of the computer system 100 may be spread across any number of networked computers and / or reside on each of several networked computers and / or virtual machines. One of skill in the art will appreciate that a wide array of different computer topologies are possible for the computer system 100 and all such topologies are within the scope of the present disclosure.

[0053] Turning to Figure 7 with the foregoing in mind, the computer system 100 comprises one or more processing units (CPUs) 59, a network or other communications interface 84, a user interface 78 (e.g., including an optional display 82 and optional keyboard 80 or other form of input device), a memory 92 (e.g., random access memory, persistent memory, or combination thereof), one or more magnetic disk storage and / or persistent devices 90 optionally accessed by one or more controllers 88, one or more communication busses 12 for interconnecting the aforementioned components, and a power supply 79 for powering the aforementioned components. To the extent that components of memory 92 are not persistent, data in memory 92 can be seamlessly shared with non-volatile memory 90 or portions of memory 92 that are nonvolatile / persistent using known computing techniques such as caching. Memory 92 and / or memory 90 can include mass storage that is remotely located with respect to the central processing unit(s) 59. In other words, some data stored in memory 92 and / or memory 90 may in fact be hosted on computers that are external to computer system 100 but that can be electronically accessed by the computer system 100 over an Internet, intranet, or other form of network or electronic cable using network interface 84. In some embodiments, the computer system 100 makes use of algorithms that are run from the memory associated with one or more graphical processing units in order to improve the speed and performance of the system. In some alternative embodiments, the computer system 100 makes use of algorithms that are run from memory 92 rather than memory associated with a graphical processing unit.

[0054] In some embodiments, the memory 92 of the computer system 100 stores:• a scoring module that performs any of the methods of the present disclosure;• a plurality of three-dimensional reference structures 102-1, .. . , 102-N, where N is a positive integer of two or greater (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or greater), each comprising (i) an initial molecule element 104, with associated three-dimensional (3-D) coordinates 106 and (ii) a different interacting partnerelement 108 with associated three-dimensional (3-D) coordinates 110;• an identity of a plurality of candidate molecules 112-1, . . . , 112-M, where M is a positive integer of two or greater (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or greater), that is the same or different than N, where each candidate molecule in the plurality of candidate molecules is within a defined evolutionary chemical space relative to the initial molecule element as indicated by it corresponding distance to initial molecule element 114-M;• a plurality of three-dimensional (3-D) testing complexes 116-1, ..., 116-N, ..., 116-N+l, . . . , 116-K, where K is a positive integer greater than N, each three-dimensional (3-D) testing complex 116 comprising one of the candidate molecules 112 / 118 and one of the interacting partner elements 108 / 122 complexed to each other along with the 3-D coordinates 120 of the candidate molecule 112 / 118 within the 3-D testing complex and the 3-D coordinates 124 of the interacting partner element 108 / 122 within the 3-D testing complex 116.

[0055] In some implementations, one or more of the above identified data elements or modules of the computer system 100 are stored in one or more of the previously mentioned memory devices, and correspond to a set of instructions for performing a function described above. The above identified data, modules or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures or modules, and thus various subsets of these modules may be combined or otherwise re-arranged in various implementations. In some implementations, the memory 92 and / or 90 optionally stores a subset of the modules and data structures identified above. Furthermore, in some embodiments the memory 92 and / or 90 stores additional modules and data structures not described above.

[0056] Features of the methods and systems provided by the invention can include one or more of the following enumerated embodiments.

[0057] Embodiment 1. With reference to Figure 6, one aspect of the present disclosure provides a method of identifying a multi-targeting molecule. The method comprises providing a plurality of reference structures. Each member of the plurality of reference structures comprises a subject molecule element in complex with an interacting partner element as illustrated at element 602 of Figure. 6. In element 602 of Figure 6, each of the complexes are to the same initial subject molecule element (e.g., initial peptide). The method further comprises providing an in silicolibrary of candidate molecules as illustrated in element 604 of Figure. 6 In some embodiments, each of the candidate molecules in the in silica library of candidate molecules illustrated in element 604 of Figure 6 includes two or more mutations relative to the initial subject molecule element (e.g., initial peptide) of element 602 of Figure 6. The candidate molecules of the in silica library are evaluated for interaction with two or more of the interacting partner elements of the plurality of structures as illustrated in element 606 of Figure 6. An in silica computed interaction between the candidate molecule and the two or more of the interacting partner elements of the plurality of structures identifies the candidate molecule as a multi-targeting molecule.

[0058] Embodiment 2. The method of embodiment 1, wherein the library of candidate molecules is a library of polypeptide molecules.

[0059] Embodiment 3. The method of embodiment 1 or 2, in which: the subject molecule is a polypeptide, such as a polypeptide ligand; the interacting partner elements are polypeptides, such as a receptor polypeptide; the library of candidate molecules is a library of polypeptide molecules; the evaluating of the candidate molecules of the library for interacting with two or more of the interacting partner elements of the plurality of structures is by evaluating the candidate molecules of the library for binding two or more of the interacting partner elements of the plurality of structures; wherein the interaction between the candidate molecule and the two or more of the interacting partner elements of the plurality of structures satisfying a reference threshold identifies the candidate molecule as a multi-targeting molecule.

[0060] Embodiment 4. The method of any one of the preceding embodiments, wherein the library of candidate molecules is a library of polypeptides and comprises one or more of at least about: 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of all single point mutations of a portion (e.g, about: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100%) of a reference amino acid sequence, including, in some embodiments, wherein the library is a DMS library; a sample (e.g, about: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40%) of all possible double point mutations of a reference amino acid sequence; a sample (e.g, about: 1, 2, 3, 4, 5%) of any other possible high-order mutation (e.g, triple or quadruple point mutations, etc.) of a reference amino acid sequence.\ 6) \Embodiment 5. The method of any one of the preceding embodiments, wherein the evaluation of the interaction with the two or more of the interacting partners elements comprisespredicting one or more structures of the candidate molecule by threading the candidate molecule through the structure of the subject molecule element in the reference structure (e.g., using AlphaFold, OmegaFold, ESMFold, molecular docking, molecular dynamics, RosettaFold, FoldX, or dTERMen etc.) as illustrated in element 608 of Figure 6.

[0062] Embodiment 6. The method of any one of the preceding embodiments, wherein the plurality of reference structures comprise structures determined by one or more of: NMR, cryo- EM, X-ray diffraction, or computationally, optionally wherein the structure is an RMSD resolution of less than about: 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 Angstrom, or less, e.g., less than about 5 angstroms, such as about 3-1 Angstrom resolution.

[0063] Embodiment 7. The method of any one of the preceding embodiments, wherein the plurality of reference structures comprises at least about: 2, 3, 4, 5, 6, 7, 8, 9, 10 structures, or more, e.g, 20, 30, 40, 50, 100, 200, 400, 500, 1000, 2000, 4000, 5000, 10000 structures.

[0064] Embodiment 8. The method of any one of the preceding embodiments, wherein the multi-targeting molecule comprises one or more of: agonist activity for at least one interacting partner in a reference structure; antagonist activity for at least one interacting partner in a reference structure; neutral binding activity for at least one interacting partner in a reference structure; non-binding activity for at least one interacting partner in a reference structure; any combination thereof including partial agonist activity, partial antagonist activity, as well as condition-specific activity of any of the foregoing, such as pH dependent activity.

[0065] Embodiment 9. The method of any one of the preceding embodiments, wherein the plurality of reference structures comprises pairs of subject sequence and interacting partner are polypeptide ligand and a receptor (including cell-surface receptors, such as GPCRs and int egrins), respectively.

[0066] Embodiment 10. The method of any one of the preceding embodiments, further comprising identifying an ensemble of sequence variants of the multi-targeting molecule that substantially preserve the structure, e.g., to an RMSD resolution of less than about: 5, 4, 3, 2, 1 Angstrom, or less, e.g., less than about 5 Angstroms, such as about 3-1 Angstrom.

[0067] Embodiment 11. The method of embodiment 10, wherein the identification of the ensemble of sequence variants is by AlphaFold, OmegaFold, ESMFold, molecular docking, molecular dynamics, RosettaFold, FoldX, RFDiffusion, ProteinMPNN, or dTERMen.

[0068] Embodiment 12. The method of any one of the preceding embodiments, further comprising experimentally (e.g., in vitro, in vivo, or ex vivo) testing the interactions of a multitargeting molecule identified by the method, optionally wherein the multi -targeting molecule is a polypeptide, the one or more interacting partners are polypeptides (e.g., receptors, such as GPCR receptors) and the interaction comprises binding.

[0069] Embodiment 13. The method of any one of the preceding embodiments, further comprising synthesizing the multi-targeting molecule.

[0070] Embodiment 14. A system comprising a microprocessor and computer-readable instructions that, when executed by the microprocessor, causes the microprocessor to perform the method of any one of embodiments 1-13.

[0071] Embodiment 15. A multi-targeting molecule produced or producible by the method of any one of the embodiments of the present disclosure including embodiments 1-13.

[0072] Additional disclosure on methods for identifying a multi -targeting molecule is detailed with reference to Figure 8 and discussed below. Figure 10 illustrates pseudocode in accordance with one embodiment of the present disclosure that is consistent with the disclosure of Figure 8.

[0073] Referring to block 800, in some embodiments, a method of identifying a multi-targeting molecule comprises providing a plurality of three-dimensional reference structures, each member of the plurality of three-dimensional reference structures comprising an initial molecule element 104 in complex with a different interacting partner element 108 in a plurality of interacting partner elements.

[0074] Referring to block 802, in some embodiments, each three-dimensional reference structure 102 in the plurality of three-dimensional reference structures comprises a different interacting partner element 108 in the plurality of interacting partner elements determined by one or more of: NMR, cryo-EM, or X-ray diffraction.

[0075] Referring to block 804, in some embodiments, each three-dimensional reference structure 102 in the plurality of three-dimensional reference structures comprises a different interacting partner element 108 in the plurality of interacting partner elements determined by one or more of: NMR, cryo-EM, or X-ray diffraction.

[0076] Referring to block 806, in some embodiments, each three-dimensional reference structure 102 in the plurality of three-dimensional reference structures comprises a differentinteracting partner element 108 in the plurality of interacting partner elements determined by NMR with an RMSD resolution of less than about: 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 Angstrom, or less.

[0077] Referring to block 808, in some embodiments, each three-dimensional reference structure 102 in the plurality of three-dimensional reference structures comprises a different interacting partner element 108 in the plurality of interacting partner elements determined by X- ray diffraction with a crystallographic resolution of less than about: 3.5, 3.4, 3.3, 3.2, 3.1, 3.0, 2.9, 2.8, 2.7, 2.6, 2.5, 2.4, 2.3, 2.2, 2.1, 2.0, 1.9, 1.8, 1.7, 1.6, Angstroms or less.

[0078] Referring to block 810, in some embodiments, a three-dimensional reference structure 102 in plurality of three-dimensional reference structures is formed by computationally docking the initial molecule element 104 to an interacting partner element 108 in the plurality of interacting partner elements.

[0079] Referring to block 812, in some embodiments, the plurality of three-dimensional reference structures comprises at least about 2, 3, 4, 5, 6, 7, 8, 9, 10 structures, or more.

[0080] Referring to block 814, in some embodiments, the plurality of three-dimensional reference structures comprises at least 20, 30, 40, 50, 100, 200, 400, 500, 1000, 2000, 4000, 5000, or 10000 structures. In some embodiments, the plurality of three-dimensional reference structures comprises at least about: 2, 3, 4, 5, 6, 7, 8, 9, 10 structures, or more, e.g., 20, 30, 40, 50, 100, 200, 400, 500, 1000, 2000, 4000, 5000, 10000 structures.

[0081] Referring to block 816, in some embodiments, a plurality of candidate molecules is provided in silico. Each candidate molecule 112 in the plurality of candidate molecules is within a defined evolutionary chemical space relative to the initial molecule element 104. Examples of a defined evolutionary chemical space are context dependent. For instance, in the case where each candidate molecule and the initial molecule element is a polymer, the defined evolutionary chemical space can be defined by an allowed number of mutations or an allowed percentage of mutations in the sequence of the candidate molecule relative to the sequence of the initial molecule element. In the case where each candidate molecule and the initial molecule element is a small molecule, the defined evolutionary chemical space can be defined by an allowed range of graph edit distances in the candidate molecule relative to the initial molecule element.

[0082] In some embodiments, each candidate molecule 112 is a polymer.

[0083] In some embodiments, the initial molecule element 104 is a polymer.

[0084] In some embodiments, each interacting partner element is a polymer.

[0085] In some embodiments a polymer is a molecule composed of repeating structural units. These repeating structural units are termed residues herein. In some embodiments, each residue pi in the set of {pi, . . ., p / / } residues represents a single different residue in the polymer. To illustrate, consider the case where the polymer comprises 100 residues. In this instance, the set of {pi, .. ., p / . J comprises 100 residues, with each residue in {pi, .. ., p / . j representing a different one of the 100 residues.

[0086] In some embodiments, a polymer is a natural material. In some embodiments, a polymer is a synthetic material. In some embodiments, a polymer is an elastomer, shellac, amber, natural or synthetic rubber, cellulose, Bakelite, nylon, polystyrene, polyethylene, polypropylene, or polyacrylonitrile, polyethylene glycol, or polysaccharide.

[0087] In some embodiments, a polymer is a heteropolymer (copolymer). A copolymer is a polymer derived from two (or more) monomeric species, as opposed to a homopolymer where only one monomer is used. Copolymerization refers to methods used to chemically synthesize a copolymer. Examples of copolymers include, but are not limited to, ABS plastic, SBR, nitrile rubber, styrene-acrylonitrile, styrene-isoprene-styrene (SIS) and ethylene-vinyl acetate. Since a copolymer consists of at least two types of constituent units (also structural units, or particles), copolymers can be classified based on how these units are arranged along the chain. These include alternating copolymers with regular alternating A and B units. See, for example, Jenkins, 1996, “Glossary of Basic Terms in Polymer Science,” Pure Appl. Chem. 68 (12): 2287-2311, which is hereby incorporated herein by reference in its entirety. Additional examples of copolymers are periodic copolymers with A and B units arranged in a repeating sequence (e.g. (A-B-A-B-B-A-A-A-A-B-B-B)n).

[0088] In some embodiments, the polymer is a polypeptide. As used herein, the term “polypeptide” means two or more amino acids or residues linked by a peptide bond. The terms “polypeptide” and “protein” are used interchangeably herein and include oligopeptides and peptides. An “amino acid,” “residue” or “peptide” refers to any of the twenty standard structural units of proteins as known in the art, which include imino acids, such as proline and hydroxyproline. The designation of an amino acid isomer may include D, L, R and S. The definition of amino acid includes nonnatural amino acids. Thus, selenocysteine, pyrrolysine, lanthionine, 2-aminoisobutyric acid, gamma-aminobutyric acid, dehydroalanine, ornithine, citrulline and homocysteine are all considered amino acids. Other variants or analogs of theamino acids are known in the art. Thus, a polypeptide may include synthetic peptidomimetic structures such as peptoids. See Simon et al., 1992, Proceedings of the National Academy of Sciences USA, 89, 9367, which is hereby incorporated by reference herein in its entirety. See also Chin et al., 2003, Science 301, 964; and Chin et al., 2003, Chemistry & Biology 10, 511, each of which is incorporated by reference herein in its entirety.

[0089] The polymer evaluated in accordance with some embodiments of the disclosed systems and methods may also have any number of posttranslational modifications. Thus, a polypeptide includes those that are modified by acylation, alkylation, amidation, biotinylation, formylation, y-carboxylation, glutamylation, glycosylation, glycylation, hydroxylation, iodination, isoprenylation, lipoylation, cofactor addition (for example, of a heme, flavin, metal, etc I), addition of nucleosides and their derivatives, oxidation, reduction, pegylation, phosphatidylinositol addition, phosphopantetheinylation, phosphorylation, pyroglutamate formation, racemization, addition of amino acids by tRNA (for example, arginylation), sulfation, selenoylation, ISGylation, SUMOylation, ubiquitination, chemical modifications (for example, citrullination and deamidation), and treatment with other enzymes (for example, proteases, phosphotases and kinases). Other types of posttranslational modifications are known in the art and are also included.

[0090] In some embodiments, a polymer is an organometallic complex. An organometallic complex is chemical compound containing bonds between carbon and metal. In some instances, organometallic compounds are distinguished by the prefix “organo-“ e.g. organopalladium compounds. Examples of such organometallic compounds include all Gilman reagents, which contain lithium and copper. Tetracarbonyl nickel, and ferrocene are examples of organometallic compounds containing transition metals. Other examples include organomagnesium compounds like iodo(methyl)magnesium MeMgl, diethylmagnesium (Et2Mg), and all Grignard reagents; organolithium compounds such as n-butyllithium (n-BuLi), organozinc compounds such as diethylzinc (Et2Zn) and chloro(ethoxycarbonylmethyl)zinc (ClZnCH2C(=O)OEt); and organocopper compounds such as lithium dimethylcuprate (Li [CuMe2] ). In addition to the traditional metals, lanthanides, actinides, and semimetals, elements such as boron, silicon, arsenic, and selenium are considered form organometallic compounds, e.g. organoborane compounds such as triethylborane (EtsB).

[0091] Referring to block 818, in some embodiments, the plurality of candidate molecules is a plurality of polypeptides.

[0092] Referring to block 820, in some embodiments, the initial molecule element 104 is a polymer, the plurality of candidate molecules is a plurality of polypeptides comprising one or more of: a) a first collection of candidate molecules that collectively represent at least about 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of all single point mutations of a portion (e.g., about: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100%) of a sequence of the initial molecule element 104, including, in some embodiments; b) a second collection of candidate molecules that collectively represent at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, or 40% of all possible double point mutations of the sequence of the initial molecule element 104; c) a third collection of candidate molecules that collectively represent at least 1%, 2%, 3%, 4%, 5% of all possible 3-order mutations through M-order mutations of the initial molecule element 104, in which M is a length of the sequence of the initial molecule element 104; or d) any combination of any of the foregoing a), b), and c).

[0093] Referring to block 820, in some embodiments, a plurality of three-dimensional testing complexes is formed. Each respective three-dimensional testing complex 116 comprises a candidate molecule 112 three-dimensionally threaded into a place of the initial molecule element 104 in a corresponding three-dimensional reference structure 102 in the plurality of three- dimensional reference structures thereby replacing the initial molecule element in the corresponding three-dimensional reference structure 102 to form the respective three- dimensional testing complex 116. In some embodiments, a program such as AlphaFold, OmegaFold, ESMFold, molecular docking, molecular dynamics, RosettaFold, FoldX, RFDiffusion, ProteinMPNN, or dTERMen is used to accomplish this threading.

[0094] Referring to block 822, in some embodiments, the plurality of candidate molecules is evaluated for interaction with the plurality of interacting partner elements using a scoring function that jointly scores an interaction between a candidate molecule 112 of the plurality of candidate molecules and each respective three-dimensional testing complex 116 in the plurality of testing complexes that includes the candidate molecule 112 to form a single score.Satisfaction of a threshold by the single score satisfies identifies the candidate molecule 112 as a multi-targeting molecule.

[0095] An example, in embodiments where the initial molecule element and the interacting partner element are polymers, is given by Equation 1 in the section below entitled “DMS-based scoring function to identify high-order mutated binders '' Equation 1 uses three metrics: AG, AF2 RMSD, AF2 pAE interaction.

[0096] In some embodiments where the initial molecule element 104 is a polymer, rather than computing AF2 RMSD using AlphaFold2 as described for Equation 1, root mean squared deviation (RMSD) is computed using a program such as PyMold, visual molecular dynamics (VMD), Chimera / ChimeraX, BioPython / MDTraj / ProDy, TM-align, GROMACS, or related RMSD functions.

[0097] In some embodiments where the initial molecule element 104 is a polymer rather than computing AF2 pAE using AlphaFold2 as described for Equation 1, other measures of confidence in structural alignment between the candidate molecule 112 / 118 and the interacting partner element 108 / 122 are used instead of AF2 pAE, such as the confidence matrices provided by RoseTTAFold, RoseTTAFold2, or OmegaFold, which are typically given as given as per- residue RMSD estimates or inter-residue distance error distributions. Such matrices can be evaluated to give a mean confidence analogous to AF2 pAE.

[0098] In embodiments in which the initial molecule element 104 and the candidate molecule are small molecules, rather than computing AF2 pAE using AlphaFold2, confidence in the interaction between the candidate molecule 112 / 118 (e.g., small molecule) and the interacting partner element 108 / 122 (e.g., receptor) is garnered by focusing on the residues in the interacting partner element 108 / 122 that interact with the atoms in the small molecule. For instance, a distance threshold (e.g., <5-8 A) can be used to focus on the interacting partner element 108 / 122 - candidate molecule 112 / 118 interactions that are physically meaningful. In fact, the AlphaFol d2 model can still provide a PAE matrix where inter-chain (or inter-molecular) PAE can be extracted between the interacting partner element 108 / 122’ s residues and the candidate molecule 112 / 118’s atoms. For example, the PAE for receptor-small molecule pairs would still reflect how confidently AlphaFold2 predicts their relative spatial orientation.

[0099] In some embodiments, rather than using AF2 pAE, a different measure of confidence in the interaction between the interacting partner element 108 / 122’s residues and the candidate molecule 112 / 118 ’ s atoms is used such as those provided by AutoDoc / AutoDock Vina, RosettaLigan, GNINA, or HADDOCK.

[0100] In some embodiments where the candidate molecule 1 12 / 118, rather than determining RMSD of the testing complex 116 relative to a reference structure 102 using AF2 RMSD, programs such as PyMOL, Chimera, Chimera X, RDKit, Visual Molecular Dynamics (VMD), or Open Babel can be used.

[0101] Referring to block 824, in some embodiments, any one or any combination of a) the initial molecule element 104 being a first polypeptide, b) interacting partner element 108 in the plurality of interacting partner elements being a polypeptide; c) the plurality of candidate molecules being a plurality of polypeptide molecules; and d) the evaluating of the plurality of candidate molecules for interaction with the plurality of interacting partner elements of the plurality of three-dimensional testing complexes comprises evaluating the plurality of candidate molecules for binding against each of the interacting partner elements in the plurality of interacting partner elements.

[0102] Referring to block 826, in some embodiments, the first polypeptide is a polypeptide ligand.

[0103] Referring to block 828, in some embodiments, an interacting partner element 108 in the plurality of interacting partner elements is a receptor polypeptide.

[0104] Referring to block 830, in some embodiments, the multi -targeting molecule comprises one or more of: a) agonist activity for at least one interacting partner in the plurality of interacting partner elements; b) antagonist activity for at least one interacting partner in the plurality of interacting partner elements; c) neutral binding activity for at least one interacting partner in the plurality of interacting partner elements; d) non-binding activity for at least one interacting partner in the plurality of interacting partner elements; e) any combination of a), b), c), or d) including partial agonist activity, partial antagonist activity, as well as condition-specific activity of any of the foregoing, such as pH dependent activity for an interacting partner element 108 in the plurality of interacting partner elements.

[0105] Referring to block 834, in some embodiments, an ensemble of sequence variants of the multi-targeting molecule that preserve the structure of the multi-targeting molecule to an RMSD resolution of less than about: 5, 4, 3, 2, 1, 0.5, or 0.3 Angstroms, or less is identified.

[0106] Referring to block 836, in some embodiments, the identification of the ensemble of sequence variants is by AlphaFold, OmegaFold, ESMFold, molecular docking, molecular dynamics, RosettaFold, FoldX, RFDiffusion, ProteinMPNN, or dTERMen.

[0107] Referring to block 838, in some embodiments the interactions of a multi-targeting molecule identified by the method are tested in vitro, in vivo, or ex vivo.

[0108] Referring to block 840, in some embodiments, the multi-targeting molecule is a polypeptide, the plurality of interacting partner elements are polypeptides and the interaction comprises binding. In some embodiments the binding is quantified by an IC50, EC50, Kd, KI, or pKI for the multi-targeting molecule with respect to a reference structure. IC50, EC 50, Kd, KI, and pKI are generally described in Huser ed., 2006, High-Throughput-Screening in Drug Discovery, Methods and Principles in Medicinal Chemistry 35; and Chen ed., 2019, A Practical Guide to Assay Development and High-Throughput Screening in Drug Discovery, each of which is hereby incorporated by reference.

[0109] Referring to block 842, in some embodiments, a first interacting partner element 108 in the plurality of interacting partner elements is a receptor e.g., a G protein-coupled receptor).

[0110] Referring to block 844, in some embodiments, the initial molecule element 104 is a first small molecule, the first candidate molecule 112 of the plurality of candidate molecules is a second small molecule, and the scoring function normalizes a score of the interaction between the first candidate molecule 112 and a three-dimensional testing complex 116 by a graph edit distance (GED) between the first candidate molecule 112 and the initial molecule element 104.

[0111] In some embodiments a small molecule satisfies two or more rules, three or more rules, or all four rules of the Lipinski's rule of Five; (i) not more than five hydrogen bond donors, (ii) not more than ten hydrogen bond acceptors, (iii) a molecular weight under 500 Daltons, and (iv) a LogP under 5. See, Lipinski, 1997, Adv. Drug Del. Rev. 23, 3, which is hereby incorporated herein by reference in its entirety. In some embodiments, a small molecule satisfies one or more criteria in addition to Lipinski's Rule of Five. For example, in some embodiments, a small molecule has five or fewer aromatic rings, four or fewer aromatic rings, three or fewer aromatic rings, or two or fewer aromatic rings.

[0112] In some embodiments, a small molecule is an organic compound having a molecular weight of less than 500 Daltons, less than 1000 Daltons, less than 2000 Daltons, less than 4000 Daltons, less than 6000 Daltons, less than 8000 Daltons, less than 10000 Daltons, or less than 20000 Daltons. In some embodiments, a small molecule is an organic compound having a molecular weight of between 400 Daltons and 10000 Daltons.

[0113] GED distance serves as a basis for quantifying differences in molecule structure. See, for example, Garcia-Hernandez, 2020, “Learning the Edit Costs of Graph Edit Distance Applied to Ligand-Based Virtual Screening,” Current Topics in Medicinal Chemistry 20, 1582-1592, which is hereby incorporated by reference. In some embodiments, the graph edit distance is a chemical graph edit distance. In some embodiments, the graph edit distance is an exact graph edit distance. Figure 9 illustrates the difference between chemical and exact graph edit distance. As seen in Figure 9, the chemical graph edit distance considers chemical changes whereas the exact GED considers the number of elementary changes described above.

[0114] Referring to block 846, in some embodiments, the initial molecule element 104 is a first polypeptide, the first candidate molecule 112 of the plurality of candidate molecules is a second polypeptide, and the scoring function normalizes a score of the interaction between the first candidate molecule 112 and a three-dimensional testing complex 116 by a mutation distance between the first candidate molecule 112 and the initial molecule element 104.

[0115] Referring to block 848, in some embodiments, the initial molecule element 104 is a first polypeptide, each candidate molecule 112 of the plurality of candidate molecules is different polypeptide, and a candidate molecule 112 in the plurality of candidate molecules is within the defined evolutionary chemical space when a sequence of the candidate molecule 112 and a sequence of the initial molecule element 104 share at least seventy percent sequence identity with the proviso that the sequence of the candidate molecule 112 includes at least two or more sequence substitutions relative of the sequence of the initial molecule element 104.

[0116] In some embodiments, the initial molecule element 104 is a first polypeptide, each candidate molecule 112 of the plurality of candidate molecules is different polypeptide, and a candidate molecule 112 in the plurality of candidate molecules is within the defined evolutionary chemical space when a sequence of the candidate molecule 112 and a sequence of the initial molecule element 104 share at least 80%, 85%, or 90% sequence identity with the proviso that the sequence of the candidate molecule 112 includes at least two or more sequence substitutions relative of the sequence of the initial molecule element 104.

[0117] In some embodiments, the initial molecule element 104 is a first polypeptide, each candidate molecule 112 of the plurality of candidate molecules is different polypeptide, and a candidate molecule 112 in the plurality of candidate molecules is within the defined evolutionary chemical space when a sequence of the candidate molecule 112 and a sequence of the initialmolecule element 104 share at least 40%, 45%, 50%, 55%, 60%, or 65% sequence identity with the proviso that the sequence of the candidate molecule 112 includes at least two or more sequence substitutions relative of the sequence of the initial molecule element 104.

[0118] Referring to block 850, in some embodiments, the initial molecule element 104 is a first small molecule, each candidate molecule 112 of the plurality of candidate molecules is a different small molecule, and a candidate molecule 112 in the plurality of candidate molecules is within the defined evolutionary chemical space when a graph edit distance (GED) of the candidate molecule 112 and the initial molecule element 104 is between 2 and 5. In some such embodiments, the GED is chemical graph edit distance. In some such embodiments, the GED is exact graph edit distance.

[0119] Referring to block 852, in some embodiments, the initial molecule element 104 is a first small molecule, each candidate molecule 112 of the plurality of candidate molecules is a different small molecule, and a candidate molecule 112 in the plurality of candidate molecules is within the defined evolutionary chemical space when a graph edit distance of the candidate molecule 112 and the initial molecule element 104 is between 4 and 10. In some such embodiments, the GED is chemical graph edit distance. In some such embodiments, the GED is exact graph edit distance.

[0120] Referring to block 854, in some embodiments, the initial molecule element 104 is a first small molecule, each candidate molecule 112 of the plurality of candidate molecules is a different small molecule, and a candidate molecule 112 in the plurality of candidate molecules is within the defined evolutionary chemical space when a graph edit distance of the candidate molecule 112 and the initial molecule element 104 is between 6 and 14. In some such embodiments, the GED is chemical graph edit distance. In some such embodiments, the GED is exact graph edit distance.

[0121] Referring to block 856, in some embodiments, the method further comprises synthesizing the multi-targeting molecule.

[0122] Referring to block 858, in some embodiments, a multi-targeting molecule identified by any method of the present disclosure is provided.

[0123] EXEMPLIFICATION

[0124] Methods.

[0125] Backbone extraction

[0126] Backbones can be extracted programmatically by directly fdtering side chain atoms or using a tool that provides a built-in functionality with that intent such as Pymol and UCSF Chimera. An ensemble of backbones can also be obtained with RFDiffusion in cases where additional conformation variability is required. In this work, a backbone per complex was obtained with RFDiffusion by preserving initial atomic positions as close as possible to native ones.

[0127] Mutagenesis landscape exploration

[0128] For each provided peptide, all single-point mutations for its sequence were generated. In other embodiments, all high-order mutations (e.g., double- or triple-point mutations) could also be covered; however, due to their usual large mutational landscape, a random fraction of such high-order mutations can be generated instead as a purpose to identify potential epistatic effects.

[0129] Generated variants were then threaded onto the original binder backbone in the complex using Rosetta. Only modeled backbones were relaxed using Rosetta FastRelax because relaxing sidechains was not efficient as they are not used as input for AlphaFold2.

[0130] Binding prediction

[0131] Using certain methods of Bennett et al (2023), for each protein-ligand complex, AlphaFold2 was employed for binding prediction by evaluating whether refolding each target and peptide sequences produced a bound complex structure similar to the natural complex. This refolding was performed using the target structure as a template and atom positions were used to initiate AlphaFold2’s initial guess. Resultant structures were then minimized using OpenMM in triplicates.

[0132] Binding energy calculation

[0133] Binding energy (AG) was calculated for minimized structures for each protein-ligand complex using FoldX. Then, the relative free energy change (AAG) was calculated on respect to the natural complex. As the experiments were carried out in triplicates, the average and standard deviation were taken for each refolded structure.

[0134] Multi-target binders ’ identification

[0135] After predicting binding and calculating binding energy, a molecule variant (candidate molecule 112) was identified as a binder if the refolded structure (3-D testing complex 116) obtained with AlphaFold2 presented an RMSD < 2 A with respect to its originating natural complex (3-D reference structure 102), a pAE interaction < 10 A (Bennett et al., 2023), and AAG< 0. Here, pAE stands for the average predicted aligned error between inter-chain residue pairs, e.g., between the binder (candidate molecule 112 / 118) and the receptor (interacting partner element 108 / 124).

[0136] Consequently, a candidate molecule 112 was considered a multi-target binder if it matches these restrictions for two or more targets.

[0137] In some embodiments, a different criterion is used for RMSD, such < 3 A, < 2.5 A, < 2.0 A, or < 1.8 A with respect to its originating natural complex.

[0138] In some embodiments, a different criterion is used for pAE interaction, such < 15 A, < 12 A, < 8 A, or < 7 A. For instance, requiring pAE interaction to be < 7 A requires that the spatial alignment between the candidate molecule 112 / 118 and the interacting partner element 108 / 124 be spatially much closer than the case in which pAE interaction are < 10 A.

[0139] DMS-based scoring function to identify high-order mutated binders.

[0140] With the purpose of assessing if in silico DMS (Deep Mutational Scanning) single-point mutations could be used for high-order mutated binders’ identification, a novel high advantageous scoring function was devised that takes into account the contribution of each mutation on top of the initial peptide. In other words, for a given high-order variant, its score was calculated as the weighted sum contribution of each single mutation considering the following equation:where i is protein target, j is {AAG, AF2 RMSD, AF2 pAE interaction} as described in Section Multi-target binders ’ identification, m is the mutations from the initial peptide used to create the in silico DMS, and N is the number of mutations in the higher-order variant relative to the initial peptide. To see how Equation 1 is used to score a high-order mutated binder using information from component single-order mutations of the initial peptide, consider the case of a thirty residue binder that bears four mutations relative to the initial peptide, at positions 5, 13, 18, and 23. In this instance, the “N” in equation 1 will be “4” since there are four mutations. Further, j is summed over the three thresholds {AAG, AF2 RMSD, AF2 pAE interaction} for each of the four mutations in the high-order mutated binder relative to the initial peptide.

[0141] Consider the case in which the mutation at position 5 satisfied the threshold for AAG but not the thresholds for AF2 RMSD and AF2 pAE interaction against a first target. In this case, the summation of position 5 is 1 + 0 + 0, for a total of 1 against the first target.

[0142] Further the mutation at position 5 satisfied the threshold for AAG but not the thresholds for AF2 RMSD and AF2 pAE interaction against a second target. Again, in this case, the summation of position 5 is 1 + 0 + 0, for a total of 1 against the second target.

[0143] Consider the case in which the mutation at position 13 satisfied the threshold for all three metrics {AAG, AF2 RMSD, AF2 pAE interaction} against the first target. In this case, the summation of position 13 is 1 + 1 + 1, for a total of 3 against the first target.

[0144] Further the mutation at position 13 satisfied the threshold for metrics AAG and AF2 RMSD but not AF2 pAE interaction against the second target. In this case, the summation of position 13 is 1 + 1 + 0, for a total of 2 against the second target.

[0145] Consider the case in which the mutation at position 18 satisfied the threshold for two of the metrics AF2 RMSD and AF2 pAE interaction, but not AAG against both the first and second targets. In this case, the summation of position 13 is 0 + 1 + 1, for a total of 2 against the first target and 0 + 1 + 1 , for a total of 2 against the second target.

[0146] Finally, consider the case in which the mutation at position 23 did not satisfy the threshold for any of the metrics against either the first target or the second target. In this case, the summation of position 23 is 0 + 0 + 0, for a total of 0.

[0147] In this example, the summation of positions 5, 13, 18, and 23 across the AAG, AF2 RMSD, and AF2 pAE interaction thresholds is 1+1 (scores for position 5 against the first and second targets) + 3+2 (scores for position 13 against the first and second targets) + 2 +2 (scores for position 18 against the first and second targets), + 0 + 0 (scores for position 23 against the first and second targets) for a grand total of 11. Thus,has the value “11” in this instance, where i is the set {1, 2} because there are two targets in this example. The value 11 is then normalized by N in accordance with Equation 1, which is “4” in this example, and thus the score provided by Equation 1 for this example is 11 / 4 or 2.75.

[0148] It will be appreciated that variations of Equation 1 can be used to score a higher-order peptide variant relative to the initial peptide. For instance, in some embodiments, each of themetrics is weighted differently. For instance, instead of assigning each metric {AAG, AF2 RMSD, AF2 pAE interaction} a “1” if the threshold for the metric is satisfied, each metric contributes to the final score differently. For instance, satisfaction of the AAG threshold gamers wi * 1, rather than “1”, satisfaction of the AF2 RMSD threshold garners W2 * 1, rather than “1”, and satisfaction of the AF2 pAE threshold garners W3 * 1, rather than “1”, where, in some embodiments, wi, W2, and W3 are real numbers greater than zero that may or may not equal each other. In one instance wi is 1.5, W2 is 1, and W3 is 1.75. In some embodiments, wi, W2, and W3 are each independently values between 0.01 and 10. In some embodiments, wi, W2, and W3 are each independently values between 0.01 and 100. In some embodiments, wi, W2, and W3 are each independently values between 1 and 1000.

[0149] Departing from these in silico DMS profiles, mutations that improved binding energy to different target combinations were identified by stacking profiles onto the same heatmap as shown in Figure 2. This elegant and insightful visualization revealed variants that are predicted to improve binding across all targets. It can also be used for identifying variants that improve binding for all but one or more targets. Such outcomes can also be used for sequence library design planning, guide computational modeling, and, therefore, reduce time and cost for engineering multi-target binders. Figures 1A, IB, and 1C provide AAG values for each of the possible naturally occurring single substitutions of the initial peptide at each of the 30 positions. In Figure 2, this information is combined into a single heat map. In Figure 2, there are eight possible values that each pixel (particular substation as a particle position) can have: “all”, “target 1 & target 3,” “target 1 & target 2”, “target 2 & target 3”, “target 1, “target 3”, “target 2”, “none” in accordance with table 1 below.

[0150] Table 1 Further detail on Figure 2 legend.

[0151] Thus, in the case where AAG for a given mutation at a given position is less than zero against all three targets as summarized in Figures 1A, IB, and 1C, the mutation at the given position will be assigned “All” and shaded with the shade for “All” in Figure 2. In the case where AAG for a given mutation at a given position is less than zero against targets 1 and 3, but not target 2, as summarized in Figures 1A, IB, and 1C, the mutation at the given position will be assigned “Target 1 & Target 2” and shaded with the shade for “Target 1 & target 2” in Figure 2. In some embodiments, each of the possible categories in Table 1 is assigned a different color, rather than a different shade.

[0152] In some embodiments, a different threshold is used. For instance, in some embodiments, the threshold is -0.5 kcal / mol is used. In the case where the threshold for AAG is -0.5 kcal / mol, and a given mutation at a given position has a AAG that is less than -0.5 kcal / mol against all three targets as summarized in Figures 1A, IB, and 1C, the mutation at the given position will be assigned “All” and shaded with the shade for “All” in a plot analogous to Figure 2. In the case where the threshold for AAG is -0.5 kcal / mol, but for the given mutation at the given position the AAG is greater than -0.5 kcal / mol against all three targets as summarized in Figures 1A, IB, and 1C, the mutation at the given position will be assigned “None” and not shaded (or colored) in a plot analogous to Figure 2.

[0153] In some embodiments, a plot such as Figure 2 is used to combine more than 1 metric. For instance, in some embodiments a plot analogous to Figure 2 combines all three metrics {AAG, AF2 RMSD, AF2 pAE interaction}. Thus, in the case where a given mutation at a given position has satisfied the threshold for all three metrics for all three targets as summarized in Figures 1 A, IB, and 1 C (in the case of AAG and analogous plots for AF2 RMSD, AF2 pAE interaction), the mutation at the given position will be assigned “All” and shaded with the shade for “All” in Figure 2. Thus, in the case where a given mutation at a given position has satisfied the threshold for two of the three metrics for all three targets as summarized in Figures 1A, IB, and 1C (in the case of AAG and analogous plots for AF2 RMSD, AF2 pAE interaction), the mutation at the given position will be assigned “None” since it did not satisfy all three metrics for any of the targets.

[0154] Plots such as Figure 2 allow for the rapid identification of mutation combinations that have desired specificity. For instance, Figure 2 identifies four mutations: {position 7, mutation Q}, {position 19, mutation N}, {position 23, mutation Q}, and {position 27, mutation K} that have negative AAG for targets 3 but not targets 1 and 2, indicating that these mutations are selective for target 3.

[0155] Scores of experimentally validate multi-binders.

[0156] The effectiveness of these approaches was further illustrated by retrospectively assessing if in silico DMS profiles obtained for the three receptor-peptide complexes can identify multitarget binders from a pool of 33 peptides synthesized and tested for binding experimentally. The average mutational distance of these sequences from the reference sequence was 10-13 mutations. All 33 peptides were found to bind at least one target, while 6 were found to bind all targets.

[0157] These selected sequences were then scored and ranked using the exemplary scoring function (Equation 1), which was devised to account for the contribution of each mutation on top of the initial peptide. After applying the filters described in Section ‘Multi-target binders’ identification’, 5 out of the 6 full multi-target binders in the top 15 sequences were identified (Figure 3).

[0158] A total of 5,000 peptide sequences with similar mutational load to the above experimentally synthesized peptides (10-13 mutations) were generated on a random basis and scored by the above in silico DMS-based scoring function (Figure 4). Applying the cut-off threshold (dashed line in Figure 4) identified for the experimental validated peptides, revealed that out that only 5% of the simulated sequences scored above it, suggesting that in silico DMS- based prediction can be used to enrich for multi-target binders.

[0159] Results

[0160] In silico DMS reveals unique target-peptide profiles.

[0161] To assess the pipeline proficiency at revealing positions prone to mutation that leads to increased binding energy against a set of targets, three receptors were selected and complexed with the same peptide available in RCSB. An in silico DMS profile was made for each by first generating all single-point mutations. Then, mutated sequences were thread onto the initial complex backbones and in silico refolding was employed to predict binding. Finally, binding energy was measured as described in Section ‘Binding energy calculation’ and variants predictednot to bind at the receptor binding site (RMSD > 5 A) or to have a low binding energy (AAG > 0) were filtered out.

[0162] Figures 1A, IB, and 1C show the unique landscape of promising mutations across the different targets. Interestingly, the Target 1 profile revealed that only a few mutations would be acceptable on top of the initial target-peptide conformation, while Target 3, on the other hand, showed the highest propensity to accommodate an extensive number of single-point mutations. For the Target 3 receptor, most of the highly mutable positions were in the initial 12 residues. Without intending to be limited by any particular theory, it is believed this is caused by the initial complex conformations’ that may influence on side-chain refolding and packing by AlphaFold2.

[0163] These profiles also confirmed several mutations that keep physicochemical properties as the original amino acid as promising (e.g., DxE and ExD across all targets, or a neutral L4A mutation in Targets 1 and 3), while it also shed light on unexpected mutations that change such properties. One example is found in the first heatmap, where mutations E— >K and K— >V showed improved peptide binding.

[0164] To test the ability of the novel enrichment function embodied by Equation 1, peptides were synthesized. Each of the peptide includes two or more mutations relative to the same initial peptide. The synthesized peptides were tested for binding against Targets 1, 2, and 3 and grouped into two sets. Set 1 peptides consisted of those peptides that did achieve binding below a set threshold to any of the targets. Set 1 consisted of 96 peptides. Set 2 also consisted of 96 peptides. Set 2 peptides consisted of those peptides that bound below a set threshold to at least one of the three targets. In set 2, sixteen peptides bound below the set threshold to all three targets. The single mutation data relative to the initial peptide for AAG, AF2 RMSD, AF2 pAE interaction was generated as described above, for instance in conjunction with Figures 1A, IB, and 1C and used to score each of the peptides in sets 1 and 2 using Equation 1. Using this scoring technique, all peptides in set 1 failed to satisfy the binding threshold against Targets 1, 2 and 3. Figure 5 illustrates. In Figure 5, the Equation 1 score for each of the peptides in set 1 (the nonbinder peptides) against Targets 1, 2 and 3 is plotted in light grey. By contrast, the Equation 1 score for each peptides in set 2 satisfied the threshold. For instance, all 16 of the peptides in the second set of peptides that had been found to bind to all three targets below the target binding threshold in vitro, achieved a high score using Equation 1. This is illustrated in Figure 5. In Figure 5, the Equation 1 score for each of the sixteen peptides in set 2 that werefound to bind to all three targets in vitro is plotted in dark grey. It is seen that each of these peptides has a score exceeding the threshold illustrated in Figure 5, indicating that they likely bind to all three targets.

[0165] It should be understood that for all numerical bounds describing some parameter in this application, such as “about,” “at least,” “less than,” and “more than,” the description also necessarily encompasses any range bounded by the recited values. Accordingly, for example, the description “at least 1, 2, 3, 4, or 5” also describes, inter alia, the ranges 1-2, 1-3, 1-4, 1-5, 23, 2-4, 2-5, 3-4, 3-5, and 4-5, et cetera.

[0166] For all patents, applications, or other reference cited herein, such as non-patent literature and reference sequence information, it should be understood that they are incorporated by reference in their entirety for all purposes as well as for the proposition that is recited. Where any conflict exists between a document incorporated by reference and the present application, this application will control. All information associated with reference gene sequences disclosed in this application, such as GenelDs or accession numbers (typically referencing NCBI accession numbers), including, for example, genomic loci, genomic sequences, functional annotations, allelic variants, and reference mRNA (including, e.g., exon boundaries or response elements) and protein sequences (such as conserved domain structures), as well as chemical references {e.g., PubChem compound, PubChem substance, or PubChem Bioassay entries, including the annotations therein, such as structures and assays, et cetera), are hereby incorporated by reference in their entirety.

[0167] Headings used in this application are for convenience only and do not affect the interpretation of this application.

[0168] Preferred features of each of the aspects provided by the invention are applicable to all of the other aspects of the invention mutatis mutandis and, without limitation, are exemplified by the dependent claims and also encompass combinations and permutations of individual features (e.g., elements, including numerical ranges and exemplary embodiments) of particular embodiments and aspects of the invention, including the working examples. For example, particular experimental parameters exemplified in the working examples can be adapted for use in the claimed invention piecemeal without departing from the invention. For example, for materials that are disclosed, while specific reference of each of the various individual and collective combinations and permutations of these compounds may not be explicitly disclosed,each is specifically contemplated and described herein. Thus, if a class of elements A, B, and C are disclosed as well as a class of elements D, E, and F and an example of a combination of elements A-D is disclosed, then, even if each is not individually recited, each is individually and collectively contemplated. Thus, in this example, each of the combinations A-E, A-F, B-D, B-E, B-F, C-D, C-E, and C-F are specifically contemplated and should be considered disclosed from disclosure of A, B, and C; D, E, and F; and the example combination A-D. Likewise, any subset or combination of these is also specifically contemplated and disclosed. Thus, for example, the sub-groups of A-E, B-F, and C-E are specifically contemplated and should be considered disclosed from disclosure of A, B, and C; D, E, and F; and the example combination A-D. This concept applies to all aspects of this application, including elements of a composition of matter and steps of method of making or using the compositions.

[0169] The forgoing aspects of the invention, as recognized by the person having ordinary skill in the art following the teachings of the specification, can be claimed in any combination or permutation to the extent that they are novel and non-obvious over the prior art — thus, to the extent an element is described in one or more references known to the person having ordinary skill in the art, they may be excluded from the claimed invention by, inter alia, a negative proviso or disclaimer of the feature or combination of features.

[0170] It should be understood that the example embodiments described above may be implemented in many different ways. In some instances, the various methods and machines described herein may be implemented by a physical, virtual, or hybrid general purpose computer, or a computer network environment.

[0171] Embodiments or aspects thereof may be implemented in the form of hardware, firmware, or software. If implemented in software, the software may be stored on any non-transient computer-readable medium that is configured to enable a processor to load the software or subsets of instructions thereof. The processor then executes the instructions and is configured to operate or cause an apparatus to operate in a manner as described herein.

[0172] Further, firmware, software, routines, or instructions may be described herein as performing certain actions and / or functions of the data processors. However, it should be appreciated that such descriptions contained herein are merely for convenience and that such actions, in fact, result from computer devices, processors, controllers, or other devices executing the firmware, software, routines, instructions, etc.

[0173] It should also be understood that the schematics may include more or fewer elements, be arranged differently, or be represented differently. But it should further be understood that certain implementation may dictate that the schematic be implemented in a particular way.

[0174] The described computer-readable implementations may be implemented in software, hardware, or a combination of hardware and software. Examples of hardware include computing or processing systems, such as personal computers, servers, laptops, mainframes, and microprocessors. In addition, one of ordinary skill in the art will appreciate that the records and fields shown in the figures may have additional or fewer fields, and may arrange fields differently than the figures illustrate. Any of the computer-readable implementations provided by the invention may, optionally, further comprise a step of providing a visual output to a user, such as a visual representation of multi-target molecules, their properties (e.g., binding properties, specificities, biophysical characteristics), including sequences, etc.

Claims

What is claimed is:

1. A method of identifying a multi-targeting molecule, the method comprising: a) providing a plurality of three-dimensional reference structures, each member of the plurality of three-dimensional reference structures comprising an initial molecule element in complex with a different interacting partner element in a plurality of interacting partner elements; b) providing a plurality of candidate molecules in silica, wherein each candidate molecule in the plurality of candidate molecules is within a defined evolutionary chemical space relative to the initial molecule element; c) forming a plurality of three-dimensional testing complexes, wherein each respective testing complex comprises a candidate molecule three-dimensionally threaded into a place of the initial molecule element in a corresponding three-dimensional reference structure in the plurality of three-dimensional reference structures thereby replacing the initial molecule element in the corresponding three-dimensional reference structure to form the respective three-dimensional testing complex; and d) evaluating the plurality of candidate molecules for interaction with the plurality of interacting partner elements using a scoring function that jointly scores an interaction between a candidate molecule of the plurality of candidate molecules and each respective three-dimensional testing complex in the plurality of testing complexes that includes the candidate molecule to form a single score, wherein the single score satisfying a threshold identifies the candidate molecule as a multi-targeting molecule.

2. The method of claim 1, wherein the plurality of candidate molecules is a plurality of polypeptides.

3. The method of claim 1 or 2, in which: a) the initial molecule element is a first polypeptide; b) interacting partner element in the plurality of interacting partner elements is a polypeptide; c) the plurality of candidate molecules is a plurality of polypeptide molecules;d) the evaluating of the plurality of candidate molecules for interaction with the plurality of interacting partner elements of the plurality of three-dimensional testing complexes comprises evaluating the plurality of candidate molecules for binding against each of the interacting partner elements in the plurality of interacting partner elements or e) any combination of the foregoing a), b), c), and d).

4. The method of claim 3, wherein the first polypeptide is a polypeptide ligand.

5. The method of claim 3, wherein an interacting partner element in the plurality of interacting partner elements is a receptor polypeptide.

6. The method of any one of claims 1-5, wherein the initial molecule element is a polymer, the plurality of candidate molecules is a plurality of polypeptides comprising one or more of: a) a first collection of candidate molecules that collectively represent at least about 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of all single point mutations of a portion (e.g., about: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100%) of a sequence of the initial molecule element, including, in some embodiments; b) a second collection of candidate molecules that collectively represent at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, or 40% of all possible double point mutations of the sequence of the initial molecule element; c) a third collection of candidate molecules that collectively represent at least 1%, 2%, 3%, 4%, 5% of all possible 3-order mutations through M-order mutations of the initial molecule element, wherein M is a length of the sequence of the initial molecule element; or d) any combination of any of the foregoing a), b), and c).

7. The method of any one of claims 1-6, wherein each three-dimensional reference structure in the plurality of three-dimensional reference structures comprises a different interacting partnerelement in the plurality of interacting partner elements determined by one or more of: TWTR., cryo-EM, or X-ray diffraction.

8. The method of any one of claims 1-6, wherein each three-dimensional reference structure in the plurality of three-dimensional reference structures comprises a different interacting partner element in the plurality of interacting partner elements determined by NMR with an RMSD resolution of less than about: 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 Angstrom, or less.

10. The method of any one of claims 1-6, wherein each three-dimensional reference structure in the plurality of three-dimensional reference structures comprises a different interacting partner element in the plurality of interacting partner elements determined by X-ray diffraction with a crystallographic resolution of less than about: 3.5, 3.4, 3.3, 3.2, 3.1, 3.0, 2.9, 2.8, 2.7, 2.6, 2.5, 2.4, 2.3, 2.2, 2.1, 2.0, 1.9, 1.8, 1.7, 1.6, Angstroms or less.

11. The method of any one of claims 1-10, the method further comprising forming a three- dimensional reference structure in plurality of three-dimensional reference structures by computationally docking the initial molecule element to an interacting partner element in the plurality of interacting partner elements.

12. The method of any one of claims 1-10, wherein the plurality of three-dimensional reference structures comprises at least about 2, 3, 4, 5, 6, 7, 8, 9, 10 structures, or more.

13. The method of any one of claims 1-10, wherein the plurality of three-dimensional reference structures comprises at least 20, 30, 40, 50, 100, 200, 400, 500, 1000, 2000, 4000, 5000, or 10000 structures.

14. The method of any one of claims 1-13, wherein the multi-targeting molecule comprises one or more of: a) agonist activity for at least one interacting partner in the plurality of interacting partner elements;b) antagonist activity for at least one interacting partner in the plurality of interacting partner elements; c) neutral binding activity for at least one interacting partner in the plurality of interacting partner elements; d) non-binding activity for at least one interacting partner in the plurality of interacting partner elements; e) any combination of a), b), c), or d) including partial agonist activity, partial antagonist activity, as well as condition-specific activity of any of the foregoing, such as pH dependent activity for an interacting partner element in the plurality of interacting partner elements.

15. The method of any one of claims 1-14, wherein the plurality of reference structures comprises pairs of subject sequence and interacting partner are polypeptide ligand and a receptor (including cell-surface receptors, such as GPCRs and integrins), respectively.

16. The method of any one claims 1-15, further comprising identifying an ensemble of sequence variants of the multi-targeting molecule that preserve the structure of the multi-targeting molecule to an RMSD resolution of less than about: 5, 4, 3, 2, 1, 0.5, or 0.3 Angstroms, or less.

17. The method of claim 16, wherein the identification of the ensemble of sequence variants is by AlphaFold, OmegaFold, ESMFold, molecular docking, molecular dynamics, RosettaFold, FoldX, RFDiffusion, ProteinMPNN, or dTERMen.

18. The method of any one of claims 1-7, the method further comprising testing, in vitro, in vivo, or ex vivo, the interactions of a multi-targeting molecule identified by the method.

19. The method of claim 18, wherein the multi -targeting molecule is a polypeptide, the plurality of interacting partner elements are polypeptides and the interaction comprises binding.

20. The method of claim 19, wherein a first interacting partner element in the plurality of interacting partner elements is a receptor.21 . The method of claim 20, wherein the receptor is a G protein-coupled receptor.

22. The method of claim 1, wherein the initial molecule element is a first small molecule, the first candidate molecule of the plurality of candidate molecules is a second small molecule, and the scoring function normalizes a score of the interaction between the first candidate molecule and a three-dimensional testing complex by a graph edit distance between the first candidate molecule and the initial molecule element.

23. The method of claim 1, wherein the initial molecule element is a first polypeptide, the first candidate molecule of the plurality of candidate molecules is a second polypeptide, and the scoring function normalizes a score of the interaction between the first candidate molecule and a three-dimensional testing complex by a mutation distance between the first candidate molecule and the initial molecule element.

24. The method of claim 1, wherein the initial molecule element is a first polypeptide, each candidate molecule of the plurality of candidate molecules is different polypeptide, and a candidate molecule in the plurality of candidate molecules is within the defined evolutionary chemical space when a sequence of the candidate molecule and a sequence of the initial molecule element share at least seventy percent sequence identity with the proviso that the sequence of the candidate molecule includes at least two or more sequence substitutions relative of the sequence of the initial molecule element.

25. The method of claim 1, wherein the initial molecule element is a first small molecule,each candidate molecule of the plurality of candidate molecules is a different small molecule, and a candidate molecule in the plurality of candidate molecules is within the defined evolutionary chemical space when a graph edit distance of the candidate molecule and the initial molecule element is between 2 and 5.

26. The method of claim 1, wherein the initial molecule element is a first small molecule, each candidate molecule of the plurality of candidate molecules is a different small molecule, and a candidate molecule in the plurality of candidate molecules is within the defined evolutionary chemical space when a graph edit distance of the candidate molecule and the initial molecule element is between 4 and 10.

27. The method of claim 1, wherein the initial molecule element is a first small molecule, each candidate molecule of the plurality of candidate molecules is a different small molecule, and a candidate molecule in the plurality of candidate molecules is within the defined evolutionary chemical space when a graph edit distance of the candidate molecule and the initial molecule element is between 6 and 14.

28. The method of any one of claims 1-27, the method further comprising synthesizing the multi-targeting molecule.

29. A system comprising a microprocessor and computer-readable instructions that, when executed by the microprocessor, causes the microprocessor to perform the method of any one of claims 1-27.

30. A multi-targeting molecule produced or producible by the method of any one of claims 1-28.

31. A non-transitory computer readable storage medium, wherein the non-transitory computer readable storage medium stores instructions, which when executed by a computer system, cause the computer system to perform a method for identifying a multi -targeting molecule, the method comprising any one of claims 1-27.