Methods and systems for characterizing structural features of a protein

The structural contribution model addresses the challenge of predicting protein stability and binding sites by performing atom-level analysis, enhancing the accuracy of drug design through precise characterization of protein features and mutation effects.

WO2025160446A1PCT designated stage Publication Date: 2025-07-31AI PROTEINS INC
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/013015
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-25
Filing Date
2025-01-24
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing methods for predicting protein structure and stability under physiological conditions are limited in accurately characterizing structural features and ensuring stability, as they often introduce artifacts due to non-physiological examination conditions, and there is a need for techniques that can predict stable conformations and binding sites for rational drug design.

Method used

A method involving a structural contribution model that performs an atom-level analysis of protein sequences, using a 3D graph to determine the structural relevance of each atom and its contribution to the overall stability, allowing for the identification of paratopes and protein-protein interfaces, and evaluating the effects of mutations.

Benefits of technology

Enables accurate characterization of protein stability and binding sites, facilitating rational drug design by identifying key structural features and assessing the impact of mutations, thus improving the precision of therapeutic development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025013015_31072025_PF_FP_ABST
    Figure US2025013015_31072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are methods and systems for characterizing one or more structural features of a molecule of interest (e.g., a protein), examples of which include determining the stability of a polypeptide, analyzing and identifying a paratope of a polypeptide, and / or examining a protein-protein interface. The methods and systems facilitate the identification and characterization of protein structures, and are useful in rational drug design.
Need to check novelty before this filing date? Find Prior Art

Description

DOCKET NO.: AIP-005WO PATENT METHODS AND SYSTEMS FOR CHARACTERIZING STRUCTURAL FEATURES OF A PROTEIN CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 624,960, filed January 25, 2024, the entire disclosure of which is incorporated by reference herein. REFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY

[0002] The instant application contains a Sequence Listing which has been submitted electronically in XML file format and is hereby incorporated by reference in its entirety. The XML copy, created on January 16, 2025, is named “AIP-005WO SL.XML” and is 775,337 bytes in size. BACKGROUND

[0003] Accurate characterization of protein sequences and the prediction of three-dimensional (3D) structures of physiologically relevant proteins can provide valuable insights into development and characterization of protein-based therapeutics. The prediction of the 3D structure of a protein that is likely to be stable under physiological conditions is extremely valuable in drug discovery. Furthermore, characterization of protein-protein interfaces can identify targets for rational drug design, whereby drugs can be designed to modulate the interface regions between proteins. Similarly, determining the binding region (paratope) of a protein that binds to a target molecule can also be helpful in designing more effective therapeutics. However, X-ray crystallographic or cryo-electron microscopic approaches may introduce structural artifacts into a protein of interest because the protein is examined under non-physiological conditions that can force the protein into adopting an unnatural conformation making the resolution of a protein under physiological conditions challenging.

[0004] As recent advances in machine learning have catalyzed a wave of new efforts to build protein-based tools and therapeutics, tools for predicting the performance of these proteins have taken on new importance. Previous efforts to predict protein structure have largely focused on determining the 3D structure into which a given sequence is likely to fold. However, relatively little attention has been paid to the orthogonal problem of ensuring that such a structure, once folded, will remain folded, and more specifically whether the structural features that facilitate its function will persist under physiological conditions.DOCKET NO.: AIP-005WO PATENT

[0005] Although significant developments have been made in predicting 3D structures of proteins of interest, there is still a need for new techniques that can accurately characterize protein sequences, such as predicting stable conformations and binding sites, to facilitate drug discovery. SUMMARY

[0006] Disclosed herein are methods and systems for characterizing one or more structural features of a molecule of interest (e.g., a polypeptide, a polynucleotide, or a small molecule), examples of which include determining the stability of the molecule (e.g., a polypeptide), analyzing and identifying a paratope of the molecule (e.g., a polypeptide, or a polynucleotide), and / or examining an interface of a molecule (e.g., a protein-protein interface, or a protein- polynucleotide interface). As described herein and as exemplified by a polypeptide, a polynucleotide, or a small molecule as a molecule of interest, the methods involve implementing a structural contribution model which performs an atom-level analysis of an element of a sequence, such as: an amino acid of a polypeptide sequence, a nucleotide of a polynucleotide sequence, or a sequence of a small molecule, to characterize one or more structural features of the sequence. The atom-level analysis models individual atoms as a fluid structure, subject to destabilizing perturbations. Modeling perturbations to the interconnected atomic structure enables determination of individual contributions of each atom, as well as individual contributions of each element of the sequence (e.g., each amino acid), to the overall structural integrity of the molecule. Thus, the methods and systems described herein facilitate the identification and characterization of molecular structures, and are useful in rational drug design. Given that the structural combination model is constructed from first principles of physical chemistry, the principles described herein can be used to identify similar features in other molecules of interest, for example, nucleic acids, lipids, ligands, molecular tags when a preliminary three-dimensional (3D) structure is available. The 3D structure can be generated via structural determination using a variety of different methods (e.g., via X-ray crystallography, cryo-electron microscopy, nuclear magnetic resonance (NMR) approaches), molecular modeling in silico, or a combination thereof.

[0007] In one aspect, the disclosure provides a method for characterizing one or more structural features of a molecule (e.g., a polypeptide, a polynucleotide, or a small molecule) of interest, wherein the method comprises, for each of one or more atoms of a sequence of the molecule of interest: (a) performing a pairwise atom to atom interaction analysis across at least one pair of atoms in the sequence to generate a 3D graph, wherein the 3D graphDOCKET NO.: AIP-005WO PATENT comprises nodes representing atoms and at least one edge between at least two nodes representing at least one interaction between the at least one pair of atoms, wherein the at least one edge comprises a weighted value based at least in part on presence or absence of an interaction between the at least one pair of atoms; (b) iteratively interrogating the 3D graph to determine a structural relevance of each of the one or more atoms of the sequence, wherein iteratively interrogating the 3D graph comprises: (i) for an atom of the 3D graph, assigning cooperativity scores representing a positional certainty to each of one or more other atoms of the sequence, wherein the cooperativity scores are assigned based at least in part on the weighted value of a corresponding edge in the 3D graph between the atom and the one or more other atoms, and (ii) combining the cooperativity scores of the atom across each of the one or more other atoms to determine a structural relevance of the atom; and for each element in the sequence of the molecule of interest, combining the structural relevance of each atom in the element to generate a measure of structural contribution for the element.

[0008] In various embodiments, the molecule of interest is a polypeptide or a polynucleotide. In various embodiments, the sequence is an amino acid sequence, or a nucleotide sequence. In various embodiments, the elements is an amino acid, or a nucleotide.

[0009] In another aspect, the disclosure provides a method for characterizing one or more structural features of a polypeptide of interest, wherein the method comprises, for each of one or more atoms of an amino acid sequence of the polypeptide of interest: (a) performing a pairwise atom to atom interaction analysis across at least one pair of atoms in the amino acid sequence to generate a 3D graph, wherein the 3D graph comprises nodes representing atoms and at least one edge between at least two nodes representing at least one interaction between the at least one pair of atoms, wherein the at least one edge comprises a weighted value based at least in part on presence or absence of an interaction between the at least one pair of atoms; (b) iteratively interrogating the 3D graph to determine a structural relevance of each of the one or more atoms of the amino acid sequence, wherein iteratively interrogating the 3D graph comprises: (i) for an atom of the 3D graph, assigning cooperativity scores representing a positional certainty to each of one or more other atoms of the amino acid sequence, wherein the cooperativity scores are assigned based at least in part on the weighted value of a corresponding edge in the 3D graph between the atom and the one or more other atoms, and (ii) combining the cooperativity scores of the atom across each of the one or more other atoms to determine a structural relevance of the atom; and for each amino acid in the amino acid sequence of the polypeptide, combining the structural relevance of each atom in the amino acid to generate a measure of structural contribution for the amino acid.DOCKET NO.: AIP-005WO PATENT

[0010] In some embodiments, the weighted value of the at least one edge is determined by performing a transform of one or more cooperativity scores corresponding to a presence of the at least one interaction between the at least one pair of atoms. In some embodiments, the at least one interaction is an interaction comprising an energy well. In some embodiments, the at least one interaction is a covalent bond, an electrostatic bond, a hydrogen bond, a disulfide bond, or a van der Waals force. In some embodiments, the transform is a logistic transform. In some embodiments, the logistic transform comprises a k parameter, an x0 parameter, and / or a covalent coefficient.

[0011] In some embodiments, the k parameter is equal to or greater than 0, equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 1, equal to or greater than 1.5, equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, equal to or greater than 4, equal to or greater than 4.5, equal to or greater than 5, equal to or greater than 5.5, equal to or greater than 6, equal to or greater than 6.5, equal to or greater than 7, equal to or greater than 7.5, equal to or greater than 8, equal to or greater than 8.5, equal to or greater than 9, equal to or greater than 9.5, equal to or greater than 10, equal to or greater than 10.5, equal to or greater than 11, equal to or greater than 11.5, or equal to or greater than 12.

[0012] In some embodiments, the x0 parameter is equal to or less than 0, equal to or less than -0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than -0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than - 0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than -0.55, equal to or less than -0.6, equal to or less than -0.65, equal to or less than -0.7, equal to or less than - 0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than -0.9, equal to or less than -0.95, equal to or less than -1, equal to or less than -1.5, equal to or less than -2, equal to or less than -2.5, or equal to or less than -3.

[0013] In some embodiments, the covalent coefficient is a covalent coefficient sigma and / or covalent coefficient pi. In some embodiments, the covalent coefficient sigma is equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to orDOCKET NO.: AIP-005WO PATENT greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, or equal to or greater than 0.99. In some embodiments, the covalent coefficient pi is equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, equal to or greater than 0.99, or equal to or greater than 1.

[0014] In some embodiments, the one or more cooperativity scores comprises a hydrogen bond score, a disulfide bond score, and / or a van der Waals force score. In some embodiments, the weighted value of the at least one edge is assigned a threshold value due to a presence of a covalent bond. In some embodiments, the weighted value of the at least one edge is between 0 and 1. In some embodiments, the weighted value of the at least one edge closer to 1 indicates higher cooperativity of two atoms connected via the at least one edge in comparison to a lower cooperativity of two atoms connected via at least one edge with the weighted value closer to 0.

[0015] In some embodiments, the method further comprises: prior to step (a), obtaining a 3D atomic structure comprising the one or more atoms of the sequence of the molecule (e.g., a polypeptide, a polynucleotide, or a small molecule). In some embodiments, assigning the cooperativity scores representing the positional certainty to each of the one or more other atoms of the sequence comprises: (a) performing additional pairwise atom to atom analyses between the atom and each of the one or more other atoms, comprising: (i) generating a plurality of pulse scores within a pulse array for the atom; (ii) for each other atom, selecting a higher of: (a) a pulse score of the pulse array corresponding to the one or more other atoms; (b) a marginal pulse score of a marginal pulse array corresponding to the one or more other atoms; and (iii) combining the selected scores in step (ii) across the one or more other atoms to generate the cooperativity scores. In some embodiments, the pulse score or the marginal score is a score based on the weighted value of a corresponding edge in the 3D graph between the atom and the one or more other atom. In some embodiments, combining the cooperativity scores across the one or more other atoms comprises summating the cooperativity scores of the one or more other atoms. In some embodiments, combining structural relevance of atoms of the element comprises summating the structural relevance of the atoms of the element to generate the measure of structural contribution for the element.DOCKET NO.: AIP-005WO PATENT

[0016] In another embodiment, the method further comprises: prior to step (a), obtaining a 3D atomic structure comprising the one or more atoms of the amino acid sequence of the polypeptide. In another embodiment, assigning the cooperativity scores representing the positional certainty to each of the one or more other atoms of the amino acid sequence comprises: (a) performing additional pairwise atom to atom analyses between the atom and each of the one or more other atoms, comprising: (i) generating a plurality of pulse scores within a pulse array for the atom; (ii) for each other atom, selecting a higher of: (a) a pulse score of the pulse array corresponding to the one or more other atoms; (b) a marginal pulse score of a marginal pulse array corresponding to the one or more other atoms; and (iii) combining the selected scores in step (ii) across the one or more other atoms to generate the cooperativity scores. In another embodiment, the pulse score or the marginal score is a score based on the weighted value of a corresponding edge in the 3D graph between the atom and the one or more other atom. In another embodiment, combining the cooperativity scores across the one or more other atoms comprises summating the cooperativity scores of the one or more other atoms. In another embodiment, combining structural relevance of atoms of the amino acid comprises summating the structural relevance of the atoms of the amino acid to generate the measure of structural contribution for the amino acid.

[0017] In some embodiments, the method further comprises characterizing the one or more structural features of the molecule (e.g., a polypeptide, a polynucleotide, or a small molecule) of interest using the measure of structural contribution for the element. In some embodiments, characterizing the one or more structural features of the molecule of interest using the measure of structural contribution comprises one or more of: (a) determining a stability of the molecule of interest; (b) examining an interface of the molecule of interest; (c) identifying a paratope of the molecule of interest; and (d) evaluating an effect of one or more point mutations on a structure of the molecule of interest.

[0018] In certain embodiments, the method further comprises characterizing the one or more structural features of the polypeptide of interest using the measure of structural contribution for the amino acid. In certain embodiments, characterizing the one or more structural features of the polypeptide of interest using the measure of structural contribution comprises one or more of: (a) determining a stability of the polypeptide of interest; (b) examining an interface of the polypeptide of interest; (c) identifying a paratope of the polypeptide of interest; and (d) evaluating an effect of one or more point mutations on a structure of the polypeptide of interest.DOCKET NO.: AIP-005WO PATENT

[0019] In some embodiments, a method for determining the stability of a molecule (e.g., a polypeptide, a polynucleotide, or a small molecule) of interest is provided, wherein the method comprises performing the method for characterizing one or more structural features of the molecule of interest. In certain embodiments, a method for determining the stability of a polypeptide of interest is provided, wherein the method comprises performing the method for characterizing one or more structural features of the polypeptide of interest. In some embodiments, the method comprises setting the k parameter to equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 1, equal to or greater than 1.5, equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, equal to or greater than 4, equal to or greater than 4.5, equal to or greater than 5, equal to or greater than 5.5, equal to or greater than 6, equal to or greater than 6.5, equal to or greater than 7, equal to or greater than 7.5, equal to or greater than 8, equal to or greater than 8.5, equal to or greater than 9, equal to or greater than 9.5, equal to or greater than 10, equal to or greater than 10.5, equal to or greater than 11, equal to or greater than 11.5, or equal to or greater than 12. In some embodiments, the method comprises setting the x0 parameter to equal to or less than - 0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than -0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than - 0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than -0.55, equal to or less than -0.6, equal to or less than -0.65, equal to or less than -0.7, equal to or less than - 0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than -0.9, equal to or less than -0.95, or equal to or less than -1. In some embodiments, the method comprises setting the covalent coefficient sigma to equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, or equal to or greater than 0.9. In some embodiments, the method comprises setting the covalent coefficient pi to equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, equal to or greater than 0.99, or equal to or greater than 1.DOCKET NO.: AIP-005WO PATENT

[0020] In some embodiments, determining the stability of the molecule (e.g., a polypeptide, a polynucleotide, or a small molecule) comprises determining at least one under-specified or disordered region of the molecule. In some embodiments, the molecule of interest is a monomeric structure.

[0021] In certain embodiments, determining the stability of the polypeptide comprises determining at least one under-specified or disordered region of the polypeptide. In certain embodiments, the polypeptide of interest is a monomeric structure.

[0022] In some embodiments, a method for examining an interface of the molecule (e.g., a polypeptide, a polynucleotide, or a small molecule) of interest is provided, wherein the method comprises performing the method for characterizing one or more structural features of the molecule of interest. In certain embodiments, a method for examining an interface of the polypeptide of interest is provided, wherein the method comprises performing the method for characterizing one or more structural features of a polypeptide of interest. In certain embodiments, the method requires selecting certain k and x0 parameters, a covalent coefficient sigma and a covalent coefficient pi. For example, in some embodiments, the method comprises setting the k parameter to equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, equal to or greater than 4, equal to or greater than 4.5, equal to or greater than 5, equal to or greater than 5.5, or equal to or greater than 6. Alternatively or in addition, in some embodiments, the method comprises setting the x0 parameter equal to or less than -0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than -0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than -0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than -0.55, equal to or less than -0.6, equal to or less than -0.65, equal to or less than -0.7, equal to or less than -0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than -0.9, equal to or less than -0.95, equal to or less than -1, equal to or less than -1.5, equal to or less than -2, equal to or less than -2.5, or equal to or less than - 3. Alternatively, or in addition to each of the foregoing k and x0 parameters, in some embodiments, the method comprises setting the covalent coefficient sigma to equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, or equal to or greater than 0.99. Alternatively, or in addition to each of the foregoing k and x0 parameters and covalent coefficient sigma, in some embodiments, the method comprises setting the covalentDOCKET NO.: AIP-005WO PATENT coefficient pi to equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, equal to or greater than 0.99, or equal to or greater than 1.

[0023] In some embodiments, examining the interface of the molecule (e.g., a polypeptide, a polynucleotide, or a small molecule) of interest comprises determining at least one interface region of the molecule of interest. In some embodiments, examining the interface of the molecule comprises determining structurally important interactions. In some embodiments, the interface of the molecule is an interface between at least two molecule regions. In some embodiments, the at least two molecule regions are on the same molecule or more than one molecule (e.g., two, or more polypeptides; two, or more polynucleotides; at least one polypeptide, and at least one polynucleotide; and the like).

[0024] In exemplary embodiments, examining the interface of the polypeptide of interest comprises determining at least one interface region of the polypeptide of interest. In exemplary embodiments, examining the interface of the polypeptide comprises determining structurally important interactions. In exemplary embodiments, the interface of a polypeptide is an interface between at least two polypeptide regions. In some exemplary, the at least two polypeptide regions are on the same polypeptide or more than one polypeptide (e.g., two, or more polypeptides).

[0025] In some embodiments, examining the interface of the molecule (e.g., a polypeptide, a polynucleotide, or a small molecule) comprises, prior to performing a pairwise atom to atom interaction analysis across at least one pair of atoms in the sequence to generate a 3D graph, the steps of: (a) selecting a focal chain; and (b) selecting a medial chain. In some embodiments, the focal chain and the medial chain comprise a region of the molecule, and wherein the region of the focal chain is the same as, or different than the region of the medial chain. In some embodiments, the region comprises a fragment, or full length sequence of the molecule. In some embodiments, the focal chain and the medial chain are on the same or different molecule. In some embodiments, examining the interface of the molecule comprises performing the method of the foregoing aspect across the medial chain of the molecule to generate the measure of structural contribution for the element (e.g., an amino acid, a nucleotide, or an element of a small molecule) of the focal chain of the molecule.

[0026] In exemplary embodiments, examining the interface of the polypeptide comprises, prior to performing a pairwise atom to atom interaction analysis across at least one pair of atoms in the amino acid sequence to generate a 3D graph, the steps of: (a) selecting a focalDOCKET NO.: AIP-005WO PATENT chain; and (b) selecting a medial chain. In exemplary embodiments, the focal chain and the medial chain comprise a polypeptide region of the polypeptide, and wherein the polypeptide region of the focal chain is the same as, or different than the polypeptide region of the medial chain. In exemplary embodiments, the polypeptide region comprises a fragment, or full length amino acid sequence of the polypeptide. In exemplary embodiments, the focal chain and the medial chain are on the same or different polypeptide. In exemplary embodiments, examining the interface of the polypeptide comprises performing the method of the foregoing aspect across the medial chain of the polypeptide to generate the measure of structural contribution for the amino acid of the focal chain of the polypeptide.

[0027] In certain embodiments, a method for identifying a paratope of the molecule (e.g., a polypeptide, a polynucleotide, or a small molecule) of interest is provided, wherein the method comprises performing the method for characterizing one or more structural features of the molecule of interest. In exemplary embodiments, a method for identifying a paratope of the polypeptide of interest is provided, wherein the method comprises performing the method for characterizing one or more structural features of a polypeptide of interest. In certain embodiments, the method requires selecting certain k and x0 parameters, a covalent coefficient sigma and a covalent coefficient pi. For example, in some embodiments, the method comprises setting the k parameter to equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 1, equal to or greater than 1.5, equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, or equal to or greater than 4. Alternatively or in addition, in some embodiments, the method comprises setting the x0 parameter to equal to or less than -0.05, equal to or less than -0.1, equal to or less than - 0.15, equal to or less than -0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than -0.4, equal to or less than -0.45, equal to or less than - 0.5, equal to or less than -0.55, equal to or less than -0.6, equal to or less than -0.65, equal to or less than -0.7, equal to or less than -0.75, equal to or less than -0.8, equal to or less than - 0.85, equal to or less than -0.9, equal to or less than -0.95, equal to or less than -1, equal to or less than -1.5, or equal to or less than -2. Alternatively or in addition to selecting a k and / or a x0 parameter, in some embodiments, the method comprises setting the covalent coefficient sigma to equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, or equal to or greater than 0.5. Alternatively or in addition to selecting a k parameter, an x0 parameter, and / or a covalent coefficient sigma, in some embodiments, the method comprises setting the covalent coefficient pi to 1.DOCKET NO.: AIP-005WO PATENT

[0028] In some embodiments, a method for evaluating an effect of one or more point mutations on a structure of a molecule (e.g., a polypeptide, a polynucleotide, or a small molecule) of interest is provided, wherein the method comprises performing the method for characterizing one or more structural features of the molecule of interest. In some embodiments, evaluating an effect of one or more point mutations on a structure of the molecule of interest comprises characterizing a tolerability of the structure of the molecule of interest to the one or more point mutations.

[0029] In exemplary embodiments, a method for evaluating an effect of one or more point mutations on a structure of the polypeptide of interest is provided, wherein the method comprises performing the method for characterizing one or more structural features of a polypeptide of interest. In exemplary embodiments, evaluating an effect of one or more point mutations on a structure of the polypeptide of interest comprises characterizing a tolerability of the structure of the polypeptide of interest to the one or more point mutations.

[0030] In another aspect, the disclosure provides a method for identifying a paratope of a molecule (e.g., a polypeptide, a polynucleotide, or a small molecule), wherein the method comprises: (a) obtaining or having obtained a plurality of binding affinity values between a target antigen and a plurality of molecule variant sequences based on the molecule, wherein the molecule variant sequences of the molecule differ from one another by point mutations at predetermined positions; (b) characterizing one or more structural features of the molecule by performing a structural analysis of individual atoms of a sequence of the molecule; (c) performing element-level (e.g., an amino acid, a nucleotide, or an element of a small molecule) combinations of binding affinity values and measures of structural contribution to generate a measure of non-structural binding relevance for each of the one or more element of each of the one or more molecule variant sequences; and (d) identifying the paratope comprising one or more element in the molecule associated with measures of non-structural binding relevance that indicate contributions to binding to the target antigen.

[0031] In exemplary embodiments, the disclosure provides a method for identifying a paratope of a polypeptide, wherein the method comprises: (a) obtaining or having obtained a plurality of binding affinity values between a target antigen and a plurality of polypeptide variant sequences based on the polypeptide, wherein the polypeptide variant sequences of the polypeptide differ from one another by point mutations at predetermined positions; (b) characterizing one or more structural features of the polypeptide by performing a structural analysis of individual atoms of an amino acid sequence of the polypeptide; (c) performingDOCKET NO.: AIP-005WO PATENT amino acid residue-level combinations of binding affinity values and measures of structural contribution to generate a measure of non-structural binding relevance for each of the one or more amino acids of each of the one or more polypeptide variant sequences; and (d) identifying the paratope comprising one or more amino acids in the polypeptide associated with measures of non-structural binding relevance that indicate contributions to binding to the target antigen.

[0032] In some embodiments, the foregoing paratope identification method further comprises: (e) validating the paratope, wherein the validation comprises: (i) generating a modified molecule variant sequence by inserting or removing one or more elements associated with measures of non-structural binding relevance that indicate a lack of contribution to binding to the target antigen. In exemplary embodiments, the foregoing paratope identification method further comprises: (e) validating the paratope, wherein the validation comprises: (i) generating a modified polypeptide variant sequence by inserting or removing one or more amino acids associated with measures of non-structural binding relevance that indicate a lack of contribution to binding to the target antigen.

[0033] In some embodiments, the validation further comprises: (ii) determining a change in measures of non-structural binding relevance for one or more remaining elements in the modified molecule variant sequence. In some embodiments, (ii) determining the change in measures of non-structural binding relevance comprises re-performing the structural analysis of individual atoms of elements of the modified molecule variant sequence. In exemplary embodiments, the validation further comprises: (ii) determining a change in measures of non- structural binding relevance for one or more remaining amino acids in the modified polypeptide variant sequence. In exemplary embodiments, (ii) determining the change in measures of non-structural binding relevance comprises re-performing the structural analysis of individual atoms of amino acids of the modified polypeptide variant sequence.

[0034] In some embodiments, the validation further comprises: (iii) expressing the modified molecule variant sequence; (iv) measuring binding affinity of the modified molecule variant sequence; and (v) comparing the measured binding affinity of the modified molecule variant sequence to a binding affinity of the molecule prior to insertion or removal of the one or more elements associated with measures of non-structural binding relevance indicative of a lack of contribution to binding. In some embodiments, (iv) measuring binding affinity of the modified molecule variant sequence to the target is determined by surface plasmon resonance. In exemplary embodiments, the validation further comprises: (iii) expressing the modifiedDOCKET NO.: AIP-005WO PATENT polypeptide variant sequence; (iv) measuring binding affinity of the modified polypeptide variant sequence; and (v) comparing the measured binding affinity of the modified polypeptide variant sequence to a binding affinity of the polypeptide prior to insertion or removal of the one or more amino acids associated with measures of non-structural binding relevance indicative of a lack of contribution to binding. In exemplary embodiments, (iv) measuring binding affinity of the modified polypeptide variant sequence to the target is determined by surface plasmon resonance.

[0035] In some embodiments, step (a) obtaining or having obtained a plurality of binding affinity values between a target antigen and a plurality of molecule variant sequences comprises: (i) generating a sequence library of the plurality of molecule variant sequences in which molecule variant sequences of the plurality differ from one another by point mutations at predetermined positions; (ii) expressing the plurality of molecule variant sequences; and (iii) measuring the plurality of binding affinity values between the target antigen and the expressed plurality of molecule variant sequences. In exemplary embodiments, step (a) obtaining or having obtained a plurality of binding affinity values between a target antigen and a plurality of polypeptide variant sequences comprises: (i) generating a sequence library of the plurality of polypeptide variant sequences in which polypeptide variant sequences of the plurality differ from one another by point mutations at predetermined positions; (ii) expressing the plurality of polypeptide variant sequences; and (iii) measuring the plurality of binding affinity values between the target antigen and the expressed plurality of polypeptide variant sequences.

[0036] In some embodiments, step (a) obtaining or having obtained a plurality of binding affinity values between a target antigen and a plurality of molecule variant sequences further comprises (iv) sequencing the plurality of molecule variant sequences to determine corresponding sequences of the plurality of molecule variant sequences. In exemplary embodiments, step (a) obtaining or having obtained a plurality of binding affinity values between a target antigen and a plurality of polypeptide variant sequences further comprises (iv) sequencing the plurality of polypeptide variant sequences to determine corresponding sequences of the plurality of polypeptide variant sequences.

[0037] In some embodiments, step (i) generating the sequence library of the plurality of molecule variant sequences comprises: (i) obtaining a reference molecule sequence; and (ii) iteratively substituting one element of the reference molecule sequence with a different element to generate the plurality of molecule variant sequences. In some embodiments, theDOCKET NO.: AIP-005WO PATENT plurality of molecule variant sequences comprises at least 50% of all possible molecule variant sequences that differ from the reference molecule sequence by a single element. In some embodiments, the plurality of molecule variant sequences comprises at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, a least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of all possible molecule variant sequences that differ from the reference molecule sequence by a single element. In exemplary embodiments, step (i) generating the sequence library of the plurality of polypeptide variant sequences comprises: (i) obtaining a reference polypeptide sequence; and (ii) iteratively substituting one amino acid of the reference polypeptide sequence with a different amino acid to generate the plurality of polypeptide variant sequences. In exemplary embodiments, the plurality of polypeptide variant sequences comprises at least 50% of all possible polypeptide variant sequences that differ from the reference polypeptide sequence by a single amino acid. In exemplary embodiments, the plurality of polypeptide variant sequences comprises at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, a least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of all possible polypeptide variant sequences that differ from the reference polypeptide sequence by a single amino acid.

[0038] In some embodiments, step (c) performing element-level combinations of binding affinity values and measures of structural relevance comprises: for each of the one or more elements of the molecule variant sequence, determining a difference between the binding affinity value of the molecule variant sequence comprising the element and the measure of structural relevance of the element. In some embodiments, the measure of non-structural binding relevance for an elements represents a binding contribution of the element to the binding affinity without a corresponding structural contribution of the element. In exemplary embodiments, step (c) performing amino acid residue-level combinations of binding affinity values and measures of structural relevance comprises: for each of the one or more amino acids of the polypeptide variant sequence, determining a difference between the binding affinity value of the polypeptide variant sequence comprising the amino acid and the measure of structural relevance of the amino acid. In exemplary embodiments, the measure of non- structural binding relevance for an amino acids represents a binding contribution of the amino acid to the binding affinity without a corresponding structural contribution of the amino acid.

[0039] In some embodiments, step (d) identifying the paratope comprising one or more elements comprises selecting one or more elements with measures of non-structural bindingDOCKET NO.: AIP-005WO PATENT relevance that are above a threshold score. In some embodiments, each of the molecule variant sequences in the plurality of molecule variant sequences have the same length. In some embodiments, the molecule variant sequences in the plurality of molecule variant sequences are between 30 and 60 elements in length. In exemplary embodiments, step (d) identifying the paratope comprising one or more amino acids comprises selecting one or more amino acids with measures of non-structural binding relevance that are above a threshold score. In exemplary embodiments, each of the polypeptide variant sequences in the plurality of polypeptide variant sequences have the same length. In exemplary embodiments, the polypeptide variant sequences in the plurality of polypeptide variant sequences are between 30 and 60 amino acids in length.

[0040] In some embodiments, characterizing the element sequence of the molecule by performing the structural analysis of individual atoms of the element sequence comprises, for each of one or more atoms of each element of the molecule variant sequence: (a) performing a pairwise atom to atom interaction analysis across at least one pair of atoms in the element sequence to generate a 3D graph, wherein the 3D graph comprises nodes representing atoms and at least one edge between at least two nodes representing at least one interaction between the at least one pair of atoms, wherein the at least one edge comprises a weighted value based at least in part on presence or absence of an interaction between the at least one pair of atoms; (b) iteratively interrogating the 3D graph to determine a structural relevance of each of the one or more atoms of the element sequence, wherein iteratively interrogating the 3D graph comprises: (i) for an atom of the 3D graph, assigning cooperativity scores representing a positional certainty to each of one or more other atoms of the element sequence, wherein the cooperativity scores are assigned based at least in part on the weighted value of a corresponding edge in the 3D graph between the atom and the one or more other atom, and (ii) combining the cooperativity scores of the atom across the one or more other atoms to determine a structural relevance of the atom; and for each element in the one or more elements of the molecule, combining the structural relevance of each atom in the element to generate a measure of structural contribution for the element.

[0041] In exemplary embodiments, characterizing the amino acid sequence of the polypeptide by performing the structural analysis of individual atoms of the amino acid sequence comprises, for each of one or more atoms of each amino acid of the polypeptide variant sequence: (a) performing a pairwise atom to atom interaction analysis across at least one pair of atoms in the amino acid sequence to generate a 3D graph, wherein the 3D graph comprisesDOCKET NO.: AIP-005WO PATENT nodes representing atoms and at least one edge between at least two nodes representing at least one interaction between the at least one pair of atoms, wherein the at least one edge comprises a weighted value based at least in part on presence or absence of an interaction between the at least one pair of atoms; (b) iteratively interrogating the 3D graph to determine a structural relevance of each of the one or more atoms of the amino acid sequence, wherein iteratively interrogating the 3D graph comprises: (i) for an atom of the 3D graph, assigning cooperativity scores representing a positional certainty to each of one or more other atoms of the amino acid sequence, wherein the cooperativity scores are assigned based at least in part on the weighted value of a corresponding edge in the 3D graph between the atom and the one or more other atom, and (ii) combining the cooperativity scores of the atom across the one or more other atoms to determine a structural relevance of the atom; and for each amino acid in the one or more amino acids of the polypeptide, combining the structural relevance of each atom in the amino acid to generate a measure of structural contribution for the amino acid.

[0042] In some embodiments, the weighted value of the at least one edge is determined by performing a transform of one or more cooperativity scores corresponding to a presence of the at least one interaction between the at least one pair of atoms. In some embodiments, the at least one interaction is an interaction comprising an energy well. In some embodiments, the at least one interaction is a covalent bond, an electrostatic bond, a hydrogen bond, a disulfide bond, or a van der Waals force. In some embodiments, the transform is a logistic transform. In some embodiments, the logistic transform comprises a parameter k, a parameter x0, and / or a covalent coefficient. In some embodiments, the k parameter is equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 1, equal to or greater than 1.5, equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, or equal to or greater than 4. In some embodiments, the x0 parameter is equal to or less than - 0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than -0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than - 0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than -0.55, equal to or less than -0.6, equal to or less than -0.65, equal to or less than -0.7, equal to or less than - 0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than -0.9, equal to or less than -0.95, equal to or less than -1, equal to or less than -1.5, or equal to or less than -2. In some embodiments, the covalent coefficient is a covalent coefficient sigma and / or covalent coefficient pi. In some embodiments, the covalent coefficient sigma is equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, orDOCKET NO.: AIP-005WO PATENT equal to or greater than 0.5. In some embodiments, the covalent coefficient pi is equal to 1. In some embodiments, the one or more cooperativity scores comprises a hydrogen bond score, a disulfide bond score, and / or a van der Waals force score. In some embodiments, the weighted value of the at least one edge is assigned a threshold value due to a presence of a covalent bond. In some embodiments, the weighted value of the at least one edge is between 0 and 1. In some embodiments, the weighted value of the at least one edge closer to 1 indicates higher cooperativity of two atoms connected via the at least one edge in comparison to a lower cooperativity of two atoms connected via at least one edge with the weighted value closer to 0.

[0043] In some embodiments, the method comprises: prior to step (a), obtaining a 3D atomic structure comprising the one or more atoms of the element sequence of the molecule (e.g., a polypeptide, a polynucleotide, or a small molecule) of interest. In exemplary embodiments, the method comprises: prior to step (a), obtaining a 3D atomic structure comprising the one or more atoms of the amino acid sequence of the polypeptide of interest.

[0044] In some embodiments, assigning the cooperativity score representing positional certainty to each of one or more other atoms of the element comprises: (a) performing additional pairwise atom to atom analyses between the atom and each of the one or more other atoms, comprising: (i) generating a plurality of pulse scores within a pulse array for the atom; (ii) for each other atom, selecting a higher of: (a) a pulse score of the pulse array, corresponding to the one or more other atoms; (b) a marginal pulse score of a marginal pulse array, corresponding to the one or more other atoms; and (iii) combining the selected scores in step (ii) across the one or more other atoms to generate the cooperativity scores. In some embodiments, the pulse score or the marginal score is a score based on the weighted value of a corresponding edge in the 3D graph between the atom and the one or more other atom. In some embodiments, combining the cooperativity scores across the one or more other atoms comprises summating the cooperativity scores of the one or more other atoms. In some embodiments, combining structural relevance of atoms of the element comprises summating the structural relevance of the atoms of the element to generate the measure of structural contribution for the element.

[0045] In exemplary embodiments, assigning the cooperativity score representing positional certainty to each of one or more other atoms of the amino acid comprises: (a) performing additional pairwise atom to atom analyses between the atom and each of the one or more other atoms, comprising: (i) generating a plurality of pulse scores within a pulse array for the atom; (ii) for each other atom, selecting a higher of: (a) a pulse score of the pulse array,DOCKET NO.: AIP-005WO PATENT corresponding to the one or more other atoms; (b) a marginal pulse score of a marginal pulse array, corresponding to the one or more other atoms; and (iii) combining the selected scores in step (ii) across the one or more other atoms to generate the cooperativity scores. In exemplary embodiments, combining structural relevance of atoms of the amino acid comprises summating the structural relevance of the atoms of the amino acid to generate the measure of structural contribution for the amino acid.

[0046] In another aspect, the disclosure provides a non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to perform any one of the methods described herein.

[0047] Further details of the nature and advantages of the present disclosure can be found in the following detailed description taken in conjunction with the accompanying figures. The present disclosure is capable of modification in various respects without departing from the spirit and scope of the present disclosure. Accordingly, the figures and description of these embodiments are not restrictive. BRIEF DESCRIPTION OF THE FIGURES

[0048] A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized in combination with the accompanying drawings.

[0049] Wherever practicable, similar or like reference numbers may be used in the figures and may indicate similar or like functionality. For example, a letter after a reference numeral, such as “third party 110A,” indicates that the text refers specifically to the element having that particular reference numeral. A reference numeral in the text without a following letter, such as “third party entity 110,” refers to any or all of the elements in the figures bearing that reference numeral (e.g. “third party entity 110” in the text refers to reference numerals “third party entity 110A” and / or “third party entity 110B” in the figures).

[0050] FIG.1A is an example system overview for characterizing a protein sequence, in accordance with an embodiment. FIG.1B depicts an example block diagram for characterizing a protein sequence, in accordance with an embodiment.

[0051] FIGs.2A-2D depict a process for building a 3D graph, in accordance with an embodiment.DOCKET NO.: AIP-005WO PATENT

[0052] FIGs.3A-3E depicts a process for interrogating a protein structure to create an interaction graph based upon the interaction of individual atoms in the protein structure. FIGs.3A-3B depict a process of using a 3D graph for a first pulse cycle, in accordance with the embodiment shown in FIGs.2A-2D. FIGs.3C-3E depict a process of using a 3D graph for a subsequent pulse cycle, in accordance with the embodiments shown in FIGs.2A-2D and FIGs.3A-3B.

[0053] FIGs.4A-4D depict a process for determining the structural contributions of atoms in an exemplary protein of interest. FIG.4A is a schematic of an exemplary interaction graph, in accordance with an embodiment. FIG.4B-4C are exemplary pulse arrays and marginal pulse arrays, respectively, for exemplary atoms, in accordance with an embodiment. FIG.4D provides exemplary scores representing a measure of structural contribution for exemplary atoms, in accordance with the embodiments shown in FIGs.4A-4C.

[0054] FIGs 5A-5C depict process steps for identifying and characterizing a paratope region in an exemplary protein of interest. FIG.5A provides a flowchart with steps to identify a paratope of a protein of interest, in accordance with an embodiment. FIG.5B depicts a flowchart for characterizing a protein sequence, in accordance with an embodiment. FIG.5C depicts a flowchart for validating a paratope of a protein of interest, in accordance with an embodiment.

[0055] FIG.6 illustrates an exemplary computing device 600 for implementing system and methods described in FIGs.1A-1B, 2A-2D, 3A-3D, 4A-4D, and 5A-5C.

[0056] FIG.7 shows improved performance of the disclosed structural contribution model as a structural classifier versus that of RoseTTAFold’s Rosetta Energy Units per residue (REU per res), buried nonpolar surface area per residue (Buried NPSA per res), or mean Alphafold2 pLDDT score (mean PLDDT).

[0057] FIGs.8A and 8B depict a use case of the structural contribution model for evaluating effects of point mutations on a given protein structure.

[0058] FIGs.9A-9C depict a use case of the structural contribution model for examining protein-protein interfaces for structurally important interactions.

[0059] FIGs.10A-10C depict a use case of the structural contribution model for performing a paratope salience analysis.

[0060] FIGs.11A-11D show polypeptide residue characterization (including paratope residues) using the structural contribution model.DOCKET NO.: AIP-005WO PATENT

[0061] FIGs.12A-12C depict a use case of the structural contribution model for performing a protein-DNA interface analysis. DETAILED DESCRIPTION I. Definitions

[0062] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. The mention of techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques or substitutions of equivalent techniques that would be apparent to one of skill in the art. While the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the presently disclosed subject matter.

[0063] As used herein, "about" will be understood by persons of ordinary skill and will vary to some extent depending on the context in which it is used. If there are uses of the term which are not clear to persons of ordinary skill given the context in which it is used, "about" will mean up to plus or minus 10% of the particular value.

[0064] The articles “a” and “an” are used in this disclosure to refer to one or more than one (i.e., to at least one) of the grammatical object of the article, unless the context is inappropriate. By way of example, “an element” means one element or more than one element.

[0065] The term “and / or” is used in this disclosure to mean either “and” or “or” unless indicated otherwise.

[0066] It should be understood that the expression “at least one of” includes individually each of the recited objects after the expression and the various combinations of two or more of the recited objects unless otherwise understood from the context and use. The expression “and / or” in connection with three or more recited objects should be understood to have the same meaning unless otherwise understood from the context.

[0067] The use of the term “include,” “includes,” “including,” “have,” “has,” “having,” “contain,” “contains,” or “containing,” including grammatical equivalents thereof, should be understood generally as open-ended and non-limiting, for example, not excluding additional unrecited elements or steps, unless otherwise specifically stated or understood from the context.DOCKET NO.: AIP-005WO PATENT

[0068] It should be understood that the order of steps or order for performing certain actions is immaterial so long as the present invention remain operable. Moreover, two or more steps or actions may be conducted simultaneously.

[0069] At various places in the present specification, variable or parameters are disclosed in groups or in ranges. It is specifically intended that the description include each and every individual subcombination of the members of such groups and ranges. For example, an integer in the range of 0 to 5 is specifically intended to individually disclose 0, 1, 2, 3, 4, 5, and an integer in the range of 1 to 3 is specifically intended to individually disclose 1, 2, and 3.

[0070] The use of any and all examples, or exemplary language herein, for example, “such as” or “including,” is intended merely to illustrate better the present disclosure and does not pose a limitation on the scope of any invention(s) unless claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of that provided by the present disclosure.

[0071] As used herein, the terms “protein” and “polypeptide” are used interchangeably and generally refer to a macromolecule that includes one or more linked chains of amino acid residues, which can be natural amino acids, unnatural amino acids or both. In various embodiments, the terms “protein” or “polypeptide” further encompass a “miniprotein” which represents a protein or a portion thereof. In various embodiments, a miniprotein may include between 15 and 75 amino acids. In various embodiments, a miniprotein may include between 20 and 70 amino acids, between 25 and 65 amino acids, between 30 and 60 amino acids, between 35 and 55 amino acids, or between 40 and 50 amino acids.

[0072] As used herein, the term “epitope” refers to a region of a protein that is specifically recognized by a binding partner, such as an antibody or another binding protein. The epitope may generally span a portion of the protein. Often, proteins may have multiple such regions where binding partners can attach. Epitopes typically fall into two classes: continuous epitopes (also known as linear epitopes), which are epitopes defined by linear sequences of consecutive amino acids, and discontinuous epitopes (also known as conformational epitopes), which are epitopes defined by discontinuous amino acids that are brought together into spatial proximity when a protein is in its folded state.

[0073] As used herein, the term "paratope" refers to the specific region of a binding molecule (e.g., binding protein) that recognizes and binds an epitope of a target molecule. The paratope typically comprises 5-20 amino acids.DOCKET NO.: AIP-005WO PATENT

[0074] As used herein, the phrase “structural relevance” is generally used in the context of atoms, e.g., atoms of an amino acid residue present in a polypeptide. The structural relevance of an atom refers to positional certainty of the atom, which contributes to the overall structure (e.g., structural integrity of the corresponding amino acid and / or the polypeptide). For example, the structural relevance of an atom can be defined according to the rototranslational certainty of the atom relative to other atoms. A highly rototranslationally affixed atom has a higher structural relevance in comparison to a less rototranslationally affixed atom that has a lower structural relevance.

[0075] As used herein, the term "measure of structural contribution" is used in the context of individual amino acid residues of a polypeptide. The measure of structural contribution of an amino acid refers to a quantitative or qualitative assessment of how the amino acid influences or contributes towards the structural integrity of the protein. This measure of structural contribution can help determine which regions or motifs within the protein structure are valuable for its interaction with a binding partner and may provide insights into designing better binding partners for therapeutic or diagnostic purposes. In various embodiments, the determination of a measure of structural contribution of an amino acid involves an iterative pairwise atom analysis of the structural relevance of the atoms of the amino acid. For example, a first amino acid with highly structurally relevant atoms will have a higher measure of structural contribution than a second amino acid with less structurally relevant atoms. In this example, the first amino acid is more likely to contribute towards the structural integrity of the protein in comparison to the second amino acid.

[0076] As used herein, the term "non-structural binding relevance" is used in the context of individual amino acid residues of a polypeptide. Generally, the non-structural binding relevance of an amino acid refers to the amino acids contribution to the binding activity of the polypeptide. An amino acid’s non-structural binding relevance does not encompass the amino acids contribution to the structural integrity of the polypeptide (e.g., the amino acids contribution to the polypeptides primary, secondary, or tertiary structures). Instead, examples of non-structural binding relevance for an amino acid may involve, without limitation, factors such as post-translational modifications, temporal or spatial availability, or other non- structural factors that influence binding interactions. In various embodiments, the non- structural binding relevance of an amino acid can be determined as a difference between the amino acid’s measure of structural contribution and a binding affinity value (e.g., binding affinity between the amino acids of a protein of interest and its binding partner).DOCKET NO.: AIP-005WO PATENT

[0077] As used herein, the term “cooperativity” is used to refer to a phenomenon where the binding of a ligand to one region of a binding surface influences the binding properties of adjacent or nearby regions on that same surface. Without wishing to be bound by a particular theory, this can be due to conformational changes, altered electrostatics, or other effects localized to the binding interface. For example, when one ligand binds to a specific region of the binding surface, it may induce a subtle change in the shape or charge distribution of nearby binding regions. This can either enhance (positive cooperativity) or hinder (negative cooperativity) the binding of other ligands to those adjacent regions. Thus, in this context, cooperativity is meant to refer to how a localized binding event can influence other binding events specifically at the binding surface level, without necessarily invoking larger-scale structural changes in the entire molecule or complex.

[0078] As used herein, the term “rototranslation” is used to refer to a rigid translation between coordinate frames. Without wishing to be bound by theory, in structural biology, because atoms occupy three-dimensional space and orientation matters for bond geometry, the term rototranslation is meant to refer to the six-dimensional vectors incorporating both the three translational components (x,y,z) and the three Euler angles ( , , ) required to define theposition and orientation of two objects relative to each other. For example, rototranslation can also refer to the rigidity between two atoms of a protein.

[0079] As used herein, the term “inpainting” is meant to refer to the generation of a complete object from an incomplete set of parts by generating compatible replacements for the missing components without altering the components already present. In protein design, inpainting is the generation of a more complete protein from a provided set of structural fragments such that the original fragments are not moved relative to each other but are instead embedded in a structurally complete protein.

[0080] Throughout the description, where systems and compositions are described as having, including, or comprising specific components, or where processes and methods are described as having, including, or comprising specific steps, it is contemplated that, additionally, there are systems and compositions and kits of the present disclosure that consist essentially of, or consist of, the recited components, and that there are processes and methods according to the present disclosure that consist essentially of, or consist of, the recited processing steps.

[0081] In the disclosure, where an element or component is said to be included in and / or selected from a list of recited elements or components, it should be understood that the element or component can be any one of the recited elements or components, or the element orDOCKET NO.: AIP-005WO PATENT component can be selected from a group consisting of two or more of the recited elements or components.

[0082] Further, it should be understood that elements and / or features of a system or a method provided and described herein can be combined in a variety of ways without departing from the spirit and scope of the present disclosure and invention(s) herein, whether explicit or implicit herein. For example, where reference is made to a particular system, that system can be used in various embodiments of systems of the present disclosure and / or in methods of the present disclosure, unless otherwise understood from the context. In other words, within this application, embodiments have been described and depicted in a way that enables a clear and concise application to be written and drawn, but it is intended and will be appreciated that embodiments may be variously combined or separated without parting from the present teachings and invention(s). For example, it will be appreciated that all features described and depicted herein can be applicable to all aspects of invention(s) provided, described, and depicted herein. II. Introduction

[0083] Disclosed herein is a novel approach to assessing protein structural cooperativity at the atomic level for purposes of determining, for example, (i) the stability of a protein, (ii) the contribution of at least one atom to stability of a protein, and (iii) the contribution of an amino acid, or a group of amino acids, to stability of a protein, identifying a paratope of a protein, examining a protein-protein interface, or evaluating the effects of mutations, such as, but not limited to, point mutations, substitutions, insertions, or deletions, on structure of the protein.

[0084] By modeling a protein as a series of independently mobile atoms and measuring the relative positional certainty between them, it is possible to generate a structural profile that can identify and localize specific structural deficiencies and quantify functionally relevant stability of both monomers and complexes. The approach yields a generally applicable metric of protein structural cooperativity that customizable based on functional relevance, enabling an automated assessment of protein stability and function.

[0085] When designing structures at the molecular level, it is valuable to understand the fundamental positional uncertainty under which the designed structures operate. Without wishing to be bound by theory, structural disorder is widespread in proteins and thought to be a determinant of protein function. (Uversky, et al. (2019) PROGRESS IN MOLECULAR BIOLOGY AND TRANSLATIONAL SCIENCE, 166:1–17; Dass, et al. (2020) SCIENTIFIC REPORTS, 10 (1): 14780). Even if quantum-scale effects are disregarded, proteins exist in solution underDOCKET NO.: AIP-005WO PATENT conditions that effectively obviate inertia due to the low Reynolds number (Needleman, et al. (2019) PHYSICSTODAY, 72 (9): 32–38), as Brownian motion and interactions with solvent effectively randomize the local environment (Mo, et al. (2019) ANNUALREVIEW OFFLUIDMECHANICS, 51 (1): 403–28). While this is often considered in terms of the position of entire molecules, the same phenomena also occur on intramolecular scales, as random molecular motion displaces part of a molecule relative to the rest of it, and as individual moieties displace themselves relative to their past locations. (Levchuk (1993) UKRAINS’KYI BIOKHIMICHNYI ZHURNAL, 65 (2): 3–16).

[0086] These phenomena make the design of macromolecular structures quite unlike the design of macroscopic structures. On the bulk material level, as when designing a building, it is sufficient to consider the balance of forces impacting each part of the structure. This carries over in molecular terms to assessing the energetic environment of each atom, effectively measuring the balance of forces imposed on it by the rest of the structure. (Xiong, et al. (2014) NATURE COMMUNICATIONS, 5: 5330; Alford, et al. (2017) J. CHEMICAL THEORY AND COMPUTATION, 13(6): 3031–48; Brooks, et al. (2009) J. COMPUTATIONAL CHEMISTRY, 30(10): 1545–1614; Opuu, et al. (2020) SCIENTIFIC REPORTS, 10(1): 11150). However, this analysis necessarily presupposes that the relative positions of the structural members are going to remain consistent over time, which is not physical. This problem compounds when considering the application of energy-based scoring systems in determining molecular stability, archetypally by simply weighting and summing all interactions to produce a global G of folding: a given moiety may be “stabilized” by interactions that are themselves not persistent across time, producing a chimerical sense of stability.

[0087] If energetic totality alone is insufficient to establish stability, then additional parameters need to be considered. One of those is self-reinforcement of stabilizing interactions, effectively a resistance to unraveling as opposed to unfolding, in that it is effectively assumed that all destabilizing perturbations will occur continuously, and protein unfolding is thus a progressive breakdown of stabilizing interaction networks. Under this paradigm, a stable structure is one for which any displacement of a structural element creates a locally energetically suboptimal conformation that will in turn bias the molecule toward returning to the original conformation as the local minimum energy state. The salient parameter, then, is not the total energy of the conformation but rather the degree to which the relative position in six-dimensional space (x,y,z, , , ) of one set of atoms specifies the relative position of another set of atoms.DOCKET NO.: AIP-005WO PATENT

[0088] At the local level, that of individual atomic interactions, this reduces to a transform of the energy of interactions. Atoms which do not interact do not specify relative positions at all. (Tenny, et al. (2022) IN: STATPEARLS[INTERNET], StatPearls Publishing) Atoms which interact strongly, by contrast, may have a single set of rototranslational parameters specifying the minimum energy state between them, and thus have strong relative rototranslational certainty: given knowledge of the position of one atom, the position of the other may be confidently predicted. At the level of an entire structure, this property is also representative of structural cooperativity; rototranslations are informative between rigid bodies, but it is still possible to quantitate how strongly motions in one flexible moiety are coupled to another.

[0089] Disclosed herein are methods, systems, and non-transitory computer readable media for characterizing one or more structural features of a protein sequence. The approaches described herein involve characterizing the structural features of a protein sequence at an atomic level that involves modeling the structure of atoms of the protein sequence as a 3D graph. The 3D graph structure enables modeling of pulses (e.g., local perturbations) that can propagate throughout the network of connections within the 3D graph structure. Altogether, this atomic level analysis enables more accurate identification of relevant residues, e.g., paratope residues, mutable residues, and salient residues (e.g., residues salient to the structural integrity of a protein, but not important to the binding of the protein to a binding partner).

[0090] FIG.1A is an exemplary system overview for characterizing a protein sequence of interest. The system overview includes a protein analysis system 130 and one or more third party entities 110A and / or 110B in communication with one another through a network 120. FIG.1A depicts one embodiment of the overall system environment. In other embodiments, additional or fewer third party entities 110 in communication with the protein analysis system 130 can be included. Generally, the protein analysis system 130 implements methods disclosed herein for characterizing one or more structural features of a protein sequence. As referenced herein, methods for characterizing one or more structural features of a protein sequence can include one or more of: determining a stability of the polypeptide of interest, examining an interface of the polypeptide of interest, identifying a paratope of the polypeptide of interest, and evaluating an effect of one or more point mutations on a structure of the polypeptide of interest. The third party entities 110 communicate with the protein analysis system 130 for purposes associated with using the characterized one or more structural features of the protein sequence e.g., for rational drug discovery.DOCKET NO.: AIP-005WO PATENT

[0091] In various embodiments, the third party entity 110 represents a partner entity of the protein analysis system 130 that operates either upstream or downstream of the protein analysis system 130. As one example, the third party entity 110 operates upstream of the protein analysis system 130 and provide information to the protein analysis system 130 to enable the characterization of one or more structural features of a protein sequence. In this scenario, the protein analysis system 130 receives data from the third party entity 110, examples of which protein sequence(s) and / or characteristics of protein sequence(s) (e.g., a protein interface of a protein sequence). The protein analysis system 130 performs methods disclosed herein to characterize one or more structural features of the protein sequence. As another example, the third party entity 110 operates downstream of the protein analysis system 130. In this scenario, the protein analysis system 130 characterizes one or more structural features of a protein sequence and provides the characterized structural features to the third party entity 110. The third party entity 110 can subsequently use the information for their own purposes. For example, the third party entity 110 may be an entity that performs rational drug design and / or drug discovery. Therefore, given the characterized structural features of a protein sequence, the third party entity 110 may design new therapeutic molecules. For example, the characterized structural features of a protein sequence may be a characterization of a protein-protein interface involving the protein sequence. If the protein-protein interface involves a target protein implicated in a disease, the third party entity 110 may receive the characterization and design new therapeutics that seek to disrupt the protein-protein interface, and ameliorate one or more symptoms of the disease.

[0092] Referring to the network 120, any suitable network 120 may be implemented that enables connection between the protein analysis system 130 and third party entities 110. The network 120 may comprise any combination of local area and / or wide area networks, using both wired and / or wireless communication systems. In one embodiment, the network 120 uses standard communications technologies and / or protocols. For example, the network 120 includes communication links using technologies such as Ethernet, 802.11, worldwide interoperability for microwave access (WiMAX), 3G, 4G, code division multiple access (CDMA), digital subscriber line (DSL), etc. Examples of networking protocols used for communicating via the network 704 include multiprotocol label switching (MPLS), transmission control protocol / Internet protocol (TCP / IP), hypertext transport protocol (HTTP), simple mail transfer protocol (SMTP), and file transfer protocol (FTP). Data exchanged over the network 704 may be represented using any suitable format, such as hypertext markup language (HTML) or extensible markup language (XML). In some embodiments, all or someDOCKET NO.: AIP-005WO PATENT of the communication links of the network 120 may be encrypted using any suitable technique or techniques. III. Overview of the Structural Contribution Model

[0093] Disclosed herein is an approach that performs atom-level analysis of elements of a molecule of interest to characterize one or more structural features of the molecule of interest. In some embodiments, an “element” of a molecule of interest is, without limitation, an amino acid (e.g., for a protein), a nucleotide (e.g., for a nucleic acid or polynucleotide), or a chemical element such as a substructure of a small molecule compound. In some embodiments, a molecule of interest is a protein or a nucleic acid and therefore, elements can include a sequence such as, without limitation, a polypeptide sequence, or a polynucleotide sequence.

[0094] Disclosed herein is an approach that performs atom-level analysis of amino acids of a protein sequence to characterize one or more structural features of the protein sequence. However, the structural contribution model described herein may be used to perform atom- level analysis of any molecule, such as, without limitation, an amino acid, a nucleic acid, or a chemical element of a compound.

[0095] The atom-level analysis models individual atoms of the amino acids as a fluid structure, subject to destabilizing perturbations. Modeling perturbations to the interconnected atomic structure enables determination of individual contributions of each atom, as well as individual contributions of each amino acid, to the overall structural integrity of the protein structure. An understanding of the contributions of each amino acid to the structural integrity of the protein structure can be useful for e.g., rational drug design. For example, depending on whether the stability of the protein structure is to be maintained or destabilized, the amino acids that most heavily contribute to the structural integrity can be included or substituted.

[0096] Generally, the disclosed methodology that determines contributions of amino acids of a protein sequence, also referred to herein as the structural contribution model, involves using an initial 3D structure of the protein sequence (determined for example, via X-ray crystallography, cryo-electron microscopy, NMR, in silico molecular modeling or a combination thereof) to generate a series of graphs with weighted edges and nodes. In each case, the nodes of the graph correspond to atoms in the protein structure, while the edges correspond to structural linkages (e.g., bonds) between atoms. The strength of the structural linkages between atoms are represented by assigned weights of the edges.DOCKET NO.: AIP-005WO PATENT

[0097] The process begins with a protein structure and a set of interatomic scores of that protein. Using the protein structure, an edgeless version of the undirected graph is generated where every atom in the protein structure is listed in arbitrary order and assigned an index. Next, a two-dimensional map is created such that pairs of indices are mapped to a floating- point number. With this edgeless graph generated, edges are added by examining the structural linkages between all pairs of atoms.

[0098] Values are assigned to each pair of atoms according to the interactions between the pair of atoms. For example, a pair of atoms can be assigned (i) a covalent coefficient (referring to a covalent bond / interaction, if present, between the pair of atoms), (ii) a hydrogen bond term (referring to a covalent bond / interaction, if present between the pair of atoms), and (iii) a van der Waals force term (referring to a van der Waals force / interaction, if present, between the pair of atoms). Generally, two covalently-bonded atoms constrain themselves to remain in a consistent orientation, so what stabilizes the position of one stabilizes the position of the other. By contrast, the van der Waals and hydrogen bond terms explicitly relate the stability of individual atom pairs to the relative certainty of the rototranslation between them. An ideal gas, for example, would have a value of 0 between all pairs of atoms. A value of 1 indicates that a given pair of atoms is completely fixed, that the position of one atom specifies the other with absolute certainty. Since each atom pair can have at most one interaction of each type between them, a separate graph is made for each score term. These graphs are then overlaid and the maximum value between each pair kept for the overall interaction graph. This process is repeated for every atom pair.

[0099] Unlike the other terms, the covalent term (i.e., covalent bond) is not a measure of relative rototranslational certainty, at least not directly. Here, the persistence of covalent bonds is taken for granted; in effect, the floor of relative rototranslational certainty is a hypothetical molecule in which all covalent bonds are intact and all bond lengths and angles ideal but the torsional angles between them randomly assigned. Thus, the covalent coefficients will be subtracted later from the final graph.

[0100] At this point, the protein structure itself is no longer relevant, as the interaction graph contains all relevant information about the relative positions between atoms. The methodology can therefore transition from building the 3D graph to scoring. For example, all edge weights are fixed and scoring involves manipulating the node weights. Here, the scoring process is intended to capture a destabilizing perturbation, and the corresponding propagation of the perturbation through the 3D graph. Scoring is an iterative process involving theDOCKET NO.: AIP-005WO PATENT propagating of node weights along the graph. First, the scoring process proceeds in two nested loops: an outer “atom loop” that iteratively assigns a final score to each atom, and an inner “pulse loop” that repeatedly updates the 3D graph until a score can be computed. The pulse loop’s graph updates can be discretized into “pulse cycles” in which every atom is considered, all updates processed, and all atoms subsequently updated.

[0101] One or more nodes will not require an update every pulse cycle. A node is updated when one of its neighbors is updated, and therefore, the 3D graph can be updated in parts. First, every node changed in the last cycle has its neighbors updated in a separate graph, referred to herein as a “marginal” graph. Then, the marginal array is compared to the 3D graph, and any node with a higher value in the marginal array is updated in the 3D graph. The node is added to the list of updated neighbors for the next cycle.

[0102] The atom loop iterates by setting all node weights to zero except for the single atom being scored (hereafter the “scoring atom”), which receives a node weight of 1. This includes both the current and the marginal array. The update list is set to include only the scoring atom initially, and the pulse loop begins.

[0103] Every iteration of the pulse loop itself iterates over the update list and comprises a third loop, the update loop. In the update loop, a single atom from the update list has its current node weight pulled from the graph, as well as a list of atoms that share an edge with it. Each of these atoms has a value for its new node weight that is calculated by multiplying the atom’s node weight by the edge weight of the edge between them. If the neighboring atom has a zero value in the marginal array, the updated value is added; otherwise, the atom’s value is updated only if it is greater than the previous value. The update loop executes once for every atom in the update list, then finishes.

[0104] During a pulse loop, the values in the marginal array are compared to values in the pulse array. Every atom with a higher node weight in the marginal array has its node weight in the pulse array updated to the node weight in the marginal array, and it is also added to the update list. This process proceeds for every atom with a nonzero value in the marginal array, at which point the marginal array is emptied and the pulse loop repeats with its newly updated pulse array and update list.

[0105] The pulse loop continues until one of two end points is reached e.g., when no nodes are updated or when a maximum cycle count is encountered. At this point the graph’s node weights reflect the total degree to which every atom of the protein sequence specifies the position of the focal atom. Covalent-only weights can then be subtracted from the full graphDOCKET NO.: AIP-005WO PATENT weights, producing a graph in which the node weights indicate the degree to which the position of the scoring atom is specified by every other atom in the molecule. These scores are then summed to produce an atom’s total score (also referred to herein as a structural relevance of an atom) and stored in a 1-D map of atom indices to score values. The process is repeated for the next atom and structural relevance of every atom can be determined. In total, the score values of all of the atoms within an amino acid can be combined to determine the amino acid’s contribution to the protein’s structure (also referred to herein as the measure of structural contribution of an amino acid).

[0106] As discussed herein, an understanding of the contributions of each amino acid to the structural integrity of the protein structure can be useful for e.g., rational drug design. For example, depending on whether the stability of the protein structure is to be maintained or destabilized, the amino acids that have the highest measures of structural contribution can be included or substituted. IV. Implementation of Structural Contribution Model

[0107] The structural contribution model is designed to discriminate between structures that, while physically possible, are not robust to perturbation. This is in contrast to methodologies that treat molecules, e.g., proteins, as effectively stable / rigid structures. It is, for example, possible for a machine learning algorithm to return a protein structure similar to a portion of a native protein but without necessary supporting structures, as when trying to design proteins of limited length. Using the structural contribution model as a rapid post hoc filter allows for the evaluation of protein structures without recourse to similar structures for comparison, as well, which is particularly relevant for de novo protein design. Since the structural contribution model is constructed from first principles of physical chemistry, it is equally applicable to proteins, nucleic acids, ligand molecules, tags, and indeed any molecular system composed of atoms held in place by intramolecular interactions, making it applicable even to wholly new classes of molecules for which training data may not yet be available.

[0108] Accordingly, in various embodiments, the structural contribution model may be used to identify a paratope, determine stability of a protein, determine contribution of at least one atom to stability of a polypeptide, determine contribution of an amino acid to stability of a polypeptide, examine a protein-protein interface, or evaluate effects of point mutations on structure of the protein.DOCKET NO.: AIP-005WO PATENT A. Methods for Identifying a Paratope

[0109] The implementation of the structural contribution model may be discussed in the context of a paratope analysis of a molecule of interest, wherein the molecule of interest is any molecule of interest, such as, but not limited to, a protein, a polynucleotide, or a chemical compound. In this section, the implementation of the structural contribution model is discussed in the context of a paratope analysis of a protein of interest.

[0110] The detection of molecular paratopes has been a longstanding problem in the field of protein design, especially antibody design, which can impact both the computational design of functional molecules and the modification of existing molecules with new substitutions or other moieties. (Ghanbarpour, et al. (2023) ISCIENCE, 26(2): 106036; Bondarenko, et al. (2021) MABS, 13(1): 1887629; Nguyen, et al. (2017) BIOINFORMATICS, 33(19): 2971–76; Chinery, et al. (2023) BIOINFORMATICS, 39(1): btac732; Del Vecchio, et al. (2021) ARXIV:2106-00757 [Q-BIO.QM]; Liberis et al. (2018) BIOINFORMATICS, 34(17): 2944–50; Daberdaku (2018) In 15TH INTERNATIONAL CONFERENCE ON COMPUTATIONAL INTELLIGENCE METHODS FORBIOINFORMATICS ANDBIOSTATISTICS(CIBB2018); Rojas (2022) ANTIBODIES(BASEL, SWITZERLAND), 11(3): 48; Zhang, et al. (2021) In 2021 IEEE INTERNATIONALCONFERENCE ONBIOINFORMATICS ANDBIOMEDICINE(BIBM), 118–24; Lu, et al. (2021) BIORXIV2021.07.17.452531). Although the structural resolution of a given binding complex, for example, by X-ray crystallography or cryo-electron microscopy, remains the standard method for determining paratope residues, the time and expense involved in obtaining such structures renders methods based on mutational analysis more suitable for high-throughput applications. Analysis of the effects of amino acid substitutions on affinity is one such approach, but confounds the structural effects of mutations with the actual effects on the binding site – particularly in cases in which a residue may be involved in both.

[0111] The structural contribution model may, as a non-limiting example, be designed to operate as part of a comprehensive protein structure design system. With reference to FIG. 1B, the methods for characterizing a protein structure use structural information of a polypeptide (e.g., an initial set of atomic co-ordinates), such as those provided by Third Party Entity 110A or 110B, and uploaded to Network 120. Network 120 can be a private network or a public network. Alternatively or in addition, Network 120 may also relate to a database, such as any private database, or any public database. Depending upon the circumstances, the structural information of a protein of interest can be either publicly available or not publicly available.DOCKET NO.: AIP-005WO PATENT

[0112] Reference is now made to FIG.1B, which depicts an exemplary block diagram for characterizing a protein sequence, namely a paratope analysis in accordance with an embodiment. As one example, FIG.1B introduces a protein of interest 150, a Single Site Mutation (1SSM) module 155, a structural contribution model module 160, a paratope analysis module 165, a validation module 175, and a sequence envelope module 170.

[0113] In various embodiments, a protein of interest 150 can be a naturally occurring protein or a synthetic protein, e.g., a miniprotein). In various embodiments, the protein of interest can be synthesized using any methods of protein synthesis including without limitation solid-phase peptide synthesis (SPPS), liquid-phase peptide synthesis (LPPS), recombinant DNA technology, cell-free protein synthesis (CFPS), native chemical ligation (NCL), expressed protein ligation (EPL), mirror image phage display, and any other methods of protein synthesis known to one skilled in the art.

[0114] By way of example, the protein of interest 150 is analyzed e.g., by the mutation library module 155 by mutating one or more residues in the protein sequence relative to a non- native amino acid sequence to generate a 1SSM library. In certain embodiments, the protein of interest 150 is analyzed by mutating each residue in the protein sequence, one position at a time, relative to a non-native amino acid to generate a 1SSM library. In other embodiments, the protein of interest 150 is analyzed by mutating every position in the sequence of the protein of interest relative to a non-native amino acid to generate a 1SSM library. The mutation library module 155 can generate a sequence library of the plurality of protein sequences in which protein sequences of the plurality differ from one another by point mutations at predetermined positions. Further analysis can involve expressing the plurality of protein sequences, measuring the plurality of binding affinity values between the target antigen and the expressed plurality of protein sequences.

[0115] Depending upon the circumstances, generating the sequence library of the plurality of protein sequences comprises obtaining a reference protein sequence, and iteratively substituting one amino acid of the reference protein sequence with a different amino acid to generate the plurality of protein sequences. Once made, the binding affinities of the resulting mutated protein sequences (muteins) to a target protein are measured to identify amino acid residues that contribute to the structural integrity of the mutein, the binding affinity of the mutein or both. In various embodiments, the output from the mutation library module 155 include the plurality of binding affinities for the proteins sequences of the 1SSM library.DOCKET NO.: AIP-005WO PATENT

[0116] The structural contribution module 160 generally analyzes the proteins sequence of a protein of interest 150 to determine the residues of the protein sequence that contribute to the structural integrity of the protein of interest 150. For example, the structural contribution module 160 may implement a structural contribution model that determines measures of structural contributions for one or more amino acids of the protein sequence. In particular embodiments, the structural contribution module 160 implements a structural contribution model that determines measures of structural contributions for every amino acid of the protein sequence. As described herein, a measure of structural contribution of an amino acid refers to a quantitative or qualitative assessment of how the amino acid influences or contributes towards the structural integrity of the protein.

[0117] The paratope analysis module 165 performs amino acid residue-level combinations of binding affinity values determined by the mutation library module 155 and measures of structural contribution determined by the structural contribution module 160 to generate measures of non-structural binding relevance for one or more amino acid residues of the protein sequence of the protein of interest 150. The paratope analysis module 165 may identify the paratope comprising one or more amino acid residues in the protein of interest associated with measures of non-structural binding relevance e.g., to determine a sequence envelope 170. In particular embodiments, the sequence envelope 170 is an annotated representation of the paratope, showing what residues are important and wherein the physical space of the protein of interest, insertions or other mutations may be made.

[0118] In various embodiments, the paratope determined by the paratope analysis module 165 may further be subjected to a validation module 175, e.g., for purposes of validating the paratope. The validation process can, in various embodiments, be further used to determine the sequence envelope module 170.

[0119] Accordingly, the paratope analysis problem can be decomposed into two parts: (i) quantifying how important each residue is for binding, and (ii) subsequently quantifying how much of that importance is structural. The structural contribution model can be used to generate a set of scores using the holistic settings outlined above by normalizing its results to the same scale. For example, the structural contribution model can determine a difference between binding affinity values (e.g., determined by the mutation library module 155 shown in FIG.1B) and measures of structural contribution (determined by the structural contribution module 160 shown in FIG.1B). In various embodiments, the difference between the binding affinity values and the measures of structural contribution may result in a range of scores fromDOCKET NO.: AIP-005WO PATENT -1 to 1, in which more positively scoring residues are more likely to be more important components of the paratope. With this list in hand, subsequent analyses (such as producing a vector with its tail at the center of mass of the protein and its head at the center of mass when weighted by paratope score to determine the likely binding face of the protein) can be performed.

[0120] The paratope analysis module 165 in FIG.1B may conduct the steps for identifying a paratope. For example, the paratope analysis module 165 may perform amino acid residue- level combinations of binding affinity values and measures of structural contribution to generate a measure of non-structural binding relevance for each of the one or more amino acids of each of the one or more polypeptide variant sequences. Specifically, for each residue, the paratope analysis module 165 may subtract the measure of structural contribution for the residue from the binding affinity value corresponding to the residue (e.g., binding affinity value of a peptide include the residue). In various embodiments, a large difference value for a residue would indicate that the residue contributes to the binding of the protein. In contrast, a small difference value for a residue would indicate that the residue does not strongly contribute to the binding of the protein. The paratope analysis module 165 effectively highlights which residues are more important to binding than their structural importance would indicate, separating residues directly involved in forming the binding surface from those involved in stabilizing that surface.

[0121] In various embodiments, the residues directly involved in forming the binding surface are identified as contributing towards binding in the polypeptide of interest’s current structural context. For example, the residues directly involved in forming the binding surface may become more mutable if that context is changed by, for example, replacing secondary structural elements not involved in binding. Furthermore, subtracting the measures of structural contribution from the binding affinity values effectively controls for this change in context, resulting in a score for each residue that indicates its importance to binding over and above that which its structural context would indicate is important to maintaining the protein’s overall shape.

[0122] In various embodiments, by determining the difference between the binding affinity values and measures of structural contribution, the paratope analysis module 165 generates measures of non-structural binding relevance for the one or more amino acids. Here, the amino acids associated with measures of non-structural binding relevance can be identified as the paratope, also termed “sequence envelope.”DOCKET NO.: AIP-005WO PATENT

[0123] Reference is made to FIG.5A, which provides a flowchart with steps to identify a paratope of a protein of interest, in accordance with an embodiment.

[0124] With reference to FIG.5A, step 505 comprises obtaining a plurality of binding affinity values between a target antigen and a plurality of protein sequences based on the protein of interest. Step 510 comprises determining a measure of structural contribution for each of one or more amino acid residues of the protein of interest by performing a structural analysis of individual atoms, as provided herein, for example in FIG.5B. Step 515 comprises performing amino acid residue-level combinations of binding affinity values and measures of structural contribution to generate a measure of non-structural binding relevance for each of the one or more amino acid residues of each of the one or more protein sequences, as provided herein. Step 520 comprises identifying the paratope comprising one or more amino acid residues in the protein of interest associated with measures of non-structural binding relevance, as provided herein. I. Generating and Interrogating a Site Saturation Mutagenesis (SSM) Library

[0125] Disclosed herein are methods for generating a site saturation mutagenesis (SSM) library. Generally, such methods are performed by system 130 shown in FIG.1A and specifically the mutation library module 155 shown in FIG.1B. For example, beginning with a protein sequence of a protein of interest, the mutation library module 155 may mutate one or more residues of the protein sequence to generate additional sequences in the 1SSM library. For example, the mutation library module 155 may mutate two or more residues, three or more residues, four or more residues, five or more residues, six or more residues, seven or more residues, eight or more residues, nine or more residues, ten or more residues, fifteen or more residues, twenty or more residues, thirty or more residues, forty or more residues, or fifty or more residues in the protein sequence.

[0126] In various embodiments, the 1SSM library is equal to 19x the length of the construct, one for every possible amino acid mutation at every possible position. For example, given a protein sequence of N amino acids of length, the mutation library module 155 may mutate the residue at position 1 to one of the other 19 natural amino acids. Thus, the 1SSM library includes at least these 20 protein sequences (including the initial protein sequence), each of which differs from another by 1 amino acid at position 1. The mutation library module 155 may further mutate the residue at position 2 to one of the other 19 natural amino acids, and can continue this process for every other position up until the residue at position NDOCKET NO.: AIP-005WO PATENT of the protein sequence. Thus, in such embodiments, the 1SSM library includes can include 20 * N protein sequences.

[0127] In various embodiments, the protein sequences of the 1SSM library are generated (e.g., physically generated).

[0128] The mutated proteins can be generated recombinantly, where, for example, nucleic acids of interest are transformed into a host cell for expression using standard molecular biology approaches. Suitable host cells for cloning and / or expressing nucleic acids encoding the modified protein sequences include prokaryote, yeast, or higher eukaryote cells. Exemplary prokaryotic cells useful for this purpose include E. coli. In addition to prokaryotes, eukaryotic microorganism such as filamentous fungi or yeast can be used as cloning or expression hosts for polypeptide encoding vectors. Plant cell cultures of cotton, corn, potato, soybean, petunia, tomato, and tobacco can also be utilized as hosts. Suitable mammalian host cells for expressing the mutated proteins include Chinese Hamster Ovary (CHO cells), human embryonic kidney (HEK) cells, NSO myeloma cells, COS cells, and SP2 cells.

[0129] In various embodiments, generating the protein sequences of the 1SSM library include transforming into yeast and sorting the various protein sequences against a range of concentrations of the target antigen. Depending upon the circumstances, the sorting may include at least one sorting step without antigen. In various embodiments, the sorting may further include a sorting step that sorts only on display. In other embodiments, the sorting may further include a sorting step that sorts for positive antigen signal in the absence of antigen, thereby identifying nonspecific binding or binding to the signal reagent. In various embodiments, the sorting may include at least one sorting step without antigen, a sorting step that sorts only on display, and a sorting step that sorts for positive antigen signal in the absence of antigen, thereby identifying nonspecific binding or binding to the signal reagent.

[0130] In various embodiments, the 1SSM sorted libraries are then recovered and lysed using standard protocols. In various embodiments, the recovered protein sequences are analyzed using mass spectrometry, surface plasmon resonance, fluorescence-activated cell sorting, high-pressure liquid chromatography, biolayer interferometry, or other methods.

[0131] In various embodiments, a quantity of an expression host cell (e.g., a yeast cell) engineered to express a protein that binds the target labeling reagent rather than the target itself is added to each sorted group. The quantity of the expression host cell (e.g., yeast cell) can be present at the same concentration or at a different concentration throughout the sorted groups. In some embodiments, the quantity of the expression host cell (e.g., yeast cell)DOCKET NO.: AIP-005WO PATENT present at a constant concentration across sorted group allows normalization of the observed counts in each sort by the counts of the reagent binder sequences. Thus, in some embodiments, the normalization allows direct comparison of the counts of each sequence across target concentrations. In various embodiments, the comparison data of the counts of each sequence across target concentrations may be used to fit binding curves, such as but not limited to Hill curves, Monod-Wyman-Changeux (MWC) curves, Koshland-Nemethy-Filmer (KNF) curves, curves generated using a Gaddum-equation, or any other binding curve, through the residual sum of squares method. In some embodiments, affinity concentrations of each mutant may be calculated using the binding curves. The affinity concentrations can be mathematically related to compute a fold change in affinity for every mutation at every position of the polypeptide of interest. In various embodiments, the relative importance of each position to binding can be determined by using the average fold change as a measure of mutability. II. Determining Structural Contribution of Individual Amino Acids of a Protein

[0132] Methods disclosed herein involve determining structural contributions of amino acids of a protein sequence. As discussed herein, methods for determining structural contributions of amino acids involve the implementation of a structural contribution model. In various embodiments, the structural contribution model examines amino acid sequences on an atomic scale to ascertain the contribution of individual amino acid residues to the structural integrity of the protein (in terms of measures of structural contributions).

[0133] The structural contribution model operates on a series of undirected graphs with weighted edges and nodes. In some embodiments, the nodes of the graph correspond to atoms in a structure, while the edges correspond to structural linkages between those atoms. The graphs may be stored as matrices.

[0134] In various embodiments, the structural contribution model conducts a systematic pairwise assessment of atom pairs to identify which residues primarily influence target binding (for instance, rather than maintaining the structural integrity of the protein). The structural contribution model can evaluate the synergistic interactions between atoms and connections within a 3D graph. This can be accomplished as the structural contribution model can simulate a stimulus originating from an origin atom and observes the consequential effects on neighboring or interconnected atoms.

[0135] In various embodiments, the structural contribution model may be used on any amino acid sequence or structure. In various embodiments, the amino acid sequence isDOCKET NO.: AIP-005WO PATENT reducible to a single primary sequence, i.e., one without degenerate amino acids. Thus, in various embodiments, the amino acid sequence that is reducible to the single primary sequence, i.e., one without degenerate amino acids, is termed “input sequence.”

[0136] The structural contribution model is robust to changes in the manner in which a structure of a molecule is generated, however the structural contribution model does not per se generate the initial structure of a molecule of interest. Rather, the initial structure of the molecule of interest is obtained from other sources, e.g., a public or non-public database of molecular structures. Thus, in various embodiments, the structural contribution model does not comprise generating a structure of a molecule of interest.

[0137] The structural contribution model may be used to solve problems related to the stability of a macromolecule or a complex. In various embodiments, a “focal set” comprises a set of atoms for which a score will be calculated. In various embodiments, a “medial set” comprises a set of atoms through which interactions will be examined. In a non-limiting example, to assess the strength of an interface, a pair of binding atoms may comprise one binding partner’s atoms assigned to the focal set, and the other binding partner’s atoms assigned to the medial set, allowing for a score to be calculated that is reflective of the structural cooperativity of the binding surface rather than the structure as a whole. Thus, by default, the focal and medial sets both contain all atoms in the structure of interest.

[0138] In various embodiments, the focal set and medial set are the same, being the entire set of atoms in the molecule. Alternatively, it is also possible to focus the algorithm on different parts of the molecule (e.g., protein) by manipulating these two sets. In certain embodiments, when the molecule is a protein, only the side chains of the protein are set as the focal set. However, in other embodiments, any fragment of the molecule (e.g., protein) may be set as the focal set.

[0139] In general, the structural contribution model involves assigning scores that are at least pairwise-decomposable at the atomic level. It is understood that the structural contribution model may rely on a reduced score function containing at a minimum an electrostatic interaction term (hydrogen bonding) and a van der Waals term. In certain embodiments, the pairwise-decomposable scores compose a “pairwise score matrix.” Further details for assigning the pairwise-decomposable scores are described hereinbelow.

[0140] In various embodiments, determining structural contribution of amino acids of a protein sequence comprises generating a 3D graph, and pulse cycling using the 3D graph, which are described in detail hereinbelow.DOCKET NO.: AIP-005WO PATENT i. Methods for Generating a 3D Graph

[0141] This section describes the initial phase of deploying the structural contribution model, which involves generation of a 3D graph representing atoms of interest in a molecule of interest.

[0142] The structural contribution model accesses an atomic structure and a set of interatomic scores of that structure. In some embodiments, the structural contribution model generates an edgeless version of an undirected graph comprising nodes and edges using the atomic structure and set of interatomic scores. Depending upon the circumstances, the structural contribution model lists every atom in the atomic structure assigning atom indices, then creates a two-dimensional map such that pairs of indices are mapped to a floating-point number.

[0143] FIGs.2A-2D depict an exemplary process for building a 3D graph. For example, a full-atom structure of a protein (FIG.2A) is turned into a graph in which each atom occupies a node (FIG.2B) and the edges are weighted. An edge represents an interaction between two nodes (e.g., a pair of atoms). For example, two atoms that do not interact define nothing about the geometric relationship between them, and therefore, no edge exists between the two nodes representing the two atoms. The lack of an edge can be represented by assigning a cooperativity score of 0. In contrast, two atoms represented by two nodes may interact. Non- limiting examples of interactions between two atoms include a covalent bond, an electrostatic bond, a hydrogen bond, or van der Waal forces. Thus, in various embodiments, two atoms that interact with each other may comprise a specified rototranslation, resulting in a cooperativity score between 0 and 1. Thus, a score of close to 1 can be assigned to a highly specified cooperativity of two atoms. For example, a covalent bond indicates rototranslational certainty between two atoms, a cooperativity score closer to 1 can be assigned. Similarly, two atoms having a strong hydrogen bond or packing interaction specifying the location of the bonded pairs can be assigned a higher cooperativity score. In contrast, a weak van der Waal interaction between two atoms may be assigned a lower cooperativity score closer to 0.

[0144] In various embodiments, pairs of atoms can be assigned multiple cooperativity scores reflective of the multiple different interactions that can be present between the two atoms. In some embodiments, where individual atom pairs receive multiple scores, the score assigned to their interaction is the maximum score, as it is assumed that the optima, e.g., hydrogen bonds, are reflective of any considerations of optimal distances in general. Without wishing to be bound by theory, hydrogen bond energetic optima are reflective of van derDOCKET NO.: AIP-005WO PATENT Waals optima due to the energetic unfavourability of steric clashes. Thus, the optimal score of a hydrogen bond comprises the relevant atoms at some minimal distance that is not within the sum of their van der Waals radii. In various embodiments, the weight of an edge is determined according to a logistic transform of the computed energetic strength of the most structurally specifying interaction between the two nodes connected by the edge. Further details for assigning cooperativity scores are described hereinbelow.

[0145] FIG.2C depicts an exemplary 3D graph with nodes and edges overlaid upon the atomic structure. As shown in FIG.2D, the atomic structure can be removed, at which point the 3D graph of nodes and edges remains. ii. Assigning Cooperativity Scores

[0146] This section describes generation of the 3D graph and assignment of cooperativity scores between the atomic nodes and corresponding edges.

[0147] In various embodiments, the generation of the 3D graph involves assigning weights (e.g., cooperativity scores) to the edges by interrogating all pairs of atoms in the molecule. Generally, different cooperativity scores are assigned based on the type of interaction between the two atoms represented by the two nodes in the 3D graph. It is possible for multiple cooperativity scores to be assigned to two atoms if multiple interactions are present.

[0148] Covalently bonded atom pairs are assigned a single value, termed the “covalent coefficient.” The covalent value may be calculated by combining two values, the sigma and pi bond coefficients, in a ratio reflective of the pi-bond character of the individual bond. Without wishing to be bound by theory, most covalent bonds in proteins are sigma bonds, and so the atom pairs may be simply assigned the sigma bond coefficient. Notable exceptions include peptide bonds, which may be assigned 40% of the pi bond coefficient and 60% of the sigma bond coefficient, or ring structures, which may be assigned 100% of the pi bond coefficient regardless of the actual resonance structure involved. In some embodiments, the pi bond coefficient is close to 1, in some embodiments, the sigma coefficient is less than the pi bond, and in some embodiments, the sigma coefficient is close to 1.

[0149] In certain embodiments, the covalent coefficient sigma is equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to or greater thanDOCKET NO.: AIP-005WO PATENT 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, or equal to or greater than 0.99.

[0150] In certain embodiments, the covalent coefficient pi is equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, equal to or greater than 0.99, or equal to or greater than 1.

[0151] In certain embodiments, the covalent coefficient is not a direct measure of relative rototranslational certainty. Because the structural contribution model is designed to measure structural cooperativity, the persistence of covalent bonds is assumed; in effect, the floor of relative rototranslational certainty is a hypothetical molecule in which all covalent bonds are intact and all bond lengths and angles ideal but the torsional angles between them randomly assigned. Thus, in various embodiments, the covalent bond-only map may be subtracted from the overall interaction graph later. In some embodiments, the covalent coefficient is analogous to conductivity. Without wishing to be bound by theory, two covalently-bonded atoms constrain themselves to remain in a consistent orientation, thus what stabilizes the position of one stabilizes the position of the other. A long chain of covalently bonded atoms specifies their relative position. Thus, the covalent coefficient is a tunable parameter, as it controls the degree to which distant interactions can stabilize each other.

[0152] In certain embodiments, the van der Waals force and hydrogen bond terms explicitly relate the stability of individual atom pairs to the relative certainty of the rototranslation between them. This is achieved by computing a coefficient from 0 to 1 by applying a logistic transform to the raw energy value attained by whatever energy function is evaluating the particular bond in question. In some embodiments, a value of 0 implies that no property of that atom pair makes its present orientation more likely than any other. An ideal gas, for example, would have a value of 0 between all atom pairs. Thus, a value of 1 indicates that a given pair is completely fixed, and that the position of one atom specifies the other with absolute certainty.

[0153] In certain embodiments, the coefficient between two atoms reflects both the energy involved in perturbing it and the degree to which rotational components affect the score. For this reason, ideal hydrogen bonds tend to have coefficients closer to 1 than ideal van derDOCKET NO.: AIP-005WO PATENT Waals force interactions, as the van der Waals force term effectively only specifies distance but a hydrogen bond weakens as the donor-H-acceptor angle decreases from pi radians. Thus, while two atoms interacting via van der Waals forces could hypothetically “roll” over each other, changing angles while maintaining distance, and reach a degenerate but geometrically distinct state, a hydrogen bond can only rotate along the donor-H-acceptor axis without weakening.

[0154] In certain embodiments, the computations of the structural contribution model module 130 are repeated for every atom pair. If the atoms are covalently bonded, they are assigned the covalent coefficient; if not, the value of each interaction between them is computed and converted to an interaction coefficient based on one or more parameters. The one or more parameters represent primary settings that can adjust the structural contribution model results to highlight different properties of tested polypeptides. Exemplary parameters include the k and x0 parameters of logistic transforms that convert the structural properties of atom pairs into edge weights, as well as the equivalent coefficients for covalent bonds. As an example, the k and x0 parameters can be used in the following equation to provide an edge weight: 1 wherein the x0 parameter sets the value of x at which “e” is raised to the 0 power and is 1 regardless of the value of the k parameter.

[0155] As used herein, the “x0” parameter defines the score value for which the resultant edge weight is 0.5. Depending upon the circumstances, the x0 parameter is equal to or less than 0, equal to or less than -0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than -0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than - 0.35, equal to or less than -0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than -0.55, equal to or less than -0.6, equal to or less than -0.65, equal to or less than - 0.7, equal to or less than -0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than -0.9, equal to or less than -0.95, equal to or less than -1, equal to or less than -1.5, equal to or less than -2, equal to or less than -2.5, or equal to or less than -3.

[0156] As used herein, the “k” parameter defines how rapidly the edge weights trends toward 0 or 1 when the score value is higher or lower, respectively. Depending upon the circumstances, the k parameter is equal to or greater than 0, equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to orDOCKET NO.: AIP-005WO PATENT greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 1, equal to or greater than 1.5, equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, equal to or greater than 4, equal to or greater than 4.5, equal to or greater than 5, equal to or greater than 5.5, equal to or greater than 6, equal to or greater than 6.5, equal to or greater than 7, equal to or greater than 7.5, equal to or greater than 8, equal to or greater than 8.5, equal to or greater than 9, equal to or greater than 9.5, equal to or greater than 10, equal to or greater than 10.5, equal to or greater than 11, equal to or greater than 11.5, or equal to or greater than 12.

[0157] Parameters x0 and k act primarily to scale the results of different score functions relative to each other, but can also be used to tune the strictness and locality of the algorithm. The covalent coefficients control how local the results are by effectively establishing the conductivity of the covalent graph. When designing hyperstable molecules or searching for structural flaws, this may be set to a low value (~0.9), as it is assumed that successful designs will be locally stable. For natural proteins, large proteins, and a more holistic evaluation of the entire structure, the covalent coefficient should be set to a high value (for example, about 0.95) to sample broader areas of the protein. Setting to a value greater than 0.999 can lead to uninformative results as covalent interactions dominate. Van der Waals forces and hydrogen bond values will vary by score function, but assuming a normal Lennard-Jones computation, the Van der Waals forces parameters may generally be around k equal to 1, x0 equal to -0.1 for use as a classifier, while values of k greater than 8, x0 equal to -0.2 may be more appropriate for pinpointing structural weaknesses. Thus, in certain embodiments, the x0 value is set by evaluating the individual score function, with the goal of achieving a maximum hydrogen bond value of ~0.8 and a maximum Van der Waals forces value of ~0.6, while the k value is used to tune the algorithm’s sensitivity; higher k finds local flaws more efficiently, but also introduces numerical noise as slightly suboptimal regions of the structure become less conductive of the pulse chains.

[0158] In various embodiments, each atom pair has at most one interaction of each type between them. A separate graph may be made for each score term. These graphs can then be overlaid and the maximum value between each pair is kept for the overall interaction graph. In some embodiments, a score function computes a van der Waals force score for atoms interacting electrostatically, but it is assumed that, when the hydrogen bond score is computed, its score minima balance the electronic and steric forces, producing a score reflective ofDOCKET NO.: AIP-005WO PATENT optima in the van der Waals force packing. In other words, because ideal hydrogen bonds do not involve steric clashes, it is assumed that the van der Waals term is redundant with the hydrogen bond term, and so the only greater of the two is retained for a given pair of atoms.

[0159] In certain embodiments, all edge weights of the structural contribution model module 130 are fixed, and scoring involves manipulating only the node weights. Scoring can be an iterative process involving the propagating of node weights along the graph by computing for each node the value of all nodes adjacent to it, multiplied by the weight of the edge between them. The maximum value from this set is assigned to the node if it is greater than the node’s current value; otherwise, the node’s value remains unchanged. iii. Pulse Cycles using the 3D Graph

[0160] This section describes the use of pulse cycles to score and update the 3D graph. In various embodiments, a pulse cycle is meant to mimic a destabilizing perturbation, such that the resulting pulse through the 3D graph represents the effects of the perturbation on the movement of atoms within the 3D graph.

[0161] To increase efficiency of scoring, two factors may be exploited to improve the efficiency and interpretability as the pulse propagates through the 3D graph.

[0162] First, the scoring process can proceed in two nested loops: an outer “atom loop” that iteratively assigns a final score to each atom, and an inner “pulse loop” that repeatedly updates the graph until a score can be computed. The pulse loop’s graph updates can be discretized into “cycles” in which every atom is considered, all updates processed, and all atoms subsequently updated exactly once. This prevents race conditions in which, for example, atoms further down the queue get updated more than atoms at the beginning.

[0163] Second, most nodes do not require an update every cycle. A node only needs to be updated when one of its neighbors updates, and so the graph can be updated in parts. First, every node changed in the last cycle has its neighbors updated in a separate graph, the “marginal” graph. Then, the marginal array is compared to the pulse array, and any node with a higher value in the marginal array is updated in the pulse array, and added to the list of updated neighbors for the next cycle.

[0164] In some embodiments, the atom loop iterates by setting all node weights to zero except for the single atom being scored (hereafter the “scoring atom”), which receives a node weight of 1. This includes both the current and the marginal array. The update list is set to include only the scoring atom initially, and the pulse loop begins.DOCKET NO.: AIP-005WO PATENT

[0165] In some embodiments, every iteration of the pulse loop itself iterates over the update list and comprises a third loop, the update loop. In the update loop, a single atom from the update list has its current node weight pulled from the graph, as well as a list of atoms that share an edge with it. Each of these atoms has a value for its new node weight calculated by multiplying the updating atom’s node weight by the edge weight of the edge between them. If the neighboring atom has a zero value in the marginal array, it is now added with this new value; otherwise, the atom’s value is updated only if it is greater.

[0166] The update loop of the structural contribution model module 130 executes once for every atom in the update list, then finishes. In certain embodiments, the pulse loop then empties the update list and compares the marginal array to the pulse array. Every atom with a higher node weight in the marginal array has its node weight in the pulse array updated to the node weight in the marginal array, and it is also added to the update list. This process proceeds for every atom with a nonzero value in the marginal array, at which point the marginal array is emptied and the pulse loop repeats with its newly updated pulse array and update list.

[0167] In some embodiments, the pulse loop is repeated until one of the two end points is reached. In some embodiments, the graph may reach a state in which no nodes are updated. However, in some embodiments, it is contemplated that the graph may reach a state in which no nodes are updated, and since only updating nodes can spawn further updates, this state effectively locks the graph in its present state, thus the pulse loop ends.

[0168] In some embodiments, the graph’s node weights reflect the total degree to which every component of the protein specifies the position of the scoring atom. In some embodiments, the covalent-only weights are subtracted from the full graph weights, producing a graph in which the node weights indicate the degree to which the position of the scoring atom is specified by every other atom in the molecule. These scores are then summed to produce that atom’s total score and stored in a 1-D map of atom indices to score values, called the score array, and the atom loop moves to the next atom. The atom’s total score represents the structural relevance of the atom.

[0169] In some embodiments, the atom loop executes exactly once per atom in the molecule, returning the pulse array and marginal array to the initial state at the start of each set of pulse loop runs. In some embodiments, one of the functions of the atom loop is to inject an artificial cycle highlighting the scoring atom into the otherwise self-updating pulse loop, while the other function is to collect the scores at the end. Once the atom loop has iterated overDOCKET NO.: AIP-005WO PATENT every atom in the molecule, the score array is complete, and may now be either returned as a map directly or embedded in the initial structure’s B factor columns as is convenient. This latter approach is often particularly useful when dealing with classical protein modeling software that does not have an extant way to directly query per-atom values, as is often the case for residue-based approaches.

[0170] By way of example, FIGs.3A-3B depict an exemplary process of using a 3D graph for a first pulse cycle, in accordance with the embodiment shown in FIGs.2A-2D.

[0171] Each atom of the structure is interrogated sequentially. For example, a single atom (shown as the filled in node in FIG.3A) is deemed the focal atom (also referred to herein as an origin atom) and the value of that atom’s node is set to 1 while all others are set to 0 (FIG. 3A). The values of all atoms are then simultaneously updated to be the greatest of their own value or the product of the value of any adjacent atom and the edge weight between them. As shown in FIG.3B, the three nodes that are linked to the focal atom through an edge are now assigned a score other than zero (as those nodes are now filled in). This terminates the first pulse cycle.

[0172] FIG.3C depicts a process of using a 3D graph for subsequent pulse cycles, in accordance with the embodiments shown in FIGs.2A-2D and FIGs.3A-3B. Subsequent to FIG.3B, subsequent pulse cycles are performed until the graph reaches a state where further cycles are unnecessary (FIG.3C).

[0173] Additional pulse chains can then be run consisting only of covalent edges (FIG.3D) and the values of the nodes in that graph are subtracted from their equivalents in the main graph to produce the degree of relative rototranslational certainty between that atom and any others (FIG.3E). The sum of these differences is the atom’s total score. In various embodiments, the atom’s total score represents the structural relevance of the atom. The process then repeats over every other atom (e.g., setting each other atom as the focal atom).

[0174] In various embodiments, methods disclosed involve determining a structural contribution of an amino acid, which can be generated by combining the structural relevance of each atom in the amino acid. For example, given N atoms within an amino acid, each of the N atoms may have an associated structural relevance (e.g., a total score of the atom). Thus, the structural contribution of the amino acid can be the summation of the N structural relevance of the N atoms of the amino acid. iv. Exemplary 3D Graphs and Pulse Arrays for Determining Structural ContributionDOCKET NO.: AIP-005WO PATENT

[0175] This sections provides non-limiting examples of 3D graphs and pulse arrays for determining a structural contribution of an atom.

[0176] As explained herein, and with reference to FIG.2A, a 3D structure of a molecule and pairwise score matrix of a molecule may be jointly used to construct an interaction matrix between all pairs of atoms in the structure of a molecule.

[0177] Reference is now made to FIG.4A, which is a schematic representation of an exemplary interaction graph. In various embodiments, a given pair of atoms, e.g., atom 410A and atom 410B, of the molecule may have a score of 0. In various embodiments, the score can range from 0 to 1. In various embodiments, all covalently bonded atom pairs of a molecule, e.g., atom 410A and atom 410B, atom 410A and atom 410C, atom 410A and atom 410D, atom 410D ant atom 410E, atom 410E and atom 410F, may be assigned a constant score term to reflect the high degree of cooperativity between bonded atoms. In various embodiments, all covalently bonded atom pairs of the molecule assigned the constant score term may further comprise a score term to permit generalization to structure with nonideal bond geometry. In various embodiments, a non-bonded atom pairs receive a score derived by applying a logistic transform to the raw value of the score assigned to any hydrogen bond or packing interactions in which they participate; the k and x0 parameters of this transform are naturally dependent on the score function employed. Importantly, where individual atom pairs receive multiple scores, the score assigned to their interaction is simply the arithmetic maximum, as it is assumed that the optima for, for example, hydrogen bonds are reflective of any considerations of optimal atomic distances in general. For convenience, values in this matrix are called “interaction scores”. In various embodiments, the interaction matrix may be iterated over every atom in the focal set of a molecule. With reference to FIG.4A, the interaction matrix is iterated over every atom, e.g., atom 410A, atom 410B, atom 410C, atom 410D, atom 410E, and atom 410F.

[0178] With reference to FIG.4B, a pulse array is constructed in which the origin atom, e.g., atom 410A, is assigned a value of 1 and all other atoms in the medial set, e.g., atom 410B, atom 410C, atom 410D, atom 410E, or atom 410F are assigned a value of 0. These values are termed “pulse scores”. In various embodiments, the pulse array initially comprises only the origin atom.

[0179] In further reference to FIG.4B, a marginal pulse array can be constructed. For example, the marginal pulse array can initially contain all zero marginal pulse scores, and every atom in the pulse origin set may be individually examined in the interaction matrix.DOCKET NO.: AIP-005WO PATENT Thus, in certain embodiments, and with reference to FIG.4B, atom 410A is examined. In various embodiments, the marginal pulse score of every atom which interacts with the origin atom, e.g., atom 410B, atom 410C, and atom 410D, may be updated in the marginal pulse array to be the greater of its current value or the product of the focal atom’s pulse score and the interaction score of the atom pair. Thus, in exemplary embodiments, atom 410B comprises the marginal pulse score of 0.9, atom 410C comprises the marginal pulse score of 0.9, and atom 410D comprises the marginal pulse score of 0.9.

[0180] With reference to FIG.4B, following the initial cycle, e.g., cycle 1, the origin set is emptied. At this stage, the marginal pulse array and pulse array of the initial cycle may be compared. In various embodiments, all atoms with a marginal pulse score higher than their pulse score in cycle 1 may have their pulse score updated to match their marginal pulse score and are then assigned to the pulse origin set. Accordingly, and with reference to FIG.4B cycle 4, the pulse array may comprise atom 410B comprising the pulse score of 0.9, atom 410C comprising the pulse score of 0.9, and atom 410D comprising the pulse score of 0.9. With reference to FIG.4B, the marginal pulse array comprising atom 410E is then assigned the marginal pulse score of 0.81.

[0181] In various embodiments, if the pulse origin set is empty, the pulse chain may be concluded. In other embodiments, the pulse chain may be concluded prematurely based on either a fixed pulse chain maximum, or the number of atoms with nonzero values exceeding asset fraction of the total atom count of the structure. Without wishing to be bound by theory, as the nature of the algorithm calculates the most significant contributions to the pulse score first, these cutoffs typically accelerate the computation with minimal effect on the results. With reference to FIG.4B, if the pulse origin set is not empty, the updated pulse chain may be calculated. Accordingly, and with reference to FIG.4B cycle 3, the pulse array may comprise atom 410A comprising the pulse score of 1, atom 410B comprising the pulse score of 0.9, atom 410C comprising the pulse score of 0.9, atom 410D comprising the pulse score of 0.9, and atom 410E comprising the pulse score of 0.81. With reference to FIG.4B, the marginal pulse array comprising atom 410F is then assigned the marginal pulse score of 0.729. This process can be iterated over every atom of the molecule.

[0182] With reference to FIG.4C, a pulse array is constructed in which the origin atom, e.g., atom 410B, is assigned a value of 1 and all other atoms in the medial set, e.g., atom 410A, atom 410C, atom 410D, atom 410E, or atom 410F, are assigned a value of 0. TheseDOCKET NO.: AIP-005WO PATENT values are termed “pulse scores”. In various embodiments, the pulse array initially comprises only the origin atom.

[0183] In further reference to FIG.4C, a marginal pulse array is constructed. For example, the marginal pulse score of every atom which interacts with the origin atom, e.g., atom 410A, is updated in the marginal pulse array to be the greater of its current value or the product of the focal atom’s pulse score and the interaction score of the atom pair. Thus, in exemplary embodiments, atom 410B comprises the marginal pulse score of 1, and atom 410A comprises the marginal pulse score of 0.9.

[0184] In further reference to FIG.4C, following the initial cycle, e.g., cycle 1, the origin set is emptied. The marginal pulse array and pulse array of the initial cycle can be compared. Accordingly, and in reference to FIG.4C cycle 2, the pulse array may comprise atom 410B comprising the pulse score of 1, and atom 410A comprising the pulse score of 0.9. Also in reference to FIG.4C cycle 2, the marginal pulse array is then calculated and comprises atom 410C with the marginal pulse score of 0.81, atom 410D with the marginal pulse score of 0.81, and atom 410E with the marginal pulse score of 0.81.

[0185] With further reference to FIG.4B, since the pulse origin set, termed pulse array, is not empty, the updated pulse chain is calculated. Accordingly, and in reference to FIG.4C cycle 3, the pulse array comprises atom 410B comprising the pulse score of 1, atom 410A comprising the pulse score of 0.9, atom 410C comprising the pulse score of 0.81, atom 410D comprising the pulse score of 0.81, and atom 410E comprising the pulse score of 0.81. In reference to FIG.4C cycle 3, the marginal pulse array comprising atom 410F is then assigned the marginal pulse score of 0.729.

[0186] At the conclusion of the pulse chain, the origin atom’s total score is the sum of all values in the pulse array, less the origin atom’s original pulse score. Thus, in reference to FIG.4D, the structural contribution model score of atom 410A is 2.59081. Furthermore, in reference to FIG.4D, the total score of atom 410B is 2.33173. Here, the total score of each atom represents the structural relevance of each atom. The higher the total score and / or structural relevance of an atom, the more the atom contributes to the structural integrity of the amino acid (and corresponding protein). v. Flow Diagram for Determining Structural Relevance of Atoms

[0187] This section provides non-limiting flow diagrams showing steps for determining structural relevance of an atom.DOCKET NO.: AIP-005WO PATENT

[0188] Reference is now made to FIG.5B, which depicts a flowchart for characterizing a protein sequence, in accordance with an embodiment.

[0189] Step 530 involves performing a pairwise atom interaction analysis across pairs of atoms in the protein sequence to generate a 3D graph. In various embodiments, generating the 3D graph includes assigning pairwise-decomposable scores to every atom pair based on the degree to which their interactions control their relative positions, referred to herein as rototranslational certainty. Thus, in various embodiments, rototranslational certainty may be used to construct the 3D graph in which atoms are represented as nodes. In some embodiments, the 3D graph may further comprise edges weighted by the strength of the specification between the nodes.

[0190] At step 535, the 3D graph is iteratively interrogated to determine structural relevance of each of one or more atoms in the protein sequence. Step 535 may include steps 540 and 545, which are performed at each iteration. At step 540, during an iteration, a pulse array may be constructed in which the origin atom is assigned a score of 1 and all other atoms in the protein of interest’s medial set are assigned a score of 0. These values are referred to as pulse scores. In various embodiments, the purse array may include a pulse origin set, which comprises sets from which pulses may originate. One or more pulse chains can be propagated across pulse cycles. Step 545 involves combining the cooperativity scores across the one or more other atoms to determine the structural relevance of the atom. In various embodiments, step 545 is performed at the conclusion of a pulse chain. At step 550 the structural relevance of atoms in the amino acid residue is combined to generate a measure of structural contribution for the amino acid residue. III. Subtraction / determination of structural relevance of atoms

[0191] The paratope is determined by subtracting the mutability of each residue, obtained in the 1SSM, from its structural relevance, as determined by the structural contribution model. This effectively highlights which residues are more important to binding than their structural importance would indicate, separating residues directly involved in forming the binding surface from those involved in stabilizing that surface. The latter group of residues are identified as being important for binding in the polypeptide of interest’s current structural context, but they may become more mutable if that context is changed by, e.g., replacing secondary structural elements not involved in binding. Subtracting the structural contribution model scores from the mutability scores effectively controls for this, and the result is a score for each residue that indicates its importance to binding over and above that which itsDOCKET NO.: AIP-005WO PATENT structural context would indicate is important to maintaining the protein’s overall shape. In some embodiments, a residue may be important to the binding and structure of the protein. In such embodiments, the residue may be retained for the purpose of paratope analysis. In various embodiments, paratope residues (e.g., residues that more heavily contribute to the binding of the protein to its binding partner) may be substitutable, but may be restricted to any equivalent residue (a residue of the same kind, such as, but not limited to, an alike charge residue, an alike polar residue, an alike hydrophobic residue).

[0192] Thus, in various embodiments, the paratope analysis module 165 (as shown in FIG. 1B) involves subtracting the mutability scores determined by the mutation library module 155 from the structural contribution scores obtained by the structural contribution module 160. IV. Methods for Validating a Paratope

[0193] In various embodiments, the paratope identified by the paratope analysis module 165 (shown in FIG.1B) can be further validated e.g., by the validation module 175 shown in FIG.1B. The validation module 175 may include introducing mutations, such as substitutions and insertions, into the peptide sequence and then testing whether the substitutions compromise the binding activity of the mutein beyond the point of utility. In various embodiments, the sequence envelope noted above in section VB(iii) is a sequence logo, with the set of substitutions listed at each position and insertion points indicated by ellipses.

[0194] In certain embodiments, the method of validating a paratope comprises generating a modified protein sequence by inserting or removing amino acid residues associated with measures of non-structural binding relevance. The protein sequence can be inpainted with new, larger sections of a protein structure.

[0195] In some embodiments, inpainting can be logarithmic inpainting, which can be accomplished using, for example and without limitation, RoseTTAFold, DiffSDS, Protein Loop Modeling Using Deep Generative Adversarial Network, ConLooper, and HyLooper. Thus, in various embodiments, inpainting comprises generating a modified protein sequence by inserting or removing one or more amino acid residues associated with measures of non- structural binding relevance that indicate a lack of contribution to binding to the target antigen.

[0196] Depending upon the circumstances, validation of a paratope comprises determining a change in measures of non-structural binding relevance for remining residues in the modified protein sequence. It is possible to further validate the paratope by re-performing the structural analysis of individual atoms of amino acid residues of the modified protein sequence.DOCKET NO.: AIP-005WO PATENT

[0197] Validation of the paratope can comprise expressing the modified protein sequence, measuring binding affinity of the modified protein sequence, and comparing the measured binding affinity of the modified protein sequence to a binding affinity of the protein sequence prior to insertion or removal of the one or more amino acid residues associated with measures of non-structural binding relevance indicative of a lack of contribution to binding. Suitable host cells for cloning and / or expressing nucleic acids encoding the modified protein sequences include prokaryote, yeast, or higher eukaryote cells. Exemplary prokaryotic cells useful for this purpose include E. coli. In addition to prokaryotes, eukaryotic microorganism such as filamentous fungi or yeast can be used as cloning or expression hosts for polypeptide encoding vectors. Plant cell cultures of cotton, corn, potato, soybean, petunia, tomato, and tobacco can also be utilized as hosts. Suitable mammalian host cells for expressing the mutated proteins include Chinese Hamster Ovary (CHO cells), human embryonic kidney (HEK) cells, NSO myeloma cells, COS cells, and SP2 cells.

[0198] Various approaches can be used for validating a paratope. These can include, for example, determining a change in one or more measures of non-structural binding relevance for one or more residues in the modified protein sequence, such as any modified or non- modified residue in the modified protein sequence, and / or re-performing the structural analysis of individual atoms of amino acid residues of the modified protein sequence. Alternatively or in addition, paratope validation can include measuring the binding affinity to the target assessed by, for example, surface plasmon resonance, wherein the measured binding affinity of the modified protein sequence is compared to the binding affinity of the protein sequence prior to insertion or removal of the one or more amino acid residues associated with measures of non-structural binding relevance. The comparison can identify or otherwise be indicative of a lack of contribution to binding.

[0199] In various embodiments, methods for identifying a paratope can be useful in, for example, the rational design of therapeutics, for example, therapeutic proteins. For example, the identified paratope indicates the location at which a binding partner binds to the target protein. If the target protein is associated with or otherwise plays a role in a disease or disorder then the binding of the protein whose paratope has been identified or a variant thereof may result in inactivation or destabilization of the target protein, which can be used to treat the disease or disorder. As a result, therapeutics can be designed to bind to the target protein via the identified paratope.DOCKET NO.: AIP-005WO PATENT

[0200] Reference is now made to FIG.5C, which depicts a flowchart for validating a paratope of a protein of interest, in accordance with an embodiment. In various embodiments, the validation module 175 of FIG.1B may perform step 565 for validating the paratope.

[0201] In reference to step 565 of FIG.5C, validating the paratope comprises step 570 for generating a modified protein sequence by inserting or removing amino acid residues associated with measures of non-structural binding relevance.

[0202] In reference to step 565 of FIG.5C, validating the paratope also comprises step 575 for determining a change in measures of non-structural binding relevance for remining residues in the modified protein sequence. The step 575 further comprises step 580 of re- performing the structural analysis of individual atoms of amino acid residues of the modified protein sequence.

[0203] In reference to step 565 of FIG.5C, validating the paratope comprises step 585 for expressing the modified protein sequence, measuring binding affinity of the modified protein sequence, and comparing the measured binding affinity of the modified protein sequence to a binding affinity of the protein sequence prior to insertion or removal of the one or more amino acid residues associated with measures of non-structural binding relevance indicative of a lack of contribution to binding. B. Methods for Determining Stability of a Molecule

[0204] In various embodiments, the structural contribution model can be used to predict the stability of any molecule of interest. In some embodiments, the molecule of interest may be, without limitation, a polypeptide, a polynucleotide, or a chemical compound.

[0205] For example, referring to FIG.1B, the structural contribution module 160 can be deployed to analyze any molecule of interest, e.g., a protein of interest 150 to predict the stability of the molecule. In such embodiments, the structural contribution module 160 can be deployed without having to implement the other elements shown in FIG.1B (e.g., without having to implement the mutation library module 155, the paratope analysis module 165, or the validation module 175).

[0206] Methods for determining the stability of a molecule involve obtaining one or more elements of the molecule of interest and executing the structural contribution model to determine measures of structural contribution of each of the elements of the molecule. For example, the structural contribution model can be used to identify a subset of elements that contribute more heavily to the structural integrity of the molecule relative to a different subsetDOCKET NO.: AIP-005WO PATENT of elements that do not contribute to or contribute less to the structural integrity of the molecule. The subset of elements that more heavily contribute to the structural integrity of the molecule are assigned higher measures of structural contribution in comparison to lower measures of structural contribution assigned to the different subset of elements that do not contribute or contribute less to the structural integrity of the molecule. Thus, the structural contribution model may be useful for determining which elements are more important for contributing to the overall stability of the molecule. For example, in the context of a protein, the structural contribution model may be useful for determining which amino acids are more important for contributing to the overall stability of the protein. In the context of a nucleic acid, the structural contribution model may be useful for determining which nucleotides are more important for contributing to the overall stability of the nucleic acid. In the context of a small molecule compound, the structural contribution model may be useful for determining which chemical substructures are more important for contributing to the overall stability of the small molecule compound.

[0207] The methods provided for herein facilitate examination of the stability of elements over the entire molecule. For example, by determining the average measure of structural contribution across elements of the entire molecule, the structural contribution model can determine the overall stability of the molecule. Additionally, the structural contribution model can identify that a molecule is highly stable if a majority of elements contribute to the stability of the molecule. As another example, the structural contribution model can identify whether a molecule is less stable if a majority of elements do not contribute to the stability of the molecule.

[0208] In various embodiments, the structural contribution model can be implemented to determine the measures of structural contribution of only alpha carbons. Measures of structural contribution of alpha carbons can be sufficient for determining the overall stability of a molecule, since surface side chains are expected to be flexible in solvent. Alternatively, or in addition, the structural contribution model can be implemented to determine the measures of structural contribution of only a subset of elements of a molecule. In such embodiments, low measures of structural contribution of the subset of elements can indicate that the low-scoring region is not sufficiently linked to the rest of the molecule. Thus, a protein with a region associated with low measures of structural contribution is more likely to unravel.DOCKET NO.: AIP-005WO PATENT

[0209] Using such methods, it is possible to differentiate between a subset of elements that contribute to the structural integrity of the molecule (herein referred to as salient elements), elements that contribute to the binding of the molecule to a binding partner (herein referred to as paratope elements), and other elements (herein referred to as other elements or mutable elements).

[0210] In general, elements that contribute to the structural integrity of the molecule (e.g., salient elements) have larger assigned measures of structural contribution. Salient elements may also have corresponding binding affinity values, for example, as determined through a SSM library analysis. Therefore, when taking the difference between the binding affinity values and the measures of structural contribution, salient elements may have measures of non-structural binding relevance that are not near zero (e.g., greater than 1, greater than 2, greater than 3, greater than 4, greater than 5, greater than 10, greater than 15, or greater than 20).

[0211] The other elements (herein referred to as other elements or mutable elements) may have a measure of structural contribution that is zero or near zero, and also have a corresponding binding affinity value as determined, for example, through the SSM library analysis, that is similarly zero or near zero. Accordingly, these other elements do not contribute or contribute little to both the stability of the molecule as well as the binding of the molecule to its binding partner. When taking the difference between the binding affinity values and the measures of structural contribution, paratope elements typically have measures of non-structural binding relevance that are near zero.

[0212] As another example, elements that contribute to the binding of the molecule to a binding partner (herein referred to as paratope elements) may have a measure of structural contribution and further have a larger corresponding binding affinity values as determined, for example, through an SSM library analysis. Therefore, when taking the difference between the binding affinity values and the measures of structural contribution, paratope residues may have larger measures of non-structural binding relevance (thereby indicating that such paratope elements contribute towards binding as opposed to merely contributing to structural integrity of the molecule). C. Methods for Determining Stability of a Polypeptide

[0213] In various embodiments, the structural contribution model can be used to predict the stability of a protein of interest. The importance of understanding stability of a protein ofDOCKET NO.: AIP-005WO PATENT interest for the purpose of drug design is two-fold. First, a protein of interest can be designed and evaluated to sure that the protein remains stable at particular temperatures e.g., at physiologically-relevant temperatures. Protein therapeutics that remain stable at physiological temperatures are more likely to exhibit desired activity. If a protein is insufficiently thermally stable to adopt a functional conformation at the body temperature of the patient, it will not function. Second, without wishing to be bound by theory, thermal stability is correlated with resistance to many of the ways protein therapeutics are denatured, including acids and proteases. Thus, a computational model of overall stability allows for protein therapeutics to be designed and modified more efficiently by screening out modifications that would otherwise reduce their therapeutic utility.

[0214] For example, referring to FIG.1B, the structural contribution module 160 can be deployed to analyze a protein of interest 150 to predict the stability of the protein of interest 150. In such embodiments, the structural contribution module 160 can be deployed without having to implement the other elements shown in FIG.1B (e.g., without having to implement the mutation library module 155, the paratope analysis module 165, or the validation module 175).

[0215] Methods for determining the stability of a protein involve obtaining a protein sequence and executing the structural contribution model to determine measures of structural contribution of each of the amino acids of the protein. For example, the structural contribution model can be used to identify a subset of amino acids that contribute more heavily to the structural integrity of the protein relative to a different subset of amino acids that do not contribute to or contribute less to the structural integrity of the protein. The subset of amino acids that more heavily contribute to the structural integrity of the protein are assigned higher measures of structural contribution in comparison to lower measures of structural contribution assigned to the different subset of amino acids that do not contribute or contribute less to the structural integrity of the protein. Thus, the structural contribution model may be useful for determining which amino acids are more important for contributing to the overall stability of the protein. The methods provided for herein facilitate examination of the stability of amino acids over the entire protein. For example, by determining the average measure of structural contribution across amino acids of the entire protein, the structural contribution model can determine the overall stability of the protein. Additionally, the structural contribution model can identify that a protein is highly stable if a majority of amino acids contribute to the stability of the protein. As another example, the structural contribution model can identifyDOCKET NO.: AIP-005WO PATENT that a protein is less stable if a majority of amino acids do not contribute to the stability of the protein.

[0216] In various embodiments, the structural contribution model can be implemented to determine the measures of structural contribution of only alpha carbons. Measures of structural contribution of alpha carbons can be sufficient for determining the overall stability of a protein, since surface side chains are expected to be flexible in solvent. Alternatively, or in addition, the structural contribution model can be implemented to determine the measures of structural contribution of only a subset of amino acids of a protein. In such embodiments, low measures of structural contribution of the subset of amino acids can indicate that the low- scoring region is not sufficiently linked to the rest of the protein. Thus, a protein with a region associated with low measures of structural contribution is more likely to unravel.

[0217] Using such methods, it is possible to differentiate between a subset of amino acids that contribute to the structural integrity of the protein (herein referred to as salient residues), amino acids that contribute to the binding of the protein to a binding partner (herein referred to as paratope residues), and other amino acids (herein referred to as other residues or mutable residues).

[0218] In general, amino acids that contribute to the structural integrity of the protein (e.g., salient residues) have larger assigned measures of structural contribution. Salient residues may also have corresponding binding affinity values, for example, as determined through a SSM library analysis. Therefore, when taking the difference between the binding affinity values and the measures of structural contribution, salient residues may have measures of non- structural binding relevance that are not near zero (e.g., greater than 1, greater than 2, greater than 3, greater than 4, greater than 5, greater than 10, greater than 15, or greater than 20).

[0219] The other amino acids (herein referred to as other residues or mutable residues) may have a measure of structural contribution that is zero or near zero, and also have a corresponding binding affinity value as determined, for example, through the SSM library analysis, that is similarly zero or near zero. Accordingly, these other amino acids do not contribute or contribute little to both the stability of the protein as well as the binding of the protein to its binding partner. When taking the difference between the binding affinity values and the measures of structural contribution, paratope residues typically have measures of non- structural binding relevance that are near zero.

[0220] As another example, amino acids that contribute to the binding of the protein to a binding partner (herein referred to as paratope residues) may have a measure of structuralDOCKET NO.: AIP-005WO PATENT contribution and further have a larger corresponding binding affinity values as determined, for example, through an SSM library analysis. Therefore, when taking the difference between the binding affinity values and the measures of structural contribution, paratope residues may have larger measures of non-structural binding relevance (thereby indicating that such paratope residues contribute towards binding as opposed to merely contributing to structural integrity of the protein).

[0221] In various embodiments, methods for determining a stability of a protein can be useful for e.g., rational therapeutic design. For example, determining the stability of the protein can result in identification of one or more paratope residues (e.g., residues contributing to the binding of the protein to a binding partner) and / or one or more residues that may contribute towards the structural integrity of the protein (e.g., referred to herein as salient residues) but may not contribute towards the binding of the protein to a binding partner. Therefore, if the goal of a therapeutic is to destabilize the protein without abrogating its ability to bind to a binding partner, the one or more salient residues can be substituted while maintaining the one or more paratope residues. For example, such a therapeutic can be useful to competitively bind to a target (e.g., thereby outcompeting binding of the target to a native ligand, which may be implicated in disease). As another example, if the goal of a therapeutic is to maintain the stability of the protein, but to abrogate the affinity to a binding partner, one or more paratope residues can be substituted while maintaining the one or more salient residues. As another example, if the goal of a therapeutic is to maintain the stability of the protein and further maintain the binding to a binding partner, various therapeutics can be designed by introducing substitutions in amino acids other than the one or more paratope residues and one or more salient residues (e.g., such amino acids other than the one or more paratope residues and one or more salient residues may be referred to herein as mutable residues). D. Methods for Evaluating an Effect of a Point Mutation

[0222] In various embodiments, the structural contribution model can be used to evaluate or predict the effect of one or more point mutations on a structure of a molecule of interest. In some embodiments, the molecule of interest may be, without limitation, a polypeptide, or a polynucleotide.

[0223] In some embodiments, the structural contribution model is used to evaluate or predict the effect of one or more point mutations on a structure of a protein. When the one or more point mutations are introduced into the protein or peptide of interest (e.g., in vitro or inDOCKET NO.: AIP-005WO PATENT silico), the structural contribution model can be employed to characterize the tolerability of the structure of the protein or peptide to the one or more point mutations. Certain point mutations can be characterized as tolerable (e.g., can be introduced without substantively impacting the structural integrity of the protein) whereas other point mutations can be characterized as intolerable (e.g., cannot be introduced without substantively impacting the structural integrity of the protein). Referring to FIG.1B, the structural contribution module 160 can be deployed to analyze a protein of interest 150 to evaluate or predict the effect of one or more point mutations on a structure of a protein of interest 150. In such embodiments, the structural contribution module 160 can be deployed without having to implement the other elements shown in FIG.1B (e.g., without having to implement the mutation library module 155, the paratope analysis module 165, or the validation module 175).

[0224] In various embodiments, the method for evaluating an effect of one or more point mutations can comprise introducing a point mutation into an amino acid sequence of interest, where the point mutation can be selected from a particular type of amino acid, including, for example, one or more of a nonpolar amino acid, a polar amino acid, a positively charged amino acid, or a negatively charged amino acid. Therefore, methods for evaluating the effect of one or more point mutations can involve determining whether a particular type of amino acid can be better tolerated in comparison to another type of amino acid. Exemplary nonpolar amino acids include alanine (A), glycine (G), isoleucine (I), leucine (L), methionine (M), tryptophan (W), phenylalanine (F), proline (P), or valine (V). Exemplary polar amino acids include cysteine (C), serine (S), threonine (T), tyrosine (Y), asparagine (N), or glutamine (Q). Exemplary positively charged amino acids include histidine (H), lysine (K), or arginine (R). Exemplary negatively charged amino acids include aspartate (D) or glutamate (E).

[0225] Similarly, methods for evaluating an effect of one or more point mutations can further involve determining the measures of structural contribution of the amino acids, including the introduced point mutation, of the amino acid sequence (e.g., using the methods disclosed herein for determining structural contributions of amino acids).

[0226] In various embodiments, methods for evaluating an effect of one or more point mutations can comprise a repeated process that involves (i) introducing a point mutation into an amino acid sequence, and (ii) determining measures of structural contribution of the amino acids of the amino acid sequence, including the point mutation. Thus, at each repetition, the structural contribution of the amino acids can be determined to understand how the introduced point mutation affects the measures of structural contribution of the other amino acids. ForDOCKET NO.: AIP-005WO PATENT example, introduction of a point mutation into an amino acid sequence may increase the measures of structural contribution of the other amino acids of the amino acid sequence, in which case, the introduced point mutation can be categorized as a substitution that increases the structural integrity of the protein. This point mutation can be a tolerable point mutation. As another example, introduction of a point mutation into an amino acid sequence may not result in significant changes in the measures of structural contribution of the other amino acids, in which case, the introduced point mutation can be categorized as not destabilizing the protein. This point mutation can be a tolerable point mutation. As another example, introduction of a point mutation into an amino acid sequence may cause significant reduction in the measures of structural contribution of the other amino acids, in which case the introduced point mutation can be categorized as a destabilizing substitution. Thus, this point mutation can be an intolerable point mutation.

[0227] In various embodiments, methods for evaluating an effect of one or more point mutations can be useful for, for example, in the rational design of a therapeutic. For example, the results of evaluating the effect of one or more point mutations may be that certain substitutions (or types of amino acids) can be tolerable or intolerable at certain positions of the amino acid sequence. For example, assuming that, at position X of a protein sequence, non- polar amino acids are well tolerated but polar amino acids are not well tolerated (e.g., cause destabilization of the protein), then, with the understanding of tolerable / intolerable amino acids at position X of the protein sequence, rational design of a therapeutic can be accomplished to stabilize or destabilize the protein, depending on the protein’s role in disease, by modulating the amino acid at position X. For example, if a protein plays a role in a disease or disorder, a the protein can be modified by rational design to change a non-polar amino acid at position X to a more polar group, thereby destabilizing the protein and minimizing its contribution to the disease or disorder. E. Methods for Examining Protein-Protein Interfaces

[0228] In various embodiments, the structural contribution model can be used to identify and interrogate a protein-protein interface. Referring to FIG.1B, the structural contribution module 160 can be deployed to analyze a protein of interest 150 to identify and interrogate a protein-protein interface involving a protein of interest 150. In such embodiments, the structural contribution module 160 can be deployed without having to implement the other elements shown in FIG.1B (e.g., without having to implement the mutation library module 155, the paratope analysis module 165, or the validation module 175).DOCKET NO.: AIP-005WO PATENT

[0229] The analysis of protein-protein interfaces is a pressing problem in computational protein design with implications for the design of therapeutic macromolecules as well as the elucidation of the biophysical underpinnings of macromolecular interactions. (Jones, et al. (1997) JOURNAL OFMOLECULARBIOLOGY, 272(1): 133–43; Bryant, et al. (2022) NATURECOMMUNICATIONS, 13(1): 1–11; Das, et al. (2021) SCIENTIFIC REPORTS, 11(1): 1–12; Daberdaku (2018) supra; Xue, et al. (2015) FEBS LETTERS, 589(23): 3516–26; Stringer, et al. (2022) BIOINFORMATICS, 38(8): 2111–18; Wu, et al. (2022) ADVANCES IN KNOWLEDGE DISCOVERY AND DATA MINING, 365–78; Malhotra, et al. (2021) NATURE COMMUNICATIONS, 12(1): 1–12; Rao, et al. (2014) INTERNATIONAL JOURNAL OF PROTEOMICS, (February): 147648; Sriramulu, et al. (2023) JOURNAL OF MOLECULAR GRAPHICS & MODELLING, 122(July): 108461; Mohseni, et al. (2023) BIOINFORMATICS, 39 (Supplement_1): i544–52; Bai, et al. (2016) PROC. NAT. ACAD. SCIENCES, 113(50): E8051–58; Meytin (2021) In 2021 IEEE MIT UNDERGRADUATE RESEARCH TECHNOLOGY CONFERENCE (URTC), 1–5; Stites (1997) CHEMICAL REVIEWS, 97(5): 1233–50).

[0230] In various embodiments, examining a protein-protein interface comprises analyzing a protein-protein interface within a single protein. Thus, examining a protein-protein interface can involve determining at least one interface region of the single protein of interest. Here, the interface region can play a role in one or more structurally important interactions of the protein. For example, a structurally important interaction can involve an interaction that maintains the stability of the protein. As another example, a structurally important interaction can involve a binding interaction with e.g., a binding partner.

[0231] In various embodiments, examining a protein-protein interface comprises analyzing a protein-protein interface between two proteins. Thus, examining a protein-protein interface may involve analyzing an interface between at least two polypeptide regions of two different proteins. The polypeptide region can comprise a fragment of a protein. Alternatively, the polypeptide region can comprise a full length amino acid sequence of a protein.

[0232] In various embodiments, examining a protein-protein interface comprises setting a first portion of an amino acid sequence as a focal chain and setting a second portion of the amino acid sequence as a medial chain. Differentiating between the focal chain and the medial chain of the amino acid sequence enables performing the methods for determining structural contribution of an amino acid provided for herein, for only one of the chains. For example, the methods for determining structural contribution of amino acids can be performed for only amino acids of the focal chain or for only amino acids of the medial chain. Thus,DOCKET NO.: AIP-005WO PATENT pulses can be propagated across amino acids of the focal chain, and the resulting scores (e.g., measures of structural contribution) can be used to assess interface quality. Altogether, differentiating between focal and medial chains enables a localized analysis to determine how the individual amino acids of the focal chain contribute to the stability of the portion of the amino acid at the protein-protein interface (e.g., a binding surface).

[0233] In certain embodiments, examining the interface of the polypeptide comprises, prior to performing a pairwise atom to atom interaction analysis across at least one pair of atoms in the amino acid sequence to generate a 3D graph, the steps of: (a) selecting a focal chain; and (b) selecting a medial chain. In some embodiments, the focal chain and the medial chain comprise a polypeptide region of the polypeptide, and wherein the polypeptide region of the focal chain is the same as, or different than the polypeptide region of the medial chain. In some embodiments, the method for examining the interface of the polypeptide comprises determining measures of structural contribution (using the methods disclosed herein) for the amino acids of the focal chain of the polypeptide. In some embodiments, the method for examining the interface of the polypeptide comprises determining measures of structural contribution (using the methods disclosed herein) for the amino acids of the medial chain of the polypeptide. In some embodiments, the method for examining the interface of the polypeptide comprises determining measures of structural contribution (using the methods disclosed herein) for both the amino acids of the focal chain and amino acids of the medial chain of the polypeptide.

[0234] In certain embodiments, methods for examining a protein-protein interface can be useful for e.g., rational drug design. For example, assuming that a given protein-protein interface is implicated in disease, the results of the protein-protein interface analysis can identify one or more amino acids (e.g., of the focal chain or of the medial chain) that are important for maintaining the stability of the protein-protein interface. Thus, methods for rational drug design may seek to generate therapeutics that effectively target or destabilize one or more of the identified amino acids. Such therapeutics can therefore destabilize the protein- protein interface and ameliorate or treat the disease. F. Methods for Examining Protein-Nucleic Acid Interfaces

[0235] In various embodiments, the structural contribution model can be used to identify and interrogate a protein-DNA interface, or an interface between a protein and any nucleic acid (e.g., DNA, RNA, cDNA, and the like).DOCKET NO.: AIP-005WO PATENT

[0236] In various embodiments, examining a protein-nucleic acid interface comprises analyzing an interface within a nucleic acid that is involved in a protein-nucleic acid interaction. Thus, examining a protein-DNA interface can involve determining at least one interface region of the nucleic acid. Here, the interface region can play a role in one or more structurally important interactions of the nucleic acid. For example, a structurally important interaction can involve an interaction that maintains the stability of the nucleic acid. As another example, a structurally important interaction can involve a binding interaction with e.g., a binding partner such as the protein.

[0237] In various embodiments, examining a protein-nucleic acid interface comprises analyzing a protein-nucleic acid interface between a protein and a nucleic acid. Thus, examining a protein-nucleic acid interface may involve analyzing an interface between at least two regions of two different molecules, such as a protein and nucleic acid. The polypeptide region can comprise a fragment of a protein. The nucleic acid region can comprise a fragment of a DNA, RNA, cDNA, or any other polynucleotide sequence. Alternatively, the polypeptide region can comprise a full length amino acid sequence of a protein, and the nucleic acid region can comprise a full length DNA, RNA, cDNA, or any other polynucleotide sequence.

[0238] In various embodiments, examining a protein-nucleic acid interface comprises setting a portion of an amino acid sequence as a focal chain and setting a portion of the nucleotide sequence as a medial chain. In some embodiments, examining a protein-nucleic acid interface comprises setting a portion of an amino acid sequence as a medial chain and setting a portion of the nucleotide sequence as a focal chain. Differentiating between the focal chain and the medial chain of the sequence enables performing the methods for determining structural contribution of an amino acid or a nucleic acid provided for herein, for only one of the chains. For example, the methods for determining structural contribution of amino acids can be performed for only amino acids of the focal chain or for only amino acids of the medial chain. Thus, pulses can be propagated across amino acids of the focal chain, and the resulting scores (e.g., measures of structural contribution) can be used to assess interface quality. In other embodiments, the methods for determining structural contribution of nucleic acids can be performed for only nucleic acids of the focal chain or for only nucleic acids of the medial chain. Accordingly, in said embodiments, pulses can be propagated across nucleic acids of the focal chain, and the resulting scores (e.g., measures of structural contribution) can be used to assess interface quality. Altogether, differentiating between focal and medial chains enables a localized analysis to determine how the individual amino acids or nucleic acids of the focalDOCKET NO.: AIP-005WO PATENT chain contribute to the stability of the portion of the amino acid at the protein-nucleic acid interface (e.g., a binding surface).

[0239] In certain embodiments, examining the interface of the protein-nucleic acid comprises, prior to performing a pairwise atom to atom interaction analysis across at least one pair of atoms in the amino acid sequence or the nucleic acids sequence to generate a 3D graph, the steps of: (a) selecting a focal chain; and (b) selecting a medial chain. In some embodiments, the focal chain comprises a polypeptide region of the polypeptide, and the medial chain comprise a nucleic acid region of the polynucleotide. In some embodiments, the medial chain comprises a polypeptide region of the polypeptide, and the focal chain comprise a nucleic acid region of the nucleic acid. In some embodiments, the method for examining the interface of the protein-nucleic acid comprises determining measures of structural contribution (using the methods disclosed herein) for the amino acids of the focal chain of the polypeptide. In some embodiments, the method for examining the interface of the protein-nucleic acid comprises determining measures of structural contribution (using the methods disclosed herein) for the nucleic acids of the focal chain of the nucleotide. In some embodiments, the method for examining the interface of the protein-nucleic acid comprises determining measures of structural contribution (using the methods disclosed herein) for the amino acids of the medial chain of the polypeptide. In some embodiments, the method for examining the interface of the protein-nucleic acid comprises determining measures of structural contribution (using the methods disclosed herein) for the nucleic acids of the medial chain of the nucleic acid. In some embodiments, the method for examining the interface of the protein-nucleic acid comprises determining measures of structural contribution (using the methods disclosed herein) for both the amino acids of the focal chain and nucleic acids of the medial chain of the nucleic acid, or the amino acids of the medial chain and nucleic acids of the focal chain of the nucleic acid. V. Computer Embodiments

[0240] Also provided herein is a computer readable medium comprising computer executable instructions configured to implement any of the methods described herein. In various embodiments, the computer readable medium is a non-transitory computer readable medium. In some embodiments, the computer readable medium is a part of a computer system (e.g., a memory of a computer system). Examples of a computing device can include a personal computer, desktop computer laptop, server computer, a computing node within a cluster, message processors, hand-held devices, multi-processor systems, microprocessor-DOCKET NO.: AIP-005WO PATENT based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like.

[0241] FIG.6 illustrates an exemplary computing device 600 for implementing system and methods described in FIGs.1A-1B, 2A-2D, 3A-3D, 4A-4D, and 5A-5C.

[0242] In some embodiments, the computing device 600 includes at least one processor 602 coupled to a chipset 604. The chipset 604 includes a memory controller hub 620 and an input / output (I / O) controller hub 622. A memory 606 and a graphics adapter 612 are coupled to the memory controller hub 620, and a display 618 is coupled to the graphics adapter 612. A storage device 608, an input interface 614, and network adapter 616 are coupled to the I / O controller hub 622. Other embodiments of the computing device 600 have different architectures.

[0243] The storage device 608 is a non-transitory computer-readable storage medium such as a hard drive, compact disk read-only memory (CD-ROM), DVD, or a solid-state memory device. The memory 606 holds instructions and data used by the processor 602. The input interface 614 is a touch-screen interface, a mouse, track ball, or other type of input interface, a keyboard, or some combination thereof, and is used to input data into the computing device 600. In some embodiments, the computing device 600 may be configured to receive input (e.g., commands) from the input interface 614 via gestures from the user. The graphics adapter 612 displays images and other information on the display 618. As an example, the display 618 can show visualizations of molecular interface. The network adapter 616 couples the computing device 600 to one or more computer networks.

[0244] The computing device 600 is adapted to execute computer program modules for providing functionality described herein. As used herein, the term “module” refers to computer program logic used to provide the specified functionality. Thus, a module can be implemented in hardware, firmware, and / or software. In one embodiment, program modules are stored on the storage device 608, loaded into the memory 606, and executed by the processor 602.

[0245] The types of computing devices 600 can vary from the embodiments described herein. For example, the computing device 600 can lack some of the components described above, such as graphics adapters 612, input interface 614, and displays 618. In some embodiments, a computing device 600 can include a processor 602 for executing instructions stored on a memory 606.DOCKET NO.: AIP-005WO PATENT

[0246] Each program can be implemented in a high level procedural or object oriented programming language to communicate with a computer system. However, the programs can be implemented in assembly or machine language, if desired. In any case, the language can be a compiled or interpreted language. Each such computer program is preferably stored on a storage media or device (e.g., ROM or magnetic diskette) readable by a general or special purpose programmable computer, for configuring and operating the computer when the storage media or device is read by the computer to perform the procedures described herein. The system can also be considered to be implemented as a computer-readable storage medium, configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.

[0247] The signature patterns and databases thereof can be provided in a variety of media to facilitate their use. “Media” refers to a manufacture that contains the signature pattern information of the present invention. The databases of the present invention can be recorded on computer readable media, e.g., any medium that can be read and accessed directly by a computer. Such media include, but are not limited to: magnetic storage media, such as floppy discs, hard disc storage medium, and magnetic tape; optical storage media such as CD-ROM; electrical storage media such as RAM and ROM; and hybrids of these categories such as magnetic / optical storage media. One of skill in the art can readily appreciate how any of the presently known computer readable mediums can be used to create a manufacture comprising a recording of the present database information. "Recorded" refers to a process for storing information on computer readable medium, using any such methods as known in the art. Any convenient data storage structure can be chosen, based on the means used to access the stored information. A variety of data processor programs and formats can be used for storage, e.g., word processing text file, database format, etc. EXAMPLES

[0248] The following Examples are merely illustrative and are not intended to limit the scope or content of the invention in any way. EXAMPLE 1 – THE STRUCTURAL CONTRIBUTION MODEL PROVIDES AN IMPROVED APPROACH FOR PROVIDING STRUCTURAL CLASSIFIERS

[0249] This example describes the use of the structural model approach to identify improved metrics of protein structure relative to other current approaches.DOCKET NO.: AIP-005WO PATENT

[0250] The selection of candidate protein sequences is a valuable step in many protein design workflows, and high throughput approaches typically involve the use of computational approaches to predict which designs will likely fold most readily. To examine the performance of the structural contribution model in this role, a large and topologically diverse set of miniproteins (Rocklin et al.2017) was used to evaluate a set of metrics to classify protein structures as stable or unstable, using a scalar cutoff for the stability score to determine stable structures. Given that hyperstable structures are generally desirable, the performance of each metric was assessed at multiple stability cutoffs. Each parameter was independently fit to maximize F , which is a harmonic mean between precision and recall for each stabilitycutoff, with a low beta value of 0.1 to account for the increased importance of maximized true positive rate rather than true negative rate for this task, because generating protein structures computationally is less resource-intensive than physically producing the proteins and evaluating their properties. Candidate metrics tested included Rosetta energy units per residue (a standard method for evaluating protein stability) as well as the buried nonpolar surface area per residue (a means of assessing core packing). (Rocklin et al.2017 supra). Additionally, the pLDDT score of an AlphaFold2 (AF2) prediction of the design sequence was included as a metric for filtration. It is routine to filter designs by AF2’s mean pLDDT score to assess the probability of the sequence folding into the designed structure. For this reason, analysis was restricted to sequences with AF2 mean pLDDT score more than 70. Under these conditions, cutoffs exist for which the mean structural contribution model serves as a more effective determinant of stable proteins and unstable protein bins across a wide range of measures of stability, with the structural contribution model’s performance increasing relative to all extant methods as the standards for stability increase.

[0251] A head-to-head comparison of the structural contribution model with the current state of the art approaches is shown in FIG.7. The performance of the structural contribution model as a determinant of protein stability on the entirety of the Rocklin miniprotein protease resistance dataset was compared to that of Rosetta’s Rosetta Energy Units per residue (REU per residue), buried nonpolar surface area per residue (Buried NPSA per residue), mean Alphafold2 pLDDT score (mean PLDDT), and the mean and maximum values of the structural contribution model. Cutoff values were used to convert a continuous spectrum of the predictions of the structural contribution model into a binary prediction as to whether a given protein was stable. Here, given a numerical cutoff, a protein is deemed stable if it scores above the cutoff. Conversely, a protein is deemed unstable if it scores below the cutoff. Since this score is a scalar value, the large array of scores (e.g., measures of structuralDOCKET NO.: AIP-005WO PATENT contribution) determined by the structural contribution model was converted to a single scalar value for comparison. Specifically, measures of structural contribution were averaged to a scalar value, which was then compared to the cutoff.

[0252] Generally, a cutoff that is too restrictive may mark stable designs as unstable, resulting in false negatives. Conversely, a cutoff that is too permissive may mark unstable designs as stable, resulting in false positives. This relationship between the undesirability of false positives and false negatives is termed the value. If false positives and false negatives are equally undesirable, = 1. was set to less than one to reflect that false positives are much worse than false negatives. Here, the value equal to 0.1 was chosen. The performance of each metric according to the F parameter was then measured. The F parameter measureshow well a given predictor performs on two metrics: precision and recall. “Precision,” as used herein, refers to the number of true positives divided by the sum of true and false positives. Precision indicates how effective a given predictor is at weeding out negative results, here defined as unstable proteins. “Recall,” as used herein, refers to the number of true positives divided by the sum of true positives and false negatives. Recall indicates how well a given predictor detects positive results, here defined as stable proteins. The F equation is asfollows:

[0253] The value was set to 0.1 to reflect the increased importance of precision over recall. The maximum attainable F value, which is the harmonic mean between precision andrecall of each model for each stability cutoff, was determined across all possible cutoffs for each metric. FIG. 7 shows the maximum possible F obtainable by setting the cutoff for eachmetric to any value. In other words, FIG.7 compares how effective each predictor could possibly be if the perfect cutoff was known ahead of time. This can be found by assigning one cutoff to stability and thereby converting the experimentally determined stability values into real positives and negatives. The performance of each metric shown was then examined by testing a range of cutoffs for each and computing F for all of them. The maximum Freported over the entire range is shown in FIG.7.

[0254] As shown in FIG.7, the structural contribution model significantly outperformed all other metrics where the desired stability cutoff is very high, without losing performance where the cutoff is closer to 1.DOCKET NO.: AIP-005WO PATENT EXAMPLE 2 – PARATOPE IDENTIFICATION AND VERIFICATION FOR A TNFR1-BINDING MINIPROTEIN

[0255] This Example describes an exemplary process for identifying a paratope of an exemplary TNFR1-binding miniprotein comprising the amino acid sequence of SARDYLERLRDEGYISDVLEGQLNDLLDRGEDEQAVIDYANDFIESR (SEQ ID NO:1).

[0256] Briefly, in a first step, amino acid residues involved in maintaining the structure and binding activity were identified from the affinity values of each miniprotein in a 1SSM Library based on the TNFR1-binding miniprotein. In a second step, amino acid residues involved in maintaining the structure of the miniprotein were determined by the structural contribution model algorithm. The paratope was then identified by comparing the amino acid residues identified by the second step against the amino acid residues identified by the first step. Once identified, the paratope was verified by the process of inpainting. After that, a sequence envelope of the paratope was created. Phase 1. 1SSM Library Construction and Determination of affinity values for a TNFR1- binding Miniprotein and Single Mutant Variants Thereof.

[0257] This section describes a process constructing an iterative, site-specific, saturation mutagenesis 1SSM library, and affinity analysis of each polypeptide in the 1SSM library. The resulting information was indicative of amino acid residues in the miniprotein that contribute to the structure and stability of the miniprotein as well as the amino acids that contribute to the binding paratope of the miniprotein.

[0258] At the outset, the amino acid sequence of the TNFR1-binding minprotein polypeptide of interest was expanded into a 1SSM library as shown in Table 1, where every position in the sequence was individually mutated to every non-native amino acid except for cysteine (C), generating the sequence library shown in Table 1. The 1SSM sequence library contained sequences equal to 19X the length of the template miniprotein, one for every possible mutation at every possible position. After library construction, the affinity values of each polypeptide in the library was measured. The affinity values are shown in Table 1.

[0259] Table 1 shows the 1SSM library including sequences, affinity values, and mutations as compared to SEQ ID NO: 1.E E E E E E E E E G G GG GGGG E E YE GG GGGG E E E G GG GG G G G R GR R GG R R R R R GGG R R R R R R R R R R R R R R R R R R R R D D D D D D D D L D D L D D D D D D D D L D L D D L D D D L D D D D L L L L L L L L L L L L L L R L L L L L L L L L L L L L L R L L L L L L L L L L L L L L L L L L L D D L D D D D D D D D D D D D D N D D D D D D D D N D D D D D N N N N N N N N N N N N N L N N N N N N N N N N N N N L L L L H L L L L L L L L L L L L L L L WL L L L R QQ L Q QQQQQ QQQQQQQQQQQQQQQQQQ GG Q Q E GC GGGG GGGGGGGGGGGGGGGGG L E E GGE E E E E E G E E E E E E E E E E E E E E E L L E L E L L L L L E E L L L L L L L L L AL L L L L L L L L VVVVVVVVV VVVVVVVVVVVVVVVVVV D D V D D D D D D D D D D D D D D D D D D SS S S S S S SD D SS SSS S S S S SD D S D D D D II I I I I I ISISI I I IK YI IYY YYYI I I ISISISI ISISISISIY YYYYYY YYY YYYYYYYYYYYY GGGGGGGGC E E E GGGGGGGGGGGGGGGGGGG E E E E E E E E E E E E E E E E E E E E E E E E D D E D D D D D D D E D D D D D D D D D D D D D D D D D D R R R R R R R R R R R L R R R R R R L R R R R R R R R R R R L R L L L L L L L L L L L L L R R R R R R C R R R L L L L L L L N L L L L R R R R R R R R R R R R R R R R R E E E E E E E E E E E E E E E E E E E E E E E E E E E E L D L L L L L L L L L L L L L K L L L L L L L L L L L L YYYYYYYYYYYYYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R R R R R R A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S N S A SGGGE GGE GGGGGGGGGE G GGGGGG G R G G R R R R R R GGR G R R R G R R R R R R R R D R R R D R R D R R D R D R D D D D D D D D D D D D D D D D D L D D L D D L D D L QL L L L L L L L L L K L WL L L L L L L L D L L L L L L L L L L L T L L L L L L L L L D L L L L WL L K D D D L D D D D D D D D D D N N D D D D D D N N N N N N D D D D D D N N D N N N N L L L L L L L L G N N N N N N N N N N N L L N L N L L N L L L L L L L L L L QQQLIQQL GGGGGGQQQQN VQQQQL L QQQQQQQQQQQQ GGGG G GGG E GGGGGGP GE GE E GGG E GG E E E E E E E E E E E E E E E E E L E L E L E E L E E L L L L L L L L L H L L L L L L L F VL YV L L L L VVVVVVVVVVVVVVVVVVD VV VVVVVV D D D D D D D D D D D D D D D D D D D D D D D D D D SISISISISISID SSISISISISISISISISISISISSISISSISISISSISIYY YYYIN T S Y YYYYYYYYYYY YYYY Y GY YYY YY GGGGG GG GGGGGGG GGGG G E E E G G GGGG E E E E E E E E E E E GG E E E E E D D D D E E E E E E E E E D D D D D D D D D D D D D R R D R R R D D D D D D D D R R R R R R R R R R R R R R R R R D D D R L R R R R L L L L L L L L L L L L L L L L L L L L L H L L L L R R R R R R L R R R R R R R R R R R R R R R R R R R E E E E E E E E E E E E E E R R R E E E E E E E E E E L L L L L L L L L L L L L L L L E E E E L L L L L L L L L L L L YYYYYYP YYYYYYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R D R R R R A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SGGGGGGE GGGGGGE GGGGGG GGG GGE R R GR R R R R G G R R R R R R R G GG R R R R R R R R R R R D D D D D D D D D D D D D D R R D D R D R D D L L D L D L D D D D D D P D L L L L L L L L L L L L L L L L L L L L L L L L L L L L L TL L L L L L L L D L L QL L L L L L L L L L L L E D D D D D D D D D D D D D D D D D D D D D D D D D D N N N N N N D N N N N N N D N N N N N N L L L L L L L L TN N N N N N N N L L L N L QN L A N L L L L N L L L QQQQQQLP QQQQQQL GGGGGGGGGQQQQ L L L QQQQQQQTQQQQ GGGG GGGGGG GGGGGGGG E E E E E E E E E E E E E G L L L L L L L L L L L L E E E E E E E G E E E E E E E E E L L L L VVVVVV VVVVVVL L L L L L L L L L L L D VVVV L V VVVVVVVVVVVV D S D ISD D ISISD D D D D ISISISISISD ISD ISD ISD ISID D D SISISD ISD ISD ISD D IS SD ISD ISD D ISISD D ISISD ISYY Y Y YYSIY YI IY Y YYYY Y Y YYY YY Y Y GGGGGGGY Y YY Y Y GGGGGG GGGGGG GGG GG E E E E E E E E E E E G G GG E E E E L G E E E E E D D D D D D D D D D E E E E E E E E D D D D QD D R R D D D D D D VD D D D D R R R R R R R R R R R R R R R R R R L L L L L R R R R R R R R R L L L L L L L L L L L L L L S L L L L L L L L R R R R R R R R R R R R RLIR R R R R R R R R R R R R R R E E E E E E E E E E E E E E E E E E E E E E E E E E E E E L L E L L L L L L L L L L L L L H L L L L L L L L L L L L YYYYYYYYYYYYYYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R R R R R R R A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SGE GGGGGGGGGGGGGGG GGGE G G GG R G R R R R R D R R R R R R R R R R R GR R R G G GR R D D D D D R R R R R R D D D D D D D D D D D D D L D L L L L L L L L D L L L L L D L D D D D D D D L L L L L L L L L L L L YL L L L L L L D D D D D D D L P L L L L L L L L L L L D D D L L P L L D L D P D D D D D D D D D D D D D D D N N N N N N L L L L N N N N N N N N N N N D L N L L N N N N N N N N N N L L L L L L L QQQQQQL L L L L L Q L L L L L L L L GGGGGGQQQQQQQQQ QQQQ Q Q E GGGGGGGG GG Q QQ G Q QQ E GE E GE GGGGGGGGG E E E E E L E E E E E E E L E L L E E E E E E E E E E L L L L L VL L L L L L L L V L GL L L L L L L L L VVVVV VVVVVVVVVD VVVVVVVVVVVV D D D D D D D D D D D D D D D D D D D D D D D D D D SISISISISIS ESISISISISISISISISISISISISISISIDSISISIS SITIY YYY YY Y Y YYYYSI IY Y Y Y YYY YY Y GGGGGGGGG GG YY Y Y Y GGGG G GGGGGYGGG GG E E E E E E E E E E E E E H E E E E E GE E G D D D D D D D D D D D D D D D E E E P E E D R D R D D ADEIGD D D D D R R R R R R R R R R R R R R R L L L L R R R R R R R R R R R L L L L L L L L L L L L R R R R R R R R R R L R L L L L L L L L L L L R R R R R R R R R R R E E E E E E E L E R R R R R R E E E E E E E E E L E E E E E E E E E E E L L L L L L L L L L L L L L L L L L L L L L L L L L YYYYYYYYYYYYYYYY D Y D YYYYYYYYYYY D D D D VD D D D D D D D D D D D D D D D D D D D D R R R R R R R R R R R R R R R R AR AR R R R R R R R R R R A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S W YA S A S A S A S A S A S A S A S A S A S A SE GG GGGE GE G GGGGG GGGGE GGGGGGG R R G R R R RGIR GR GP R R R R R GR R R GR R R R D D D D D D L L L D D R D D D D D D D R D D D D R R R R R L D D D D D D D D D L D L L L L L L L L L L N L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L D D D D D D D D D D D D D D D D D D D D L D D D D D D D N N N N N N D N N N D N N N L L N N N N N N N N L E L L L L N L L L L L L L N L L N N S L N L N N N L L L L N L QQQQQQQL L L L S GG QQQQGQQQQQQQQ QQQQQQQ G G G QG G G E E GGG E E E E GGGGG GGGGGGE GGG GGGE GE L E E E E E E E E E E E E L E E E E E E E L E L L L L L L L L L L L L L L L L L L L L L L L VVVVVVVVL L L VVVVVVVVTVVVVVVVVVVVV D D IID D D D D D D D D D D D D D D D D D D D D D D D D SS S SISISISISIDSISISISISISISISISISISISIDSISISISISISISIYYIYYYYSYYIYYYYYYYYYYYSIY YY Y GG G G Y Y YY Y G P GGGGGGG G YGGGGG G E E G G E E E E G E E E E E E GG GG E E E E E E N E EGIE G E E E AE E D D D D D D D D D D D D D D Y D D R R D R R R R D D D D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R L L L L GL L L L L L L L L L L L L L L L L L L L L L L R R L R R R R R R R R R R R R R R R R R R R R VR R R R R R E E E E E E E E E E E E E E E E E E E E E E E E L L E L L L L L L L L L L E E E E L L L L L L L L L L L L L L L L L YYYYYYYYYYYYYYYYYYYYYYYYYYYYY D D L D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R R R R R R R A A A S A S A S A S A S A S A S A S A S A S A S E A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S RGGGGGGGGGGGGGGGGGGGG E GGG E G R R R R R R R R R R R R R R R R R G R R R G G R GGR D D D D D D D D D D D D D D D D D R D R R R R R D R D L L L L L L L L L L L L L L D D L L L D D D D L D D D L L L L L L L L L L L L L L L L L L L L GL L L L L L E L L L F L L D D D D D D D D D D D D D D D D D D D D L L L L L D D L D N D D D D D N N L L L N N N N N N N N N N D N L N L N N N N N N N L L L L L L L N N N N L N N N L QQQ L L L L L L L L L L L Q Q QQQQ Q L L L Q L Q GGG QGQQQQ Q QQQQQF QQ QQ GGGG G QG E E E GGE GGGG L L L E L E L L E E E E E E E E GE GGGGGGGGGGGGE L L L L L L L E E E E E L L E E E E E E E L L E L VVV L VVVVVVVVVVV L L L L L L L L L VVVVVVVVVVVVVL V D D D D D MD D D D D D D D D D D D D D D D D D D D D VD SISISISISISISISISISISISISISISISISISISISIS SISISISIS SIDSIYYYYI IYYYYYYYYYWYYYYY Y Y YSIY GGG YY YY Y GGGG GGG GGG GGGGGGYE E E E E E E EIG E E E GGG E E E E E G Y E E GGG E G E E E E E E E E E D D MD D D D D D D D D D D D D D D D D D D R R R R R R R R R R X R R R R R K D D D D D D R R R R R R L L L L L L E L R R R R R R R L L L R L L L L L L L L L L L L TL L L R R R R R R R R R R R R R R R R L R R R R R R R R R R E E E E E E E E E E E E QE R R R E E E E E E E E E E L L L L L L L L L L L E E E E L L L L L L L L L ES L L L R L L L W YYYYYYYYYYYYYYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R R R R R R R AA A S S A S A S A S A S A S A S A S A S A S TA S A S A S A S A S A S A S K S A S A S A S A S A S A S A S A S A SGC GGGGG E GGGGGGGGGGGGGG GGGG R R R R R WR GGR R R R R R R R R R R R R GR R R R G D R R D D D D R R L D D D L L L L D D D D R D D D D D D D D D D D D D D L L L D D E L L L L L L L L L L L L L L L L LLL L D D D D L L L L L L L L L L L L L L L L LIL L L L L L L L L L L L D D D D D D D D D D D D D D D D D D D D D D D D D N N N N N N N L L N N N N N N N N N N N L L L L L N N N N N N N N N N L L L L L L L L L L L L L L N L L L L L QD Q L QQL QQQQ G GG GGGGGGGQ L L QQ QQQQQR K QQQQQQQQQ G G GGGGGGGGGGG G G E E E G G GG E E E E E L E E E E E L L L E K E E E E E E E E E E E L E L E E L L L L L L L L L L L L L L L L L L L L L L L VVVVVVVVVVVV D D VVVVVVVVVVVVVVVVV D D SSD S D D SSID S D S D S D S D D S SSISID S D S D D SSD S D S D D SSD S D S D D S D D S D D II I I I I I I IHI I I I I I I I I ISISI IS SIY YYYYYYYYYYYYYY YYI IY Y Y Y Y Y Y G GGGGGGG Y Y L Y Y GGGG GGGGGGGGGGGGGGGGG E E E E QE E E E E E E E E E E E E E E E E E E D D D D D D D D D D D E E E E E D D D D D D D D R R R R R R R R R R R R L R D D D D D D D D D D R R R R R R R R R R R R R R R R L L L L L L L L L L L L L L L L L L L L L L L L L L L L R R R R R R R R R R R WR R R R R R R R R R R R QR K R R E E E E E E E E E E E E GE E E E E E E E E E E E E E E E L L L L L L L L L L L L L L L L L L L L L L L L L L L L L YYYYYYYYYYYYYYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D D D D D D D D GD R R R R R R R R R R R R R R R R R R R R R R R R R R R R R A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SGGEP GGGGGGGGE G G GGGGG GE GGG G R R R R R R R R R R R GR GR GR R R R GR GR R R G D D R D R R D R R R D R M D L L D D D D D D D D D D D D D L L L L L L D D D L D L D L L D L L D L D L L L D D L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L D D D D D L L L L H L N D D D D D D L L L D D D D D D D D D D D D D D D D N N N N N D N N N N N D N N L L N N L N N L L N NNL N L N L L N N N N N L N L N N L L L L QQ QQ QL LIL L QQQ Q QMQL L L L L QQQQQL L L L QQ Q Q Q Q QQQWQQ GG G G G G G E E GGGGGGE GGGE GE GGE GGGGGG GGGG L E E E E E E L E E E L E E E L E E E E E E E E E E E N L L L L L L VL L L L L L L VL L L L L L L L L L L VVVVVVVVD VVVVVVVV VVVVVVVVVVV D D D D D SSIS S SD S D D SSISID S D S D D S S D D SSID S D D D D D SSISD D D D S D ISD D D II I I I I I I I I I ISISI IS SIIIS SISIYYYYYYYYY F YYY Y YYYYYI IY YIY YY GG YY Y Y Y Y E E GGG E GGGG E GGGGGG E GGGGGG E GGGG E GGGG D E E E E K D E E E E E E E E E E E R D D D D D W E E D D D D D D E E E D D D D D D D D D D D D R R R R R R R R R R R D D D D R R R R R R R R R R R R R R R R R L L R R L L L R L L L L R L L L L L L L L R L L L L L L R L L L R L L L R R R R R R R E E L L E R R R R R R R R WE E E E E E R R R R R E R R E E E E E L L L L L L L L E E E E E E E E E E E E E L L AL L L L L L L N L L L L L L L L YYYY Y YY D YYYYYYYYYY D D D D D D D D D D D YYYYYYYYY YYY D D D D D D D D D D D D MD D D D R R R R AAR R R R R R R R R R R R AR R R R R R R R R R R R R A S GA S A S A S A S A S A S A S A S A S A S A S A S AA S A S A S A S A S A S A S A S A S A S L A S A S A SGE GGG R R GG R GGE GGG R GGGGG GGG E R GR R GGG R GGGGR G R L R R R R R R R R R R R R D D D D R R D R R D R R D D D L L L D L D D D D L D D D D D D D D L D D D D D D L D L L L L S L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L D L D D D D D D D D D D D D D D D E D D D D D D D D D N D N N N N N N N D N N N N N N N N N N N D N N N N N VL N L N L L L L L L L N L V L L L L N L QQQ QQQL L L L L L L L L L L L QL Q Q QQQQQQQQQQQQQQQQQQQGQ GGGG G E GE GGGGGGG G G G E GGGG GE GGE GGGGE G E E E L E L E E E E E E E E E E E E L E E L E E E E L E L L L L E L L L L L L L L L L P L L L L L V V L L L V VVV V VVVVVVVVVVVVVVVP VVVVVD V D S D D D D D SSIS SID S D F D D D SS S SD S D D SSD S D D SSD D SSID D D SS SID S D D D S D ISI I I I I I ISI I I I I I I I I ISISISISIYI IW F YYYYYYYYYYYYYYYYYYYYYYYYYYY GGGGGGGGGGG GGGGGGGGGGGGGGGGG E E E E E E E E E E E G E E E E E E E E E E E D E E E E E E D D D D R R D D R D D D D D D D D D D D D D D D D D D D D D D H R R R R R R R R R R R R R R R R R R R R R R R R R R L L L L L L L L L L L L L L R R L L L L L L L L L L L L L L R R L R R R R R R R R R R R R H R R R R R R R R R E E H E E E E E R R R E E Y S E L E E E E E E VE E E E E E E E E L L L L L L L L L L L L L L L L L L L L L L L L L L L L YYYYYYYYYYYYYYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R R R R R R R K R R R R R R R R R R R R R R R R A S A S A S A S A S A S A S A F A S A S A S A S A S Q S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SGGGGGGGGGE GGE GGGGG GGGGGG E G R R R R R R R K R GR R GR R R G G R R R R R R R GG D D D D D D D R R R R R R D D D D R D D D D D D D R D L L L L L L L L L D L L D YL L L D L D D D D L D D D L L L L L L L L L L L D L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L D D D D D D D D D D D D D D N N N D D D D D D D D D D D D N N N N N N N N D N N N N N N N N D N L L L L L L L L D N L L N L L L L L NP L N N N N L L L L L N N L QQQQQQQL L L G GG GGGGQQQQ QQQQQ QQ Q QL L QQ Q Q Q Q QE Q G G G G G G G E E E E E E E GGG G GGG GG G G GG E E E E GE E E E E E E E E E E E E GE E L L L L D L L L L E E L L L E VVVVVVVVVL L L L L L L L R L L L L L L L D D D D D D D D D VVVVVVVVVVVVVVVVVL VV II I I I I I I ID D D D D D D D D D D D D D D D D VD D SS S S S S S S S S SISIS SISISISISISISISISISISISISIDSISIYYYYYYYYYIYKIYYYYYYYYYY YSIY G G GGY GY Y Y GG G G GGG GYY GGG GG G G GG E E E E E GG E E G GG G E E E E E E E E E E E E E E TE E GE E D D D D D D D D D E D D E D D D D D D E D R R R R R R R R D D D D D D D D D R R R D R R R R R D R R R R R R R R VR L R R L L L L L L L L L L L R R R L L L L L L L L L L L L L L L L R R R R R R R R R R R R R R R R R R R R R R L R R E E E E E E E E E R E R E E E E E E E E E E E E R E E L L L L L L L E E E L L L L L E L L L YL L L L L L L L L L E L L L YYYYYYYYYF YY D D D D D D YYYYYYYYYYYYYY YY D D D D D D S D D D D D D D D D D D D D Y D D D R R R R R R R R R R R R R R R R R R R R R R R R R R R R AAA R S S S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SAIA S A SGGGGGGGGGGGGGG GG GGGGGGGGGG R R R R R R R R QR R R GR R GR R VR R R R R R R G R R D R R L D D D D D D D D R D D D D D D D D D D D D D D L L D L L L L L L L D L D L L D L L L L L L L D D L L L L L L L L L L L L L L L L L L L L L L L L L L L L L D D D D D D D D D D L L L L L L D D D D D L L D L D D N D D D D D D D D D D D N N N N N N N N N N N N N L N L L L N L L L L L L L L N N L L N N L L N N N N N N N N L L L L L L L L L L N Q QQ QQQ Q L QL Q QQ L Q Q QQ Q Q Q QQQQQ QQ G G G G G Q Q GQ G GG Q E G GG G GG G G G GGGGG G E E E G G E G E E GG L E E E L L L E E L L L L E E L L L E L E E S L E E E L L E E E E E E L L L L L L L L E E L L E VVVVVVVVVVV L L VV L D D VVVVVVVVVVVVVD VVV S D ISD ISD ISD D ISISD D ISISD D D D D ISISISISISISD D D IS SISID D SSD ISD ISD ISD ISD D ISISPSID S Y ISD ISYYYY YYYYYMYYYIYYIYYYYYYYYYIY G Y Y Y E GG GGG G G Y G GGGGGG G G GGGG V E E E G E E E E E E E E E E G E E E G G G E E E E E E E E G E E G D D D D E D E D D D D D D D D D D D D D D D D D D D D R D R R R R R R R R R R R D R K D R R R D L R L L L R L L L L L L L L L R L R R R R R R L R R R R L R R R R R R R R R R R L L L L L L L L L L L L R R L E R R R F R R R R R R R R S L E E E R E E E E E E E E E R E E E E E E E E E K E E L L E L L L L E E E L L L GL L L L L L L L L L L L L L L L L L YYYYYYYYYYYYYYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D D D QD D D D D H R R R R D R R R R R R R R R R R R R R R R R R R R R R R R A S A S A S A S R S A S A S A S A S A S A S A S A S A S E S A S A S A S A S A S A S A S A S A S A S A S A S A S A SE GG R R GGE GGVGGGGGGEIGGGGGG G GG R GR G G G R R R R G R R R R R R R R R R R R R D D R D R R R R R L D D D R D D S D D D D D D D D D D L L L L L L L L L L D D D D D D D D L L L L L L D L L L L L L L L L L L L L L L L L L L L L L L L L L L D D L L L L L L L L L L L L D D D D D D D D D D D D D N N D D V D D D D D D D D D D L N N N N N N N N N N K N N N N N N N N N N N N N L L L L L L L L L L N N L QQL L QL L L L L L L L L L K L L L GGQQQ QQQQQQQQQAQQQQQQQQQQQQ E GGGG E GGGG G G E GE GGGGGGGGGGE GGGGG E L E E E L E E E L E E E E E E E E E E E L E E E E E L VL L L VL L L L K L L L L L L L L L L L L L L L V VVVD VVVVVVVVVVVVVVVVVVVVVV D D SIG S D D SSS D S D D D SS SD S D S D S D D SSD S D D D SS SD D SSD D SSD D D SS SD D SIYI I IAI I I I I I I I I I I I I I I I I I I ISISY YYY YYYY YY YYYYYIP YY YY Y Y YYYYY GGGGGGGGGGGGGGGGGGGGGGGGGG E E G E E E E E E E E E E E E E E E E E E E E E E E E E GS D D R D D D D D D D D D D D D D D D D D D D D D L R R R R R R R R D D D D D R R R R R R R R R R R R R R R R R R R L R L L L L L L L L L L L L L L L L L L L L L L L L L E R R R R GR R R R R R R R R R R R R L R R R R R R L E E E E E E TE E E E E E R R R E E E E E E E E E E E E E E L L L L L L L L L L L L L L L L L L L L L QL L L L L YY D YYYYYYYYYYYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D AD D D D D D R R AR R R R R R R R R R R R R R R R R R R R R R R R R R A S MA S A S A S A S A S A S A S A S A S A S A S A P A S A S A S A S A S A S A S A S A S A S A S A S A S A SE E G GGGGG GGGGE GGGG GG GGGGGGG R GG G R R R R R R R R R R R GR R G R R GR R R R R R R D L D D D D R R R R D D D D D R R D D D N TD D D D D D D L D AD D L L L L D L L D D D L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L D D D D D D D D D D D D D D D D D D D D D D D D D D D D N L N F N N N N N D N N N N N N N L L L N L L NF L N L L N N L L L N L L N N N N N N N N L L L L L L L L L L L L QQQQQQQQQQQQQL L QQQQQQQQQQQQQQQQ GGGGGGGGGG G GG G E GAGGG G E GGG GGGE E GGE G E E E E P E E E E L E E E E E E E E E E E E L L E E L E L L L L L L L L L L L L L L L L L L L L L L V L L VL VVVVVVVVVVVVVVVVVVVVVVV VVVD V D D D D D D D D D D D D D D D R D D D D D D D D D D D SIS S S S S S S S S S S SDS S S S S S S V SIS S SISYI I I I I I I I I I I ISI I IS SI I I I I I I IYYYYYYYYYYYYI I IYYY YYYYYYVYYY G E GG G GG G GY YY G G G G GGG GG G GGGGG E E E E E G E E E E E E GG G G E E E G E E E E E E E E E E D R D D D D D D D D D D D E E E E D D D D D D D D D R R R R R TAR R R D D D D D D R R R R R D N R L R R R R R R R R R R R R L L L L L L L L L L L L L L L L P L L L L L L L L L L L R R R R R R R R R R R L R AR R R R R R R R R R MR E R E E E E E E E E E E E R E E E R E E E E E E E E E E L L L L L L L L L L L L E E L E E L L L L L L L L L L L L L L L L YYYYYYYYYYYYY YYYYYYYYYYYYYYY D F D D D D D D D D D D D YP D D D R D D D D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R R R AR R R R A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S H A S A S A S A SGGGE GGG GGG GGGGGGGGGGGGGG GG R R R GR R R GR R R GR R R R R R R R R R R R R G R D D R R R R R R D D D D D D D D D D D D D D AD D D D L L L D L L L D D D D L L L L L L L L D D L D L L L L L L L L L L L L L L L L L L L L L L L H L VL L L L L L D D D L D D D L L L L L L L L L L D D D D D D D D D D D D D D D H D D D D D D D N N N N N N N N N N N N N N N N N N N N N N N N N L L L L L L L L L L L N N N L L L N L L L L L L L L L L L L QQQL QQ L L Q QQQ QQQQQQQQH QQQQQ Q GGGQF GG Q Q QQ G G G G G E E Q E GGG GGG GGGGGGGGGG GG E E E L L L E E E E E E E TE E E E E E E E E E E E E L L L L E L L L L L L L L L L L L L L L L L L L L L L VVVVVV V V D D VVVVVVVV VVVVVVVVVVVVVD S D D D ISISISD ISISID S L ISD ISD D ISISID D SSD D ISISID D SISV ISD ISD ISD ISD ISD ISD ISD D ISISD D IS SIS YYYYYYIYYYIR YYYY Y YYYQYYYY Y Y GG Y YY YY GGGG GGGGG GGGGG GGGGG G E E E E G GGGG R G E E E E E E E E E E E E E E E E E D D D E E E E E E E E D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R L R R R R R D R R R MR R R R R R R R R R R R R R L L L L L L L L L L L L L L L L L L L L L L L L L L L L R R R R R R R R R R R R R R E E E R R R R R R R R R R R R R R R E L E E ME E E E E E E E E E E E E E E E E E E E L L L L L L L L L L L E L L L L L L L L L L L L L L L L L L YYYYYYYYYYYYYYYYYYYYYYYYYYYGY D D D D D D D D D N D E D D D D D D D D D D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R R R R R R R A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SE GGGE G R R R GR GGGGGGGGGE GGGE GGGGGGGGG R R GG G R R R G R R R R R R R R R R R R R R D D D R D L L L D L D D D D D D D D D R R D D D R D D R R R D D D L L L L L L L L L D D L L L D D L L D D D L D L L L L L L L L L L L L L L L L L L D D D D L L L L L L L L L L L GL L L L L L L L L L L D D D D D D D D D D D D D S D D D D D D D D D D N N N D N L L L N L N N N N L L L N N N L L L L N N D N N N N N N N N N N N N N N L L N N L L L L L L L L L L L L L L QQQL QQYQQQQQQQQL L GG G GG QQQQQQQQQQQQQQQQ G G G GGG G E GN GE GGGGGD MG G E E GG GGGE G E E AE VE E L L L L L L L L E E E L L L L E E E E E E E E E E E E E L E E L VVVVVVVVVVVV L VVL L L L L L L L L L L L L L L VVVVVVVVVVVVVVVV D D D S D D S D S D D SSD S D S D D D D SSIIS S SD D S D S D D SSD S D S D D SSD S D S D D D SSISD SISI ISI I I I I I I I I ISI I I I I I I I I I I ISIYYYIYYYYYYYYYYYYIYYYYYYYYYYYYYYY GGGF GGGGGGGGGGG GGGGGGGGK N GGGG E E E E E E E E E E E E E E E G E E E E E E E E E E E E E E E S D D D D D D D D D D D R D D D D D D D D D D D D D D D D R D D R S R R R R R H R R R R R R L L R R R R R R R R R R R R R R L L L L L L L L L L L L L L L L L L L L L L L L L L R R R R R R R R R R R R L L R R R P R R R R R R R R R R R R R R E E E E E E E E E E E E L E E E E E E E E E E E E E E E N E E L L L L L L L L L L L L L L L L L L L L L L L L L L L L L YYYYYYYYYYYY D YYYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D H D D D D D D D D D D D D D R R R R R R R R R R R R AR R R R R R R R R R R R R R R R R R A S A S A S A S A S A S A S A S A S A S A S VA S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SGGGG GGGGGE GGGGGGGG GGGGGGG G R R R G D R R R R R GR R R R R R R R GR R R R R R R GR L D R R D D DRD D D L L L D R D R D D D D DID D D D L L L L L L L L D D L L D D D D D L L L D D L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L D L L L D D D D L L L L L L L L D D D D D D D D D D D D D D D D D D D D D D D D N N N N N N YN N N N N N MN N N N N N N N N N N N N N L L L L L L L L L L L L L N Q L L L QQ L L L L L L L L L L L QL L Q Q QQQ Q Q G QQ Q Q G GQQQQQ QQ QQQQQQQ G G G G G G G E G GGG G GGGGG GG GGGGGGG E E E E E E E E R E E E E E E E E E E E E E E E E E E L L L L L L E V L L L L L L L L VWL L L L L L L L L L L L VL V L V V D VVVVVVVVD VVVVVV VV VVVVVVV D D SINISD ISD D IS SD ISD ISD D ISISD ISISID S D ISD ISD ISD ISD D ISISID D D SSISD D D D D ISISISISISD Q IS SIYYYIE YYY YYYYYYYYYYIYYYYAYYIY Y G GG GGG G Y YY G G GG G GGGGGGG GGGGGG E E E E E E E E E E E E VGG E E E E E GG E E E E E E E R D D D D D D D D E E E E D D D D D D D D D D D D D D D D D L D R R R R R R R R L R R R L R R R R R R R D R R R R R R R R R R L L L L L L L L L L L L L L L L L L L L D L L L L L L L R R N R R R R R R R R R R R E E E R R R R R E E E E E R R R R R R R R R R E E E E E E E E E E E E E E E E E E E E L L L L L L L L L L L L E YL L L L L L L L L L L L L L L L L YYYYYYYYYYY YYYYYYYYYYYYYYYYY D D D D D D D D D D D D R D D YD D D D TD D D D D D D D D R R R R R R R R R R R AR R R R R R R R R R R R R R R R R A S A S A S A S A S H S A S A S A S A S A S D A S A S A S A S A S A S A S A S A S A S D S A S A S A S A S A S A SGE GGGE GGGGGGGGG GG GE GGGGGG GG R R GR R R R R R R G G G A R R R R R R GR R R R R R D D R D D L L L L D R R D D R D R R R D D F D D D D D D D D D D D L D D D D D D D L L L L L L L L L L L L L L L L D L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L D D D D D D N D D D D D D D D D D L D D D AD D D D D N DS N N N N N N N N AN N N N N N D N N N N N N QN N L L L L L L L L N N N L L L L QQQ L L L L L L L L L L L L L GGGQQQQQQQQQ L L QQ QQ QL QQ L QQ QQ GG G G Q Q Q G QQQ G E E E G E GE GGG GGGGGGGGGE GGGGGG L E E L E L E E E H E E E E E E E E E L E E E E E E E L L L L VL VL L L L L L L L L L L L VL L L L L L L VVMVVD V VVVVVVVVVVVLIV VVVVVVV D D D D D D D D D D D D D TD D D D D D D D D D D D D SISISISISIWISISISISISISISISISISISD IS SISISISISISIS S S SISIYYY YYYIY YYYYYYYYY YI I IY Y Y YY YYYYYYY GGGGGGH GGGGGGGGGG GGGGGGGGGGG E E E E E E E E E E E E E E E E E G E E E E ME E E E E E E D D D D D D D D D D D D D D D D D D R R R R R R R R D D D D D D D D D D D R R R R R R R R R R L L L L L R R R R R R R R R R R L L L L L L L L L L L L L L L L L L L L L L L R R R R R R R R L R R R R R R R R R E R R R R R R R R R R R E E E E E E E E L E E E E E E E E E E E E E E E E E E E E D L L L L L L L L L L L L L L L L L L L L L L L L L L L L YYYYYYYY D YYYYYYYYVYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R R R R R R R A S A S A S A S A S A S A S W S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SE GGGGGGG GE GE G GGGGGGE E E GGGGE GR R R R R R GR GR GGR G R R R R R GGGR R R R G R R D M D D D R D R D R R D R R D D QD R R R D D D D R D L L D L D L L D L D L D D L D D L R L L L D DDIL L D L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L D QD D D D D D D D L L L L L L L L L D L D D D D D D D D D L L D D D D D LF N N N N N N D NDL L L N L L N L L L E N L L NIN N N N N N D D N L L N L N N L L L L L N P N N N N L L L L L L N L QQQQQQQQQQL L L L QQQQQQQQQQQL QQQQQQQ GGGGG G Q G G Q G E GGE GGGGGG GE GGGGGGGGGGE G E E E E E E L E E E E E E E E E E E E E E E E E E E L L E L L L L L L L L L L L L L L L L L L L L L L L L L L L L VVVVVVVVVVVVVVVVVVVVVV VVVVVVV S D D I D D D D WD D D D D D D D D D D D VD D D D D D D SSISIS SISISISISISID SSISIS SID SSISISISISISIDSISISISISIS GSIYYYIYYYYYYNIYYIYIYYYYYYSIYYYY Y G GGGY Y Y YY G G G G GG GL GGG Y GGGG G E E G G GGG G G E E E E E E E E GE E GE E E E E E E E GE E E E E E E D D D TD D D D D E E P D D D D D D D D D D E D D D D D D D R R R R R D R R R R R R R R D QR R R D L R R R R R L L L L L L R R R R R R R R L L L L L L L L L L L L L L L L L L L L L R R R R R R R R R R R R L R R R R R R R R L R R R R R R R E E E E E R L E E E E E E E E R L L L L L L E E E E E E E E E R L L L L L L L LEIE E E E E E E L L L L F L L L L TL L L L YYYYYYYYYYYYYYYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R AR R R R R R R R R R R R QR R R R R R R R R R R R A S A S A S A S A S A S A S A S A S A S A S A S A S T S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SGGGG GGGGG GGGGGE E E G GGGGGG R R R R GR R R R R GR R R R R G GGR R R R R G R D D R R R R R D R D R D D D D D D D D D D D D D L L L L D L L L D L D L L D L D D D D L D L L L L D L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L D D D WD D D L L L L L L L D D D D D D D D D D D D D D N N D R D D N N N N N N N D D N N N N L L L L N L L N L L N N T L N L N N N N N L N N L L L N QQQQQQ QL L L L QL L Q L L L L QL L L QQQ GQGQQQQGQ Q GQIQQQGQQ G Q GGGGL E E E E E E GE GGGGE G G GGGE GGGE GG L L L L L L E L L E E E E E E E E E L L L L L L L L L L E E L L L E E E L E E VL L L VL L VVVVVVVVVVVVVVVVVVVVD VVV VV D S D D SSD D D D D SS K S SID D S D D D SS SID D S D S D S D D DSID D D D D SS S SISE II I I I I ISI I I ISI I I ISIS SI I I ISIYYYYYYYY YYYY YYYI IS Y YYYY AGGGGGGGY E E GGGGGY E GG Y E GG Y GS GGY Y E GGGG E GG E E E E E E E E E E E E E D D D E E E E E D D D D D D D D E E D E D D D D D D D D D D R R R D D D R D D R R R R R R R R R R R R R L R R R R R R R R YR L L L L L L L L L L L L AL L L L L ML L L L L L R R R R R R R L R R R R R R R R R E E T R E E E E E R R R R E E R R R E E E E E E R R E E E E E E E E E L L L L L L L L L L L L L E E E L L L L L L L L L L L ML L YYYYYYYYYYYYMYYYY YYYYYYYYY D D D D D D D D Y D D D D D D D D D D D D D D D D D D D R R R R R R R R AR R R R R E R R R TR R R R R VR R R A S A S A S A S A S A S A S QA S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SGYGGGGE GGGGGGGG GG GGGGE GGGG R R R R R G R R R R R R G G S R R R G G R R R R R R R R D D D D R L L L L D D D D D D R R R R R D D D D D D D D D D D D D L L L L L L L L D D D D D D L L L L L L L L L L L L L L L L L L L L L L L L L L L L D L L L L L L L L S L L L D D D D D D D D L L L L L L D D D D D D M D D D D D D D N W N N N N N N N H D D D D D N N N N N N D N N N N N N N N N N N L L L L L L N L L L L L L N L QQQ L L L L L L QL L L L L GGGGG QQQ QQQ QQL L L L L QQ Q Q Q Q GQQQQ Q QQQQQQ G G G G GG G E GGGE GGQ E GGGGE GE E GGG GG E E E E E E E E L E E E E L E E E E L E L L E E E E E E L L L L L L L L L L L L L L L L L L L L L V VVL L L VVVVVVVVVVVVVVWVVVV V V D VVV VV D S D S D S D S D D SSD D D D SS SD S D S D D SSD S D S D S D D D SS SID D SG SID D S D D D SSA II I I I I I I I I I I I I I I I I I I I I ISISIYYD YYYSIYYYYYYYYYYYYYY YSIYY GGGG Y Y GG GGG GGGGG GG YYY Y G G G G G G GGGG E E E E E E E E E E E E E G G G E E G E E E E D D D D D E E E E D D D D D D E E E E E D D D D F D D R R R R R D D D D D D D D R R R R R R R R R R R R R R R R R D D D R R R R R R R L L L L L L L L L L L L L L L L L L L L L L L L L L L L R R R R R D L R R X R R R R R R R R R R R R R R R R R E E E R Y R E E E E E E E E E E R E E E E E E E E E L L L L L L L L L L L L L L L E E E E E E L L L L L L L L L L L L L L YYYYYY YYYYYYYYYYYYYYYYYYD YYY D D D D D D Y D R R R D D D D D D D D D D D D D D D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R R R ARIA S A S A S A S A S A S S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S K A S A S A S A S A S A SGGE GGN G GGGGGE TGGGGGE EF H GGGGG R R R R G R R R R GR R R R R G G R R R R R G R R R R R D D D D D R D R R D R R R D D D D D L L L L D D D D D D D L D D L D D D D L K D D L L L L L L L L L L L L L L L L L L D D L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L D D D D D D D D D D D D D D D L D D D D D D D D D N N D N N N N N N N N D N N N N N N D N N N N N N N N N L L N L L L L N L L L L L L N L L L L L L L N L L GGQQQQQ QQ L L L L L QL L L QQ QQQ QQQQQQQ QQQQQ QQQ GGG G Q G G E E GGGGG E F E E E E VG GGGG G G GGY GGG L E E G G E L E E E E E E E ME E E E E E E L E E E L L L L L L L L L L L L L L L L L L L L L L L VL L L VVVVVVVVVVVVVVVVVVV VVVVVD VVV D S D ISID D SSD ISD ISD D ISISD D ISISID S D ISID D S D ISD ISD ISD D V ISISID D D SISD ISK ISD ISIS D VSD ISD ISIYYIYYY YYYYYSIYYYYYYSIYYYYYY G Y Y E GG GY YYY G GG GGG GGGY GYGGGGG G E G G G G GG E E E E E E E GE E E E E GE E E E E E D D E E E E E E E D D D D D D R R D D R D D E D D D D D D D E D D D D D D D D D D R R R R R F R R R R R R R R L L R R R R R R R R R R R L L L L L L VL L L L L L L L L L R L L L L L L L L L R R R R R R R R R R R R R R R R R L R R R R R R R R R E E R E E E E E E E E ERIE E E E E E R E E E E E E E E E L L E L L L L L L L L L L L L L L L L L E Y YWYH YYYYYYYL L L L L L L L L L YY YYYYY YYYYYYYYY D K D D D D D D D D D D D D D D D D DYID D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R R R R R R A S AR S P S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SG GGGGGGGGGGGGGGGG GGGG GGGGG R G R H R R R R R R R R R R R R R R R GR R R R GR R R D D D R R R R D D D D D D D D D D D D D D D H D D D D L L ML L L L L L L L L D D D D D L L L L L L L L L D L L L L L L L L L L L L L L L L L L L L L L L L L L L L L L ML L L L D D YD L L D L D D D D D D D D D D D D TD D D D D D D D GD N N N N N N N N N N N N N N N N N N N N N N N N N N N N L L L L L L L L L L L L L L L L L L L L L L L N L L L L L L QQQQQQQQQQQQQQQQQQQQQQQ Q GGGGGGGGGGH G GGG QQQQQ G GGGGGGG G E E E E E E E E N E E E E E E E GGGGG E E E E E E E E E E E E L L L L L L L L L L L L L L L L L L L L L L L E L L L L L L VVVVVVVVVVVVVVVV D D D D D D D D D D D D D D D VVVVVVVL VVVVV D D D D D D D D D D D D D D SISISISISISISISISISISISISISIAISISISISISISISISISISISISISISIYYYYYYYYYYYYYYYYYYYYYYYYYYYYY GGGGGGGGGGGGGGGGGGGGGD GGGGGGG E E E E E E E E E E E E E E E E E E E E E E E E E E E E E D D D D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R R R R R R R R R R R L R R R R R R R R R R R R R L L L L L L L L L L L L L L L L L L L L L L L L L L L L R R R R R R R R R R R R R R R R R R E E E E E E E R R R R R R R R R R R E E E E E E E E E L E E E E E E E E E E E E E L L L L L L L L L L L L L L L L L L L L L L L L L L L L YYYYYYYYYYYYYYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R R R R R R R R R R GR R R R R R R R R R R R R A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SGGGGGGE R GG R GGGGGGGGGGGGGG GGGG R R R R R R R R R R D GR R R R GG R R R R Y R R D D R R R R R D D D D D D D D D D D D D D D D D D D L L D L L L D L L L L L L L L L L L D L L L D D D L D D L L L L L L L L L L L L L L L L L L L L L L L L L L L L L D D D L L L L L L L D D D D D D D D D D L L D D D D D D D D D D D D D D D D N N N N N N N GN N N N N N L L N N N N N L L L N L L L L L L N N N L N N N N L N N L L Q GQQ L L L L L L L L L L L L Q Q L Q L QQ QQQ GQQQQQ Q QQQQQQQQQQ GGGGG G GG GG GG GQQ GG GG G G G E E GE E E GE E E E E E E E E E E GE E E G G E E GG L L E L L L L E L L L L E VL L L L L L L L L L L L E E E L L L L E E L V V L V L N GVVVVVD VVVVV V VVVVVVVVVVD VR D S D ISD D D D ISISISISID SYD ISD ISD ISD ISD D ISISD D N D D D D D ISIDISISISISISISD D IS SD ISD ISISID D SSIYYYYYYIYYYYYYYYYY YYYIYYYYIY R G GY Y Y GGGG G G G GG GGGG Y G GG G GG G E G E E GE E E E E E E E G G E E E E G GG E E E E E E E E E E E E D D D D D D E D D D D D D D E D D D D D D D D D D D D D D R R R R R R D R R R R R R R R R R R D R L R R R R R R R R R L L L L L L L L L L L L R R L L L L L L L L L L L L L L L L R R R R R R R R R R R R R R E R R R R R E E E E E E R R R R R R R R E P E E R E E E E E E F E E E R E E E E E L L L L L L L L L L L L L L L L L L L L E E L E L L L L L L D YYYYYYY YYYYYYYYYYYYYL L YYYYYY Y YY D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R D R R R R R R N R R R R R R R R DS R R R MR R AAA R S S S A S A S A S F S A S A S A S A S A S A S A S V S A S A S A S A S A S A S A S A S A S A S A S A S A S A SE GGGGGGGGG GGE G ES GL GGG GGGGGG R R R R R R R R R G R R R GR G G R R R R R R R R G R R R R R R R D D D D D D D YD D D R D D D D D D D D D L L L D D D D L L L L L L D L D D VD L L L L L L L L L L L L L L L L L L L L L L I L L L L L L L L L L L L L L D D D D D D D D L D D D L L L L L L L L L L L D D D D D D D D D D D D D D D N N N N D N N N D D N L L N N N N N R N N N N N N N N N N N L L L L L L L L N N L L N L QQ QQL L L L L L L L L L L L L L QQQ GGGG G QQQ QQ L L QQQQ Q Q QQQQQQQQQQQ GG G G G G G E E E E E E GGG G G GGGGT GGGGG E E E E E GE GE E E E E E E E YE E E E E L L L L L L L L L E L E E L VVVVVVVVVL L L L L L L L L L L L L L L L L L D VVQVVVVVVVVVVVVVVVVV D S D S D S D S D S D D SSD SSID D S D P SSD D S D S D D D SS SD S D D D SS SD S D S D S D S D S D II I I I I I ISI I I ISI I I I I I I I I I ISIYYYYY YYIT Y YIY YYI IY Y Y Y Y YYYYYYYYYYY GGGGGGGGG E E GGGGGTGGG E GGGGGGGGGGG E E E E E E E E E E E E E E E E E E E E E E E E E E D D D D D D D D D R D D D D D D D D D R D D D D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R R R R R L L L L L L L L L L L L L L L L L YL L L L L L L L L L R R R R R R R L R R R R R R R R R R R R R R R R R R R R R R E E E E E E E E E L E E E E E E E E E E E E E E E E E E E E L L L L L L L L L L L L L L L L L L L L L L L L L L L L YYYYYYYYY D YYYYYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R R R WR R R R R R R R R R R R R R R R R R R R A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SE GWGGGE QGGGE GG GE E GGGGGGE GG G R R R R R G D D R R GR R GG R R R G R E R G R G R R R N R R R G G R R R R R D D D D D D D D D G D D D D D L L L L L D L L L L D D L D L L D D D D L L L D D D D L L L L L L L L L L L L L L L L L L L L L L L L L L L L D D K D D L L L D D D D L D L L L L L L L D D L L L L D N N N N N D D D D D N N D N D D D D D D D D D N N D N L L L L L N N N L L L L N N L N N N L L L N N N N N N N N L L L L L L N L L N L L QQQQQL L L GGQQQQQQQQQQQL L QQQQQQL QQ Q GGGE E GGGGGGGGG Q GG Q G G Q QG E E E L E E E E E E E E E E E EGIGGGGG E E E E E E GK GG L L L M L L L L L L L L L L L L L L L L L L LEIE E L L E E L L VVVVVF VVVV VVAVVVVVVVVVVVVVVV D D SISD ISD D D ISISISD ISD D ISISD V ISID D D SSD ISD ISD ISID D D SSISD ISD ISD D ISISD ISD ISD ISD ISD D ISISIYYYYY S YYYYYII IYYYYYIYYYYYYGYYYYY GGGGGGGGGGG GGGGYGGGGGGGGGG G E E E E E E E E E GE E E G E E E E E E E GE D D D D D E E E E R D D D D D D D D D D D E E E E E D D D D D D D D D D D D D R R R R R R R R R R GR R R R R L L L L L L L L L L L R R L L R R R R R R R R R L R L R R R R R R R R R R R L L L R R R R L L L L L L L L L L L R R R R R R R R R R R E E E E E E E R E E E R R E E E E E E E E E E E E E E A L L L L L L L L L E E E E LIL L L L L L L L L L L L L YYYY L YY L L L L L YYYYE YYYY YYQYYYYYYN R Y D D D D WD D D D D D D D D D D Y D D D D D D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R R R R R R AAA R S S S A S A S A S A S A S A S A S A S A S A S A S A S A S S S A S A S A S A S A S A S A S A S A S A S A S A SG R GGGGGGGGGE GGGGGG R R R R R R R R GR R R R R R GMG GGGGGG R R GF R R R R R GG D R L D D D D D R R D D R R R D D D D D D D D D D D D D D L L L L L L L L D L L L L L L D D D L L L L L D D D L L L L L L L L L L L L L L L L L L L L L L L L L L L D L D D D D D D D D L L L L L L L L L L D D D D D D D N D N N N N N N N N D D D D D D D D D D D D N N N N N L N N N N L L L L L L L L N L L N N N N N N N L L N N N QL L L L L L L L L L L L L QQQQQQQ L L Q QQ L Q QQQQ Q QQQQQQ Q GQ GG GGG Q Q QQ G GG E G GG GS GGGG G GGGG VE E E E E E E E E E E E E E E E G GG E E E E E GGG L L TL L L L L E E E E L V L E E E VL L L L L L L L L L L L L L L VVL L L VVVVVVVD D VVVVVVVD VVVVVH V VVV D S D D S D S D S D D SL SISID D SSD S D S D S D D SSD S D D S D S D S D S D S D D D SSIHD S D D ISI I I I I I I I I I I I I ISI I I I I I I IS SY YYYYYYYYYYYYYYYIYYYYYYYYI IY G GG Y YYY GGGG GGGGGGGGGE GGGG G E G GG WG E E E E E E E E E E E E E E E E E E E E E E GG D D D D D D D E E E E E E D D D D D D D D D D D D D D D D D D D D R E R R R R R R WR R R R R R R R R R R R R R R R R R D D L L L L L L L WL L L L L L L L L L L L L L L L L L L R R R R R R R R R R R R R R R R R R R R R R R R R R R R R L L E E E E E E E E E E E E E E E E E E E E E E E E E E E R R L L L L L L L L L L L L L L L L L L L L L L L L L L L E E L L YYYYYYYYYYYK YYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R R R R R R R R R R R R R R R R R R R R R R D D R F A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S L S A SGGGGE GGGGGGGGG GGGGGGGE G GGGG R R R R R GR R R R R G R R R R R R R R R R R R L D D R R R T R R D D D D D D D D L D L D D D D D D L D D D D L D D D D D D D L L L L L L L VL L L L L L L L L L L L L D D D L L L L L L L L L L L L L L L L L L L L L L L D D D D D L L L L L L L D D D D D D D D D D D D D D D D D D D D N N N N N N N N N N N N N N N N N N N N L L L L L ML N N N N N N L L L L N L L L L L L L N L L L L L L L QQQQL L QQQQQQ GG GGQ L QQQQQQQQQQQQQQQQQ GGGG G GGGGGGGGGGGGR GGGGG E E E G E E E L E E E E E W E E E E E E E E E E L L L E L E E E E E L L L L L L L L L L L L L L L L L L L L L L VVVVLS VVVVVVVV VVVVVVVVVVVVYV D D D V D SISD D D ISISISID D D SSISD ISD ISISD D ISISID D SSD D ISISD D D D D ISISISISISISD D IS SD ISD ISD ISIYYYIY YYYIYYYYIYY YYY Y YYYYYYYYYYY GGGGG GGGGGGG GGGGGG G GG E G G G E E E E E E E E E E E E E E G GG E E E E E G E E E E E E D D E E D D D D D D D D D D D D D D D D N R R R R R R R R R R D D D D D D D D D R D R R R R R L L R R R R R R QL L L L L L R R R R R L L L L L L L L L L L L L F L L L L R R R R R R R R R R R R L R R R R R R R R R R R R R R E E E E E R E E E E E E R E L L L E L L E E E E E E E E E E E E E E E L L L L L L L L L L L L L L L L L L L L L L L YYYYYL YYYYYYY YYYYYYYYYYYYYY D Y D D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R R R R R R R YR R P R R R R R R R R R R R R R R A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A SG R K G G R GG D R GGGE GGGG R R GGGGGGGGGGGR GE G R R GR D R R R R R R R R R R R R R R L D D R R L L D D D L L L L D D D R R R R D D D D L L D D D D D D D D D D D D D R D D D L L L L L L L L L L L L L L D L L L L L L L L L L L L L L L L L L L L L L L L AL L L L L L L L L L N D D D D D D D D D D D D D D D D L N N N N N N D D D D D D D D D D D N N N N N D N QL L N N L N N N N N N N N N N N N L L L N L L L L L L L L L L L L L L L L N L L GQQL L Q L Q Q QQ E GGQQ QQQ QQ Q Q QQQ QQQQL Q E GGGGGGQ Q GGGG E GG GG Q Q G E G E GGGGGGGGGE G L E L E E GE E E E E E L E E E L E E E E E E E E E L E VL VL L L L L L L L L L L L L L L L L L L L L L V L D V VVVVVVVVV VVVVVK VVVVVVVVVV S D D LSISID D D SSISD ISD ISD ISID D D D SSISISD D D D D D D D D D D D D D D ISIEISISIRISISISISISISISIPIDSISIYYYIYYYYYYIYYYYYYYYYYYYYYYSIY GGG GGY Y GG GGG GG GGGGY Y GGG G G G G E E E E E G GGG G E E E E E E E E E E R E E E E E E E E E GE E D D D D D D D E D D D D D D D D D D E D R D D D R R R D D D D D DDD R R R R R R R R R R R RIR L L R R R R R R R R R R R L L L L L L L L L K L L L L L L L L L L L L L L L L R R R R R R R L R R R R R R R R R R R R R R R R R R R E E E R R E E E E R E E E E E E E E E E E E E E E E E L E E E L L E L L L L L L E L L L L L L L L L L L L L L L L L L VL YYY D TYYYYYS YYYYYYYYYYYYYYYYYYY D D D D D D D D D D D D D D D D D D D D D D D D D D D D R R R R L R R R R R R R R R R R R R R R R R R R R R R R R A S A S M S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S A S G S A S A S A S A S A S)8 077 07 07 07 09 07 08 077M(- E0- -E - E - E - E - E- -0-0-yti91 E281 9628 E6 E1 E E 535 741 88582998ni579 1 391 0 9ff4130 78 69 65 34 29 07 30 03 A. .3.1.1.1.5.2.1.1.44.2DIQ:E O04142434567890 S N 8884848484848484858 R R R S S R R R R R R R S S S S S S S R S SEEIEIEIEIEIEIEIE REIEI IF F F F F F F F F MF D D D D D D D D D D D N N N N N N N N N N N AA A AAAAAAAAY YYYYYYYYYY DIDID A D DIDID D DIVVI I I I IDL VVVVVVIV VA AN AAAAAAAA QQQQQQQQQ Q QE E E E E E E E E E E D D D D D D D D D D D E E E E E E E E E E E GGGGAGGGG G G R R R R R R R R R R R D D D D D D D D D D L L L L L L L L L L D L L L L L L L L L L L L D D D D D D D D D D D N N N N N N N N N N L L L L L L L L N L L QQQQQQ L QQ Q QGGGGGG Q GG G GE E E E E E E E GE E L L L L L L L L E L L VVVVVV L VV V VD D D D D D E D D D DS S S S M SISD S SI I I I I IY YISISIM YYYYY YYYY GGGGGGMGGGG E E E E E E E E E E E D D D D D D D D D D D R R R R R R R R R R R L L L L L L L L L L L R R R R R R R R R R RecE E E E E E E E E E E L L L L L L L L L L LneYYYYYYYYYYYu D D D DqeDID D D D D D R R R R R R R R R R S A S A S A S A S A S A S A S A S A S A S A SDOCKET NO.: AIP-005WO PATENT

[0260] Briefly, the 1SSM library was transformed into yeast cells and sorted against a range of concentrations of a target antigen bound by the parental miniprotein. The sorting phase included at least two selections without antigen, one that sorted only via yeast display and at least one more that sorted for positive signal in the absence of antigen, which identified nonspecific binding or binding to a signaling reagent. The sorted libraries were recovered, lysed, and the DNA was amplified and sequenced via next-generation sequencing (NGS). The binding affinities values of the protein variants were measured as follows. The number of times each unique sequence was seen in each NGS result (the “counts”) was first normalized by the counts of a control sequence within each sort. These normalized values were then combined across sorts for each unique sequence, normalized to the maximum result (usually at the highest antigen concentration), and a Hill curve was fitted to these normalized values to determine an EC50 value.

[0261] The observed affinity values were used to determine the fold change in affinity for every mutation at every position of the TNFR1-binding miniprotein. From this, the relative importance of each position to structural integrity and binding was determined by using the average fold change as a measure of mutability. Phase 2. Determination of Structural Contribution Model Score for the TNFR1-binding Miniprotein.

[0262] This section describes the analysis of the structure of the polypeptide of interest with the structural contribution model system.

[0263] The structural contribution model mapped the strength of interactions between amino acid residues and used the product of these interactions to establish the degree of stability in the relative position of all pairs of amino acid residues in the protein. In this way, the structural contribution of each residue to the overall structural integrity was calculated.

[0264] Pairwise-decomposable scores were assigned to every atom pair in the focal set of the miniprotein, which comprised every atom in the amino acid sequence of SARDYLERLRDEGYISDVLEGQLNDLLDRGEDEQAVIDYANDFIESR (SEQ ID NO.1). The pairwise-decomposable scores were assigned to atom pairs based on the degree to which their interactions control their relative positions - their rototranslational certainty. The atom pairs were assigned vector values as shown in Table 2. With reference to Table 2, atom “N” is a nitrogen atom, atom “CA” is an alpha carbon atom, atom “C” is a carbon atom, atom “O” is an oxygen atom, atom “OXT” is an XX atom, atom “CB” is a beta carbon atom, atom “OG” is an XX atom, and atom “H” is a hydrogen atom.DOCKET NO.: AIP-005WO PATENT

[0265] Table 2 shows vector assignments for each atom in the analyzed structure. TABLE 2

[0266] The resulting scores were then used to construct a 3D graph (such as an exemplary 3D graph shown in FIG.2D) in which the nodes are atoms and the edges were weighted by the strength of the rototranslational certainty between them. Although the entire 3D graph is not shown here due to size limitations, the resulting exemplary contact map between each atom is shown in the Table 3. Two disconnected atoms that did not interact and thus did not inform the geometric relationship between them were scored at 0, and not specified in the cooperativity interaction matrix. A score of close to 1 was assigned due to a highly specified cooperativity of two atoms having a strong covalent bond, hydrogen bond or packing interaction specifying the location of the bonded pairs, as this indicates likelihood for the two atoms to form the interaction. The assigned covalent scores are shown in Table 4, the hydrogen bond scores are shown in Table 5, the assigned Van der Waals scores are shown in Table 6.DOCKET NO.: AIP-005WO PATENT

[0267] Table 3 shows the contact map for each origin atom within indicated focal chain set. The contact map provided herein is exemplary and a part of a large set. The combined pulse score reflect all interactions between the particular atom pair.T NETA PN A 1 N 2 H H H O C 1 O O H C 2 C H C C O H C N H N H deni99 3 0 1 3 86 4 5 2 0 7 8 28 9 5 9 4 2 9 8 9 7 3 8 7 1 5 4 7 4 58 68 12 80 37 60 37 17 24 8 0 4 5 5 9 6 8 5 be7 95 7 4 9 6 92 9 6 6 0 6 3 9 0 0 76 63 07 57 3 7 2 0 9 merosloC Pc2 S12 6 012 91 1 0.0.101.4 01.4 01.9 03.5 00.2 03.2 01.2 2 4 5 2 7 6 5 4 04 53 51 91 24 u. .101.01.01.02.02.02.01.01.02.03.01.02.0.202.0 l#acnioah 7 7 7 6 O F C 4 4 74 34 4 4 64 64 74 74 54 64 54 34 51 41 74 74 64 74 64 74 64 54 74W50edm 1 otH 2 D B 1 T B 20- o N A H B G E 2 C C H H C C C H 1 N H H G E A X 2 C H H B A EPIH H O C O 1 N N H C C NA:.#OlaNcnioah F C 7 74 74 74 74 74 74 74 74 74 74 74 7 7 7 7 7 7 7 7 7 7 7 7 7T34 4 4 4 4 4 4 4 4 4 4 4 4 4EK EC LniB gm 2 2 1 1 H H 2 1 TO AirotH H O A H 2 H 1 H 2 H E D D G G B B 1 H H 2 H 1 H 2 H A H H Z E D G B X A 1 H 2 H D T 1 H H N N C N C C C O O C C NDOCKET NO.: AIP-005WO PATENT

[0268] Table 4 shows the covalent contact maps for each origin atom within indicated focal chain set. The contact map provided herein is exemplary and a part of a large set.eTNETA PesleruoPcS.90.90.90.90.90.90.90.90.90.90.90.90.90.90.90 1 1.909.0.90.90.90.90.90.90 l#acnioah F C74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74O We5 m2 2 2 1 d 1 1 0 ootH H H H E D D G G B B A H H 2 T H Z D G B X B0- N A N N N N N C C C C C C C N H 2 H 1 N C H 1 H 2 H 2 C C O C HPIA:. l#aOcnioahNF C74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74T E4K EC LniB g O Airm2 2 1 1 otH H H H 2 E D D G G B B 1 T A H H E O AH 2 H 1 H 2 H H H H H H H Z D G B X A D T 1 H 2 1 2 1 2 1 H H N N C N C C C O O C C NDOCKET NO.: AIP-005WO PATENT

[0269] Table 5 provides exemplary portion of a large data set and includes the hydrogen bond map for each origin atom within indicated focal chain set. TABLE 5

[0270] Table 6 provides exemplary portion of a large data set and includes the Van der Waals forces map for each origin atom within indicated focal chain set.T NETA PootG H H G G G H H G D H H H H H H H E H H A H A H H Z N A 1 2 1 2 1 1 2 H 2 1 H 2 H 1 H 1 H 2 1 H 2 H 2 N H H 2 H 2 C H 2 eroc88 88 99 46 61 9 74 97 2 12 34 5 8 63 18 36 6 1 3 9 9 5 3 1 Se4 4 0 9 0 17 8 6 8 6 7 34 3 2 9 9 30 11 83 48 63 69 44 27 34 sl90 9 1 0 7 1 2 9 1 1 0 1 2 0 1 2 0 1 2 1 1 2 9 1 1 9 1 1 0 1 2 2 1 5 6 7 1 0 4 2 1 8 6 3 9 1 2 2 1 8 .38 3 4 2 5 .24 1 2 1 3 1 4 3 3 4 2 u P.0.0.0.0.0.0.0.0.0.0.0.10.0 0.0.0 0.0.0.0.101.0.10.101.0 l#acnioah F C74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 74 7 7 7 7 7 7 7O 4 4 4 4 4 4 4W50em2 2 2 2 2 2 2 1 2 2 1 1 2 2 1 d H H H D H G G G G0- ootH H H H H H H H N AH 1 H 2 H 2 H 2 H 2 H 2 H 2 H 1 H 2 H 2 H E 1 H H E H H H H H H H H H H E E EPIH 2 1 H 1 2 2 2 1 2 2 H H HA:.#OlaNcnioah F 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7T C 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4E6K Eni2 2 1 1C L B gO Airm otH H H H D D G G B B 2 1 T O AH 2 H 1 H 2 H E A H H Z E D G B X A 1 H H 2 H 1 H 2 H H H D T 1 2 1 H H N N C N C C C O O C C NDOCKET NO.: AIP-005WO PATENT

[0271] Where individual atom pairs received multiple scores, the score assigned to their interaction was taken to be the maximum score, as it was assumed that the optima, e.g., hydrogen bonds, were reflective of any considerations of optimal distances in general.

[0272] After the creation of an interaction matrix and iteration over every atom in the focal set, a pulse array was constructed as shown in Table 7. A marginal pulse array that contained all zero values was then constructed. Every atom in the pulse origin set was individually examined in the interaction matrix, the score of every atom with which it interacts was updated in the marginal pulse array to be the greater of its current value, or the product of the focal atom’s pulse score and the interaction score of the atom pair, as shown in Table 7.

[0273] Table 7 provides an exemplary portion of a large data set and includes the origin atom and the corresponding focal chain set and pulse array scores. TABLE 7DOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENT

[0274] Next, the origin set was emptied, and the marginal pulse array and the pulse array were compared. All atoms with a marginal pulse score higher than their pulse score had their pulse score updated to match their marginal pulse score and were then assigned to the pulse origin set, as shown in Table 7. For clarity, Table 7 shows two pulse cycles for a representative set of atoms, but additional pulse cycles are not shown.

[0275] When the pulse origin set was empty, the pulse chain was concluded. This step was iterated over every atom in the focal set of the miniprotein, as shown in Table 7, and the calculated interaction scores within the interaction matrix following the pulse chain are shown in Table 3. Covalent-only weights were then subtracted from the full graph weights, producing a graph in which the node weights indicated the degree to which the position of the scoring atom is specified by every other atom in the molecule, as explained herein.

[0276] At the conclusion of the pulse chain, the origin atom’s 1D array structural contribution model score was the sum of all values in the pulse array, less the origin atom’s original pulse score, which normally has a value of 1. This 1D array is shown in Table 8. As shown in Table 8, the higher the score for particular atom, the higher its contribution (e.g., structural relevance) to the stability of the full-length miniprotein.

[0277] Table 8 provides exemplary portion of a large data set and includes the 1D score array. TABLE 8DOCKET NO.: AIP-005WO PATENTPhase 3. Determination of the Sequence Envelope for the TNFR1-binding Miniprotein.

[0278] This section describes the identification and verification of a paratope to provide a sequence envelope that defines the amino acid residues that contribute to the binding affinity of the test miniprotein.

[0279] The paratope was determined by subtracting the binding affinity of protein sequences including each residue (as determined from the 1SSM library construction and interrogation of phase 1) from the determined measures of structural contribution (as determined by the structural contribution model score in phase 2). The observed affinity values from phase 1 detailed the impact on binding of every mutation, but these effects could have resulted from two phenomena: the actual binding surface being changed, or the stability of that surface being affected. Thus, in analyzing single mutant affinity values alone, the core is identified as important to binding, but its importance is context-sensitive. Thus, subtracting the structural contribution model score from Phase 2 from the binding affinity values of Phase 1 produces a score, shown in Tables 9, 10, 11, and 12, that indicates the non-structural bindingDOCKET NO.: AIP-005WO PATENT relevance of every amino acid to binding activity, thereby allowing for specific identification of the paratope residues shown below. Tables 30-33 show amino acid substitutions and associated scores for each wild-type amino acid residue of the polypeptide sequence SARDYLERLRDEGYISDVLEGQLNDLLDRGEDEQAVIDYANDFIESR (SEQ ID NO:1).

[0280] Table 9 shows amino acid substitutions and associated scores for each wild-type residue 1-12 of SEQ ID NO: 1.TNETA P5 Y4-4-9-1-3-3-9-1-1-2-4-2-1-8-3-1- 02- 6 8 1 2 5 6 5 8 7 6 3 3 2 8 6 15 49 05 76 9 9 5 5 6 2 1 7 6 4 4 3 7 9 9 62 81 9 1 11 0 6 8 5 6 9 3.8 4 3 9 5 16 111 9 2 3 8 2 4 5 0 4 D0. . . . . .-0-2- 02-0-2-0-2.03. .02. . .-0-2-2-1.08.00.0 16 42 36 92 15 63 3 26 30 87 40 9 3 2 6 4 3 4 1 8 6 5 6 7 8 6 4 7 64 66 1 4 9 0 3 5 2 2 6 5 2 7 3 1 6 3 4 10 92 01 2 5.23.68.17.74.3.6.1.3.5.8.4.7.1.7.8.9 1.3 R - 0- - -4-1-4-7-2-3-4-7-0-2-3-1-3- 7 7 8 3 6 3 O 4 4 4 6 9 5 17 60 56 24 6 6 9 2 4 1 4 3 0 9 3 7 2 1 5 0 5 42 27 67 1W9.8 5 065 8 2 6 1 .6 2.0 .10 .09 .9 2.0 4 5 0 9 82.7 3 3 7 . 01.02 A 0 -9.-4.-6-2- 0 1 -4-8-7- 0 1. . .-3-3-2- 0 1.-8-21--PIA8 1 1 3 1 2 3 8 6 1 82 3 0 8 1 8 5 7 1 4 6 2 71:.n 7 o 7 3 7 .55 9 3 0 .26 9 4 .01 8 6 3 2 4 7 4 .71 6 9 5 9 8 8Oitu1- .0 1- .4 2.-2-42.-81- .5.32.-20.-22.8N1 St0 0 0 - 0 0T9iE Et_nsb K Leu oiousC B ditisni_ O AsD Tero T p WmadicaA A A R N D C Q E G HIL K M F P S T W Y VDOCKET NO.: AIP-005WO PATENT

[0281] Table 10 shows amino acid substitutions and associated scores for each wild-type residue 13-25 of SEQ ID NO: 1.87 98 07 96 86 2 4 0 1 5 5 1 0 8 7 5 9 9 9 5 3 8 24 41 5 69 92 4 36 5 5 7 8 8 74 10 4 9 1 9 3 7 1 6 TNETA P19 64 78 84 14 84 02 27 5 7 7 8 3 4 4 1 1 .8.5.3.5.9.5.8 1 4 5 1 9 6 1 6 72.5.5.0.0.6.7.0.3.3.51 D0-2-1- 02-0-2-2-1-2-2-2-0-0-1- 02-2- 6 3 3 9 2 1 2 8 3 5 2 3 4 34 35 53 99 2 8 5 48 2 4 8 8 4 7 8 0 7 5 7 9663 89 01 0 80 26 9 5 .0.9.4.4.8.0 7. .4.1 1. .364 7 2 2 6 5 50.7.5.8. .8.00 0 1 S- - - - - - -7-1-7-8-9-9-7- 09-1-1- 47 7 9 1 0 99 7 2 6 7 3 2 0 15 14 40 39 10 .9.42 72.4 6 6 53 72.9.52.65.680.8.4.41I - - - -0- -1-2-4-6-O 18 10 55 73 19 07 17 53 82 5 1 4 7 1 7 0 8 2 2 7 0 2 6 2 8 7 26 26 82 33 17 8 3 1W.0 8.3.9.5.4.2.5.7.9 2 350 40- 1 Y3.02-3.-3-4. .-2.-6.-5.-5-3.- - -2-2-2-3-2-0 2 2 2 2 2 0 102-PIA8:.n 7 73 11 47 76 79 12 80 52 00 9 3 9 1 4 1 70 o 5 O 3it0.2 .86 4.8 2.5 0.6 0.4 5.3 8.1 1.5 8 9 4.0 .96 2 6.8 2 5.1 5 5 1.6 .52 2.1 3.u2-2-2-3-2 4 2 2 2 2 1 1 1 7 3 20N0 1 Gti - -0- - - - - - - - - -t1-T1E E_K Lenso b u C B udito ns_ O AisiD TesTirodiA p WmacaA A R N D C Q E G HIL K M F P S T W Y VDOCKET NO.: AIP-005WO PATENT

[0282] Table 11 shows amino acid substitutions and associated scores for each wildtype residue 26-38 of SEQ ID NO: 1.87 37 66 2 5 88 1 2 6 3 4 6 8 6 6 3 9 8 21 84 8 .357 35 44 88 35 73 9 4.17 31 TNETA P84 25 10 9 6 2 2 9 1 0 5 2 4 7 9 1 4 7 2 .1.1.81 .36 .96 .06 .35 .5 1.6 .49 .49 0 9 9.4.08 4 03 G8-8-2-2-5-4-4- 03- 0 1 -7-7.-94.-35.-83.-76- 0 1 - 0 1.-75.-24- 2 9 44 6 8 1 9 2 0 1 4 2 6 4 5 0 3 0 8 8 9 7 1 77 45 19 3 6 0 7 8 7 1 4 3 2 .4.9.1.2 9.7 2 9 2 5 9 8 4 0 6 92 R2- 01-2-2.-1-02-3.1.-90.-52-2.2.-81 6. .-2-72.-22.-31- .00 40 36 8 7 0 4 7 7 6 5 8 50 9 0 7 39 0 4 3 58 44 00 71 2 8 4 0 0 9 8 9 8 .5.9.7.8 5 6 5 7 6 2 4 0 4 9 7 81 1 1 22.52.21.41.32.70.52.32.32.3.4.6. .72 D- - -0- - - - - - - - -2-2-1-1-1- 0 2 3 4 O 8 1 4 98 8 0 4 7 3 6 0 6 8 9 3 0 8W8.79.14 7 750 7.77 6.5.1.0 11. .6 9.0- 2 L- -6-1- 03- -8-6-2-PIA3 5 8 1 3 96:.n 7 9 3 4 5 0 31 o 1O 6it5.4.1N1 2 Lt 700.8 0.2 .-9.33 03.u -1-7-209-1-2-T1itE E_nsb K Leu oiousC B ditisni_ O AsD Tero T p WmadicaA A A R N D C Q E G HIL K M F P S T W Y VDOCKET NO.: AIP-005WO PATENT

[0283] Table 12 shows amino acid substitutions and associated scores for each wildtype residue 39-47 of SEQ ID NO: 2.T NETA P2 9 8 6 9 93.3.94.4.8 7 53.9 27.36.6.8.5.1.0 07.0 03.4.5.07.4 D -2-5- 08-7-7-1-1-9-1-1-3-4-4-1-8- 5 4 5 9 8 2 6 6 0 0 1 6 7 2 3 9 8 5 3 3 6 7 3 5 1 5 7 13 64 6 9 8 9 2 2 6 6 5 0 6 3 4 7 8 4 2 9 9 3 9 8 6 4 4 8 4 4 5 3 9 9 1 5 10 3 0 4 744 9 2 1.1 5 5 1.4 N2.-2.-02.-3. .-4-0-0. . . .2.10.6.2.12-1-2-0-1-2-2-2-1.01.0 10 62 55 99 21 3 9 8 2 4 5 4 7 7 5 1 6 8 4 2 0 0 6 3 3 8 4 1 2 0 6 1 6 1 8 0 0.4.6. . . .0.9.7.O 4 A 0- -2-3-2-8-5-1-8-W5008 6 5 3 3 2- 4I2 86 1 8 9 6 44 19P5 0 69 56 62 55 04 79A:.93 Y n4.9o2.- -45.-87.-52.-8-4-23.3.3.2 - 0OitN2 o utT1 niitsE Eeu n m b K L oitau C BdiiO Ase so TdisA D T R p WcaA A R N D C Q E G HIL K M F P S T W Y VDOCKET NO.: AIP-005WO PATENT

[0284] From an understanding of the paratope, the identity of amino acid residues unimportant for binding were identified. In this case, the amino acid residues unimportant for binding were found to be located primarily in the loops. To test this, the loop sequences were replaced with larger sequences using one of many inpainting algorithms. The inpainted structures were scored and a selection of diverse loop insertions were chosen. Next, the selected loop insertions were expressed clonally and their affinities were observed by surface plasmon resonance.

[0285] Finally, the paratope analysis and the insertion affinities were combined to produce an annotated representation of the paratope which identifies immutable residues (and what mutations they will tolerate) and where in physical space insertions can be made in the miniprotein. The annotated representation of the paratope was represented as a sequence logo, or a sequence envelope, for a given affinity range. The sequence envelope detailed the mutations, including both substitutions and insertions, that can be made to the miniprotein without compromising the affinity beyond the point of utility. The sequence envelope was defined with respect to a given range of acceptable affinities, which depends upon the particular use of the molecule. Taking a maximum fold change of 10, the following sequence envelope was generated and is presented herein as a list of the tolerable residues at each position of the full-length miniprotein, namely SARDYLERLRDEGYISDVLEGQLNDLLDRGEDEQAVIDYANDFIESR (SEQ ID NO:1): Position 1 substitutions: M, D, S, G, A, K, P, V, I, L, Q, H, F, and N; Position 2 substitutions: S, G, A, T, and V; Position 3 substitutions: K, S, A, T, Y, V, R, Q, and H; Position 4 substitutions: M, D, G, K, A, S, E, T, P, W, I, Y, R, Q, H, F, and N; Position 5 substitutions: M, T, Y, W, I, V, L, Q, H, and F; Position 6 substitutions: W, L, Q, and I; Position 7 substitutions: M, D, G, K, A, S, E, T, Y, W, I, V, L, R, H, F, and N; Position 8 substitutions: M, D, G, K, A, S, E, T, P, W, Y, L, R, Q, H, F, and N; Position 9 substitutions: M, L, and I; Position 10 substitutions: M, K, S, A, T, Y, V, L, R, Q, H, F, and N; Position 11 substitutions: D, S, T, W, L, R, H, F, and N;DOCKET NO.: AIP-005WO PATENT Position 12 substitutions: D, G, K, S, E, T, V, W, R, Q, and M; Position 13 substitutions: M, D, G, K, A, S, Y, I, L, R, Q, H, F, and N; Position 14 substitutions: M, D, G, K, A, S, E, T, Y, P, I, V, L, R, Q, H, F, and N; Position 15 substitutions: H, R, P, and I; Position 16 substitutions: D, N, and S; Position 17 substitutions: D, G, K, A, S, E, T, P, W, I, Y, V, L, R, Q, H, F, and N; Position 18 substitutions: V and I; Position 19 substitutions: D, K, L, R, and N; Position 20 substitutions: M, D, K, S, A, E, T, Y, W, I, V, L, R, Q, H, F, and N; Position 21 substitutions: M, D, G, K, A, S, E, T, Y, W, I, V, L, R, Q, H, F, and N; Position 22 substitutions: A, Y, W, L, Q, and M; Position 23 substitutions: L and Y; Position 24 substitutions: D, A, P, G, T, Y, I, Q, M, H, N, S, V, W, R, F, K, E, and L; Position 25 substitutions: M, D, S, A, E, T, Y, W, V, Q, H, F, and N; Position 26 substitutions: L, V, and I; Position 27 substitutions: M, L, V, and I; Position 28 substitutions: M, D, G, K, A, S, E, T, Y, W, I, V, L, R, Q, H, F, and N; Position 29 substitutions: M, D, G, K, A, S, E, T, Y, W, V, L, R, Q, H, F, and N; Position 30 substitutions: D, N, and G; Position 31 substitutions: M, E, R, and G; Position 32 substitutions: M, D, G, S, A, E, T, P, W, I, Y, V, L, R, Q, H, F, and N; Position 33 substitutions: M, D, S, E, T, P, W, I, Y, V, R, Q, H, F, and N; Position 34 substitutions: M, G, K, A, S, E, T, Y, W, L, R, Q, H, F, and N; Position 35 substitutions: A, H, and R; Position 36 substitutions: D, G, V, I, and M; Position 37 substitutions: K, T, V, I, L, and M;DOCKET NO.: AIP-005WO PATENT Position 38 substitutions: L, D, and Y; Position 39 substitutions: R, T, Y, and G; Position 40 substitutions: K, A, T, I, and L; Position 41 substitutions: M, K, S, A, E, T, Y, W, I, V, L, R, Q, H, F, and N; Position 42 substitutions: R and D; Position 43 substitutions: T, Y, W, V, and F; Position 44 substitutions: T, V, I, L, and M; Position 45 substitutions: M, D, G, K, A, S, E, Y, W, I, V, L, R, H, F, and N; Position 46 substitutions: M, D, S, G, K, E, T, P, W, I, Y, V, L, R, Q, H, F, and N; and Position 47 substitutions: K, E, P, W, I, Y, R, H, F, and N. EXAMPLE 3 – IDENTIFICATION OF STRUCTURAL CONTRIBUTIONS TO BINDING AFFINITY OF A FRAGMENT OF A TNFR1-BINDING MINIPROTEIN

[0286] This Example describes the process for identifying the structural contributions of a fragment of a TNFR1-binding miniprotein structure using the structural contribution model algorithm.

[0287] This section describes the analysis of the structure of the polypeptide of interest with the structural contribution model system, as explained in Example 2. As referred to herein, a position of an amino acid within a polypeptide is denoted as “xN”, where “N” refers to the numerical location of the amino acid. The structure of the polypeptide of interest analyzed herein comprised the amino acids at positions x1, x11, x18, x28, x34, and x41 of the full sequence miniprotein of Example 2, which comprises the substituted amino acids of S, D, V, D, Q, and N, respectively. The structure of the polypeptide analyzed herein comprised the amino acids found in the tertiary structure of the polypeptide, and not in the linear amino acid sequence. Thus, the amino acids of the polypeptide of interest analyzed herein comprised x1=S, x11=D, x18=V, x28=D, x34=Q, and x41=N, and are referred to in this example as S-D- V-D-Q-N.

[0288] Pairwise-decomposable scores were assigned to every atom pair between the various atoms of amino acids S, D, V, D, Q, and N in the focal set of the miniprotein based on the degree of their rototranslational certainty. The atom pairs were assigned vector values as shown in Table 13. In reference to Table 13, atom “N” is a nitrogen atom, atom “CA” is anDOCKET NO.: AIP-005WO PATENT alpha carbon atom, atom “C” is a carbon atom, atom “O” is an oxygen atom, atom “OXT” is an XX atom, atom “CB” is a beta carbon atom, atom “OG” is an XX atom, and atom “H” is a hydrogen atom.

[0289] Table 13 shows vector assignments for each atom in the analyzed structure. TABLE 13

[0290] The resulting scores were then used to construct a 3D graph (not shown) in which the nodes are atoms and the edges weighted by the strength of the rototranslational certainty between them. Although the 3D graph is not shown due to its size, the resulting contact map between each atom is shown in the Table 14. Two disconnected atoms that did not interact and thus did not inform the geometric relationship between them were scored at 0, and not specified in the cooperativity Interaction Matrix. A score of close to 1 was assigned to a highly specified cooperativity of two atoms having a strong covalent bond, hydrogen bond or packing interaction specifying the location of the bonded pairs, as this indicates likelihood for the two atoms to form the interaction. The assigned covalent scores are shown in Table 15, the hydrogen bond scores are shown in Table 16, the assigned Van der Waals scores are shown in Table 17. Where individual atom pairs received multiple scores, the score assigned to their interaction is the maximum score, as it was assumed that the optima, e.g., hydrogen bonds, were reflective of any considerations of optimal distances in general.

[0291] Table 14 provides exemplary portion of a large data set and includes the contact map for each origin atom within indicated focal chain set. The combined pulse score reflects all interactions between the particular atom pair.den 8 i 9 81 74 84 51 84 52 41 23 43 T NE ootX G B A G G X B G XT N A O H C C H 2 O O H 1 O C O H 1 O OA Pden 2 i 4 1 2 0 5 3 7 7 3 1 7 3 2 5 7 35 65 2 8 6 0 3 2 6 2 b meer2 oslo5 44 94 2 5 5 2 3 1 9 5 0 0 0 8 0 4 2 25 64 9 9 8 C PcS1030902015 5 4 7 u. . . . .02.09.06.01.01.01.04.01.01.0 l#an cioah F C 1 1 1 1 1 1 1 1 1 1 1 1 1 1 ed m T T T ootB B G A X X X G G N A O H 1 C H 2 H 1 H C O O O O O H O den 1 7 4 7 7 4 3 i 68 52 81 86 52 12 86 81 0 8 9 1 8 4 0 1 5 9 07 3 b meosler96 70 74 16 70 90 16 74 67 05 89 4 5 4 94 CuoPcS0.06.09.03.05.06.02.05.03.02.02.02.03.02.0 l#an cioah F C 1 1 1 1 1 1 1 1 1 1 1 1 1 1O W5ed m 0oT o0tX G - N A N O N H 1 O N H 2 O C O O O N CPIA #:. lOan ciNoah F C 1 1 1 1 1 1 1 1 1 1 1 1 1 1T 4 E1K EC Lnim T B giroB BO AtG A X B G O A H H H A H H D T H N C C O O C O 1 2 3 H 1 2DOCKET NO.: AIP-005WO PATENT

[0292] Table 15 provides exemplary portion of a large data set and includes the covalent contact maps for each origin atom within indicated focal chain set.TN Eed m T ootA G A N A H 1 C C O P e sleruoPcS9.09.09.09.09.0 l#an cioah F C 1 1 1 1 1 ed m ootB H B H B N A 2 C O 1 C e sleruoPcS9.09.09.09.09.09.09.09.09.09.09.09.09.09.0 l#an cioah F C 1 1 1 1 1 1 1 1 1 1 1 1 1 1O Wem 5 d T 0 ootG H A X B H G A B B0- N A O 3 H O C C 2 H N N N C C CPIA:l#.an OcioahN5 F C 1 1 1 1 1 1 1 1 1 1 1 1 1 1T E1K EniC L B gOim T B B ArotG A X B G H H H A H H D T O A H N C C O O C O 1 2 3 H 1 2DOCKET NO.: AIP-005WO PATENT

[0293] Table 16 provides exemplary portion of a large data set and includes the hydrogen bond map for each origin atom within indicated focal chain set. TABLE 16

[0294] Table 17 provides exemplary portion of a large data set and includes the Van der Waals forces map for each origin atom within indicated focal chain set.4 8 1 5 6 4 4 6 2 4TeN dm TE ootA X B B B B N A H O H 1 H 3 H 2 H 1 H 1 H H H HT 1 2 1 2A P 26 10 8 52 23 76 43 38 20 42 46 20 e 89 74 1 0 2 6 9 6 4 5 6 4 sleroc4 .14 74 87 43 51 82 72 34 24 22 34 u P S 0.30.30.10.10.40.10.10.10.10.10.10 l#acnioah F C 1 1 1 1 1 1 1 1 1 1 1 1 edm ootB B B B B B B N AH 1 H 1 H 1 H 2 H 1 H 3 H 2 H 2 H 2 H 3 H 2 H 3 96 16 53 24 97 71 32 89 17 4 6 5 9 8 8 5 2 9 2 0 3 0 05 20 2 6 e 9 9 9 2 3 9 9 0 0 8 1 95 8 sleroc4 .16 .04 .25 .19 1.0 5 2 2 2 3 2 94 u P S 0 0 0 0 0.204.01.01.01.01.0.101.0 l#acnioah F C 1 1 1 1 1 1 1 1 1 1 1 1 1O Wedm50 ootB G B G B A G G G G G N AH 2 H H 2 H H 2 H 2 H H0- H H H H 3 HPIA:. l#aniOcoahNF C 1 1 1 1 1 1 1 1 1 1 1 1 1 1T 7 E1K EC LniB girm otT B BO A G A X B G O A H H H A H H D T H N C C O O C O 1 2 3 H 1 2DOCKET NO.: AIP-005WO PATENT

[0295] After the creation of an interaction matrix and iteration over every atom in the focal set, a pulse array was constructed as shown in Table 18. A marginal pulse array that contained all zero values was then constructed. Every atom in the pulse origin set was individually examined in the interaction matrix, the score of every atom with which it interacts was updated in the marginal pulse array to be the greater of its current value, or the product of the focal atom’s pulse score and the interaction score of the atom pair, as shown in Table 18.

[0296] Table 18 provides an exemplary portion of a large data set and includes the origin atom and the corresponding focal chain set and pulse array scores. TABLE 18DOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENT

[0297] Next, the origin set was emptied, and the marginal pulse array and the pulse array were compared. All atoms with a marginal pulse score higher than their pulse score had their pulse score updated to match their marginal pulse score and are then assigned to the pulse origin set, as shown in Table 18.

[0298] When the pulse origin set was empty, the pulse chain was concluded. This step was iterated over every atom in the focal set of the miniprotein, as shown in Table 18. Covalent- only weights were then subtracted from the full graph weights, producing a graph in which the node weights indicated the degree to which the position of the scoring atom was specified by every other atom in the molecule, as explained herein. The calculated interaction scores within the interaction matrix following the pulse chain are shown in Table 18.DOCKET NO.: AIP-005WO PATENT

[0299] At the conclusion of the pulse chain, the origin atom’s 1D array structural contribution model score was the sum of all values in the pulse array, less the origin atom’s original pulse score, which normally has a value of 1. This 1D array is shown in Table 19. Accordingly, as shown in Table 19, the higher the score for particular atom, the higher its contribution to the stability of the full-length miniprotein of Example 2. Conversely, the closer the score is to 0, the lower the atom’s contribution to the stability of the full-length miniprotein of Example 2.

[0300] Table 19 provides exemplary portion of a large data set and includes the 1D score array. TABLE 19EXAMPLE 4 – IDENTIFICATION OF STRUCTURAL CONTRIBUTIONS TO BINDING AFFINITY OF A FRAGMENT OF A TNFR1-BINDING MINIPROTEIN

[0301] This Example describes the process for identifying the structural contributions of a fragment of a TNFR1-binding miniprotein structure using the structural contribution model algorithm.

[0302] This section describes the analysis of the structure of the polypeptide of interest with the structural contribution model system, as explained in Example 2. The structure of the polypeptide of interest analyzed herein comprised the amino acids x6, x23, x27, x36, and x40 of the full sequence miniprotein of Example 2, which comprises the amino acid L, L, L, V, and A, respectively. The structure of the polypeptide analyzed herein comprised the amino acids found in the tertiary structure of the polypeptide, and not in the linear amino acidDOCKET NO.: AIP-005WO PATENT sequence. Thus, the amino acids of the polypeptide of interest analyzed herein comprised x6=L, x23=L, x27=L, x36=V, and x40=A, and are referred to in this example as L-L-L-V-A.

[0303] Pairwise-decomposable scores were assigned to every atom pair between the various atoms of amino acids L, L, L, V, A in the focal set of the miniprotein based on the degree of their rototranslational certainty. The atom pairs were assigned vector values as shown in Table 20. In reference to Table 20, atom “N” is a nitrogen atom, atom “CA” is an alpha carbon atom, atom “C” is a carbon atom, atom “O” is an oxygen atom, atom “OXT” is an XX atom, atom “CB” is a beta carbon atom, atom “OG” is an XX atom, and atom “H” is a hydrogen atom.

[0304] Table 20 shows vector assignments for each atom in the analyzed structure. TABLE 20

[0305] The resulting scores were then used to construct a 3D graph (not shown) in which the nodes are atoms and the edges weighted by the strength of the rototranslational certainty between them. Although the 3D graph is not shown due to its size, the resulting contact map between each atom is shown in the Table 21.DOCKET NO.: AIP-005WO PATENT

[0306] Table 21 provides exemplary portion of a large data set and includes the contact map for each origin atom within indicated focal chain set. The combined pulse score reflects all interactions between the particular atom pair.deni38 7 b 0 7 3 7 7 5 5 9 7 0 92 40 71 52 meer2 osloc9 9 1 3 3 3 3 4 9 1 5 2 9 4 3 5 Cu P S.0.0.01.01.01.02.0 l#acnioah F C 4 3 3 3 3 3 2TN Ee1 1 1 1T dm ootG D 1 D DA H H D H H H H P N A 3 1 C 1 2 1 2 deni69 12 50 91 28 5 75 b meer39 3 91 48 2 6 31 osloCu Pc7 5 S2.0 02.3 01.4 1 01.3 4 02.4 01.8 0.10 l#acnioah F C 2 4 3 3 3 3 2 em T 1 1 dootX D D N A O C H 1 H 2 H 3 H 2 H 3 deni3 64 3 3 1 5 54 5 64 56 9 b me8 osler2 2 2 28 34 03 82 o 6 7 2 5 4 8 2 Cu PcS.201.01.02.0.102.02.0 l#acnioah O F C 5 4 3 3 3 4 2W50e2 dm 1 ootA D0- N H H G HPIA N O 2 3 H C 1A:.#OlaniNc aT 1 oh F C 2 2 2 2 2 2 2E2K EniC L B gO Airm otT G B X A D T O A C C O O C C NDOCKET NO.: AIP-005WO PATENT

[0307] Two disconnected atoms that did not interact and thus did not inform the geometric relationship between them were scored at 0, and not specified in the cooperativity Interaction Matrix. A score of close to 1 was assigned to a highly specified cooperativity of two atoms having a strong covalent bond, hydrogen bond or packing interaction specifying the location of the bonded pairs, as this indicates likelihood for the two atoms to form the interaction. The assigned covalent scores are shown in Table 22, the hydrogen bond scores are shown in Table 23, and the assigned Van der Waals scores are shown in Table 24. Where individual atom pairs received multiple scores, the score assigned to their interaction is the maximum score, as it was assumed that the optima, e.g., hydrogen bonds, were reflective of any considerations of optimal distances in general.

[0308] Table 22 provides exemplary portion of a large data set and includes the covalent contact maps for each origin atom within indicated focal chain set.e sleruoPcS9.09.09.09.09.0 l#an cioah F C 2 2 2 2 2T NeE d m1 T ootD B A N A C H A 2 C C H 1 P e sleruoPcS9.09.09.09.09.0 l#an cioah F C 2 2 2 2 2 ed m ot2 o D G B N A C C O C H 2 e sleruoPcS9.09.09.09.09.09.09.0 l#an cioah F C 2 2 2 2 2 2 2O Wem T5 d0 ootG B X N A H H A 1 C H0- C O H 3PIA:l#.aniOcoahN2 F C 2 2 2 2 2 2 2T E2K EC LnigOim T B ArotG B X A D T O A C C O O C C NDOCKET NO.: AIP-005WO PATENT

[0309] Table 23 provides exemplary portion of a large data set and includes the hydrogen bond map for each origin atom within indicated focal chain set. TABLE 23

[0310] Table 24 provides exemplary portion of a large data set and includes the Van der Waals forces map for each origin atom within indicated focal chain set.ootA N A H H 1 H 3 H 2 H 2 H 1 H 1 75 7 1 1 9 7 7 esler1 9 7 3 37 92 44 13 63 oc4 2 .14 8 .12 1 .13 37 21 62 u P S 0 0 0.20.30.20.10 l#acniaT oh F C 5 4 4 4 4 4 3NE Tem 2 1 2 1 1A d B G G G G G P ootN AH 1 H 2 H 2 H 1 H 3 H 3 H 3 87 6 9 5 2 7 7 e 0 0 8 3 9 0 6 3 8 7 4 0 9 sleruoPc3 9 S2.7 0.13 2 01.7 0 01.0 4 02.4 55 0.101.0 l#acnioah F C 5 5 4 4 4 4 4 e 1 2 2 2 1 dm ootB G G G G G N AH B 2 C H 3 H 2 H 1 H 1 H 1 94 70 42 70 8 8 1 7 8 7 5 7 9 3 esler6 8 7 1 54 00 50 oc3 1 6 1 2 1 3 7 5 8 u P S.0.0.01.01.01.01.0 l#acnioah O F C 5 5 4 4 4 4 4W50e2 2 2 2 1 dmB G G G G0- ootB G H H H H H H HPIN A 3 2 1 3 2 2 3A:.#OlaniNc aT 4 oh 2 F C 2 2 2 2 2 2 2EK EniC L B gO Airm otT G B X A D T O A C C O O C C NDOCKET NO.: AIP-005WO PATENT

[0311] After the creation of an interaction matrix and iteration over every atom in the focal set, a pulse array was constructed as shown in Table 25. A marginal pulse array that contained all zero values was then constructed. Every atom in the pulse origin set was individually examined in the interaction matrix, the score of every atom with which it interacts was updated in the marginal pulse array to be the greater of its current value, or the product of the focal atom’s pulse score and the interaction score of the atom pair, as shown in Table 25.

[0312] Table 25 provides an exemplary portion of a large data set and includes the origin atom and the corresponding focal chain set and pulse array scores. TABLE 25DOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENT

[0313] Next, the origin set was emptied, and the marginal pulse array and the pulse array were compared. All atoms with a marginal pulse score higher than their pulse score had their pulse score updated to match their marginal pulse score and are then assigned to the pulse origin set, as shown in Table 25.

[0314] When the pulse origin set was empty, the pulse chain was concluded. This step was iterated over every atom in the focal set of the miniprotein, as shown in Table 25. Covalent- only weights were then subtracted from the full graph weights, producing a graph in which the node weights indicated the degree to which the position of the scoring atom is specified by every other atom in the molecule, as explained herein. The calculated interaction scores within the interaction matrix following the pulse chain are shown in Table 21.

[0315] At the conclusion of the pulse chain, the origin atom’s 1D array structural contribution model score was the sum of all values in the pulse array, less the origin atom’s original pulse score, which normally has a value of 1. This 1D array is shown in Table 26. Accordingly, as shown in Table 26, the higher the score for particular atom, the higher its contribution to the stability of the full-length miniprotein of Example 2. Conversely, the closer the score is to 0, the lower the atom’s contribution to the stability of the full-length miniprotein of Example 2. Here, as shown in Table 26, all atoms had scores indicating high contribution to the stability of the full-length miniprotein of Example 2.

[0316] Table 26 provides exemplary portion of a large data set and includes the 1D score array. TABLE 26DOCKET NO.: AIP-005WO PATENT EXAMPLE 5– IDENTIFICATION OF STRUCTURAL CONTRIBUTIONS TO BINDING AFFINITY OF A FRAGMENT OF A TNFR1-BINDING MINIPROTEIN

[0317] This Example describes the process for identifying the structural contributions of a fragment of a TNFR1-binding miniprotein structure using the structural contribution model algorithm.

[0318] This section describes the analysis of the structure of the polypeptide of interest with the structural contribution model system, as explained in Example 2. The structure of the polypeptide of interest analyzed herein comprised the amino acids x19, x23, x27, x33, x36, and x40 of the full sequence miniprotein of Example 2, which comprises the amino acid G, Y, I, S, F, I, respectively. The structure of the polypeptide analyzed herein comprised the amino acids found in the tertiary structure of the polypeptide, and not in the linear amino acid sequence. Thus, the amino acids of the polypeptide of interest analyzed herein comprised x19=G, x23=Y, x27=I, x33=S, x36=F, and x40=I, and are referred to in this example as G-Y- I-S-F-I.

[0319] Pairwise-decomposable scores were assigned to every atom pair between the various atoms of amino acids G, Y, I, S, F, and I in the focal set of the miniprotein based on the degree of their rototranslational certainty. The atom pairs were assigned vector values as shown in Table 27. In reference to Table 27, atom “N” is a nitrogen atom, atom “CA” is an alpha carbon atom, atom “C” is a carbon atom, atom “O” is an oxygen atom, atom “OXT” is an XX atom, atom “CB” is a beta carbon atom, atom “OG” is an XX atom, and atom “H” is a hydrogen atom.

[0320] Table 27 shows vector assignments for each atom in the analyzed structure. TABLE 27DOCKET NO.: AIP-005WO PATENT

[0321] The resulting scores were then used to construct a 3D graph (not shown) in which the nodes are atoms and the edges weighted by the strength of the rototranslational certainty between them. Although the 3D graph is not shown due to its size, the resulting contact map between each atom is shown in the Table 28. Two disconnected atoms that did not interact and thus did not inform the geometric relationship between them were scored at 0, and not specified in the cooperativity Interaction Matrix. A score of close to 1 was assigned to a highly specified cooperativity of two atoms having a strong covalent bond, hydrogen bond or packing interaction specifying the location of the bonded pairs, as this indicates likelihood for the two atoms to form the interaction. The assigned covalent scores are shown in Table 29, the hydrogen bond scores are shown in Table 30, the assigned Van der Waals scores are shown in Table 31. Where individual atom pairs received multiple scores, the score assigned to their interaction is the maximum score, as it was assumed that the optima, e.g., hydrogen bonds, were reflective of any considerations of optimal distances in general.

[0322] Table 28 provides exemplary portion of a large data set and includes the contact map for each origin atom within indicated focal chain set. The combined pulse score reflects all interactions between the particular atom pair.l#acnioah F C 4 4 4 5 4 5 5 5 5 1 4TN Ee1 1 1 1T dm ootA G G B B D GA N A C H H A 2 H H A H H H P C 2 3 H H 2 1 3 deni59 8 1 1 2 7 9 2 8 1 8 7 94 8 3 9 7 3 7 3 8 1 6 5 4 85 2 6 be61 2 5 7 56 3 3 9 6 3 4 merosloc2 2 1 0 3 4 1 4 0 3 1 2 2 2 1 4 5 0 3 49 Cu P S.10.0.0.0.0.0.0.0.01.02.0 l#acnioah F C 4 1 5 1 5 1 5 1 5 1 4 em 1 1 1 2 2 2 dootG D D D D H H D H H A N A 3 1 O C C 3 H 3 H H 3 C deni76 85 4 b 5 4 10 47 7 3 8 9 9 1 67 98 01 0 93 me6 osler43 2 3 2 9 2 7 84 77 46 26 5 9 6 89 25 CuoPcS.101.09.01.01.01.01.01.02.01.01.0 l#acnioah O F C 4 4 5 4 5 1 4 1 4 1 1W5em 1 2 2 2 20 d0- ootA G D H B A D D D HPIN A C 1 C O H C O C O C 1A:. l#OaniN8coah T2F C 5 5 5 5 5 5 5 5 5 5 5E EK LC Bnigm B B TO AirotB H H H A B X A D T O A 3 2 1 H H C O O C C NDOCKET NO.: AIP-005WO PATENT

[0323] Table 29 provides exemplary portion of a large data set and includes the covalent contact maps for each origin atom within indicated focal chain set.l#an cioah F C 5 5 5 4TN EeT d m BA ootA N A H P 1 C C C e sleruoPcS9.09.09.09.0 l#an cioah F C 5 5 5 5 ed m ootB H B A N A 2 O C C e sleruoPcS9.09.09.09.09.09.09.09.09.09.09.0 l#an cioah F C 5 5 5 5 5 5 5 5 5 5 5O We5 d m o B T 0 otB B B A X A0- N A C C C C N H 3 C C O H HPIA #:. ln OaciaN9 o h 2 F C 5 5 5 5 5 5 5 5 5 5 5TE EK LniC B g m TOiroB B B AtH H A B X A D T O A 3 2 H 1 H H C O O C C NDOCKET NO.: AIP-005WO PATENT

[0324] Table 30 provides exemplary portion of a large data set and includes the hydrogen bond map for each origin atom within indicated focal chain set. TABLE 30

[0325] Table 31 provides exemplary portion of a large data set and includes the Van der Waals forces map for each origin atom within indicated focal chain set.1 3 7 2 7 5 3 1 5l#acnioah F C 5 5 4 1 5 4 5 5 4 4 4TN Ee2 dmT D B B BT ootX N A O O H H H H AA O 1 1 O 1 1 O C H P 36 1 5 1 1 4 21 22 0 6 2 7 2 1 4 1 6 2 5 1 1 47 8 39 8 0 8 esler6 3 2 8 6 2 3 6 56 0 95 uoPc2 S.19 0.12 0.14 0.12 6 0.101.9 01.7 4 0.100.1 05.6 01.0 l#acnioah F C5 5 4 5 5 4 5 5 5 4 4 em 1 1 1 d T ootX G B G B B G N A H O H 2 H H 2 H 2 H 2 H 2 H O H 2 51 67 72 51 36 6 2 5 29 17 3 6 1 5 1 7 5 9 8 2 e 4 9 8 2 9 4 5 sler6 oc6 6 5 6 6 6 7 5 0 6 9 1 2 0 6 2 3 9 3 6 3 5 u P S.01.03.01.01.0.10.002.02.0.103.0 l#acnioah F C 5 5 5 5 5 4 5 5 5 4 5O W5e1 10 dm0ootA B B G B B B G B- N A H H O H 3 H 3 H 3 H 3 H 3 H 2 H 2 H 1PIA:. l#OacniN1 oah T3F C 5 5 5 5 5 5 5 5 5 5 5E EK LC Bnigm B B TO AirotB H H H A B X A D T O A 3 2 1 H H C O O C C NDOCKET NO.: AIP-005WO PATENT

[0326] After the creation of an interaction matrix and iteration over every atom in the focal set, a pulse array was constructed as shown in Table 32. A marginal pulse array that contained all zero values was then constructed. Every atom in the pulse origin set was individually examined in the interaction matrix, the score of every atom with which it interacts was updated in the marginal pulse array to be the greater of its current value, or the product of the focal atom’s pulse score and the interaction score of the atom pair, as shown in Table 32.

[0327] Table 32 provides an exemplary portion of a large data set and includes the origin atom and the corresponding focal chain set and pulse array scores. TABLE 32DOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENTDOCKET NO.: AIP-005WO PATENT

[0328] Next, the origin set was emptied, and the marginal pulse array and the pulse array were compared. All atoms with a marginal pulse score higher than their pulse score had their pulse score updated to match their marginal pulse score and are then assigned to the pulse origin set, as shown in Table 32.

[0329] When the pulse origin set was empty, the pulse chain was concluded. This step was iterated over every atom in the focal set of the miniprotein, as shown in Table 32. Covalent- only weights were then subtracted from the full graph weights, producing a graph in which the node weights indicated the degree to which the position of the scoring atom was specified by every other atom in the molecule, as explained herein. The calculated interaction scores within the interaction matrix following the pulse chain are shown in Table 28.

[0330] At the conclusion of the pulse chain, the origin atom’s 1D array structural contribution model score was the sum of all values in the pulse array, less the origin atom’s original pulse score, which normally has a value of 1. This 1D array is shown in Table 33.DOCKET NO.: AIP-005WO PATENT Accordingly, as shown in Table 33, the higher the score for particular atom, the higher its contribution to the stability of the full-length miniprotein of Example 2. Conversely, the closer the score is to 0, the lower the atom’s contribution to the stability of the full-length miniprotein of Example 2.

[0331] Table 33 provides exemplary portion of a large data set and includes the 1D score array. TABLE 33EXAMPLE 6 – EVALUATING EFFECTS OF POINT MUTATIONS ON A GIVEN STRUCTURE

[0332] The structural contribution model was used to evaluate the effect of point mutations on a given structure.

[0333] In particular, residue 71 of polypeptide top7 (1gys) (Kuhlman et al. (2003) SCIENCE302: 1364-1368. DOI:10.1126 / science.1089427) was mutated from phenylalanine (F) to alanine (A). Next, the polypeptide top7 was subjected to the structural contribution model as outlined in the previous examples. Briefly, pairwise-decomposable scores were assigned to every atom pair in the focal set of the top7 protein, or the top7 F71A mutant, based on their rototranslational certainty. The atom pairs were assigned vector values. The resulting scores were then used to construct a graph in which the nodes are atoms and the edges weighted by the strength of the rototranslational certainty between them. Two disconnected atoms that do not interact and thus do not inform the geometric relationship between them were scored at 0, and are not specified in the cooperativity interaction matrix. A score of close to 1 was assigned to a highly specified cooperativity of two atoms having a strong covalent bond, hydrogen bond or packing interaction specifying the location of the bonded pairs, as this indicates likelihood for the two atoms to form the interaction.DOCKET NO.: AIP-005WO PATENT

[0334] After the creation of an interaction matrix and iteration over every atom in the focal set, a pulse array was constructed.

[0335] A marginal pulse array that containing all zero values was then constructed. Every atom in the pulse origin set was individually examined in the interaction matrix, the score of every atom with which it interacted was updated in the marginal pulse array to be the greater of its current value, or the product of the focal atom’s pulse score and the interaction score of the atom pair.

[0336] Next, the origin set was emptied, and the marginal pulse array and the pulse array were compared. All atoms with a marginal pulse score higher than their pulse score have their pulse score updated to match their marginal pulse score and were then assigned to the pulse origin set.

[0337] When the pulse origin set was empty, the pulse chain was concluded. This step was iterated over every atom in the focal set of the protein. Covalent-only weights were then subtracted from the full graph weights, producing a graph in which the node weights indicated the degree to which the position of the scoring atom is specified by every other atom in the molecule, as explained herein.

[0338] At the conclusion of the pulse chain, the origin atom’s 1D array structural contribution model score was the sum of all values in the pulse array, less the origin atom’s original pulse score, which had a value of 1.

[0339] Reference is now made to FIGs.8A and 8B, which depict the results of the use case of the structural contribution model for evaluating effects of point mutations on a top7 (1gys) polypeptide structure. As shown in FIGs.8A and 8B, residue 71 (as indicated by the solid black arrow) of top7 (1qys) was mutated from F (FIG.8A) to A (FIG.8B). FIG.8A illustrates the residues with varying levels of structural contribution to polypeptide’s stability, while FIG.8B illustrates the residues with varying levels of structural contribution to polypeptide’s stability after the F71A substitution. As shown in FIG.8B, the F71A mutation adversely effected not only residue 71 (solid black arrow) but also neighboring residues, as indicated by the hollow arrows, which have reduced scores as a consequence of the removed Van der Waals interactions propagating through the helical turn. EXAMPLE 7 – EXAMINING PROTEIN-PROTEIN INTERFACES

[0340] The structural contribution model was used to examine protein-protein interfaces for structurally important interactions.DOCKET NO.: AIP-005WO PATENT

[0341] Briefly, focal set and medial sets were selected for the analysis. Pairwise- decomposable scores were assigned to every atom pair between the focal set and the medial set of the protein based on their rototranslational certainty. The atom pairs were assigned vector values. The resulting scores were then used to construct a graph in which the nodes are atoms and the edges weighted by the strength of the rototranslational certainty between them. Two disconnected atoms that do not interact and thus do not inform the geometric relationship between them were scored at 0, and are not specified in the cooperativity interaction matrix. A score of close to 1 is assigned to a highly specified cooperativity of two atoms having a strong covalent bond, hydrogen bond or packing interaction specifying the location of the bonded pairs, as this indicates likelihood for the two atoms to form the interaction. The interaction scores within the interaction matrix are shown in a contact map.

[0342] After the creation of an interaction matrix and iteration over every atom in the focal set, a pulse array was constructed.

[0343] A marginal pulse array containing all zero values was then constructed for the focal set and the pulse chain was propagated across the medial set only. Every atom in the pulse origin set was individually examined in the interaction matrix, the score of every atom with which it interacted was updated in the marginal pulse array to be the greater of its current value, or the product of the focal atom’s pulse score and the interaction score of the atom pair.

[0344] Next, the origin set was emptied, and the marginal pulse array and the pulse array were compared. All atoms with a marginal pulse score higher than their pulse score have their pulse score updated to match their marginal pulse score and were then assigned to the pulse origin set.

[0345] When the pulse origin set was empty, the pulse chain was concluded. This step was iterated over every atom in the focal set of the protein. Covalent-only weights were then subtracted from the full graph weights, producing a graph in which the node weights indicated the degree to which the position of the scoring atom is specified by every other atom in the molecule, as explained herein.

[0346] At the conclusion of the pulse chain, the origin atom’s 1D array structural contribution model score for each atom of the focal set was the sum of all values in the pulse array, less the origin atom’s original pulse score, which had a value of 1.

[0347] Reference is now made to FIGs.9A-9C, which depict a use case of the structural contribution model for examining protein-protein interfaces for structurally importantDOCKET NO.: AIP-005WO PATENT interactions. Here, the structural contribution model was used to examine protein-protein interfaces for structurally important interactions, rather than simply moieties in proximity to an interface. This was done by setting the focal set to one amino acid chain and the medial set to another amino acid chain. Accordingly, the focal set and the medial set may be on different, or the same protein. As shown in FIGs.9A-9C, the structural contribution model was ran with the focal chain set to chain B (FIG.9A), chain E (FIG.9B), and then both together (FIG.9C). When ran with one chain set to focal and the other to medial, the structural contribution model effectively measured how well the atoms of focal chain interacted with the atoms of the medial chain, highlighting such interactions as hydrogen bonds to the backbone that are expected to persist through structural perturbation. In contrast, analysis of the entire structure (FIG.9C) allowed for sidewise contacts between interface residues to be examined, which allows further assessment of any possible mutations as for FIG.9A and FIG.9B. The most highly stabilizing residues are shown in FIG.9C (as indicated by solid arrows). Alternatively, the residues most stabilized by the interaction were determined by subtracting the structural contribution model scores of each chain from those of the pair (not shown). EXAMPLE 8 – PERFORMING PARATOPE ANALYSIS

[0348] The structural contribution model was further used for paratope salience analysis.

[0349] FIGs.10A-10C depict a use case of the structural contribution model for performing a paratope salience analysis. This involved using a source of data for residue- specific mutability. As illustrated in exemplary FIGs.10A-10C, the HUTNFR1 binder is shown with the inverse of the average fold change for all substitutions at each position used as a proxy for observed relevance to binding (FIG.10A). As can be seen, the entire core was shown as relevant to binding (exemplary residues of the core are labeled with hollow arrows), likely because mutations to the core disrupt the folding of the protein. By applying the structural contribution model to the structure, FIG.10B shows the amino acids relevant to binding. FIG.10C is the result of taking the difference between the normalized observed salience scores (FIG.10A) and the normalized observed structural contribution model scores (FIG.10B). Thus, FIG.10C shows the degree to which each residue is more important than is structurally explicable (FIG.10C). Specifically, FIG.10C, outlines both a specific face of the protein as the paratope and structural residues that are likely important for binding. These structural residues include (labeled as C) the valine and isoleucine that hold the two interface helices in a specific orientation. Furthermore, the two residues in the upper left (labeled as B)DOCKET NO.: AIP-005WO PATENT are likely to participate in a helix-capping interaction. Notably, the residue (labeled as A) was identified by both the salience score (FIG.10A) and the structural contribution model score (FIG.10B) as relevant to binding. Thus, this residue (labeled as A) was deemed as valuable for the structural integrity of the protein, but not as relevant for binding. This highlights the use of the structural contribution model for differentiating between residues important for binding (e.g., paratope residues) in contrast to residues that are important for structural integrity (and not important for binding). EXAMPLE 9 – STRUCTURAL CONTRIBUTION MODEL MORE ACCURATELY DISTINGUISHES BETWEEN PARATOPE RESIDUES AND NON-PARATOPE RESIDUES

[0350] This Example demonstrates the advantage of using the structural contribution model to distinguish paratope and non-paratope residues.

[0351] In brief, using the methods for characterizing protein stability and determining the paratope of a protein described herein, paratope residues (FIG.11D), salient or paratope residues (FIG.11C), and mutable residues (FIG.11B) were identified on the amino acid chain on Example 2, and shown herein in FIG.11A. Overall, FIGs.11A-11D show more accurate characterization of residues (including paratope residues) using the structural contribution model.

[0352] Of note, the methods disclosed herein successfully distinguished between salient residues (residues that contribute towards the structural integrity of the protein) and paratope residues shown in FIG.11D (residues that contribute towards the binding of the protein to a binding partner). For example, the residue 6 (a leucine) shown as “A” in FIG.11C was identified as a salient residue using the methods disclosed herein. This differs from traditional methods, such as RoseTTAFold, or AlphaFold which would fail to distinguish between salient residues and paratope residues and instead, would categorize both salient and paratope residues together (residues as identified in FIG.11C). Thus, the structural contribution model is capable of identifying salient and paratope residues, while the traditional methods conflate salient and paratope residues, and simply identify them as immutable without being able to discern why.

[0353] Using the structural contribution model, each residue 1-47 of the amino acid sequence SARDYLERLRDEGYISDVLEGQLNDLLDRGEDEQAVIDYANDFIESR (SEQ ID NO: 1) was classified as a salient residue, a paratope residue, or other residue (e.g., a mutable residue non-salient / non-paratope residue). The classifications are shown in the TableDOCKET NO.: AIP-005WO PATENT 34 below. In comparison, the traditional methods (such as, but not limited to, alanine scanning and similar) could not distinguish between the reasons why a residue may be immutable. Thus, using traditional methods, all salient and paratope residues would read the same (identified as “immutable” in the table below). The use of surface exposure to distinguish between salient and paratope residues relies on arbitrary cutoffs at select size ranges, thus the traditional methods are inferior to the structural contribution model, which reliably and effectively distinguishes between salient and paratope residues.

[0354] Table 34 shows residue classification. TABLE 34DOCKET NO.: AIP-005WO PATENTEXAMPLE 10 – EXAMINING PROTEIN-DNA INTERFACES

[0355] The structural contribution model was used to examine protein-DNA interfaces for structurally important interactions.

[0356] Briefly, focal set and medial sets were selected for the analysis. Pairwise- decomposable scores were assigned to every atom pair between the focal set and the medial set of the protein based on their rototranslational certainty. The atom pairs were assignedDOCKET NO.: AIP-005WO PATENT vector values. The resulting scores were then used to construct a graph in which the nodes are atoms and the edges weighted by the strength of the rototranslational certainty between them. Two disconnected atoms that do not interact and thus do not inform the geometric relationship between them were sco...

Claims

DOCKET NO.: AIP-005WO PATENT CLAIMS What is claimed is:

1. A method for characterizing one or more structural features of a molecule of interest, the method comprising: for each of one or more atoms of a sequence of the molecule of interest: (a) performing a pairwise atom to atom interaction analysis across at least one pair of atoms in the sequence to generate a 3D graph, wherein the 3D graph comprises nodes representing atoms and at least one edge between at least two nodes representing at least one interaction between the at least one pair of atoms, wherein the at least one edge comprises a weighted value based at least in part on presence or absence of an interaction between the at least one pair of atoms; (b) iteratively interrogating the 3D graph to determine a structural relevance of each of the one or more atoms of the sequence, wherein iteratively interrogating the 3D graph comprises: (i) for an atom of the 3D graph, assigning cooperativity scores representing a positional certainty to each of one or more other atoms of the sequence, wherein the cooperativity scores are assigned based at least in part on the weighted value of a corresponding edge in the 3D graph between the atom and the one or more other atoms, and (ii) combining the cooperativity scores of the atom across each of the one or more other atoms to determine a structural relevance of the atom; and for each element in the sequence of the molecule of interest, combining the structural relevance of each atom in the element to generate a measure of structural contribution for the element.

2. The method of claim 1, wherein the molecule of interest is a polypeptide or a polynucleotide.

3. The method of claim 1, wherein the sequence is an amino acid sequence, or a nucleotide sequence.

4. The method of claim 1, wherein the elements is an amino acid, or a nucleotide.DOCKET NO.: AIP-005WO PATENT 5. The method of claim 1, wherein the weighted value of the at least one edge is determined by performing a transform of one or more cooperativity scores corresponding to a presence of the at least one interaction between the at least one pair of atoms.

6. The method of claim 2, wherein the at least one interaction is an interaction comprising an energy well.

7. The method of claim 6, wherein the at least one interaction is a covalent bond, an electrostatic bond, a hydrogen bond, a disulfide bond, or a van der Waals force.

8. The method of claim 2, wherein the transform is a logistic transform.

9. The method of claim 8, wherein the logistic transform comprises a k parameter, an x0 parameter, and / or a covalent coefficient.

10. The method of claim 9, wherein the k parameter is equal to or greater than 0, equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 1, equal to or greater than 1.5, equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, equal to or greater than 4, equal to or greater than 4.5, equal to or greater than 5, equal to or greater than 5.5, equal to or greater than 6, equal to or greater than 6.5, equal to or greater than 7, equal to or greater than 7.5, equal to or greater than 8, equal to or greater than 8.5, equal to or greater than 9, equal to or greater than 9.5, equal to or greater than 10, equal to or greater than 10.5, equal to or greater than 11, equal to or greater than 11.5, or equal to or greater than 12.

11. The method of claim 9, wherein the x0 parameter is equal to or less than 0, equal to or less than -0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than -0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than -0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than -0.55, equal to or less than -0.6, equal to or less than -0.65, equal to or less than -0.7, equal to or less than - 0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than -0.9, equal to or less than -0.95, equal to or less than -1, equal to or less than -1.5, equal to or less than -2, equal to or less than -2.5, or equal to or less than -3.

12. The method of claim 9, wherein the covalent coefficient is a covalent coefficient sigma and / or covalent coefficient pi.DOCKET NO.: AIP-005WO PATENT 13. The method of claim 12, wherein the covalent coefficient sigma is equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, or equal to or greater than 0.

99.

14. The method of claim 12, wherein the covalent coefficient pi is equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, equal to or greater than 0.99, or equal to or greater than 1.

15. The method of any one of claims 1-14, wherein the one or more cooperativity scores comprises a hydrogen bond score, a disulfide bond score, and / or a van der Waals force score.

16. The method of claim 1, wherein the weighted value of the at least one edge is assigned a threshold value due to a presence of a covalent bond.

17. The method of any one of claims 1-16, wherein the weighted value of the at least one edge is between 0 and 1.

18. The method of claim 17, wherein the weighted value of the at least one edge closer to 1 indicates higher cooperativity of two atoms connected via the at least one edge in comparison to a lower cooperativity of two atoms connected via at least one edge with the weighted value closer to 0.

19. The method of claim 1, further comprising: prior to step (a), obtaining a 3D atomic structure comprising the one or more atoms of the sequence of the molecule of interest.

20. The method of any one of claims 1-19, wherein assigning the cooperativity scores representing the positional certainty to each of the one or more other atoms of the sequence comprises: (a) performing additional pairwise atom to atom analyses between the atom and each of the one or more other atoms, comprising: (i) generating a plurality of pulse scores within a pulse array for the atom;DOCKET NO.: AIP-005WO PATENT (ii) for each other atom, selecting a higher of: (a) a pulse score of the pulse array corresponding to the one or more other atoms; (b) a marginal pulse score of a marginal pulse array corresponding to the one or more other atoms; and (iii) combining the selected scores in step (ii) across the one or more other atoms to generate the cooperativity scores.

21. The method of claim 20, wherein the pulse score or the marginal score is a score based on the weighted value of a corresponding edge in the 3D graph between the atom and the one or more other atom.

22. The method of any one of claims 1-21, wherein combining the cooperativity scores across the one or more other atoms comprises summating the cooperativity scores of the one or more other atoms.

23. The method of any one of claims 1-22, wherein combining structural relevance of atoms of the element comprises summating the structural relevance of the atoms of the element to generate the measure of structural contribution for the element.

24. The method of any one of claims 1-23, further comprising characterizing the one or more structural features of the molecule of interest using the measure of structural contribution for the element.

25. The method of claim 24, wherein characterizing the one or more structural features of the molecule of interest using the measure of structural contribution comprises one or more of: (a) determining a stability of the molecule of interest; (b) examining an interface of the molecule of interest; (c) identifying a paratope of the molecule of interest; and (d) evaluating an effect of one or more point mutations on a structure of the molecule of interest.

26. A method for determining the stability of a molecule of interest, the method comprising performing the method of claim 1.

27. The method of claim 26, wherein the method comprises setting the k parameter to equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 1, equal to or greater than 1.5,DOCKET NO.: AIP-005WO PATENT equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, equal to or greater than 4, equal to or greater than 4.5, equal to or greater than 5, equal to or greater than 5.5, equal to or greater than 6, equal to or greater than 6.5, equal to or greater than 7, equal to or greater than 7.5, equal to or greater than 8, equal to or greater than 8.5, equal to or greater than 9, equal to or greater than 9.5, equal to or greater than 10, equal to or greater than 10.5, equal to or greater than 11, equal to or greater than 11.5, or equal to or greater than 12.

28. The method of claim 26, wherein the method comprises setting the x0 parameter to equal to or less than -0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than -0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than -0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than - 0.55, equal to or less than -0.6, equal to or less than -0.65, equal to or less than -0.7, equal to or less than -0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than -0.9, equal to or less than -0.95, or equal to or less than -1.

29. The method of claim 26, wherein the method comprises setting the covalent coefficient sigma to equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, or equal to or greater than 0.

9.

30. The method of claim 26, wherein the method comprises setting the covalent coefficient pi to equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, equal to or greater than 0.99, or equal to or greater than 1.

31. The method of claim 26, wherein determining the stability of the molecule comprises determining at least one under-specified or disordered region of the molecule.

32. A method for examining an interface of the molecule of interest, the method comprising performing the method of claim 1.

33. The method of claim 32, wherein the molecule of interest is a polypeptide, or a polynucleotide.DOCKET NO.: AIP-005WO PATENT 34. The method of claim 32, wherein the method comprises setting the k parameter to equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, equal to or greater than 4, equal to or greater than 4.5, equal to or greater than 5, equal to or greater than 5.5, or equal to or greater than 6.

35. The method of claim 32, wherein the method comprises setting the x0 parameter equal to or less than -0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than - 0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than -0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than -0.55, equal to or less than -0.6, equal to or less than -0.65, equal to or less than -0.7, equal to or less than -0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than -0.9, equal to or less than -0.95, equal to or less than -1, equal to or less than -1.5, equal to or less than -2, equal to or less than -2.5, or equal to or less than -3.

36. The method of claim 32, wherein the method comprises setting the covalent coefficient sigma to equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, or equal to or greater than 0.

99.

37. The method of claim 32, wherein the method comprises setting the covalent coefficient pi to equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, equal to or greater than 0.99, or equal to or greater than 1.

38. The method of claim 32, wherein examining the interface of the molecule of interest comprises determining at least one interface region of the molecule of interest.

39. The method of claim 32, wherein examining the interface of the molecule comprises determining structurally important interactions.

40. The method of claim 32, wherein the interface of the molecule of interest is an interface between at least one polypeptide region and at least one polynucleotide region.

41. The method of claim 32, wherein examining the interface of the molecule comprises, prior to performing a pairwise atom to atom interaction analysis across at least one pair of atoms in the sequence to generate a 3D graph, the steps of: (a) selecting a focal chain; and (b) selecting a medial chain.DOCKET NO.: AIP-005WO PATENT 42. The method of claim 41, wherein the focal chain comprises a polypeptide region of a polypeptide of interest, or a polynucleotide region of a polynucleotide of interest, and the medial chain comprises a polypeptide region of a polypeptide of interest, or a polynucleotide region of a polynucleotide of interest.

43. The method of claim 42, wherein the polypeptide region comprises a fragment, or full length amino acid sequence of the polypeptide.

44. The method of claim 42, wherein the polynucleotide region comprises a fragment, or full length nucleotide sequence of the polynucleotide.

45. The method of claim 41, wherein the focal chain and the medial chain are on the same or different polypeptide.

46. The method of claim 41, wherein the focal chain and the medial chain are on the same or different polynucleotide.

47. The method of any one of claims 32-36, wherein examining the interface of the molecule comprises performing the method of claim 1 across the medial chain of the molecule to generate the measure of structural contribution for the element of the focal chain of the molecule.

48. The method of claim 47, wherein: a. the medial chain of the molecule is a medial chain of the polypeptide of interest, and the focal chain of the molecule is a focal chain of polypeptide of interest; b. the medial chain of the molecule is a medial chain of the polynucleotide of interest, and the focal chain of the molecule is a focal chain of polynucleotide of interest; c. the medial chain of the molecule is a medial chain of the polypeptide of interest, and the focal chain of the molecule is a focal chain of polynucleotide of interest; or d. the medial chain of the molecule is a medial chain of the polynucleotide of interest, and the focal chain of the molecule is a focal chain of polypeptide of interest.

49. A method for identifying a paratope of the molecule of interest, the method comprising performing the method of claim 1.

50. The method of claim 44, wherein the method comprises setting the k parameter to equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 1, equal to or greater than 1.5, equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, or equal to or greater than 4.DOCKET NO.: AIP-005WO PATENT 51. The method of claim 44, wherein the method comprises setting the x0 parameter to equal to or less than -0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than -0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than -0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than - 0.55, equal to or less than -0.6, equal to or less than -0.65, equal to or less than -0.7, equal to or less than -0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than -0.9, equal to or less than -0.95, equal to or less than -1, equal to or less than -1.5, or equal to or less than -2.

52. The method of claim 44, wherein the method comprises setting the covalent coefficient sigma to equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, or equal to or greater than 0.

5.

53. The method of claim 44, wherein the method comprises setting the covalent coefficient pi to 1.

54. A method for evaluating an effect of one or more point mutations on a structure of the molecule of interest, the method comprising performing the method of claim 1.

55. The method of claim 49, wherein evaluating an effect of one or more point mutations on a structure of the molecule of interest comprises characterizing a tolerability of the structure of the molecule of interest to the one or more point mutations.

56. A method for characterizing one or more structural features of a polypeptide of interest, the method comprising: for each of one or more atoms of an amino acid sequence of the polypeptide of interest: (a) performing a pairwise atom to atom interaction analysis across at least one pair of atoms in the amino acid sequence to generate a 3D graph, wherein the 3D graph comprises nodes representing atoms and at least one edge between at least two nodes representing at least one interaction between the at least one pair of atoms, wherein the at least one edge comprises a weighted value based at least in part on presence or absence of an interaction between the at least one pair of atoms; (b) iteratively interrogating the 3D graph to determine a structural relevance of each of the one or more atoms of the amino acid sequence, wherein iteratively interrogating the 3D graph comprises: (i) for an atom of the 3D graph, assigning cooperativity scores representing a positional certainty to each of one or more other atoms of the amino acid sequence, wherein the cooperativity scores are assigned based atDOCKET NO.: AIP-005WO PATENT least in part on the weighted value of a corresponding edge in the 3D graph between the atom and the one or more other atoms, and (ii) combining the cooperativity scores of the atom across each of the one or more other atoms to determine a structural relevance of the atom; and for each amino acid in the amino acid sequence of the polypeptide, combining the structural relevance of each atom in the amino acid to generate a measure of structural contribution for the amino acid.

57. The method of claim 56, wherein the weighted value of the at least one edge is determined by performing a transform of one or more cooperativity scores corresponding to a presence of the at least one interaction between the at least one pair of atoms.

58. The method of claim 57, wherein the at least one interaction is an interaction comprising an energy well.

59. The method of claim 58, wherein the at least one interaction is a covalent bond, an electrostatic bond, a hydrogen bond, a disulfide bond, or a van der Waals force.

60. The method of claim 57, wherein the transform is a logistic transform.

61. The method of claim 60, wherein the logistic transform comprises a k parameter, an x0 parameter, and / or a covalent coefficient.

62. The method of claim 61, wherein the k parameter is equal to or greater than 0, equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 1, equal to or greater than 1.5, equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, equal to or greater than 4, equal to or greater than 4.5, equal to or greater than 5, equal to or greater than 5.5, equal to or greater than 6, equal to or greater than 6.5, equal to or greater than 7, equal to or greater than 7.5, equal to or greater than 8, equal to or greater than 8.5, equal to or greater than 9, equal to or greater than 9.5, equal to or greater than 10, equal to or greater than 10.5, equal to or greater than 11, equal to or greater than 11.5, or equal to or greater than 12.

63. The method of claim 61, wherein the x0 parameter is equal to or less than 0, equal to or less than -0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than -0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than -0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than -0.55, equalDOCKET NO.: AIP-005WO PATENT to or less than -0.6, equal to or less than -0.65, equal to or less than -0.7, equal to or less than - 0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than -0.9, equal to or less than -0.95, equal to or less than -1, equal to or less than -1.5, equal to or less than -2, equal to or less than -2.5, or equal to or less than -3.

64. The method of claim 61, wherein the covalent coefficient is a covalent coefficient sigma and / or covalent coefficient pi.

65. The method of claim 64, wherein the covalent coefficient sigma is equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, or equal to or greater than 0.

99.

66. The method of claim 64, wherein the covalent coefficient pi is equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, equal to or greater than 0.99, or equal to or greater than 1.

67. The method of any one of claims 56-66, wherein the one or more cooperativity scores comprises a hydrogen bond score, a disulfide bond score, and / or a van der Waals force score.

68. The method of claim 56, wherein the weighted value of the at least one edge is assigned a threshold value due to a presence of a covalent bond.

69. The method of any one of claims 56-68, wherein the weighted value of the at least one edge is between 0 and 1.

70. The method of claim 69, wherein the weighted value of the at least one edge closer to 1 indicates higher cooperativity of two atoms connected via the at least one edge in comparison to a lower cooperativity of two atoms connected via at least one edge with the weighted value closer to 0.

71. The method of claim 56, further comprising: prior to step (a), obtaining a 3D atomic structure comprising the one or more atoms of the amino acid sequence of the polypeptide.DOCKET NO.: AIP-005WO PATENT 72. The method of any one of claims 56-71, wherein assigning the cooperativity scores representing the positional certainty to each of the one or more other atoms of the amino acid sequence comprises: (a) performing additional pairwise atom to atom analyses between the atom and each of the one or more other atoms, comprising: (i) generating a plurality of pulse scores within a pulse array for the atom; (ii) for each other atom, selecting a higher of: (a) a pulse score of the pulse array corresponding to the one or more other atoms; (b) a marginal pulse score of a marginal pulse array corresponding to the one or more other atoms; and (iii) combining the selected scores in step (ii) across the one or more other atoms to generate the cooperativity scores.

73. The method of claim 72, wherein the pulse score or the marginal score is a score based on the weighted value of a corresponding edge in the 3D graph between the atom and the one or more other atom.

74. The method of any one of claims 56-73, wherein combining the cooperativity scores across the one or more other atoms comprises summating the cooperativity scores of the one or more other atoms.

75. The method of any one of claims 56-74, wherein combining structural relevance of atoms of the amino acid comprises summating the structural relevance of the atoms of the amino acid to generate the measure of structural contribution for the amino acid.

76. The method of any one of claims 56-75, further comprising characterizing the one or more structural features of the polypeptide of interest using the measure of structural contribution for the amino acid.

77. The method of claim 76, wherein characterizing the one or more structural features of the polypeptide of interest using the measure of structural contribution comprises one or more of: (a) determining a stability of the polypeptide of interest; (b) examining an interface of the polypeptide of interest; (c) identifying a paratope of the polypeptide of interest; and (d) evaluating an effect of one or more point mutations on a structure of the polypeptide of interest.DOCKET NO.: AIP-005WO PATENT 78. A method for determining the stability of a polypeptide of interest, the method comprising performing the method of claim 56.

79. The method of claim 78, wherein the method comprises setting the k parameter to equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 1, equal to or greater than 1.5, equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, equal to or greater than 4, equal to or greater than 4.5, equal to or greater than 5, equal to or greater than 5.5, equal to or greater than 6, equal to or greater than 6.5, equal to or greater than 7, equal to or greater than 7.5, equal to or greater than 8, equal to or greater than 8.5, equal to or greater than 9, equal to or greater than 9.5, equal to or greater than 10, equal to or greater than 10.5, equal to or greater than 11, equal to or greater than 11.5, or equal to or greater than 12.

80. The method of claim 78 or 79, wherein the method comprises setting the x0 parameter to equal to or less than -0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than -0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than -0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than -0.55, equal to or less than -0.6, equal to or less than -0.65, equal to or less than -0.7, equal to or less than -0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than - 0.9, equal to or less than -0.95, or equal to or less than -1.

81. The method of claim 78, wherein the method comprises setting the covalent coefficient sigma to equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, or equal to or greater than 0.

9.

82. The method of claim 78, wherein the method comprises setting the covalent coefficient pi to equal to or greater than 0.3, equal to or greater than 0.4, equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, equal to or greater than 0.99, or equal to or greater than 1.

83. The method of claim 78, wherein determining the stability of the polypeptide comprises determining at least one under-specified or disordered region of the polypeptide.DOCKET NO.: AIP-005WO PATENT 84. The method of claim 78, wherein the polypeptide of interest is a monomeric structure.

85. A method for examining an interface of the polypeptide of interest, the method comprising performing the method of claim 56.

86. The method of claim 85, wherein the method comprises setting the k parameter to equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, equal to or greater than 4, equal to or greater than 4.5, equal to or greater than 5, equal to or greater than 5.5, or equal to or greater than 6.

87. The method of claim 85, wherein the method comprises setting the x0 parameter equal to or less than -0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than - 0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than -0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than -0.55, equal to or less than -0.6, equal to or less than -0.65, equal to or less than -0.7, equal to or less than -0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than -0.9, equal to or less than -0.95, equal to or less than -1, equal to or less than -1.5, equal to or less than -2, equal to or less than -2.5, or equal to or less than -3.

88. The method of claim 85, wherein the method comprises setting the covalent coefficient sigma to equal to or greater than 0.9, equal to or greater than 0.91, equal to or greater than 0.92, equal to or greater than 0.93, equal to or greater than 0.94, equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, or equal to or greater than 0.

99.

89. The method of claim 85, wherein the method comprises setting the covalent coefficient pi to equal to or greater than 0.95, equal to or greater than 0.96, equal to or greater than 0.97, equal to or greater than 0.98, equal to or greater than 0.99, or equal to or greater than 1.

90. The method of claim 85, wherein examining the interface of the polypeptide of interest comprises determining at least one interface region of the polypeptide of interest.

91. The method of claim 85, wherein examining the interface of the polypeptide comprises determining structurally important interactions.

92. The method of claim 85, wherein the interface of a polypeptide is an interface between at least two polypeptide regions.

93. The method of claim 92, wherein the at least two polypeptide regions are on the same polypeptide or more than one polypeptide (e.g., two, or more polypeptides).DOCKET NO.: AIP-005WO PATENT 94. The method of claim 85, wherein examining the interface of the polypeptide comprises, prior to performing a pairwise atom to atom interaction analysis across at least one pair of atoms in the amino acid sequence to generate a 3D graph, the steps of: (a) selecting a focal chain; and (b) selecting a medial chain.

95. The method of claim 94, wherein the focal chain and the medial chain comprise a polypeptide region of the polypeptide, and wherein the polypeptide region of the focal chain is the same as, or different than the polypeptide region of the medial chain.

96. The method of claim 95, wherein the polypeptide region comprises a fragment, or full length amino acid sequence of the polypeptide.

97. The method of claim 96, wherein the focal chain and the medial chain are on the same or different polypeptide.

98. The method of any one of claims 85-97, wherein examining the interface of the polypeptide comprises performing the method of claim 56 across the medial chain of the polypeptide to generate the measure of structural contribution for the amino acid of the focal chain of the polypeptide.

99. A method for identifying a paratope of the polypeptide of interest, the method comprising performing the method of claim 56.

100. The method of claim 99, wherein the method comprises setting the k parameter to equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 1, equal to or greater than 1.5, equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, or equal to or greater than 4.

101. The method of claim 99, wherein the method comprises setting the x0 parameter to equal to or less than -0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than -0.2, equal to or less than -0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than -0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than - 0.55, equal to or less than -0.6, equal to or less than -0.65, equal to or less than -0.7, equal to or less than -0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than -0.9, equal to or less than -0.95, equal to or less than -1, equal to or less than -1.5, or equal to or less than -2.DOCKET NO.: AIP-005WO PATENT 102. The method of claim 99, wherein the method comprises setting the covalent coefficient sigma to equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, or equal to or greater than 0.

5.

103. The method of claim 99, wherein the method comprises setting the covalent coefficient pi to 1.

104. A method for evaluating an effect of one or more point mutations on a structure of the polypeptide of interest, the method comprising performing the method of claim 56.

105. The method of claim 104, wherein evaluating an effect of one or more point mutations on a structure of the polypeptide of interest comprises characterizing a tolerability of the structure of the polypeptide of interest to the one or more point mutations.

106. A method for identifying a paratope of a polypeptide, the method comprising: (a) obtaining or having obtained a plurality of binding affinity values between a target antigen and a plurality of polypeptide variant sequences based on the polypeptide, wherein the polypeptide variant sequences of the polypeptide differ from one another by point mutations at predetermined positions; (b) determining measures of structural contribution of amino acids of the polypeptide by performing a structural analysis of individual atoms of an amino acid sequence of the polypeptide; (c) performing amino acid residue-level combinations of binding affinity values and measures of structural contribution to generate a measure of non-structural binding relevance for each of the one or more amino acids of each of the one or more polypeptide variant sequences; and (d) identifying the paratope comprising one or more amino acids in the polypeptide associated with measures of non-structural binding relevance that indicate contributions to binding to the target antigen.

107. The method of claim 106, further comprising: (e) validating the paratope, wherein the validation comprises: (i) generating a modified polypeptide variant sequence by inserting or removing one or more amino acids associated with measures of non-structural binding relevance that indicate a lack of contribution to binding to the target antigen.

108. The method of claim 106, wherein the validation further comprises: (ii) determining a change in measures of non-structural binding relevance for one or more remaining amino acids in the modified polypeptide variant sequence.DOCKET NO.: AIP-005WO PATENT 109. The method of claim 108, wherein (ii) determining the change in measures of non- structural binding relevance comprises re-performing the structural analysis of individual atoms of amino acids of the modified polypeptide variant sequence.

110. The method of claim 106, wherein the validation further comprises: (iii) expressing the modified polypeptide variant sequence; (iv) measuring binding affinity of the modified polypeptide variant sequence; and (v) comparing the measured binding affinity of the modified polypeptide variant sequence to a binding affinity of the polypeptide prior to insertion or removal of the one or more amino acids associated with measures of non-structural binding relevance indicative of a lack of contribution to binding.

111. The method of claim 110, wherein (iv) measuring binding affinity of the modified polypeptide variant sequence to the target is determined by surface plasmon resonance.

112. The method of any one of claims 106-111, wherein (a) obtaining or having obtained a plurality of binding affinity values between a target antigen and a plurality of polypeptide variant sequences comprises: (i) generating a sequence library of the plurality of polypeptide variant sequences in which polypeptide variant sequences of the plurality differ from one another by point mutations at predetermined positions; (ii) expressing the plurality of polypeptide variant sequences; and (iii) measuring the plurality of binding affinity values between the target antigen and the expressed plurality of polypeptide variant sequences.

113. The method of claim 112, further comprising (iv) sequencing the plurality of polypeptide variant sequences to determine corresponding sequences of the plurality of polypeptide variant sequences.

114. The method of claim 112, wherein (i) generating the sequence library of the plurality of polypeptide variant sequences comprises: (i) obtaining a reference polypeptide sequence; and (ii) iteratively substituting one amino acid of the reference polypeptide sequence with a different amino acid to generate the plurality of polypeptide variant sequences.

115. The method of claim 114, wherein the plurality of polypeptide variant sequences comprises at least 50% of all possible polypeptide variant sequences that differ from the reference polypeptide sequence by a single amino acid.DOCKET NO.: AIP-005WO PATENT 116. The method of claim 114, wherein the plurality of polypeptide variant sequences comprises at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, a least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of all possible polypeptide variant sequences that differ from the reference polypeptide sequence by a single amino acid.

117. The method of any one of claims 106-116, wherein (c) performing amino acid residue- level combinations of binding affinity values and measures of structural relevance comprises: for each of the one or more amino acids of the polypeptide variant sequence, determining a difference between the binding affinity value of the polypeptide variant sequence comprising the amino acid and the measure of structural relevance of the amino acid.

118. The method of claim 117, wherein the measure of non-structural binding relevance for an amino acids represents a binding contribution of the amino acid to the binding affinity without a corresponding structural contribution of the amino acid.

119. The method of any one of claims 106-118, wherein (d) identifying the paratope comprising one or more amino acids comprises selecting one or more amino acids with measures of non-structural binding relevance that are above a threshold score.

120. The method of any one of claims 106-119, wherein each of the polypeptide variant sequences in the plurality of polypeptide variant sequences have the same length.

121. The method of any one of claims 106-120, wherein the polypeptide variant sequences in the plurality of polypeptide variant sequences are between 30 and 60 amino acids in length.

122. The method of any one of claims 106-121, wherein characterizing the amino acid sequence of the polypeptide by performing the structural analysis of individual atoms of the amino acid sequence comprises: for each of one or more atoms of each amino acid of the polypeptide variant sequence: (a) performing a pairwise atom to atom interaction analysis across at least one pair of atoms in the amino acid sequence to generate a 3D graph, wherein the 3D graph comprises nodes representing atoms and at least one edge between at least two nodes representing at least one interaction between the at least one pair of atoms, wherein the at least one edge comprises a weighted value based at least in part on presence or absence of an interaction between the at least one pair of atoms;DOCKET NO.: AIP-005WO PATENT (b) iteratively interrogating the 3D graph to determine a structural relevance of each of the one or more atoms of the amino acid sequence, wherein iteratively interrogating the 3D graph comprises: (i) for an atom of the 3D graph, assigning cooperativity scores representing a positional certainty to each of one or more other atoms of the amino acid sequence, wherein the cooperativity scores are assigned based at least in part on the weighted value of a corresponding edge in the 3D graph between the atom and the one or more other atom, and (ii) combining the cooperativity scores of the atom across the one or more other atoms to determine a structural relevance of the atom; and for each amino acid in the one or more amino acids of the polypeptide, combining the structural relevance of each atom in the amino acid to generate a measure of structural contribution for the amino acid.

123. The method of claim 122, wherein the weighted value of the at least one edge is determined by performing a transform of one or more cooperativity scores corresponding to a presence of the at least one interaction between the at least one pair of atoms.

124. The method of claim 123, wherein the at least one interaction is an interaction comprising an energy well.

125. The method of claim 124, wherein the at least one interaction is a covalent bond, an electrostatic bond, a hydrogen bond, a disulfide bond, or a van der Waals force.

126. The method of claim 123, wherein the transform is a logistic transform.

127. The method of claim 126, wherein the logistic transform comprises a parameter k, a parameter x0, and / or a covalent coefficient.

128. The method of claim 127, wherein the k parameter is equal to or greater than 0.5, equal to or greater than 0.6, equal to or greater than 0.7, equal to or greater than 0.8, equal to or greater than 0.9, equal to or greater than 1, equal to or greater than 1.5, equal to or greater than 2, equal to or greater than 2.5, equal to or greater than 3, equal to or greater than 3.5, or equal to or greater than 4.

129. The method of claim 127, wherein the x0 parameter is equal to or less than -0.05, equal to or less than -0.1, equal to or less than -0.15, equal to or less than -0.2, equal to or less than - 0.25, equal to or less than -0.3, equal to or less than -0.35, equal to or less than -0.4, equal to or less than -0.45, equal to or less than -0.5, equal to or less than -0.55, equal to or less than -0.6,DOCKET NO.: AIP-005WO PATENT equal to or less than -0.65, equal to or less than -0.7, equal to or less than -0.75, equal to or less than -0.8, equal to or less than -0.85, equal to or less than -0.9, equal to or less than -0.95, equal to or less than -1, equal to or less than -1.5, or equal to or less than -2.

130. The method of claim 127, wherein the covalent coefficient is a covalent coefficient sigma and / or covalent coefficient pi.

131. The method of claim 130, wherein the covalent coefficient sigma is equal to or greater than 0.1, equal to or greater than 0.2, equal to or greater than 0.3, equal to or greater than 0.4, or equal to or greater than 0.

5.

132. The method of claim 130, wherein the covalent coefficient pi is equal to 1.

133. The method of any one of claims 122-132, wherein the one or more cooperativity scores comprises a hydrogen bond score, a disulfide bond score, and / or a van der Waals force score.

134. The method of claim 122-133, wherein the weighted value of the at least one edge is assigned a threshold value due to a presence of a covalent bond.

135. The method of any one of claims 122-134, wherein the weighted value of the at least one edge is between 0 and 1.

136. The method of claim 135, wherein the weighted value of the at least one edge closer to 1 indicates higher cooperativity of two atoms connected via the at least one edge in comparison to a lower cooperativity of two atoms connected via at least one edge with the weighted value closer to 0.

137. The method of claim 122, wherein the method comprises: prior to step (a), obtaining a 3D atomic structure comprising the one or more atoms of the amino acid sequence of the polypeptide of interest.

138. The method of any one of claims 122-137, wherein assigning the cooperativity score representing positional certainty to each of one or more other atoms of the amino acid comprises: (a) performing additional pairwise atom to atom analyses between the atom and each of the one or more other atoms, comprising: (i) generating a plurality of pulse scores within a pulse array for the atom; (ii) for each other atom, selecting a higher of:DOCKET NO.: AIP-005WO PATENT (a) a pulse score of the pulse array, corresponding to the one or more other atoms; (b) a marginal pulse score of a marginal pulse array, corresponding to the one or more other atoms; and (iii) combining the selected scores in step (ii) across the one or more other atoms to generate the cooperativity scores.

139. The method of claim 138, wherein the pulse score or the marginal score is a score based on the weighted value of a corresponding edge in the 3D graph between the atom and the one or more other atom.

140. The method of any one of claims 122-139, wherein combining the cooperativity scores across the one or more other atoms comprises summating the cooperativity scores of the one or more other atoms.

141. The method of any one of claims 122-140, wherein combining structural relevance of atoms of the amino acid comprises summating the structural relevance of the atoms of the amino acid to generate the measure of structural contribution for the amino acid.

142. A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1-141.

143. A system comprising: a processor; and a non-transitory computer readable medium comprising instructions that, when executed by the processor, cause the processor to perform the method of any one of claims 1-141.

Citation Information

Patent Citations

  • Polynucleotide constructs with disulfide groups

    JP2016537027A

  • System and method for prediction of protein-ligand bioactivity using point-cloud machine learning

    US11256995B1

  • Methods for identifying ligand binding sites in a biomolecule

    US20030165915A1

  • Fragmentation-based methods and systems for de novo sequencing

    US20050009053A1

  • Compositions that bind multiple epitopes of IGF-1r

    US20090130105A1