Methods for Protein Design
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- EURO LAB FUER MOLEKULARBIOLOGIE EMBL
- Filing Date
- 2023-04-11
- Publication Date
- 2026-04-20
AI Technical Summary
The prior art is difficult to effectively predict accurate high-resolution protein structures in protein design, thus limiting the progress of protein design.
AlphaFold and other neural network methods combined with Rosetta and molecular dynamics simulation, a high-reliability protein sequence and structure were designed by optimizing amino acid sequences and protein structures.
It realizes rapid prediction of protein structure and function in a short time, improves the efficiency and accuracy of protein design, and can design stable and specific functions.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a method and apparatus for protein design. [Background technology]
[0002] Description of Related Art De novo protein design has been a long-standing fundamental goal of synthetic biology but has been hindered by the difficulty of reliably predicting accurate, high-resolution protein structures from amino acid sequences. Recent advances in the accuracy of protein structure prediction methods such as AlphaFold (AF) have facilitated proteome-scale structure prediction of monomeric proteins.
[0003] Proteins contribute significantly to most life processes at the cellular scale. They realize a variety of functions, including, among others, catalysis of biochemical reactions, mechanical functions involved in cell motility, and the formation of intracellular structures. The central paradigm for understanding protein function is based on the observation that proteins fold into complex, yet specific, three-dimensional structures that are assumed to be their lowest energy state, which varies depending on their amino acid sequence. Experimental structure determination has therefore been a major research object in biology for the past 50 years and has been a major step forward in the development of new molecular structures for over 100 years. 5 Over 100,000 well-solved structures have been obtained and deposited in the Protein Data Bank (PDB) [1]. These structures reveal a diverse array of folding topologies and geometries of individual proteins, classified into 41 structures with 1390 folding topologies, and reveal that 51% of the structures previously extracted from the PDB form quaternary structures through the formation of oligomers and complexes.
[0004] The question has been to what extent does an amino acid sequence code for a protein's structure, and whether it is possible to predict a protein's structure from its amino acid sequence. This question has driven many different computational protein structure prediction approaches through decades of effort [2]. Specifically, improved methods have been explored in an independently evaluated biennial community-wide competition (CASP) that ranks the accuracy of participants' approaches with respect to determined but unpublished structures, and prediction accuracy has steadily improved [3]. This question has recently led to the development of AlphaFold [4], hereafter referred to as AF. AF is a neural network-based approach that has achieved comparable atomic accuracy with respect to crystal structures. AF has been applied at a proteome scale across several species, giving rise to a new database of predicted structures [5]. More recently, a community-wide evaluation has revealed several important applications of AF, ranging from predicting the effects of amino acid variants to building cryo-EM models [6]. While AF's input protocol was intended to predict single-stranded structures, an unexpected result was that it can predict non-contiguous strands, thus enabling the prediction of protein complexes using existing trained networks.[7] This capability has also been applied at the proteome scale to reveal novel core eukaryotic complexes.[8] These efforts led to the release of AlphaFold-Multimer, a version of AF explicitly trained for complex prediction.[9]
[0005] Despite these advances, the problem of protein folding remains much more complex than structure prediction alone. Proteins are not rigid structures, for example in physiological conditions. Proteins exhibit a thermodynamic equilibrium between folded and unfolded forms. Many proteins also change conformation according to well-described equilibria between those states and utilize these changes to exert their functions
[10] . Furthermore, there are many intrinsically disordered proteins (IDPs), which do not exhibit a stable tertiary structure or fold transiently depending on the environmental context
[10] . One notable biomedical example is the misfolding of amyloid-α helices into β-sheet-like filaments associated with Alzheimer's disease
[11] . Denatured proteins have also been discovered. These denatured proteins form multiple stable but distinct folds
[12] .
[0006] From a biophysical perspective, protein folding can be understood using the concept of a folding energy funnel in the atomic configuration space, with multiple energy minima corresponding to alternative stable states of the protein through which it transitions with characterizable kinetics. Based on this concept, computational physics methods such as molecular dynamics (MD) simulations provide methods to characterize the dynamics, thermodynamics, and kinetics of conformational transitions
[13] , ab initio folding, disorder transitions, and protein-protein and protein-ligand binding
[14] .
[0007] Despite the complex structures and dynamics of naturally occurring proteins already mentioned, evolution has still only explored a small portion of the landscape of potential protein sequences.
[15] Therefore, there is great potential in elucidating the fundamental biophysical principles of protein folding and in designing and engineering novel proteins that can exploit this vast space. Recent examples include protein logic gates
[16] , self-assembling systems
[17] and targeted therapeutics
[18] .
[0008] Established methods of protein engineering have, until recently, focused on tailoring naturally occurring proteins through iterative experimental selection processes such as directed evolution.
[19] More recently, computational design approaches have enabled de novo protein design encompassing a range of functions spanning rules for topology selection, protein backbone construction, and amino acid sequence optimization, as well as combinations of these approaches.
[20]
[0009] One of the prior art specific methods in computational protein design is the Rosetta suite of protein design tools
[21] , specifically Rosetta Remodel
[22] . The Rosetta suite combines tools for all steps from the selection of the desired protein topology to the design and validation of the folded protein sequence of amino acids. The design of a protein amino acid sequence in Rosetta Remodel consists of a combination of three tasks: topology specification, backbone generation, and fixed backbone design
[22] . After specifying the desired protein topology, the protein backbone structure can be generated using matching fragments extracted from existing proteins. Fragments matching the desired set of secondary structures and contacts are selected from PDB structures. These selected fragments are then sampled using a Markov chain Monte Carlo method to create a pool of candidate backbone geometries. The candidate backbone structures can then be fitted to sequences using fixed backbone protein design, and the backbone can be optimized between design steps, if necessary
[23] . Once the desired backbone geometry is obtained, the goal is to find an amino acid sequence that folds the protein structure. Here, the Rosetta Design protocol
[23] starts by inputting an all-valine sequence in the desired backbone and runs a Markov chain Monte Carlo method to arrive at a low-energy sequence. The general strategy of Rosetta-based protein design has been successfully applied to a variety of design problems. These applications range from the first de novo designed proteins
[23] to synthetic vaccines
[24] , complex assemblies
[25] and enzymes
[26] . However, Markov chain Monte Carlo methods in structure space can be time consuming and computationally intensive.
[0010] To address this shortcoming of classical protein design, approaches have been made to apply neural networks to various design problems. Some studies train generative models to directly generate protein sequences with desired functions. For example, in
[27] , a language model is trained on the UniProt sequence database to generate sequences with specific functions, and the generated model is fine-tuned to a specific protein family or function. Besides language models, other types of generative models, such as variational autoencoders
[28] and generative adversarial networks
[29] , have been investigated for direct protein sequence generation. Although these approaches have been successful in generating protein sequences related to a given function, they do not explicitly consider structural information and therefore cannot be applied to the task of protein design, which includes constraints on tertiary structure.
[0011] Elsewhere in the area of neural network-based protein design, generative models have been explored for the generation of protein structures
[30] , training a generative adversarial network to generate realistic backbone distance maps and coordinates. In
[31] , the same goal is achieved using a variational autoencoder. Both approaches show that fine-grained control over the designed backbone structure can be achieved by manipulating latent variables. Although these approaches are viable alternatives to classical backbone design, these approaches do not provide any guarantees about the feasibility of their predicted backbone designs.
[0012] To bridge the gap between structure-only and sequence-only approaches,
[32] trained a neural network on the PDB database of protein structures to predict protein sequences given a fixed backbone structure. These approaches rely on network components that consider the geometry of the protein backbone along with the protein sequence, and include the use of a per-residue local coordinate system to infer protein geometry.
[33] frame fixed-backbone design as a constraint satisfaction problem and train their network as a constraint solver. These methods require neural networks trained for a specific protein design task - fixed-backbone protein design - and as such, cannot be easily extended to other design applications without retraining.
[0013] Prior art has also explored the reuse of previously trained predictors of protein structure or function as part of an in silico screening framework. In these approaches, neural networks are treated as a scoring function to evaluate the quality of protein designs
[34] . The designs are then improved using gradient-based
[35] , gradient-free
[36] , or neural network-based
[37] optimization.
[0014] In
[34] , an optimization loop incorporating neural networks was first used for structure prediction in de novo protein design. Their approach was then extended to fixed-backbone design
[35] and the generation of protein scaffolds for stabilization of protein motifs, using both Markov chain Monte Carlo and gradient descent as their means of optimization. All of these approaches utilize trRosetta
[38] as a structure prediction predictor. In
[36] , a new release of AF is used as a structure predictor for fixed-backbone protein design, using greedy optimization with sequences initialized from a trained model. Summary of the Invention
[0015] In a preferred embodiment, the invention is a computer-implemented method for protein design.
[0016] A computer-implemented method for designing at least one protein includes the steps of generating at least one amino acid sequence to be tested, some of the amino acids included in the amino acid sequence to be tested being selected according to a probability distribution, and further includes the steps of predicting structural characteristics of the at least one protein from the aligned at least one amino acid sequence, calculating a fitness function value for the at least one amino acid sequence to be tested based on the structural characteristics of the at least one protein, and selecting or deselecting the at least one amino acid sequence to be tested depending on the value of the fitness function.
[0017] The method may further comprise the step of modifying at least one amino acid sequence to be tested by altering at least one of the amino acids contained in the amino acid sequence to be tested.
[0018] The method may further comprise the step of iteratively modifying at least one amino acid sequence until the value of the fitness function exceeds a preselected threshold.
[0019] Predicting the structural properties of at least one protein contained in the tested amino acid sequence may comprise predicting the positions of the amino acids relative to each other and / or predicting the entire atomic structure of at least one protein.
[0020] The method can use optimization of amino acid sequences (also called protein sequences) by gradient-free or gradient-based optimizers. An example of a gradient-free optimizer is an evolutionary algorithm. The optimization results in backbones (i.e., amino acids or polypeptides linked by peptide bonds, from which side chains branch out) and amino acid sequences that can be designed by construction and have a high degree of confidence in AF. By using an evolutionary algorithm, the search for amino acid sequences (also called protein sequences) can be automated. The modification can include mutating and / or recombining at least one amino acid sequence to be tested according to the evolutionary algorithm.
[0021] Structural properties can be predicted using AlphaFold. AlphaFold (AF) [4] can be embedded in the design loop to predict protein structures. For example, native-like protein structures can be predicted by AF by searching amino acid sequences with high confidence predictions. A flexible family of fitness functions is designed that encodes various protein design tasks. This family of fitness functions is integrated into the design loop with extensive validation using Rosetta
[21] and / or molecular dynamics simulations.
[0022] The fitness function may be defined according to at least one biological and / or physicochemical property of the protein.
[0023] The fitness function comprises one or more fitness function components that represent biological and / or physicochemical properties of at least one protein.
[0024] The method may further comprise the step of inputting at least one amino acid sequence or known structural feature associated with the target protein, to which at least one protein can bind, into AlphaFold as a structural template.
[0025] This method allows for rapid prediction of novel protein monomers starting from random sequences. These novel protein monomers can adopt diverse folding sequences within the known protein space. This method can be used to design proteins that bind to pre-specified target proteins.
[0026] Protein monomers were found to largely maintain structural integrity, and the method demonstrates the ability to predict proteins that undergo conformational switches upon complex formation, such as the α-helix to β-sheet switch during amyloid filament formation.
[0027] The method may further comprise the step of redesigning the at least one selected amino acid sequence to be tested to make the redesigned at least one amino acid sequence more native-like and / or to improve the solubility and / or expressibility of the at least one protein.
[0028] The method may further comprise the step of statistically reconstructing at least one tested amino acid sequence from predicted structural features of at least one protein, thus allowing the calculation of a reconstructed amino acid sequence from the predicted protein structure.
[0029] Calculation of the reconstructed amino acid sequence allows further validation of the disclosed methods.
[0030] The method may further comprise a step of re-predicting at least one structural characteristic of the protein based on the redesigned at least one amino acid sequence. From the restored amino acid sequence, the sequence of the re-predicted protein can be calculated. Calculating the re-predicted protein structure allows further validation of the method of the present disclosure.
[0031] The method may further comprise the step of computationally validating the selected amino acid sequence using molecular dynamics and / or Rosetta ab-initio structure prediction.
[0032] The structural integrity of the predicted protein structures has been verified and confirmed by standard ab initio folding and structure analysis methods, and more extensively by performing rigorous all-atom molecular dynamics simulations to analyze the corresponding structural flexibility, intramonomer and interfacial amino acid contacts.
[0033] The method may further comprise the step of experimentally validating the selected amino acid sequences by measuring the solubility and / or expression of at least one protein.
[0034] The at least one protein may comprise a monomer, a binder that binds to a target protein, a homodimer, a heterodimer, a trimer, a tetramer, a pentamer, an oligomer, or a protein complex. [Brief description of the drawings]
[0035] For a more complete understanding of the present invention and its advantages, reference is now made to the following description and accompanying drawings.
[0036] [Figure 1]The method is outlined below. (A) The disclosed method includes the steps of: defining a fitness function; generating structures (candidate structures) by generating candidate amino acid sequences using evolutionary search; redesigning the amino acid sequences to be more native-like; re-predicting the structure of the redesigned amino acid sequences; filtering the redesigned amino acid sequences by likelihood, solubility and prediction confidence; optimizing the codons of the remaining amino acid sequences; and ordering the codon-optimized amino acid sequences for downstream experiments. (B) The method includes an optimization loop. A pool of amino acid sequences (sequence pool) is maintained throughout the design process. AlphaFold predicts all-atom structures (all-atom geometries) and confidence metrics (pLDDT, pAC) for each sequence. These are combined in an optimization loop into a fitness function that is optimized, e.g., by maximization or minimization. The amino acid sequences (hereinafter also referred to as protein sequences) and fitness values of the fitness functions associated with the amino acid sequences are input to an optimizer, which updates the sequence pool, e.g., by an evolutionary algorithm. If the fitness value associated with a selected one of the amino acid sequences exceeds a desired threshold, the amino acid sequence and structure are validated using Rosetta and molecular dynamics simulations. The fitness function may include one or more fitness function components. The fitness function may be any function (pLDDT, pAC) of the amino acid sequence, the predicted all-atom structure, and a confidence metric. Examples of fitness functions include the root mean square deviation (RMSD) relative to a reference structure (e.g., a native structure, an experimentally verified structure, a decoy structure, etc.), one or more confidence metrics, or constraints on complex formation and / or conformational changes upon complex formation. These and further examples of fitness function components used to construct the fitness function and the mathematical formulas of these fitness function components are described in detail below.
[0037] [Diagram 2]AlphaFold predictions of protein complexes are shown: dimer (HIV protease 1DMP, de novo designed protein 4PWW), trimer (de novo designed coiled-coil 1COI), tetramer (consensus TPR superhelix 2AVP), and pentamer (phage lambda tail terminator 3FZB). Predictions (dark grey) are aligned and overlaid on the native structure (light grey). All predictions show low RMSD relative to the native structure.
[0038] [Diagram 3]De novo monomer design is shown. (A-D) Ab initio and molecular dynamics validation of designed monomers of lengths 32-256 amino acids. Designed monomers (light grey) are overlaid with their lowest energy structures from Rosetta ab initio predictions (dark grey) or relaxed states of their predicted all-atom structures (Structures). Each monomer is shown with its RMSD relative to the predicted AlphaFold structure, as well as its value of Rosetta's all-atom energy function, also called "Rosetta score", e.g., REF2015, or energy functions based on electrostatic interactions, van der Waals interactions, and / or statistical physics. The predicted aligned confidence (pAC) of each designed monomer is shown (prediction error). For validation using Rosetta, we show the distribution of Rosetta scores and RMSDs for the top 10 relaxed structures (relaxed) starting from the ab initio predictions (dark grey) and the AlphaFold structures (light grey) for all decoys (i.e., conformations with low free energy) from the Rosetta ab initio predictions (ab initio) starting from the extended conformations (light grey) and the AlphaFold structures (dark grey) relative to the AlphaFold structures (Rosetta score distributions). For validation by molecular dynamics (MD metrics), we show the Boltzmann weighted ensemble variance (CV) distribution in a 2D landscape of the number of contacts and backbone RMSF (root mean square fluctuations) of amino acids within a protein monomer relative to the averaged MD structures (CV distributions). (E) t-distributed stochastic neighbor embedding (TSNE) of designed monomeric structures ≥ 64 amino acids in length. Agglomerative clustering is used to separate the structures into 10 clusters. (F) Representatives from each cluster, showing the diversity of structures designed using AlphaFold.
[0039] [Figure 4]For de novo dimer design, we show Rosetta and molecular dynamics validation of designed homodimers (A-C) of lengths between 32 and 128 amino acids. Designed homodimers are superimposed with their lowest energy structures obtained from Rosetta relaxation of their predicted all-atom structures (complex structures). Each relaxed dimer is shown along with its RMSD relative to the AlphaFold structure, Rosetta binding energy and packing statistics. Predicted aligned confidence (pAC) is shown for each designed dimer (prediction error). For molecular dynamics validation (MD complex metrics), we show the Boltzmann weighted CV distribution in 2D space of all amino acid contacts (intramonomer and interface) and the RMSF of the complex, along with the distribution of individual monomer contacts and interface contacts (CV distribution). (D-F) Rosetta and molecular dynamics validation of designed heterodimers of lengths between 32 and 128 amino acids. Scales displayed are the same as for (A-C). (G) Orthogonality of designed dimers. The predicted structures of the on-target and off-target complexes are shown for two different pairs of designed dimers. The prediction confidence (prediction error) of the on-target dimer is consistently higher than the same confidence for the off-target dimer for pairs between amino acid monomers.
[0040] [Diagram 5]De novo oligomer design is shown. (A-G) Rosetta and molecular dynamics validation of designed trimers, tetramers, and pentamers with monomer lengths ranging from 32 to 64 amino acids, and a hexamer with a monomer length of 32 amino acids. Designed oligomers are superimposed with their lowest energy structures derived from Rosetta relaxation of their predicted all-atom structures (complex structures). Each relaxed oligomer is shown along with its RMSD relative to the AlphaFold structure, Rosetta binding energy, and packing statistics. Predicted aligned confidence (pAC) is shown for each designed oligomer (prediction error). For molecular dynamics validation, the Boltzmann weighted CV distribution in 2D space of all amino acid contacts (intramonomer and at the interface) and the RMSF of the complex is shown (CV distribution), along with the distribution of individual monomer contacts and interface contacts. (H) Additional structures.
[0041] [Figure 6] De novo binder design is shown. (A,B) Rosetta validation of designed binders on previously de novo designed proteins. Relaxed binders bound to the target protein are superimposed with their predicted all-atom structures (complex structures). Each relaxed binder is shown with its RMSD against the AlphaFold structure, Rosetta binding energy and packing statistics. Predicted aligned confidence (pAC) is shown for each designed binder (prediction error). For molecular dynamics validation, the Boltzmann weighted CV distribution in 2D space of all amino acid contacts (intramonomer and at the interface) and the RMSF of the complex are shown (CV distribution), along with the distribution of individual monomer contacts and interface contacts.
[0042] [Figure 7]Conformational changes in sequence space of AlphaFold and de novo designed proteins are shown. (A) AlphaFold prediction of conformational changes upon complex formation. (Left) Predicted structures of monomeric circadian clock proteins KaiB (KaiB ground state, KaiBgs) and fold-switch stabilized KaiB (KaiBfs) in complex with KaiC (dark grey) superimposed with native structures of KaiBgs (1R5P) and KaiBfs (5JYT), respectively. Both predictions show low RMSD with respect to the corresponding native structures. Predicted KaiBfs (dark grey) superimposed with native KaiBgs (incorrect conformation, light grey) shows high RMSD. (Right) Predicted structures of monomeric amyloid-β and its pentamer (dark grey) superimposed with native monomer (1IYT) and oligomer (2MXU). AlphaFold predicts the conformational change from an α-helix to a parallel β-sheet characteristic of amyloid. (B) Rosetta and molecular dynamics validation of designed oligomers of monomer length 32 exhibiting the amyloid-like predicted conformational change. Designed oligomers (light grey) are superimposed with their lowest energy structures (dark grey) from Rosetta relaxation of their predicted all-atom structures (complex structures). Each relaxed oligomer is shown along with its RMSD relative to the AlphaFold structure, Rosetta binding energy and packing statistics. Monomers are shown compared to the oligomer structure (multiple monomer structures) to demonstrate the conformational change upon oligomerization. Predicted aligned confidence (pAC) is shown for each designed oligomer and its constituent monomers (prediction error). Validation by molecular dynamics shows the Boltzmann-weighted CV distribution in 2D space of all amino acid contacts (intramonomer and at the interface) and the RMSF of the complex (CV distribution), along with the distribution of individual monomer contacts and interfacial contacts. (C) Additional proteins designed using a fitness function that favors conformational changes upon complex formation.
[0043] [Figure 8]Validation of designed monomers using Rosetta and molecular dynamics. (A-D) Validation of de novo designed monomers of lengths 32, 64, 128 and 256 amino acids using Rosetta and molecular dynamics. (Structures). The lowest Rosetta energy structure is shown overlaid with the predicted AlphaFold structure, with RMSD and Rosetta all-atom scores. (Prediction error) shows the predicted aligned confidence for each monomer. (Rosetta score distribution) shows the distribution of relaxations in Rosetta energy and RMSD relative to the predicted AlphaFold structure and decoys from ab initio predictions (i.e. conformations with low free energy). (Relaxation) shows this distribution for the relaxations of the 10 lowest energy ab initio decoys (dark grey) and the 10 relaxations of the AlphaFold predicted structures (light grey). (ab initio) shows the distribution of decoys with Rosetta energy <0 REU sampled starting from the extended conformation (grey) and the AlphaFold structure (red). For molecular dynamics-based validation, (RMSD / F evolution) shows the time evolution of the RMSD (dark grey) and RMSF (light grey) for each monomer over the course of a 100 ns explicit solvent dynamics simulation. (Contact evolution) shows the time evolution of the number of intramonomer contacts over the course of 100 ns of simulation (grey). (CV distribution) shows the Boltzmann-weighted distribution in the 2D ensemble variable (CV) landscape of the backbone RMSF and all contacts (intramonomer and interface) for all snapshots in the last 50 ns of the simulation. Further randomly designed structures of monomers of length 64 and 128 amino acids.
[0044] [Figure 9]Validation of the designed dimers using Rosetta and molecular dynamics is shown. (A-F) Validation of de novo designed homo- and heterodimers of monomers of length 32, 64 and 128 amino acids using Rosetta and molecular dynamics. (Complex structure) shows the AlphaFold predicted structure (light grey) overlaid with the best of the 10 relaxed ones (dark grey) using the Rosetta all-atom scoring function. The structures are shown with their RMSD relative to the relaxed structures, as well as the Rosetta binding energy dG and packing statistics packstat. (Monomer structure) shows the monomeric AlphaFold predicted structure (dark grey) overlaid with the complex structure (light grey) and shows the aligned RMSD. (Prediction error) shows the predicted aligned confidence for each AlphaFold complex prediction. For the molecular dynamics-based validation, (RMSD / F evolution) shows the time evolution of the RMSD (dark grey) and RMSF (light grey) of the whole complex over the course of a 100 ns explicit solvent dynamics simulation. (Contact evolution) shows the time evolution of all intra-monomer contacts as well as interfacial contacts over the course of a 100 ns simulation. (CV distribution) shows the Boltzmann weighted distribution in the 2D collective variable (CV) landscape of the backbone RMSF and all contacts (intra-monomer and interfacial) for the entire complex for all snapshots in the last 50 ns of the simulation. (MD monomer metrics) shows the distribution of RMSD (monomer RMSD), RMSF (monomer RMSF), monomer contacts and interfacial contacts for all monomers and interfaces in the complex over the last 50 ns of the simulation.
[0045] [Figure 10]Validation of designed oligomers using Rosetta and molecular dynamics is shown. (A-D) Validation of de novo designed trimers, tetramers, pentamers and hexamers of monomer lengths 32 and 64 amino acids using Rosetta and molecular dynamics. (Complex structures) shows the AlphaFold predicted structure (light grey) overlaid with the best of the 10 relaxations (dark grey) using Rosetta's all-atom scoring function. The structures are shown together with their RMSD to the relaxed structures, as well as the Rosetta binding energy of a single monomer to the rest of the complex dG, and the packing statistics packstat. (Monomer structures) shows the monomeric AlphaFold predicted structure (dark grey) overlaid with the complex structure (light grey) and shows the aligned RMSD. (Prediction errors) shows the predicted aligned confidence for each AlphaFold complex prediction. For molecular dynamics-based validation, (RMSD / F evolution) shows the time evolution of the RMSD (dark grey) and RMSF (light grey) of the whole complex over the course of a 100 ns explicit solvent dynamics simulation. (Contact evolution) shows the time evolution of all intra-monomer contacts as well as interfacial contacts over the course of a 100 ns simulation. (CV distribution) shows the Boltzmann-weighted distribution in the 2D collective variable (CV) landscape of the backbone RMSF and all contacts (intra-monomer and interfacial) for the whole complex for all snapshots in the last 50 ns of the simulation. (MD monomer metrics) shows the distribution of RMSD (monomer RMSD, blue), RMSF (monomer RMSF, green), monomer contacts (grey) and interfacial contacts (orange) for all monomers and interfaces in the complex over the last 50 ns of the simulation.
[0046] [Figure 11]Validation of the designed binding protein using Rosetta and molecular dynamics. (A,B) Validation of a de novo designed binder, 64 amino acids in length, against two previously designed target proteins using Rosetta and molecular dynamics. (Target protein) shows the AlphaFold predicted structure of the target protein. (Complex structures) shows the AlphaFold predicted structure of the target protein in complex with the binder (grey) superimposed with the best of the 10 relaxed (colored) ones using the Rosetta all-atom scoring function. The target protein is in blue and all designed binders are in red. The structures are shown together with their RMSD relative to the relaxed structure, as well as the Rosetta binding energy dG and packing statistics packstat. (Monomer structures) shows the monomeric AlphaFold predicted structure (dark grey) superimposed with the complex structure (light grey) and shows the aligned RMSD. (Prediction errors) shows the predicted aligned confidence for each AlphaFold complex prediction. For molecular dynamics-based validation, (RMSD / F evolution) shows the time evolution of the RMSD (dark grey) and RMSF (light grey) of the whole complex over the course of a 100 ns explicit solvent dynamics simulation. (Contact evolution) shows the time evolution of all intra-monomer contacts as well as interfacial contacts over the course of a 100 ns simulation. (CV distribution) shows the Boltzmann-weighted distribution in the 2D collective variable (CV) landscape of the backbone RMSF and all contacts (intra-monomer and interfacial) for the whole complex for all snapshots in the last 50 ns of the simulation. (MD monomer metrics) shows the distribution of RMSD (monomer RMSD), RMSF (monomer RMSF), monomer contacts and interfacial contacts for all monomers and interfaces in the complex over the last 50 ns of the simulation.
[0047] [Figure 12]Validation of designed conformational-changing oligomers using Rosetta and molecular dynamics is shown. (A-D) Validation of de novo designed trimers, tetramers and pentamers of monomeric length 32 amino acids using Rosetta and molecular dynamics. (Complex structures) shows the AlphaFold predicted structure (light grey) overlaid with the best of the 10 relaxations (dark) using the Rosetta all-atom scoring function. The structures are shown together with their RMSD to the relaxed structures, as well as the Rosetta binding energy of the single monomer to the rest of the complex dG, and the packing statistics packstat. (Monomer structures) shows the monomeric AlphaFold predicted structure (dark grey) overlaid with the complex structure (light grey) and shows the aligned RMSD. (Prediction error) shows the predicted aligned confidence of the AlphaFold prediction for each monomer (left) and the complex (right). For molecular dynamics-based validation, (RMSD / F evolution) shows the time evolution of the RMSD (dark grey) and RMSF (light grey) of each complex over the course of a 100 ns explicit solvent molecular dynamics simulation. (Contact evolution) shows the time evolution of all intra-monomer contacts as well as interfacial contacts over the course of a 100 ns simulation. (CV distribution) shows the Boltzmann-weighted distribution in the 2D collective variable (CV) landscape of the backbone RMSF and all contacts (intra-monomer and interfacial) across the complexes for all snapshots over the last 50 ns of the simulation. (MD monomer metrics) shows the distribution of RMSD (monomer RMSD), RMSF (monomer RMSF), monomer contacts and interfacial contacts for all monomers and interfaces in the complex over the last 50 ns of the simulation.
[0048] [Figure 13] An experiment of the predictions is shown.
[0049] [Figure 14](A) Snapshots of monomers at increasing fitness function values (pAC) during optimization. As the fitness function value (pAC) increases from 0.7 to 1.5, the predicted structure (top) changes. Concurrently, the local confidence for each amino acid (pLDDT, top) and the predicted aligned confidence for each amino acid pair (pAC, bottom) increase. (B) Snapshots of heterodimers at increasing fitness function values (pAC) during optimization. As the fitness function value increases from 0.8 to 1.5, the structure of each monomer changes, and the monomers are predicted to be closer together. In parallel, the predicted aligned confidence per monomer and at the interface approaches 1.0, indicating improved confidence in heterodimer formation of the designed monomers. (C) Schematic representation of the fitness function components involved in the design of monomers and dimers. For monomer design, the predicted LDDT is maximized at each sequence position, the pAC is maximized for each pair of amino acids, and the average and maximum protein radii are constrained to ensure compactness. In the design of dimers and oligomers, the pAC is further maximized, both for each constituent monomer and for the entire complex.
[0050] [Figure 15]The results of a toxin inhibitor conjugation (TIC) screen of blockers designed in silico according to the present disclosure are shown. (A) TIC screen against RcaT-Sen2 with IPTG-inducible high copy plasmids expressing 88 designed blockers (donor-library). The arrayed donor library (384 density format) was conjugated with E. coli BW25113 recipient carrying an arabinose-inducible RcaT-Sen2 plasmid. Exconjugants carrying these two plasmids (i.e., the above E. coli recipient carrying one IPTG-inducible plasmid as well as an arabinose-inducible plasmid) were selected, and colony opacity was measured using the Iris tool
[58] to assess the fitness of the exconjugants. For each exconjugant strain, opacity was normalized to the average opacity of the GFP-expressing negative control (which does not affect toxin activity) on each plate. The mean blocking scores (see x-axis) were calculated by dividing the mean normalized opacity of each strain on experimental plates containing arabinose + IPTG (RCaT and blocker induction) (n=2 biologicals) by the mean normalized opacity of control plates containing IPTG only (blocker induction) (n=2 biologicals). Mean blocking scores (cut-off > 1.5) with SD are shown, neutral scores are shown in grey, and hits in black have p < 0.01 by t-test compared to the mean blocking scores of experimental and control plates. (B). Same results as in (A) but using recipient E. coli BW25113 carrying the RcaT-Eco9 plasmid. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0051] The method of this document is a search problem to find a set of sequences of amino acids of a protein (amino acid sequences) whose mathematical fitness function exceeds a fixed threshold for fitting a design goal of the protein. The mathematical function is a combination of the sequence of amino acids that form the protein, the predicted structure of the protein, and an AF confidence measure, such as pAE or pLDDT (described later). The design goal of the protein may simply be proper protein folding, or it may be a biological function, such as binding to a target protein, dimerization, self-assembly, etc. The method uses AF to take into account both the properties of the protein sequence of amino acids (also called "amino acid sequence") as well as the all-atom protein structure of the protein sequence of amino acids, and incorporates these into a fitness function to provide a protein structure prediction as well as a measure of prediction confidence.
[0052] The methods presented here are combined with Rosetta ab initio structure prediction
[21] and validation using molecular dynamics simulations.
[0053] It is outlined in Figure 1. The optimization loop is set up to iteratively mutate a sequence pool of amino acid sequences as shown (see Figure 1A), score the sequence pool of amino acid sequences with the output of AF to model the fitness function, and then update the sequence pool with the mutated amino acid sequences and scores. In one embodiment, the initial pool of amino acid sequences (or protein sequences of amino acids) is created by sampling fixed length amino acid sequences uniformly randomly from the 20 standard amino acids. In another embodiment, it is possible to exclude some of the standard amino acids, such as cysteine, from the creation of the initial pool of amino acid sequences to avoid designing proteins containing disulfide bonds. Alternatively, the initial pool may be started or created by sampling from a predefined probability distribution of amino acids, for example as given by the statistics of natural amino acids or natural amino acid sequences, or from a predefined probability distribution of amino acid sequences. In one embodiment, during the execution of the method of the present disclosure, these predefined probability distributions may be adjusted, for example, based on properties such as the score of the sequence pool of amino acid sequences, or based on recent scientific findings.
[0054] In contrast to the previous approach using trRosetta, we use the AF program to obtain structure predictions [4]. The optimization objective, i.e. the fitness function, is defined in terms of the protein sequence, the full-atom structure, and a mathematical function of the AF confidence. The fitness function can be composed of fitness function components, i.e. mathematical expressions. The composition of the fitness function by fitness function components allows for a flexible definition of the fitness function according to the structural properties and / or biological and / or physicochemical functions of the protein to be designed. The optimization objective represents the characteristics of the desired protein design. The properties and / or biological and / or physicochemical functions may relate to one or more of the structure of the designed protein, the binding properties of the designed protein, the switch between conformations of the designed protein, the conformational change of the designed protein in response to environmental conditions, the amino acid distribution of the designed protein (e.g., human-like amino acid distribution for immune evasion), the amino acid sequence distribution of the designed protein (e.g., human-like amino acid sequence distribution), the function of the protein (e.g., enzymatic activity), including predicted functions predicted by a protein function predictor, the secondary structure of the designed protein, the oligomeric state of the designed protein, the propensity for post-translational modifications (PTMs) of the designed protein, including predicted PTMs predicted by a PTM predictor.
[0055] The first step in predicting protein structure using AlphaFold is to search a large database of protein sequences for similar protein sequences of one or more input protein sequences. The protein sequences may be of known origin and / or metagenomic origin. Once some similarity of the input protein sequences is found, the newly found protein sequences are aligned to the input protein sequences. This alignment is done by trying to match as many amino acids in the similar protein sequences as possible with the corresponding some of the amino acids in the input protein sequences. Non-matching amino acid sequences in the input protein sequences are assigned gaps "-". This produces a set of multiple sequence alignments. At this stage, it is understood that the aligned similar protein sequences do not have structural information associated with them. In a second optional step, the structures of those aligned protein sequences for which structural information is known can be input to AlphaFold as structural templates.
[0056] The AF program processes a single protein sequence (i.e., a single amino acid sequence) from the initial sequence pool (or initial pool of amino acid sequences) as an input protein sequence (i.e., an input amino acid sequence) of amino acids together with a corresponding one of the multiple sequence alignments (as described above) to predict the structure of the single protein sequence (or single amino acid sequence) by constructing an alignment of amino acids contained only in the protein sequence using the multiple sequence alignment [7]. The multiple sequence alignment predicts which amino acids contained in the protein sequence are located close to each other (see the discussion of predicted aligned error or pAE below). The method further returns the whole-atom structure of the single protein sequence (or single amino acid sequence) and a measure of the prediction confidence (see Figure 1A, AlphaFold output). The value of the prediction confidence measure is expressed as a combination of the predicted local distance difference test (pLDDT)
[39] and the predicted aligned error (pAE) [4]. The pLDDT test measures the quality of the local model, and the pAE provides a measure of confidence for each amino acid pair in the protein sequence. The protein sequence is then optimized to maximize the output fitness function (L) (i.e., predicted local distance difference test (pLDDT) and predicted aligned error (pAE) of the protein sequence, predicted all-atom structure, and the aforementioned confidence measures). This optimization can include minimizing the root mean square deviation (RMSD) of the protein relative to a reference structure, and maximizing the AF prediction confidence. The RMSD is calculated from the difference between the structure of the protein predicted from the known specific backbone structure of the protein. Since the method is suitable for predicting complex structures [6], the fitness function can also constrain the oligomeric state or conformational changes of the protein upon binding (see FIG. 1B). In general, any function of protein structure, amino acid sequence, and prediction confidence can be used for the optimization.
[0057] For the protein sequences of amino acids in the sequence pool, the value of the fitness function of the protein sequence of the amino acids is calculated. The protein sequences are further recombined and mutated by an optimizer to explore the protein sequence space. "Mutation" means replacing one or more amino acids in the protein sequence by different amino acids randomly and uniformly sampled. "Recombination" means selecting two sequences from the sequence pool and randomly exchanging random stretches of both sequences. The sequence pool is updated with the recombined and / or mutated sequences.
[0058] An evolutionary algorithm is used to update the sequence pool, for example by mutating and / or recombining protein sequences (amino acid sequences). The following
[40] was used; however, other evolutionary algorithms or other gradient-free or gradient-based optimizers can be substituted as desired.
[0059] The evolutionary algorithm is implemented as follows: A random sequence pool of N sequences is used to evaluate the fitness function for the amino acid sequences to obtain a fitness function value t associated with the amino acid sequence. In the next step, the maximum fitness function value t max In the next step, the fitness function value t ≧ t is determined until a certain number “M” of new amino acid sequences (or protein sequences) are identified. max -tolerance·t max The protein sequences in the sequence pool with a certain tolerance of suboptimal range (tolerance t max). The selected amino acid sequences (or protein sequences) were further recombined and mutated and improved with a so-called greedy method for each recombined / mutated one of the amino acid sequences (or protein sequences). The term "greedy improvement" means that only those amino acid sequences (or protein sequences) that improve their fitness function value t over their unmutated starting amino acid sequences (or protein sequences) were selected. Finally, the top N generated amino acid sequences (or protein sequences) were selected and the algorithm was repeated afterwards.
[0060] Through optimization, the protein changes its structure to increase both a local confidence measure, e.g. based on pLDDT, and a global confidence measure, e.g. based on pAE (as shown in FIG. 1C). For a given protein design task, a threshold is selected above which the sequence of amino acids of the protein (or amino acid sequence) is considered optimized. The amino acid sequences (or protein sequences) that are above this threshold are returned from the method, and a subset of the returned amino acid sequences (or protein sequences) is provided for validation, e.g. using molecular dynamics simulations and Rosetta ab initio structure prediction
[21] . To reduce the time required for processing the subset in the Rosetta program, only the subset is provided.
[0061] Fitness Function Using gradient-free methods, any function of amino acid protein sequences (or amino acid sequences), whole-atom structures, and AF confidence can be used as fitness functions. This flexibility allows many possibilities in designing amino acid sequences (or protein sequences) and protein complexes that have some specific properties and / or biological and / or physicochemical functions. Fitness functions may be specified to represent the goal of the design task.
[0062] A protein complex is a (possibly transient) association of multiple proteins. For this method, multiple protein sequences are designed simultaneously, rather than a single protein sequence.
[0063] The methods described herein focus on fitness functions using the aforementioned confidence models (or metrics) and protein backbone geometries (or backbone structures), however, this is not intended to limit the invention and other fitness functions can be used.
[0064] The confidence model returns two main measures (or confidence metrics) of the confidence of the protein structure prediction. As mentioned above, these confidence measures are the predicted local distance difference test (pLDDT)
[39] and the predicted aligned error (pAE) [4]. pLDDT gives a local confidence measure for each amino acid in the protein sequence. pAE gives the predicted error for the position of an amino acid in the local coordinate system of other amino acids. For both confidence measures, the method predicts a binned distribution of values over the amino acids in the protein sequence. The pAE values are converted to predicted aligned confidence (pAC), where pAC=1-μ pAE / pAE max Wherein, μ pAE is the average pAE, and max is the center of the bin with the highest pAE. Alternatively, one can work with the predicted template matching score (pTM)
[41] . Optimization of pLDDT and pAE (e.g., high pLDDT and low pAE) provides a prior distribution over the amino acid sequence, based on which native-like protein structures can be designed. A native-like amino acid sequence is an amino acid sequence for which the sequence of amino acids (i.e., or a series of amino acids) is equal to or greater than the minimum probability under the probability distribution observed in nature. Such a probability distribution observed in nature can be determined based on sequences and structures deposited in the PDB. Thus, a native-like amino acid sequence has:
[0065] The fitness function may include terms that maximize the value at a confidence level or confidence metric given by:
number
[0066] AF returns the coordinates of all atoms of a protein [4], which allows for the introduction of geometric (or structural) features into the fitness function in a further embodiment of the method. Fitness function components, e.g., mathematical expressions, can be introduced into the fitness function depending on any geometric features of the protein sequence and its predicted structure. These geometric features can be any geometric features of the designed protein that need to be optimized. For example, geometric features include, but are not limited to: Arbitrarily weighted sum / product of AlphaFold confidence measures Protein secondary structure determined from predicted 3D structure -Protein radius of gyration Difference in protein structure against some reference protein structure Distribution of amino acids in a protein sequence - Protein solubility predicted from protein sequence and structure The function of the protein, e.g., physicochemical and / or biological function, predicted by a separate model from the protein sequence and structure For protein complexes: the relative positions of the centers of mass of the constituent monomers or alternatively any weighted sum / intersection of these geometric features.
[0067] The main building blocks of the fitness function components introduced into the fitness function are measures of the difference between two sets of coordinates. The structural metrics of the following proteins used within the fitness function are: The aligned error[4] is given as follows: AE ij (x,y)=||T xi x i -T yj y i || In the formula, T x,I is a vector of backbone atom x i This is the aligned error used in training AF [4].
[0068] Frame-Aligned Point Error (FAPE) is given by [4],
number
[0069] The distance RMSD is given by: dRMSD(x,y)=RMSD(d(x),d(y)) The function d(x) calculates the distance between each pair of alpha carbons in the backbone of an amino acid and is calculated as follows: Given the structure of a protein, take the 3D coordinate "x" of the alpha carbon atom of its backbone (i.e., the atom in the amino acid to which the side chain is attached). The function d(x) calculates the distance between each pair of alpha carbons. In practice, this function provides enough information to recover all the 3D coordinates of the alpha carbon atoms up to the global translation, rotation, and reflection, which is why the pairwise distances can represent the actual 3D coordinates in the fitness function. If the pairwise distances for the two structures are the same, then the actual 3D coordinates are the same up to the global translation, rotation, and reflection of the protein sequence.
[0070] The template matching score
[41] is a measure of the structural similarity between protein structures,
number
number
number
[0071] L TM (x,y) is an approximation of the template matching score that does not require structural alignment [4]. The template matching score scale is normalized to 1, providing a smooth fitness function for very different structures.
[0072] Furthermore, we use a compactness fitness function L that penalizes large protein radii. comp can be used to enforce constraints on the overall shape of the protein,
number
[0073] Monomers. To generate globular monomeric proteins, the confidence fitness function (or confidence metric) was combined with a weighted version of the compactness fitness function as follows:
number
[0074] This allows a trade-off between compactness and reliability in the case of elongated structures. However, it is understood that the method can be used for other protein design and is not limited to globular monomeric proteins. The fitness function to assess the globularity of a protein is relatively simple. This is not so simple when doing this to determine whether a part of a protein is, for example, a transmembrane domain. The method can be used to design other types of proteins, given a fitness function to determine this from the protein's sequence and / or structure.
[0075] Protein complexes. In the design of protein complexes, the same fitness function L mon were used for the constituent monomers in the protein complex. This ensures that each monomer has a stable high-confidence structure. Previous studies have identified inter-monomer pAE as a predictor of complex formation
[19] . A low inter-monomer pAE (or equivalently, a high inter-monomer pAC) corresponds to a high confidence in the complex prediction. As mentioned above, pAE is the predicted error in distance and relative orientation for each pair of amino acids in the protein sequence. Assuming that the prediction is calibrated (i.e., the predicted uncertainty corresponds well to the true uncertainty), a low predicted error pAE in inter-monomer residue pairs corresponds to a highly constrained distance and relative orientation between the monomers. This is only the case when the proteins in the protein complex interact with each other. For some of the non-interacting proteins, the uncertainty in distance / relative orientation is high because the proteins are never constrained in their distance / orientation relative to each other.
[0076] Therefore, another combination of confidence and compactness fitness functions is imposed on the complex protein, e.g.,
number
[0077] Conformational change. To guide the conformational change of the protein complex upon complex formation, we use the conformational change fitness function L CC Define the fitness function L CC Let L be the fitness function defined above. cpx (X,X m ),
number
[41] between the protein structure as part of a protein complex and its structure as a monomer.
[0078] Protein-protein interactions. The interfacial pAE (ipAE) correlates with the probability of a protein-protein interaction. By optimizing for a low ipAE between two protein monomers, it is possible to engineer protein-protein interactions. In combination with additional confidence and geometry fitness function components that impose constraints (see Table 1 below), the de novo monomeric, oligomeric fitness function (L denovo ) and the fitness function (L binder ) can be defined. The sequences for which AF predicts different states in monomers and oligomers are compared with the corresponding fitness function (L change ) can be explored.
[0079] The fitness function component of the fitness function includes multiple confidence measures (e.g., pAE, interface pAE, pLDDT). The fitness function component of the fitness function includes the geometric constraints (radius of gyration R g , residue pair X i and X j The maximum distance d(X i ,X j The confidence fitness function component is determined by whether the generated structure candidates are native-like (L conf ), the desired protein-protein interaction (L ipAE ) The geometry constraints constrain the predicted protein structure to fit a designed topology and / or geometry.
[0080] In addition to the aforementioned components of the fitness function, some further components are listed and described in the table below.
[0081] [Table 1]
[0082] Representation of the input. An amino acid sequence (or a protein sequence of amino acids) is represented as a one-hot encoding array of 20 classes, one class for each standard amino acid. This allows the use of both gradient-free and gradient-based optimizers, since gradients can be estimated through the one-hot representation using a straight-through estimator or similar
[42] . As input to AF, an additional multiple sequence alignment is constructed using only the single input sequence and a blank template feature. A blank template feature is an empty template feature. Instead of passing one or more templates as input to AlphaFold, no templates are passed to AlphaFold at all. However, AlphaFold still expects an input that contains information about the templates. The use of the term "blank" template feature refers to a template input that does not actually contain a template.
[0083] Sequence template. In order to reduce the size of the search space and implement precise sequence constraints without introducing additional fitness functions, the optimized amino acid sequence (or protein sequence of amino acids) is represented as a pair (S,T),
number
number
number
[0084] Representation of complexes. AF was originally designed to accept only a single protein sequence as input. To enable AF to predict the structure of protein complexes, the principles described in [7] were followed. In this case, chain breaks to predict the structure of complex proteins are represented by introducing an increment greater than 32 in the residue index of each additional chain beyond the first chain. This increment is used to separate residues on chains that differ by more than 32, the maximum relative residue index difference embedded separately by AF [4]. Furthermore, following [7], multiple sequence alignment features are split so that each monomer has a gap at the sequence position of the other monomer and a separate copy of its sequence alignment feature.
[0085] This can be illustrated by analogy using the sequence (in pseudo-fasta format). Assume the array is as follows: >1 MERRY CHRISTMAS >2 HAPPY NEW YEAR >3 HAPPY EASTER Suppose AlphaFold has to predict the structure of those complexes, the input to AF has the following multiple sequence alignment (in pseudo-fasta format): >1:2:3 MERRYCHRISTMASHAPPYNEWYEARHAPPYEASTER >1 MERRY CHRISTMAS --------------------------- >2 ------------HAPPYNEWYEAR---------- >3 ---------------------------HAPPYEASTER A "-" indicates a gap (i.e., unmatched amino acids) in the multiple sequence alignment. That is, for a complex of N proteins, a multiple sequence alignment of N+1 sequences is provided. The first of the sequences is the sequence that concatenates all of the single protein sequences. The k+1 sequence is the sequence of the kth input protein, with the number of left gaps ("-") equal to the total length of the 1st to k-1th proteins and the number of right gaps ("-") equal to the total length of the k+1th to Nth proteins.
[0086] design research Optimization. For optimization, an evolutionary strategy optimizer according to
[40] was used as described above. The population size was set to 10. During mutation, the population size was increased by a factor of two. Protein sequences of amino acids with a suboptimal degree of up to 10% (in other words, the tolerance value above was set to 0.1) were considered as targets for mutation and recombination. Recombination was applied by crossover with a probability of 10% at each position in the sequence of amino acids. In this context, crossover means taking two amino acid sequences, and the creation of the recombined sequence is first performed by taking an amino acid from the first sequence. At each amino acid, with a probability of 10%, the sequence is switched from the sequence in which the amino acid is currently taken to the other sequence.
[0087] The optimization is performed on those sequences by mon , L cpx For L>1.5, and L cc The fitness function including was considered complete when L>0.9. Tables II and III show further parameters of the optimization runs that were performed.
[0088] AlphaFold configuration. For optimization, AF was configured for single sequence use with ensemble templates, extra MSA features disabled, and the number of MSA features restricted to the number of modeled monomers. The number of AF iterations (recycle steps) was kept as a parameter for each optimization run (see Tables II and III below). For larger protein complexes, the number of iterations was reduced to two to speed up the computation. The parameter set model_1_ptm was used for all experiments.
[0089] [Table 2]
[0090] [Table 3]
[0091] Rosetta Verification The designed protein sequences were validated using the Rosetta suite of protein design and structure prediction tools
[21] . Ab initio structure prediction based on the fragment assembly method
[43] was used as an independent baseline (i.e., an independent method) for the designed protein sequences and protein structures.
[0092] Ab initio structure prediction is performed as follows. Protein secondary structures were predicted using PSIPRED
[44] and S4PRED
[45] to obtain protein secondary structures for fragment selection. 3-mer and 9-mer fragments were selected from the Rosetta fragment database based on secondary structure and protein sequence information. Alignment information was not used because the designed protein sequence did not have sufficient homology to natural protein sequences. Ab initio structure prediction was performed by a fragment assembly method using the Abinitio Relax protocol. Two starting conformations were evaluated. The first starting conformation started from an extended conformation with all backbone torsion angles set to 180° (in fact, this is the default starting conformation for ab initio structure prediction using the Abinitio Relax protocol), and the second starting conformation started from the AF predicted structure.
[0093] It is understood that there is no a priori information about the protein structure for ab initio structure prediction. The prediction process must start from some initial conformations that are theoretically plausible conformations of the amino acid chain. This initial conformation does not necessarily have to be an extended conformation, but an extended conformation is a good initialization because it is easier to set up. Also, in contrast to starting from, for example, a random conformation, this initial conformation does not allow the protein backbones to clash. However, it is noted that within a few steps of the ab initio structure prediction process, the structure is already changed to something closer to the actual protein structure.
[0094] In both runs of the two starting conformations, 32000 decoys (conformations with low free energy) were generated. The decoys from both runs were pooled and scored using the Rosetta all-atom energy function
[46] . Overall, the 10 lowest energy decoys were selected for relaxation. The decoys were relaxed using the Relax protocol. In cases where the ab initio structure prediction failed to find a minimum energy structure, the AF predicted structures were also relaxed to generate 10 further decoys. The decoys were rescored using the same Rosetta scoring function to calculate energies that could be directly compared. It is understood that the ab initio structure prediction and relaxation use different scoring functions in Rosetta. Therefore, to be able to compare the quality of the structures between both of these protocols, the output protein structures need to be scored again (rescored) using the same scoring function. The lowest energy relaxed decoy was selected as the final predicted structure and its Rosetta score and RMSD against the AF predicted structure were evaluated.
[0095] Protein-protein interface analysis. Because Rosetta does not allow ab initio structure prediction of protein complexes other than for symmetric homo-oligomers, interfaces of designed complexes were assessed using the Interface Analyzer protocol
[47] . Protein structures were relaxed and Rosetta binding energies were calculated by removing a single monomer from the complex and subsequent repacking. As a further measure of interface quality, packing statistics were calculated using the PackStat protocol
[48] , which detects unpacked regions within the interface that are not solvent accessible.
[0096] Verification based on molecular dynamics simulation A subset of the designed monomers and higher order protein complexes of the protein sequences were further validated by performing all-atom molecular dynamics (MD) simulations in explicit solvent to analyze the structural flexibility and the corresponding properties of the internal / intramonomer and interfacial contacts (for multimeric proteins and protein complexes). Simulations performed in explicit solvent means that the protein (or protein complex) is "put in a box" in which all the water molecules and all the salt ions are explicitly included. This term is in contrast to the "implicit solvent" method, in which the water molecules and ions are not explicitly simulated, but instead are modeled by the modification of the force field.
[0097] Initial system construction. All protein parameters were described using the standard AMBER force field (ff14sb)
[49] . Each protein or protein complex was solvated using atomistic TIP3P water with a minimum padding of 10 Å to form a cubic periodic box, then electrically neutralized with an ionic concentration of 0.15 M NaCl using standard ionic parameters.
[0098] A standardized minimization, equilibration and simulation protocol, Simulation Protocol A, containing 11 stages, was developed for the solvated protein structure and the solvated protein complex structure systems. A set of restraints (RS) was applied to each system at a specific stage of equilibration. These restraints included restraining all heavy (i.e., non-hydrogen) atoms of the protein. Each system was then minimized over four stages with 1500 steps (500 steepest descents + 1000 conjugate gradients) of minimization, with different force constants in each successive stage, i.e., 10 kcalmolA in stage 1, 10 kcalmolA in stage 2, 10 kcalmolA in stage 3, and 10 kcalmolA in stage 4. 2 , and 5 kcallmolA in stage 2. 2 , 1 kcallmolA in stage 3 2In stage 3, unconstrained RS was applied. MD simulations were performed in all subsequent stages. The SHAKE algorithm was used for all atoms covalently bonded to hydrogen atoms. A time step of 2 fs was used. Long-range Coulomb interactions were treated using a GPU implementation of the particle mesh Ewald summation method (PME)
[50] . A nonbonded cutoff distance of 10 Å was used. In stage 5, each system was simulated with RS (k = 10 kcal / molA 2 ) and heated from 10K to 300K in 1ns. Then, the temperature was increased to γ=5.0ps -1 The temperature was maintained at 300 K using a Langevin thermostat with a damping constant of τ 0.05, and in stage 6 the system was allowed to equilibrate at constant volume, and thus in the NVT ensemble (i.e., constant particle number N, constant volume V, and constant temperature T) for 1 ns. The pressure relaxation time τ p The pressure was maintained at 1 atm using a Berendsen barostat at 1.0 ps, and for each of the subsequent stages the system was run in the NPT ensemble for 100 ps, with k = 10 kcal / molA in stage 7. 2 , and in stage 8, k = 5 kcal / molA 2 , and in stage 9 k = 1 kcal / molA 2 , and k = 0.5 kcal / molA at stage 10. 2 The systems were simulated with RS of 10 ns. Finally, in stage 11, the RS constraints were removed and the systems were simulated in the NPT ensemble for another 5 ns. This was followed by 100 ns of main simulations for each system in the NPT ensemble under the same conditions as in stage 11. Coordinate snapshots from the main simulations were generated every 10 ps, resulting in 10,000 snapshot trajectories per system.
[0099] Structural flexibility and contact analysis. Structural stability and flexibility were analyzed by calculating the root mean square deviation (RMSD) of the protein backbone atoms relative to the initially predicted protein structure (after protein backbone alignment) and the root mean square fluctuation (RMSF) relative to the average protein structure in our MD. For protein complexes, this analysis was performed separately for both the whole complex protein and the individual monomeric proteins. The number of side chain-side chain contacts within a monomer was calculated for each snapshot of our MD based on a heavy atom distance threshold of 4 Å. Similarly, interface contacts were calculated for each interface in the simulated protein complex based on the same criteria. A global view of protein contacts was obtained by summing all intramonomer and interface contacts.
[0100] Intramonomer contacts are contacts within a monomer. These intramonomer contacts provide an indication of the stability of the monomer. If the number of intramonomer contacts decreases throughout the MD simulation, this indicates protein instability, i.e., the protein is being pulled apart. Interfacial contacts represent intermonomer contacts between each pair of monomers. These interfacial contacts indicate the stability of the protein-protein interaction. If the number of intermonomer contacts decreases throughout the MD simulation, this indicates that the monomers are dissociating from each other.
[0101] Finally, for all simulated systems, we computed the Boltzmann-weighted distribution (G = -k B Determine the approximate potential of mean force (G) by calculating Tln(ρ), where ρ is the normalized frequency within the binned 2D landscape.
[0102] Structural Clustering Designed protein sequences of 64 amino acids or more in length were subdivided into a dictionary of fragments using Geometricus
[51] . Fragments were collected using the k-mer method with k=16 and the radius method with a cutoff of 10 Å. A fragment is a shorter chain, or a local neighborhood of an amino acid. Fragments collected using the radius method include all amino acids around a given amino acid that have at least one atom within the cutoff radius. In practice, this is the amino acids in the chains around the given amino acid, as well as other amino acids that are close in the 3D structure but not necessarily close in the amino acid sequence.
[0103] A collection of feature representations was computed for all designed proteins and the dimensionality was reduced using non-negative matrix factorization (NMF) with 50 components
[52] . This is done as follows: We counted the number of fragments in each protein structure that correspond to each fragment cluster identified by Geometricus. We returned a vector of these counts, which is a bag of feature representations. Proteins were separated into 10 clusters using Ward linkage agglomerative clustering on their NMF components. As recently shown by results [7], AlphaFold was found to predict the structure of protein complexes, and AF trained on the structure of single-chain proteins can be used to predict the structure of protein complexes. To gain confidence in the suitability of AF for designing protein complexes using the methods of this document, we first predicted the structures of dimers, trimers, tetramers and pentamers of small proteins (Figures 2 and 8: AF predictions show low root mean square deviations (RMSDs) of native structures <3 Å for the protein complexes considered). This could be due to these structures being part of the PDB and therefore part of the training dataset of AF [4], but together with recent studies on protein complex prediction using AF [6], it gives some indication that AF may be suitable for designing protein complexes using the methods of this document.
[0104] Single sequence optimization to design diverse monomers As a first step towards de novo protein design based on AF, amino acid sequences of lengths 32, 64, 128 and 256 amino acids were generated to design monomeric globular proteins (Fig. 3). The amino acid sequences designed for the monomers showed a high predicted aligned confidence exceeding 0.82 average predicted aligned confidence (Fig. 3, prediction error) and an average pLDDT (normalized to the range [0,1]) of 0.83. This indicates that AF gives high confidence in the predicted structures, since a pLDDT of 0.7 or higher corresponds to a reliable prediction of the backbone structure [4,5]. To validate the predicted structures of AF, fragment-based ab initio structure prediction using Rosetta
[21] was performed on a randomly selected subset of designs starting from the extended amino acid chain as well as from the structure of AF (see Fig. 8). For each monomeric designed protein sequence, the 10 lowest energy Rosetta structures were extracted. We extracted these extracted structures of the monomer as well as the structure of AF using the Rosetta force field
[46] and selected the lowest energy structure. For the majority of the monomer structures, the RMSD between the AF structure (gray, structures in Figure 3(A-D)) and the lowest energy structure (red) is less than 3.0 Å. For protein structures with distances above this threshold (Figure 3(C): result 8), the larger deviations seem to be due to flexible α-helices, and the predicted structures of AF have relatively low Rosetta energies.
[0105] To visualize the energy landscape of folding for each of the designed monomers, the distribution of decoys was examined for their ab initio folding with Rosetta energy
[46] and RMSD against the structure of AF (Fig. 3(A-D), Rosetta score distribution). For monomers with sizes of 32 and 64 amino acids, it was noted that the relaxation of the predicted structure (blue) has a comparable but higher energy compared to the best 10 decoys from the ab initio prediction (red) (Fig. 3(A,B), Rosetta score distribution, relaxation). Surprisingly, for monomers with sizes of 128 and 256 amino acids, the relaxed structure of AF (blue) shows a much lower energy compared to the ab initio decoys (red) (Fig. 3(C,D)). This indicates that the ab initio structure prediction was not able to find a plausible minimum energy structure for those monomers. Indeed, the distributions of monomeric decoys with sizes of 32 and 64 amino acids show a funnel-shaped distribution with a single minimum in both low energy and low RMSD characteristic of protein folding (Figure 3(A,B), Rosetta score distributions, ab initio). In contrast, for sizes of 128 and 256 amino acids, the distributions are much more diffuse and show no clear minimum, indicating that the ab initio predictions are unable to find a plausible minimum (Figure 3(A,B), Rosetta score distributions, ab initio). However, the result that the relaxed AF structure shows a very low Rosetta energy
[46] increases our confidence that the designed structure is valid.
[0106] As a further validation step, the structural stability of the designed monomers was assessed by performing all-atom molecular dynamics (MD) simulations. Boltzmann-weighted frequency distributions in the 2D ensemble variable (CV) landscape consisting of backbone RMSF and intramonomer contacts show that most systems exhibit a clear unimodal distribution centered around low RMSF and a significant number of contacts. This suggests that the majority of conformers in our ensemble are sampled in a narrow range, consistent with the maintenance of structural stability (Figure 3(A-D), MD metrics). Furthermore, the time evolution of the backbone RMSD relative to the initial AF predicted structure and the backbone RMSF relative to the MD average structure are stable at less than 4 Å and 2 Å, respectively, for the majority of the monomer systems (see Figure 8). Similarly, the time evolution of the intramonomer contacts remains stable for the simulated monomers, suggesting the maintenance of the folded structure. As expected, the number of contacts increases with monomer size, ranging from about 50-100, about 100-200, about 300-400, and about 600-800 for monomers of length 32, 64, 128, and 256 amino acids, respectively (see FIG. 8).
[0107] To evaluate the diversity of structures generated by our method, we clustered monomers ≥ 64 amino acids in size and extracted representative structures for each cluster (Figure 3(E,F)). The structures were embedded using Geometricus
[51] and agglomerative clustering was performed to visualize their distribution (Figure 3I)). From visual inspection, representative clusters include α-only (clusters 1, 2, 5, 6, 7), β-only (cluster 3) and mixed αβ proteins (clusters 4, 8, 9, 10) (Figure 3(F)).
[0108] Sequence pairing exploration allows de novo dimer design We next applied this method to the de novo design of protein homo- and heterodimers. We optimized dimers with monomer sizes of 32, 64, and 128 amino acids. Similar to the previous monomer design, the predicted dimers show high predicted aligned confidence for the structure of the AF, with an average pAC > 0.75 and an average pLDDT > 0.83 (Figure 4 (A-D), prediction error). From visual inspection, the designed interfaces show a high level of shape complementarity. Instead of ab initio structure prediction, which is generally difficult to access for protein complexes using Rosetta, we calculated the Rosetta binding energy between monomers as a measure of structure quality, as well as a score of interface packing statistics
[48] . We applied this validation to a subset of designed dimers (see Figure 9). We relaxed the dimers using the Rosetta force field
[46] and note that all relaxed structures have low RMSD to the AF predicted structure (Figure 4, (A-F), complex structures). Furthermore, most of the dimers exhibited packing statistics scores above 0.6, which are comparable to the packing statistics of crystal structures at 2.0 Å resolution [90, 103]. Most structures exhibited Rosetta binding energies better than -40REU, suggesting that the predicted interfaces are stable in Rosetta
[46] .
[0109] MD simulations of dimeric and higher oligomeric systems allow the analysis of both intramonomer contacts as well as interfacial contacts between monomers. Boltzmann-weighted frequency distributions in CV space of global RMSF and total contacts (intramonomer and interfacial) show mostly unimodal distributions with low average RMSF (≦2 Å) and a significant number of total contacts that increases with the size of the complex (Figure 4, (A-F), MD metrics). Although significant deviations are observed in some of the dimeric systems, many of the dimeric systems show a stable temporal evolution of global RMSD and RMSF (see Figure 9). Detailed analysis of RMSD, RMSF, intramonomer and interfacial contacts shows that both monomers in either homodimeric or heterodimeric systems exhibit similar flexibility as well as the number of intramonomer contacts. Each dimeric system has a significant number of interfacial contacts, but as expected, these are less than the corresponding intramonomer contacts.
[0110] For synthetic biology applications, orthogonality is a desirable property for dimers
[16] . That is, pairs of monomers designed to dimerize should only bind with their designed partners and not with monomers of other designed dimers. As a preliminary test for the orthogonality of dimers designed using our method, we predicted all combinations of monomers for designed dimer pairs of monomers 32 amino acids in size (Figure 4(G)). On-target complexes were predicted with high confidence by AF (Figure 4, (G), on-target prediction). In contrast, for off-target complexes, i.e., complexes of monomers not designed for dimer formation, the average predicted aligned confidence for amino acid pairs between monomers dropped below 0.5. This indicates that AF cannot predict these off-target combinations as dimers with any confidence, providing preliminary evidence that dimers designed using the method outlined in this document can indeed exhibit orthogonality.
[0111] Multiple sequence optimization enables homo-oligomer design To demonstrate the feasibility of designing protein complexes beyond dimers, we designed homo-oligomers ranging from trimers to hexamers. Monomers of 32 and 64 amino acid sizes were considered for trimers, tetramers, and pentamers, and 32 amino acids for hexamers. As before, a subset of oligomers was evaluated for predicted aligned confidence in AF, Rosetta binding energy, and packing statistics and behavior under 100 ns molecular dynamics simulations (see Figure 10). Designed structures showed pAC and pLDDT > 0.7 for the majority of structures (Figure 5).
[0112] The complex structure shows high predicted aligned confidence (prediction error), low RMSD relative to the Rosetta relaxed structure of <2.0 Å, good binding energy of <-40 REU, and packing statistics of >0.59, which are comparable to native protein complexes
[53] (Figure 5(A-G) Complex structure).
[0113] MD simulations of homo-oligomers show similar properties to the dimers, generally exhibiting moderate RMSD (≦6 Å) and RMSF (≦2 Å) as well as the time evolution of a significant number of intra-monomer and interfacial contacts (see Figure 10). The Boltzmann-weighted distribution in RMSF-contact space again shows a unimodal distribution where the average RMSF remains around 2 Å and the total contacts scale with the size of the complex. Detailed analysis of the individual intra-monomer and interfacial contacts shows significant contacts in each monomer and at the interface for almost all systems (Figure 5(A-G) MD metrics).
[0114] Interestingly, inspection of the designed protein complexes reveals a subset of extended oligomers with exposed interfaces on both sides (Figure 5(G-H)). The designed sequences of amino acids can potentially oligomerize into larger assemblies. Correspondingly, MD simulations of these complexes show that, as expected, intramonomer contacts for all monomers and continuous interfacial contacts for all but one interface are maintained. This indicates the feasibility of designing and validating large assemblies in AF that go beyond simple oligomers.
[0115] Designing with fixed monomers to find binders to target proteins As a natural extension to dimer and oligomer design using AF, which has potential applications in biologics design and synthetic biology, we designed proteins that bind to a fixed target protein (e.g., a fixed monomer, also called a template) (Figure 6 and Figure 12, two proteins of 64 amino acids in length were selected). The two proteins were previously designed as target proteins and show distinctly different folds (Figure 6(A,B) target proteins). The amino acid sequence of the target protein was fixed in the design process to optimize the amino acid sequence of the binder protein to be designed. The predicted structures of the designed binder proteins show high aligned confidence of >0.81, good Rosetta binding energy of <-58 REU, and packing statistics of >0.67, which are comparable to natural protein interfaces
[48] (Figure 6(A,B)). This indicates that the designed binders form dimers with the target protein. By visual inspection, it was noted that the designed binders for each target protein all appear to bind at the same interface, indicating a preference for binding sites during the design process.
[0116] MD simulations of the two designed binders for each of the two target proteins show a unimodal distribution in the 2D CV space of global RMSF-total contacts, with an average RMSF < 2 Å and an average number of total contacts ranging from 300–400, consisting of a significant number of individual intramonomer (100–200) and interface (40–120) contacts.
[0117] For the design of binders, AF was input with a template amino acid sequence of the target protein whose structure should remain fixed throughout the design process. Inputting the template allows the design of binders without the need to input the MSA of the target protein. To reduce the computations as much as possible, the amino acid sequence of the target protein was cropped around the desired binding site. The cropped amino acid sequence was input into AF as a template. Such a template keeps the computational cost constant for the design of one or more binders that bind to the target protein, regardless of the size of the target protein.
[0118] Sequence design reveals a signature of conformational changes in AlphaFold To determine whether the sequence space of AF contains information about protein conformational changes when interacting with other proteins, we predicted the structures of two proteins known to have distinct monomeric and oligomeric states: KaiB, a circadian clock protein
[54] , and amyloid-β, which is involved in Alzheimer's disease
[55] .
[0119] As a monomer, KaiB adopts its ground-state structure, KaiBgs
[54] . In complex with KaiC, KaiB changes conformation to a fold-switch stabilized state, KaiBfs
[54] . The AF prediction of KaiB as a monomer shows good agreement with the native KaiBgs structure (RMSD 0.69 Å), while the prediction of KaiB in complex with KaiC shows good agreement with the native structure of KaiBfs (RMSD 1.92 Å). However, the structural alignment of the predicted complex of KaiB with KaiC shows a high RMSD (6.11 Å) to the structure of KaiBgs (Figure 7(A)), indicating that the AF prediction captured part of the conformational changes between KaiBgs and KaiBfs.
[0120] In its monomeric state, native amyloid-β forms an α-helical structure, which undergoes a conformational change to parallel β-sheets upon aggregation (Figure 7(A) grey)
[0105] . Although not accurate in terms of RMSD (>5 Å), AF captures well the transition between the α-helical monomeric state and the parallel β-sheet oligomers (Figure (A) red). This indicates that indeed AF may have learned in a limited way to predict conformational changes upon protein binding.
[0121] Furthermore, we identified a set of oligomer designs that exhibit conformational changes upon complex formation (Figure 7(B)), with 12 designed structures exhibiting a conformational change from a monomeric α-helix (Figure 7(B), monomeric structure) to a stack of parallel β-sheets characteristic of amyloids
[55] . Both the monomeric and oligomeric states show high predicted aligned confidence in AF (Figure 7(B), predicted error), and the oligomeric states show binding energies of <-88 REU and packing statistics of >0.65 for all oligomers, indicating that the complexes are stable in the Rosetta force field
[86] .
[0122] Similarly, MD simulations show a clear minimum in RMSF-contact space centered around low RMSF (<2 Å) and a significant number of contacts (100-500). However, these conformation-switching open-end oligomers are considerably altered compared to other oligomers, including conformation-retaining open-end oligomers, as previously described. The open-ended nature is captured by a significant number of contacts at all but one consecutive interface, while the number of interface contacts (40-80) is comparable to or greater than the number of intramonomer contacts (10-60). This confirms the extended conformation of the monomer in the oligomeric state.
[0123] We designed heterodimers and homo-oligomers that exhibit conformational changes upon complex formation. By maximizing the TM score between the monomeric and oligomeric states during optimization, we were able to find proteins that exhibit the desired conformational changes (Figure 7(C)). Interestingly, the resulting proteins were found to exhibit conformational changes that go beyond the amyloid-like transitions seen in previous design experiments, indicating that more diverse conformationally changing proteins exist in the sequence space of AF.
[0124] Restoring a structure to an array To generate improved alignments for structure candidates recovered from AF, a conditional autoregressive diffusion model (ADM) was trained on protein sequences and structures in the PDB. The conditional ADM can learn to restore a protein's predicted structure to a given protein's amino acid sequence by optimizing the following loss:
number
number
[0125] A masked sequence for ADM inference was initialized, e.g., sampled from the AF, to restore (or redesign) the sequence of the predicted structure subject to constraints. The sequence is masked at all positions that were not specified to maintain (or fix) the initial sequence. For example, in the case of designing a protein binder, the sequence of the binder is completely masked, while the sequence of the target protein is kept intact (or fixed). A random one of the masked positions of the binder sequence is sampled. For this random position, a probability distribution is calculated, and the probability of forbidden amino acids is set to 0. The probability distribution is rescaled by the inverse temperature of 10, and an amino acid is sampled or set for the random position. If the sampled random position is associated with one or more sequence constraints, an amino acid is sampled or set for the position that is affected by one or more sequence constraints. This process is repeated until no masked amino acids remain. In this way, 100 sequences are sampled per predicted protein structure. The sampled sequences are ranked based on the likelihood of each sampled sequence under the model. The sampled sequence with the highest likelihood can be used in all subsequent steps in the method.
[0126] To ensure that the redesigned (or restored) amino acid sequences still fold as expected, we re-predict the structures of the redesigned (or restored) amino acid sequences using AlphaFold for 12 recycle iterations. We calculate the average pLDDT and pTM for those re-predicted protein structures and discard amino acid sequences with the product pLDDT·pTM<0.5. We calculate the RMSD between the initial prediction and the re-predicted protein structures, as well as the CamSol solubility score (Sormanni et al., 2015) for the protein structures. We rank all the remaining re-predicted protein structures according to their pLDDT·pTM, RMSD, and CamSol solubility scores. We select some of the high ranking re-predicted protein structures for experimental validation.
[0127] Design of RcaT inhibitors The structure of RcaT-Sen2 was predicted using ColabFold with 96 recycle steps and early stopping at a tolerance of 0.1 Å. Around the putative active site, a crop of 100 amino acids was extracted as a template for fixed-target design. AlphaDesign was run to determine the fitness L binder (X|RcaT-Sen2;active-site), two recycle steps, a suboptimal degree of 0.1, and a population size of 10 were used to generate predicted (or candidate) structures of 50 and 100 amino acids. The target amino acid sequence was kept fixed, and an autoregressive diffusion model was used to generate 100 amino acid sequences per predicted (candidate) protein structure. The most likely amino acid sequence was kept and
number
[0128] In experimental validation of the designed binders' phenotypes (see Figure 15), a total of 88 proteins (50 or 100 amino acids long) were designed in silico to bind to the active site and block the activity of the growth-inhibitory toxin RcaT-Sen2
[56] , which is part of the Retron-Sen2 toxin-antitoxin system (Retron-TA). Sixteen of them were also designed to further bind to the active site of a different RcaT protein from E. coli, RcaT-Eco1 (21.5% identical to RcaT-Sen2 at the protein level). To experimentally validate the functionality, solubility, and / or expressibility of the designed blockers, we tested the ability of the designed blockers to inhibit the growth-inhibitory activity of RcaT-Sen2 using a high-throughput reverse genetic screening approach called toxin-inhibitory conjugation (TIC)
[56] . In TIC, a mobilizable gene donor plasmid is conjugated en bloc with an E. coli recipient strain transformed with a toxin plasmid on an agar plate. Normally, toxin-expressing strains do not grow, but growth is restored if the mobilizable plasmid encodes a toxin blocker.
[0129] Each of the 88 designed blockers designed according to this disclosure was cloned into an IPTG-inducible high-copy vector by Golden Gate assembly
[57] to construct a gene donor library. The gene donor strains were arrayed on agar plates in a 384-colony format, conjugated with E. coli BW25113 recipients transformed with an arabinose-inducible vector carrying RcaT-Sen2, and the exconjugants were then tested for their ability to grow while co-expressing the RcaT-Sen2 toxin with the designed blockers.
[0130] To assess the promiscuity of the designed blockers, the gene donor library was probed with RcaT-Eco9
[56] , a toxin 49% identical to RcaT-Sen2 from Escherichia coli. A quarter of the designed binders (22) restored bacterial growth in the presence of RcaT-Sen2 to various degrees (Fig. 1A). Two of the top hits, the RcaT-binders “cpx-50-nr2-run_5_0” and “50aa_1_55”, completely abolished the toxicity of RcaT-Sen2 and fully restored the growth of exconjugants. Although originally designed to target RcaT-Sen2, 15 constructs were also able to inhibit the toxicity of RcaT-Eco9 (Fig. 1B). Although the two toxins are not identical, their active sites are structurally similar, and thus functional blockers are likely to bind to both active sites (no blockers were found that inhibited only RcaT-Eco9). Notably, the top five validated blocker hits against RcaT-Eco9 were designed to bind to both the RcaT-Sen2 and RcaT-Eco1 active sites, potentially explaining their apparent promiscuity. In sum, proteins designed according to the present disclosure are able to inhibit the activity of the RcaT toxin in vivo, supporting the use of this tool to design protein binders in silico.
[0131] In silico protein structure prediction and design experiments Prediction of complex formation and conformational changes was performed using Colabfold. Five AF pre-trained parameter sets (model_1_ptm, model_2_ptm, model_3_ptm, model_4_ptm, and model_5_ptm) were used for prediction, and the model was selected by the highest predicted template matching score. Complex queries were predicted without paired MSA. The top models obtained were structurally aligned using PyMol to obtain RMSD. Parameters of the prediction run for proteins with PDB IDs of IR5P, 5JYT, 1IYT, and 2MXU are summarized in Figure 13. For optimization, an evolutionary optimizer was used. The population size was set to 10. During mutation, the population size was expanded by a factor of two. Sequences with a suboptimal degree of at most 10% were considered for mutation and recombined. Recombination was applied by crossover with a probability of 10% at each sequence position. The optimization was performed using the L mon , L_cpx for L>1.5, and L ccSequences with L>0.9 for some of the fitness functions, including were considered perfect. Tables II and III (see above) show further parameters of the optimization runs performed. AF was configured for the use of a single sequence, with ensemble, template, and extra MSA features disabled and the number of MSA features limited to the number of modeled monomers. The number of AF iterations (recycle steps) was kept as a parameter for each optimization run (see Tables II and III above). For larger protein complexes, the number of iterations was reduced to two to speed up the calculations. The AF pre-training parameter set model_1_ptm (pre-trained using as many sequences as possible in the MSA and structural templates, referred to as model_1, and fine-tuned to predict template matching scores) was used for all experiments. Finally, designed proteins of length 64 or more were subdivided into a dictionary of fragments using Geometricus. Fragments were collected using the k-mer method with k=16 and the radius method with a cutoff of 10 Å. A collection of feature representations was calculated for all designed proteins and the dimensionality was reduced using non-negative matrix factorization (NMF) with 50 components. Proteins were separated into 10 clusters using Ward linkage agglomerative clustering on their NMF components.
[0132] conclusion In this paper, we present a method for de novo protein design based on optimizing protein sequences of amino acids using an evolutionary algorithm. AlphaFold (AF) [4] is embedded in the design loop as a predictor. Optimization is facilitated by measures of AF output prediction confidence (or confidence metrics), namely predicted local distance difference test (pLDDT)
[39] and predicted aligned error (pAE)
[16] . We designed a set of flexible fitness functions (or fitness function components, see Table I and the paragraph preceding Table I) that encode an extensible platform for developing further fitness functions to solve various design tasks as well as tailor-made design problems. The method includes post-processing steps that make the sequences of the designed proteins more native-like, as well as improve solubility and improve expression when such proteins are experimentally synthesized. This enables a variety of applications including de novo design of protein monomers, dimers, oligomers, context-dependent conformational switches, and binders to target proteins.
[0133] The predicted protein structures are extensively validated using the Rosetta suite
[21] , a protein design and structure prediction tool. Fragment assembly-based ab initio structure predictions
[43] were used as an independent baseline for the designed protein structures. In addition to this, we developed a further rigorous validation protocol using all-atom molecular dynamics (MD) simulations, which go beyond the computational techniques traditionally used for structure prediction evaluation. MD simulations allow extensive exploration of the putative native state, and thus instability is detected as increased overall structural flexibility, loss of internal contacts within protein monomers, and / or loss of interfacial contacts within complexes.
[0134] We also experimentally validate a subset of the designed protein binders using a well-established phenotypic high-throughput reverse genetic screening approach called Toxin Inhibition Conjugation (TIC)
[56] . In a set of 88 proteins designed to bind to the toxin protein RcaT-Sen2, 25% restore bacterial growth to some degree, and two completely abolish RcaT-Sen2 activity. Some constructs even show a more extensive inhibition of toxicity against a related, but not identical, toxin (RcaT-Eco9). This demonstrates, first, that the binders designed using this method are sufficiently soluble and expressible in vivo to be utilized in such experiments, and, further, that the in silico design methodology affords impressive binders capable of completely inhibiting the corresponding functionality of the target protein.
[0135] The method in this paper was applied to design monomeric proteins de novo starting from a completely random sequence of amino acids ranging in size from 32 to 256 amino acids long based on the fitness function of pLDDT combined with pAE. The method generates a variety of structurally stable de novo designed monomeric proteins with diverse folds. Using the function of AF to predict protein complexes, complex predictions were made for many systems showing good structural agreement. By further specifying a function based on globular compactness along with the complex predictions, a variety of stable de novo protein complexes can be designed, including homodimers, heterodimers, and homo-oligomers ranging from trimers to hexamers. Orthogonality between pairs of dimers can be shown, and thus the method outlined in this paper may have applicability in designing mutually exclusive combinations, for example, in protein logic gates
[16] . Furthermore, many open-ended predicted complexes have been observed, providing a possible route for the design of self-assembling systems
[17] .
[0136] A particularly interesting subset of predicted protein structures exhibits significant conformational changes between their monomeric and oligomeric states. Structural predictions of existing protein systems known to change conformation and / or folding between monomeric and oligomeric forms, including amyloid-α-β switching, show that AF inherently encompasses a signature of conformational changes. Conformational switching between monomeric and oligomeric forms has been observed in several de novo designed open-end oligomeric systems. By defining an objective function to maximize the structural difference between the monomeric and oligomeric forms, the method can de novo design oligomers with this conformational switching property. Context-dependent conformational switching is a desirable feature in synthetic biology applications. For example, the design of proteins that self-assemble inside cells but not outside them may be achievable by engineering membrane-permeable α-helices that spontaneously switch to membrane-impermeable β-sheeted filaments as intracellular concentrations due to accumulation drive the equilibrium to the oligomeric state.
[0137] The method can select proteins that are not modified during the design loop, and thus, by combining this selection with de novo designed proteins starting from random sequences of amino acids, monomeric binders can be designed for target proteins that can be pre-specified. By applying this approach to a selected set of target proteins, stable binders that show significant interfacial contacts across the same interface of a given target protein were designed. The method can be further optimized for therapeutic applications in powerful biological design.
[0138] Acknowledgements The inventors would like to acknowledge the support of the German Federal Ministry of Education and Research (BMBF) (de.NBI project:031A537B) and the Volkswagen Foundation (contract 95826) and the Volkswagen Foundation's "Experiment! Funding Initiative" (grant number 93874-1). The inventors would like to thank Alessio Ling Jie Yang, Jacob Bobonis, Carlos Geert Pieter Voogdt and Athanasios Typas for the experimental validation studies.
[0139] References [1] Christine Zardecki, Chenghua Shao, Maria Voigt, and Stephen K. Burley. Protein data bank: 50 years of macromolecular structures enabling research and education. The FASEB Journal, 35, 2021. [2] Brian Kuhlman and Philip Bradley. Advances in protein structure prediction and design. Nature Reviews Molecular Cell Biology, 20:681-697, 2019. [3] Andriy Kryshtafovych, Torsten Schwede, Maya Topf, Krzysztof Fidelis, and John Moult. Critical assessment of methods of protein structure prediction (ca-p) - round xiv. Proteins, 2021. [4] John M Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvu-nakool, Russ Bates, Augustin Zidek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A Kohl, Andy Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David A. Reiman, Ellen Clancy, Michal Zielinski, Martin Steinegger, Michalina Pacholska, Tamas Berghammer, Sebastian Bodenstein, David Silver, Oriol Vinyals, Andrew W. Senior, Koray Kavukcuoglu, Pushmeet Kohli, and Demis Hassabis. Highly accurate protein structure prediction with alphafold. Nature, 596:583 - 589, 2021. [5] Kathryn Tunyasuvunakool, Jonas Adler, Zachary Wu, Tim Green, Michal Zielinski, Augustin Zidek, Alex Bridgland, Andrew Cowie, Clemens Meyer, Agata Laydon, et al. Highly accurate protein structure prediction for the human proteome. Nature, 596(7873):590-596, 2021. [6] Mehmet Akdel, Douglas Eduardo Valente Pires, Eduard Porta Pardo, Jurgen Janes, Arthur O. Zalevsky, Balint Meszaros, Patrick Bryant, Lydia L. Good, Roman A. Laskowski, Gabriele Pozzati, Aditi Shenoy, Wensi Zhu, Neera Borkakoti, Sameer Velankar, Adam Frost, Kristen Lindorff-Larsen, Alfonso Valencia, Sergey Ovchinnikov, Janani Durairaj, David B. Ascher, Janet M Thornton, Norman E. Davey, Amelie Stein, Arne Elofsson, Tristan I. Croll, and Pedro Beltrao. A structural biology community assessment of alphafold 2 applications. bioRxiv , [7] Milot Mirdita, Sergey Ovchinnikov, and Martin Steinegger. Collabfold-making protein folding accessible to all. bioRxiv , [8] Ian R. Humphreys, Jimin Pei, Minkyung Baek, Aditya Krishnakumar, Ivan Anishchenko, Sergey Ovchinnikov, Jing Zhang, Travis J. Ness, Sudeep Banjade, Saket Bagde, Viktoriya G. Stancheva, Xiao-Han Li, Kaixian Liu, Zhi Zheng, Daniel J. Barrero, Upasana Roy, Israel S. Fernandez, Barnabas Szakal, Dana Branzei, Eric C. Greene, Sue Biggins, Scott Keeney, Elizabeth A. Miller, J. Christopher Fromme, Tamara L. Hendrickson, Qian Cong, and David Baker. Structures of core eukaryotic protein complexes. bioRxiv, 2021. [9] Richard Evans, Michael O’Neill, Alexander Pritzel, Natasha Antropova, Andrew Senior, Tim Green, Augustin Zidek, Russ Bates, Sam Blackwell, Jason Yim, Olaf Ronneberger, Sebastian Bodenstein, Michal Zielinski, Alex Bridgland, Anna Potapenko, Andrew Cowie, Kathryn Tunyasuvunakool, Rishub Jain, Ellen Clancy, Pushmeet Kohli, John M Jumper, and Demis Hassabis. Protein complex prediction with alphafold-multimer. bioRxiv, 2021.
[10] David D Boehr, Ruth Nussinov, and Peter E Wright. The role of dynamic conformational ensembles in biomolecular recognition. Nature chemical biology, 5(11):789-796, 2009.
[11] Subramanian Vivekanandan, Jeffrey R Brender, Shirley Y Lee, and Ayyalusamy Ramamoorthy. A partially folded structure of amyloid-beta (1-40) in an aqueous environment. Biochemical and biophysical research communications, 411(2):312-316, 2011.
[12] Kulkarni Madhurima, Bodhisatwa Nandi, and Ashok Sekhar. Metamorphic proteins: the janus proteins of structural biology. Open biology, 11(4):210012, 2021.
[13] S Kashif Sadiq, Frank Noe, and Gianni De Fabritiis. Kinetic characterization of the critical step in hiv-1 protease maturation. Proceedings of the National Academy of Sciences, 109(50):20449-20454, 2012.
[14] Neil J Bruce, Gaurav K Ganotra, Daria B Kokh, S Kashif Sadiq, and Rebecca C Wade. New approaches for computing ligand-receptor binding kinetics. Current opinion in structural biology, 49:1-10, 2018.
[15] Po-Ssu Huang, Scott E Boyken, and David Baker. The coming of age of de novo protein design. Nature, 537(7620):320-327, 2016.
[16] Zibo Chen, Ryan D Kibler, Andrew Hunt, Florian Busch, Jocelynn Pearl, Mengxuan Jia, Zachary L VanAernum, Basile IM Wicky, Galen Dods, Hanna Liao, et al. De novo design of protein logic gates. Science, 368(6486):78-84, 2020.
[17] Zibo Chen, Matthew C Johnson, Jiajun Chen, Matthew J Bick, Scott E Boyken, Baihan Lin, James J De Yoreo, Justin M Kollman, David Baker, and Frank DiMaio. Self-assembling 2d arrays with de novo protein building blocks. Journal of the American Chemical Society, 141(22):8891-8895, 2019.
[18] Aaron Chevalier, Daniel-Adriano Silva, Gabriel J Rocklin, Derrick R Hicks, Renan Vergara, Patience Murapa, Steffen M Bernard, Lu Zhang, Kwok-Ho Lam, Guorui Yao, et al. Massively parallel de novo protein design for targeted therapeutics. Nature, 550(7674):74-79, 2017.
[19] Michael J Dougherty and Frances H Arnold. Directed evolution: new parts and optimized function. Current opinion in biotechnology, 20(4):486-491, 2009.
[20] Xingjie Pan and Tanja Kortemme. Recent advances in de novo protein design: Principles, methods, and applications. Journal of Biological Chemistry, page 100558, 2021.
[21] Andrew Leaver-Fay, Michael D. Tyka, Steven M. Lewis, Oliver F. Lange, James M Thompson, Ron Jacak, Kristian Kaufman, Paul D. Renfrew, Colin A. Smith, William Sheffler, Ian W. Davis, Seth Cooper, Adrien Treuille, Daniel J. Mandell, Florian Richter, Yih-En Andrew Ban, Sarel Jacob Fleishman, Jacob E. Corn, David E. Kim, Sergey Lyskov, Monica Berrondo, Stuart G. Mentzer, Zoran Popovic, James J Havranek, John Karanicolas, Rhiju Das, Jens Meiler, Tanja Kortemme, Jeffrey J. Gray, Brian Kuhlman, David Baker, and Philip Bradley. Rosetta3: an object-oriented software suite for the simulation and design of macromolecules. Methods in enzymology, 487:545-74, 2011.
[22] Po-Ssu Huang, Yih-En Andrew Ban, Florian Richter, Ingemar Andre, Robert M. Vernon, William R. Schief, and David Baker. Rosetta remodel: A generalized framework for flexible backbone protein design. PLoS ONE, 6, 2011.
[23] Brian Kuhlman, Gautam Dantas, Gregory C. Ireton, Gabriele Varani, Barry L. Stoddard, and David Baker. Design of a novel globular protein fold with atomic-level accuracy. Science, 302:1364 - 1368, 2003.
[24] Bruno E. Correia, John T. Bates, Rebecca J. Loomis, Gretchen Baneyx, Chris Carrico, Joseph G. Jardine, Peter Rupert, Colin E. Correnti, Oleksandr Kalyuzhniy, Vinayak Vittal, Mary J. Connell, Eric Stevens, Alexandria Schroeter, Man Chen, Skye MacPherson, Andreia M. Serra, Yumiko Adachi, Margaret A. Holmes, Yuxing Li, Rachel E. Klevit, Barney S. Graham, Richard T. Wyatt, David Baker, Roland K. Strong, James E. Crowe, Philip R. Johnson, and William R. Schief. Proof of principle for epitope-focused vaccine design. Nature, 507:201 - 206, 2014.
[25] Yang Hsia, Jacob B. Bale, Shane Gonen, Dan Shi, William Sheffler, Kimberly K. Fong, Una Nattermann, Chunfu Xu, Po-Ssu Huang, Rashmi Ravichandran, Sue Yi, Trisha N. Davis, Tamir Gonen, Neil P. King, and David Baker. Design of a hyperstable 60-subunit protein icosahedron. Nature, 535:136 - 139, 2016.
[26] Florian Richter, Andrew Leaver-Fay, Sagar D. Khare, Sinisa Bjelic, and David Baker. De novo enzyme design using rosetta3. PLoS ONE, 6, 2011.
[27] Ali Madani, Bryan McCann, Nikhil Naik, Nitish Shirish Keskar, Namrata Anand, Raphael R. Eguchi, Po-Ssu Huang, and Richard Socher. Progen: Language modeling for protein generation. bioRxiv, 2020.
[28] Alex Hawkins-Hooker, Florence Depardieu, Sebastien Baur, Guillaume Couairon, Arthur Chen, and David Bikard. Generating functional protein variants with variational autoencoders. PLoS Computational Biology, 17, 2021.
[29] Donatas Repecka, Vykintas Jauniskis, Laurynas Karpus, Elzbieta Rembeza, Jan Zrimec, Simona Poviloniene′, Irmantas Rokaitis, Audrius Laurynenas, Wissam Abuajwa, Otto Savolainen, Rolandas Meskys, Martin K. M. Engqvist, and Aleksej Zelezniak. Expanding functional protein sequence space using generative adversarial networks. bioRxiv, 2019.
[30] Namrata Anand and Possu Huang. Generative modeling for protein structures. In NeurIPS, 2018.
[31] Hao Huang, Boulbaba Ben Amor, Xichan Lin, Fan Zhu, and Yi Fang. G-vae, a geometric convolutional vae for protein structure generation. ArXiv, abs / 2106.11920, 2021.
[32] Jingxue Wang, Huali Cao, John Zeng Hui Zhang, and Yifei Qi. Computational protein design with deep learning neural networks. Scientific Reports, 8, 2018.
[33] Alexey Strokach, David Becerra, Carles Corbi-Verge, Albert Perez-Riba, and Philip M. Kim. Fast and flexible protein design using deep graph neural networks. Cell systems, 2020.
[34] Ivan Anishchenko, Tamuka Martin Chidyausiku, Sergey Ovchinnikov, Samuel J Pellock, and David Baker. De novo protein design by deep network hallucination. bioRxiv, 2020.
[35] Christoffer H Norn, Basile I. M. Wicky, David Juergens, Sirui Liu, David E. Kim, Doug K Tischer, Brian Koepnick, Ivan V. Anishchenko, David Baker, and Sergey Ovchinnikov. Protein sequence design by conformational landscape optimization. Proceedings of the National Academy of Sciences of the United States of America, 118, 2021.
[36] Lewis Moffat, Joe G Greener, and David T Jones. Using alphafold for rapid and accurate fixed backbone protein design. bioRxiv, 2021.
[37] Ethan C. Alley, Grigory Khimulya, Surojit Biswas, Mohammed AlQuraishi, and George M. Church. Unified rational protein engineering with sequence-based deep representation learning. Nature Methods, pages 1-8, 2019.
[38] Jianyi Yang, Ivan V. Anishchenko, Hahnbeom Park, Zhenling Peng, Sergey Ovchinnikov, and David Baker. Improved protein structure prediction using predicted inter residue orientations. Proceedings of the National Academy of Sciences, 117:1496 - 1503, 2020.
[39] Valerio Mariani, Marco Biasini, Alessandro Barbato, and Torsten Schwede. lddt: a local superposition-free score for comparing protein structures and models using distance difference tests. Bioinformatics, 29:2722 - 2728, 2013.
[40] S. Sinai, Richard Wang, Alexander Whatley, Stewart Slocum, Elina Locane, and Eric D. Kelsic. Adalead: A simple and robust adaptive greedy search algorithm for sequence design. ArXiv, abs / 2010.02141, 2020.
[41] Yang Zhang and Jeffrey Skolnick. Scoring function for automated assessment of protein structure template quality. Proteins: Structure, 57, 2004.
[42] Eric Jang, Shixiang Shane Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. ArXiv, abs / 1611.01144, 2017.
[43] Kim T. Simons, Richard Bonneau, Ingo Ruczinski, and David Baker. Ab initio protein structure prediction of casp iii targets using rosetta. Proteins: Structure, 37, 1999.
[44] D. T. Jones. Protein secondary structure prediction based on position-specific scoring matrices. Journal of molecular biology, 292 2:195-202, 1999.
[45] Lewis Moffat and David T. Jones. A deep semi-supervised framework for accurate modelling of orphan sequences. bioRxiv, 2020.
[46] Rebecca F. Alford, Andrew Leaver-Fay, Jeliazko R. Jeliazkov, Matthew J. O’Meara, Frank Dimaio, Hahnbeom Park, Maxim V. Shapovalov, Paul D. Renfrew, Vikram Khipple Mulligan, Kalli Kappel, Jason W. Labonte, Michael S. Pacella, Richard Bonneau, Philip Bradley, Roland L. Dunbrack, Rhiju Das, David Baker, Brian Kuhlman, Tanja Kortemme, and Jeffrey J. Gray. The rosetta all-atom energy function for macromolecular modeling and design. Journal of chemical theory and computation, 13 6:3031-3048, 2017.
[47] Steven M. Lewis and Brian Kuhlman. Anchored design of protein-protein interfaces. PLoS ONE, 6, 2011.
[48] William Sheffler and David Baker. Rosettaholes: Rapid assessment of protein core packing for structure prediction, refinement, design, and validation. Protein Science, 18, 2009.
[49] James Maier, Carmenza Martinez, Koushik Kasavajhala, Lauren Wickstrom, Kevin Hauser, and Carlos Simmerling. ff14sb: Improving the accuracy of protein side chain and backbone parameters from ff99sb. Journal of chemical theory and computation, 11 8:3696-713, 2015.
[50] Ulrich Essmann, Lalith E. Perera, Max L. Berkowitz, Thomas A. Darden, Hsing-Chou Lee, and Lee G. Pedersen. A smooth particle mesh Ewald method. Journal of Chemical Physics, 103:8577-8593, 1995.
[51] Janani Durairaj, Mehmet Akdel, Dick de Ridder, and Aalt D.J. van Dijk. Geometricus represents protein structures as shape-mers derived from moment invariants. Bioinformatics, 36 Supplement_2: i718-i725, 2020.
[52] Daniel D. Lee and H. Sebastian Seung. Learning the parts of objects by non-negative matrix factorization. Nature, 401:788-791, 1999.
[53] Nicholas F. Polizzi, Yibing Wu, Thomas Lemmin, Alison M. Maxwell, Shao-Qing Zhang, Jeff Rawson, David N. Beratan, Michael J. Therien, and William F. DeGrado. De novo design of a hyperstable non-natural protein-ligand complex with sub-a accuracy. Nature chemistry, 9 12:1157-1164, 2017.
[54] Roger Tseng, Nicolette F Goularte, Archana G. Chavan, Jansen Luu, Susan E Cohen, Yong-Gang Chang, Joel Heisler, Sheng Li, Alicia K. Michael, Sarvind Tripathi, Susan S. Golden, Andy LiWang, and Carrie L. Partch. Structural basis of the day-night transition in a bacterial circadian clock. Science, 355:1174 - 1180, 2017.
[55] Rodrigo Gallardo, Neil A. Ranson, and Sheena E. Radford. Amyloid structures: much more than just a cross-β fold. Current opinion in structural biology, 60:7-16, 2019.
[56] Bobonis, J., Mitosch, K., Mateus, A., Karcher, N., Kritikos, G., Selkrig, J., Zietek, M., Monzon, V., Pfalz, B., Garcia-Santamarina, S., Galardini, M., Sueki, A., Kobayashi, C., Stein, F., Bateman, A., Zeller, G., Savitski, M.M., Elfenbein, J.R., Andrews-Polymenis, H.L., Typas, A., 2022. Bacterial retrons encode phage-defending tripartite toxin-antitoxin systems. Nature 1-7. https: / / doi.org / 10.1038 / s41586-022-05091-4
[57] Engler, C., Kandzia, R., Marillonnet, S., 2008. A One Pot, One Step, Precision Cloning Method with High Throughput Capability. PLoS ONE 3, e3647. https: / / doi.org / 10.1371 / journal.pone.0003647
[58] Kritikos, G., Banzhaf, M., Herrera-Dominguez, L., Koumoutsi, A., Wartel, M., Zietek, M., Typas, A., 2017. A tool named Iris for versatile high-throughput phenotyping in microorganisms. Nature Microbiology 2, 17014. https: / / doi.org / 10.1038 / nmicrobiol.2017.14
Claims
1. A computer-implemented method for designing at least one protein. - A step of creating at least one amino acid sequence to be tested, comprising selecting some of the amino acids contained in the amino acid sequence to be tested according to a probability distribution, - A step of predicting the structural characteristics of the at least one protein from the at least one amino acid sequence, - A step of calculating the fitness function value for the at least one amino acid sequence to be tested based on the structural characteristics of the at least one protein, A method comprising the step of selecting or deselecting at least one amino acid sequence to be tested, according to the value of the fitness function.
2. The method according to claim 1, further comprising the step of modifying the at least one amino acid sequence to be tested by changing at least one of the amino acids contained in the amino acid sequence to be tested.
3. The method according to claim 2, further comprising the step of repeating the modification of the at least one amino acid sequence until the value of the fitness function exceeds a pre-selected threshold.
4. The method according to any one of claims 1 to 3, wherein the step of predicting the structural properties of the at least one protein contained in the amino acid sequence to be tested includes predicting the relative positions of the amino acids and / or predicting the total atomic structure of the at least one protein.
5. The method according to claim 1, wherein the modification includes mutating and / or recombining the at least one amino acid sequence being tested.
6. The method according to claim 1, wherein the structural characteristics are predicted using AlphaFold.
7. The method according to claim 1, wherein the fitness function is defined according to the biological and / or physicochemical properties of the at least one protein.
8. The method according to claim 1, wherein the fitness function comprises one or more fitness function components representing the biological and / or physicochemical properties of the at least one protein.
9. The method according to claim 1, further comprising the step of inputting into AlphaFold as a structural template at least one amino acid sequence or known structural characteristics related to the target protein to which the at least one protein can bind.
10. The method according to claim 1, further comprising the steps of redesigning the selected at least one amino acid sequence to be tested to make the redesigned at least one amino acid sequence more undenatured, and / or improving the solubility and / or expressibility of the at least one protein.
11. The method according to claim 10, further comprising the step of re-predicting the structural properties of the at least one protein based on the redesigned at least one amino acid sequence.
12. The method according to claim 1, further comprising the step of statistically reconstructing the at least one amino acid sequence to be tested from the predicted structural properties of the at least one protein.
13. The method according to claim 1, further comprising the step of computationally verifying the selected amino acid sequence using molecular dynamics and / or Rosetta ab-initio structural prediction.
14. The method according to claim 1, further comprising the step of experimentally verifying the selected amino acid sequence by determining the solubility and / or expressibility of the at least one protein.
15. The method according to claim 1, wherein the at least one protein comprises a monomer, a homodimer, a heterodimer, a trimer, a tetramer, a pentamer, an oligomer, a protein complex, a component of a protein complex, a binder that binds to a target protein, and a protein exhibiting multiple three-dimensional structures.