Means and methods for the induction of protein degradation

The FoldDelay metric identifies co-translational weak spots in proteins, enabling targeted drug screening to induce misfolding and degradation, addressing the challenges of 'undruggable' proteins in drug discovery.

WO2025157996A1PCT designated stage Publication Date: 2025-07-31VLAAMS INTERUNIVERSITAIR INST VOOR BIOTECHNOLOGIE VZW +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/051802
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-26
Filing Date
2025-01-24
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Current drug discovery methods struggle to target 'undruggable' proteins lacking binding pockets, and existing technologies for protein degradation, such as PROTACs and molecular glues, are ineffective for certain targets, while covalent drugs and RNA interference face challenges in precise tissue targeting and formulation.

Method used

A novel metric called FoldDelay (FD) is developed to identify co-translational weak spots (CWS) in proteins during translation, which are accessible and crucial for native stability, allowing selective small molecule binding to induce misfolding and degradation via the ubiquitin-proteasome system or autophagy.

Benefits of technology

FD enables the identification of CWS regions in proteins, facilitating the development of targeted drug screening methods to induce protein misfolding and degradation, overcoming the limitations of existing technologies in targeting 'undruggable' proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025051802_31072025_PF_FP_ABST
    Figure EP2025051802_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The invention provides means and methods to identify amino acids in target proteins which are crucial weak spots for the native protein stability. Such regions are useful for screening of compounds interacting with such amino acid regions. Compounds identified with this technology can induce the misfolding and subsequent cellular degradation of a target protein.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] MEANS AND METHODS FOR THE INDUCTION OF PROTEIN DEGRADATION

[0002] Field of the invention

[0003] The instant invention relates to the field of protein folding, more particularly to the field of protein translation, more particularly to the field of drug discovery. The invention provides means and methods to identify specific amino acid regions and drug pockets in target proteins which can be useful for screening of compounds interacting with such regions. Compounds identified with this technology can induce the misfolding and subsequent cellular degradation of a target protein.

[0004] Introduction to the invention

[0005] Despite continuous technological advancements that have expanded the druggable target space, the challenge of 'undruggable' proteins with significant therapeutic potential persists, defined by the absence of proper binding pockets for ligand interaction. While PROTAC-modalities and molecular glues partially mitigate this challenge by inducing proximity between the protein of interest (POI) and an E3 ligase, certain targets remain refractory to these approaches due to the absence of favorable binding pockets. Covalent drugs, interacting with reactive functional groups, present an alternative avenue for achieving potent and selective inhibition of the POI without reliance on binding pockets. However, the effectiveness of these drugs hinges on the proximity of reactive groups to the active site of the target protein. RNA interference (RNAi) and CRISPR(-like) technologies adopt a distinct strategy, targeting precursor RNA and DNA, respectively, ultimately resulting in protein knockdown. Nevertheless, challenges persist in formulating and precisely targeting tissues with these technologies.

[0006] We investigated the possibility to inactivate target proteins by small molecules before these target proteins attain their final folded state. Cellular protein production begins at translation initiation. During translation, proteins occupy non-native or native intermediate states before achieving their final folded state, a process referred to as co-translational folding. In this process, proteins transiently expose new binding sites at the ribosome or in the exit tunnel, offering previously unexplored anchor points for interaction with exogenous molecules. Engaging with these transiently exposed sites, exogenous molecules can disrupt proper folding, induce misfolding, and set off cellular cascades that culminate in the clearance of misfolded proteins via the ubiquitin-proteasome system (UPS) or autophagy. An approach to identify such transiently exposed sites would be a novel paradigm for induced protein degradation.

[0007] The understanding of protein folding mechanisms has traditionally been shaped by in vitro refolding experiments of fully denatured proteins, emphasizing the transition from denaturing to native conditions. However, protein translation is generally slower than folding, leading to partial folding reactions in incomplete nascent chains during synthesis, underscoring the divergence between in vivo and in vitro folding mechanisms. In the present invention we have developed a new metric, designated herein as FoldDelay, that integrates the topological pattern of atomic interactions of the native structure with the differential translation kinetics of various codons. We show that the proposed metric allows to identify co-translational weak spots (CWS) in co-translational protein folding. CWS are regions and drug pockets in the non-natively folded protein that are accessible during translation, crucial for native protein stability, and contain the right properties to allow selective small molecule binding. Currently, it is impossible to predict or experimentally validate what intermediate states a protein occupies during translation. Therefore, we designed an in silico prediction tool, herein further designated as FoldDelay (FD). FoldDelay is a highly valuable tool to identify CWS as it identifies regions in a protein that are unsatisfied during translation, i.e. regions that are translated relatively early (N-terminal) but have native interactions with regions that are translated relatively late (C-terminal). Co-translational weak spots (CWS) serve as an excellent starting point for the screening of molecules which can bind to such CWS sequences.

[0008] Figure legends

[0009] Figure 1 - FoldDelay measures waiting times incurred by nascent residues during translation (A) Distribution of average folding times versus estimated average translation times of 133 proteins in the Protein Folding Database (PFD2.0

[0027] ). Arrows show difference in folding and translation times for individual proteins. Yellow arrows indicate proteins for which average translation time is slower than average folding time, blue arrows indicate proteins for which average translation time is faster than average folding time. (B) Schematic representation of the globular native structure of a hypothetical protein. Amino acids are coloured in a gradient from N-term (blue) to C-term (red). (C) Contact map of hypothetical structure in (A), with contacts indicated by solid lines. (D) Contact map as residue 5 emerges from the ribosome. Dotted lines indicate interactions that are not yet accessible as not all interaction partners have been added to the polypeptide chain. (E) Contact map as residue 24 emerges from the ribosome. Solid lines indicate interactions that are now available, dotted lines indicate interactions that are not. All contacts for residue 5 have at this point become available. The FD incurred by residue 5 is 19 amino acids, spanning the point where residue 5 emerged from the ribosome, until the point where its most C-terminal interaction partner, residue 24, emerges. (F) Cartoon representation of the native structure of the E. coli peptidyl-prolyl isomerase B (PPI B, UniProt code P23869) enzyme as predicted by AlphaFold. Residues are coloured on a gradient from N-term (blue) to C-term (red). (G) Contact map of PPI B from the structure in (F). (H) Per-residue FD calculation for PPI B. (I) Mean FD of domains in the SCOPe40 dataset versus the relative residue position in the domain (scaled from 1 to 100). Error bars indicate standard deviation. (J) Domain length versus mean FD of the SCOPe40 dataset. Red points indicate domains of exactly 101 amino acids, the domain length with the most datapoints in the SCOPe40 database. (K) Violin plots showing the distribution of mean FD for the domains of exactly 101 amino acids per SCOP class. Green points indicate a representative example in each group (domains with mean FD closest to the median of their respective SCOP class). (L) FD profiles of the representative examples for each SCOP class indicated in (K). (M) Cartoon representation of the native structure of a peptidyl prolyl cis-trans isomerase from S. cerevisiae (UniProt code P14832) as predicted by AlphaFold. Residues are coloured on a gradient from N-term (blue) to C-term (red). (N) Contact map of the structure in (F). (O) Per-residue FD calculation for the structure in (F).

[0010] Figure 2 - Exploring FoldDelay across proteomes. (A) Distribution of residues with the maximum FD (log scale) for each protein in E. coli (n = 3,910) and S. cerevisiae (yeast; n = 5,812). Vertical dotted lines indicate 1 second, 10 seconds and 1 minute. (B) Number of residues for different FD bins in yeast proteins. (C) Enrichment of residues with FDs between 1 and 10 seconds (or bigger than 10 seconds) versus background for the different DSSP secondary structure categories in yeast proteins. C = coil, B = P-bridge, E = extended strand in -sheet conformation, G= 3-turn helix, H = 4-turn helix, I = 5-turn helix, S = bend and T = hydrogen bounded turn. (D) Enrichment of residues with FD between 1 and 10 seconds (or bigger than 10 seconds) versus background for all amino acid types in yeast proteins. (E-G) pLDDT scores (E), solvent accessibilities (F) and stabilities for residues in yeast proteins for different categories of FD. (H) Percentage of residues in APRs (TANGO score > 5) for different FD bins in yeast proteins. To avoid biases, residues in transmembrane domains and signal peptides were filtered out. (I) Aggregation strength (TANGO score) for residues in APRs of yeast proteins for different categories of FD.

[0011] Figure 3 - Ssb binds to regions with high FoldDelays. (A) Schematic representation of Ssb with a nascent chain during translation. Ssb is targeted to the nascent chain by the ribosome-associated complex (RAC) once the nascent chain reaches a length of around 50 aas. (B) FD in the nascent chain at the start of Ssb binding for sites with a peak width between 6-8 aas (n = 3,371), as compared to an equivalent number of randomly sampled positions (n = 4,000) from the same set of proteins. The line represents the median value at each position, while the shaded region is the 95% bootstrapped confidence interval (Cl). (C) Overlap between limbo regions and Ssb binding sites (width of 5-11 aas) (D) Difference between the median FD values of Limbo regions that are also Ssb binding sites, and Limbo regions that are not Ssb binding sites. (E) FD profile of MTAP. Limbo regions that are also Ssb binding sites are shown in orange while those that are not Ssb binding sites are shown in grey. (F) Difference between the median FD values of the aligned Ssb footprints and of the randomly sampled positions showed in B. Dotted line indicates the average peak width of Ssb binding sites in the dataset. (G) FD in the nascent chain at the start of Ssb binding for sites with a peak width of 5 aa (n = 1,412), 6 aa (n= 1,277), 7 aa (n = 1,111), 8 aa (n = 983), 9 aa (n = 945), 10 aa (n = 897) and 11 aa (n = 707). (H) Average median FDs between positions -53 and -35 (Ssb binding region) per width peak of aligned Ssb binding footprints with a peak width between 5-11 aas. Based on the linear model the average FD value at these positions increases with Ssb peak width. All experimental Ssb binding sites used in this figure are derived from

[0024] ,

[0012] Figure 4 - Proteins with high FoldDelays are associated with co-translational misfolding and aggregation. (A) Total FD of proteins that are actively translated in S. cerevisiae (translatome), which have been stratified based on whether they interact co-translationally with Ssb (n = 1,913) or not (n = 910)

[0023] , Proteins that interact with Ssb are further stratified on whether they remain soluble (n = 1,495) or aggregate (n = 418) in SSB cells. (B) FD in the nascent chain at the start of Ssb binding for sites with a peak width between 5-11 aas in proteins that aggregate or remain soluble in SSB cells. There are 1,917 and 5,415 Ssb binding sites in proteins that aggregate or remain soluble, respectively. The line represents the median value at each position, while the shaded region is the 95% bootstrapped Cl. (C) Number of APRs per lOOaa in proteins bound by Ssb that remain soluble (n = 1,495) or aggregate (n = 418) in SSB cells. (D,E) Percentage of APR starting sites in bins of equal length based on the start of Ssb binding footprints with a width between 5-11 aas in proteins that aggregate (D) or remain soluble (E) in SSB cells. Fisher exact test with FDR correction was used to compare the proportion of APR starting sites at bin -55 to -35 (Ssb binding region) against the other bins. (F) Total FD of yeast proteins under physiological conditions (n = 107) or upon exposure to arsenite stress (n = 140) compared to background (MS proteome; n = 1,179)

[0040] , (G) Total FD of proteins that are co-translationally ubiquitinated under physiological conditions (n = 600)

[0043] , As background we use the translatome (n = 1,790) reported by Willmund et at

[0023] , Statistical significance was determined by unpaired Wilcoxon test with Bonferroni correction for multiple comparisons (A, C, F and G).

[0013] Figure 5 - Compensating FoldDelay through codon optimization is an implausible evolutionary strategy. (A) Distribution of codon decoding times per amino acid as reported by Tuller et al

[0044] , (B) Minimal FD vs. actual FD calculated per protein. For each protein's furthest distance interaction, the actual and minimal translation times of the separating residues were calculated using the mean translation rates of the actual codons, and the mean translation rate for the fastest synonymous codon, respectively. (C) Gain in FD as calculated by the difference between the actual FD and the minimal FD in (B). (D) Distribution of the proportional differences, calculated as the ratio between the differences shown in (C) and the actual FD. (E-H) Identical analyses as those shown in (A-D), this time using the average decoding times reported by

[0046] based on data from

[0061] ,

[0014] Figure 6 - isolation of CWS sequences from the human PARP enzyme - The figure above represents the per-residue fold-delay (FD) calculation for PARP1 (Poly [ADP-ribose] polymerase 1. The graph shows the amino acid positions of PARP1 on the x-axis and the fold-delay values on the y-axis. The fold delay value represents the distance in number of residues between the amino acid residue and the most C-terminal amino acid interaction partner as present in the 3D fold. Three regions are highlighted with a high folddelay (at least 350), meaning that these amino acid stretches cannot be fully stabilized in its native conformation until the next 350 (or more) amino acids have been produced. This means those regions are partially exposed for a long time and are 3 different co-translational weak spots.

[0015] Detailed description of the invention

[0016] The functionality of globular proteins relies on adopting a three-dimensional shape known as the native structure. Achieving this structure involves the folding of an elongated polypeptide chain into a specific conformation. Although this process is intricate, it is widely acknowledged that all the necessary information for a protein to attain its native fold is encoded in its primary amino acid sequence [1], Additionally, proteins can fold completely in physiologically relevant timescales, generally in the range of microseconds to seconds [2], Most of these folding rates are derived from classic in vitro experiments in which the (re)folding of purified, full-length protein is monitored. These experiments yielded invaluable insights, including the realization that in vitro folding rates are partly determined by topological complexity [3], An aspect that is overlooked in such experiments, however, is protein translation. Protein translation progresses at an average rate of about 20 aas / s in prokaryotes and anywhere between 4 and 10 aas / s in eukaryotes, meaning that the complete synthesis of proteins can take seconds and even up to minutes. Hence, in vivo, local folding events often take place while a polypeptide chain is still emerging from the ribosome, i.e. co-translationally. Indeed, it is estimated that one third of the E. coli cytosolic proteome folds at least one entire domain co-translationally [4], and this fraction is likely higher in eukaryotes, given their slower translation rates. Moreover, co-translational folding can increase folding efficiencies: it was found that the folding of firefly luciferase is faster if it happens during translation compared to its post-translational (re)folding [5], Other instances of proteins folding more efficiently co- than post-translationally have been reported [6-9], Increased folding efficiencies in these examples are attributed to a smoother folding landscape, where folding of partial polypeptide chains prevents aberrant non-native interactions that could lead to kinetic traps, misfolding and aggregation [8, 10], In line with this, co-translational folding has been put forth as one of the explanations for why a large portion (one third) of the E. coli cytoplasmic proteome was found to be "non-refoldable", i.e. after cell lysis and protein denaturation, these proteins do not reassemble into their native folds

[0011] , To promote co-translational folding, the cadence of translation - which happens in bursts and pauses rather than at a uniform rate - has been evolutionarily optimized ([10, 12, 13] and reviewed in

[0014] ). For example, interdomain regions tend to be enriched in non-optimal codons, which slows translation, allowing the leading domain to fold before the lagging one. It appears then that the vectorial nature of translation is exploited in vivo to increase folding efficiency: the gradual addition of residues allows the growing polypeptide chain to sample stabilizing native interactions in a reduced conformation space, thereby avoiding kinetic traps associated with interactions with not yet formed residues towards the C-terminus in the sequence

[0015] , Here, we explore the notion that co-translational folding is a double-edged sword: the incremental emergence of the nascent chain is beneficial, especially for short range interactions, as interaction partners emerge in quick succession, effectively promoting native interactions in a reduced conformation space. The opposite might be true for long-range interactors. Vectorial polypeptide production condemns residues with long-range interaction partners to idle while the remainder of the polypeptide chain is being produced, in fact making them more vulnerable to non-native intra- and intermolecular interactions, potentially leading to "premature folding" - i.e. misfolding - and / or aggregation. Indeed, several sources report that newly synthesized proteins are more vulnerable to misfolding and aggregation than existing, matured proteins [16, 17], with topologically complex proteins being more at risk

[0017] , More evidence for the pitfalls associated with co-translational folding is the existence of a branch of the proteostasis network (PN) that acts specifically at the stage of translation, shielding regions from premature folding events. Firstly, ribosomes themselves have a holdase function as their negatively charged surface interacts with nascent chains, preferentially with basic and aromatic residues [18, 19], thereby preventing premature co- translational folding. Furthermore, ribosomes effectively act as solubility tags: the association of nascent chains with a large negatively charged ribosome creates an excluded volume in which interactions with other cellular components are disfavored. Secondly, a host of dedicated chaperones engage nascent chains at the ribosome

[0020] , The typical example of this is Trigger Factor (TF) in E. coli, which directly interacts with both the nascent chain and the ribosome, thereby preventing off-pathway interactions

[0021] , In eukaryotes, co-translational chaperones are most well-studied in S. cerevisiae, in which Nascent polypeptide Associated Complex (NAC), and Ribosome Associated Complex (RAC) directly engage the ribosome and interact with the nascent chain near the ribosome exit tunnel

[0020] , RAC recruits a Hsp70 type ribosome-associated chaperone called Ssb, which prevents premature folding through bindingrelease cycles [22-24], Many of the yeast co-translational chaperones have orthologs in mammals, and the list of co-translationally acting chaperones is expanding. Furthermore, it is becoming clear that canonical cytosolic chaperones also engage nascent chains as they are still attached to the ribosome

[0024] , Still, this mechanism is not foolproof as an estimated one third of newly synthesized polypeptides are targeted for proteasomal degradation, either through mistakes in translation or inability to attain the native fold

[0025] , Clearly, the vectorial nature of translation can benefit folding outcomes, but it also poses a risk. Residues that idle on the ribosome potentially engage in off-pathway interactions, necessitating the evolution of a dedicated co-translational proteostasis network. In the instant invention we provide a method to quantify the length of time between the translation of each residues and that of all its native interaction partners. Effectively, we calculate the delay on co-translational folding experienced by individual residues which metric is herein designated as "FoldDelay" (FD). The calculation of FD was inspired by that of Contact Order, a simple yet highly effective method of capturing topological complexity that correlates well to post-translational in vitro refolding rates. But in the instant invention we propose a variant metric that also encapsulates the vectorial nature of protein translation, thereby also capturing the in vivo co-translational folding. Importantly we show that the proposed metric allows to identify cotranslational weak spots (CWS) in co-translational protein folding. CWS are regions of between 3 to 15 amino acids in the non-natively folded protein that are accessible during translation, crucial for native protein stability, and contain the right properties to allow selective small molecule binding. Importantly a CWS can be a linear sequence present in the target protein of about 3 to about 15 amino acids or a CWS can be a drug pocket of between about 3 to about 15 amino acids present in the target protein. Indeed, a drug pocket is not necessarily a linear sequence and most often consists of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or 13 amino acids which are present in different locations in the linear sequence of the target protein but the amino acids of the drug pocket in the folded target protein interact with each other in the 3-dimensional fold. With the current technologies it is impossible to predict or experimentally validate which are the intermediate states a target protein occupies during translation. Thus our novel metric, FoldDelay, is a valuable tool to identify CWS as it identifies regions in a protein that are unsatisfied during translation, i.e. regions that are translated relatively early (N- terminal) but have native interactions with regions that are translated relatively late (C-terminal). CWS regions can be used in drug screening methods to identify molecules which can bind on these CWS regions.

[0017] Accordingly the present invention provides in a first embodiment a computer-implemented method to identify a co-translational weak spot (CWS) in a target protein comprising a) determining for each amino acid residue in the primary sequence of said target protein a FoldDelay value wherein said value corresponds with the interaction between a first amino acid residue and the interaction with its amino acid interaction partners in the 3-dimensional fold of said target protein and wherein the FoldDelay value is measured in the primary sequence of said target protein as the sequence distance between said first amino acid residue with the most downstream C-terminal amino acid interaction partner present in the primary sequence of said target protein, b) plotting the determined FoldDelay values for each of the amino acids in the primary sequence of said target protein in a graph and c) identifying a sequence between about 3 and about 15 amino acids which have at least 50% higher than the average FoldDelay values in said graph and wherein said sequence is a CWS.

[0018] In yet another embodiment the invention provides a method to identify a co-translational weak spot (CWS) in a target protein comprising a) determining for each amino acid residue present in the primary sequence of said target protein the interaction with its amino acid interaction partners in the 3- dimensional fold of said target protein, b) measuring the FoldDelay value for each amino acid residue which is the amino acid sequence distance between each amino acid residue with the furthest stabilizing downstream C-terminal amino acid interaction partner as determined in the 3D sequence of said target protein, c) plotting the determined FoldDelay parameters for each of the amino acids in the primary sequence of said target protein in a graph and c) identifying a CWS between 3 and 15 amino acids which have at least 50% higher than the average FoldDelay values in said graph .

[0019] In yet another embodiment the invention provides a computer-implemented method to identify a co- translational weak spot (CWS) in a target protein comprising the steps of: a) determining for each amino acid of a target protein the intramolecular amino acid interaction partners as determined in the 3-dimensional fold of said target protein, b) calculating the FoldDelay value for each of said amino acid residues of said target protein, c) identifying a CWS as about 3 to about 15 amino acids present in said target protein, wherein the FoldDelay values of the amino acids in said CWS are at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 85 or more amino acids as measured in the primary sequence of said target protein, wherein said FoldDelay value of an amino acid residue represents the distance between said amino acid residue and the furthest stabilizing C-terminal amino acid interaction partner as defined in the 3D fold of said target protein and wherein the distance between said amino acid and said furthest stabilizing C- terminal interaction partner is measured in the primary sequence of said target protein.

[0020] In yet another embodiment the invention provides a method to identify a co-translational weak spot sequence (CWS) in a target protein comprising a) determining for each amino acid residue in the primary sequence of said target protein a FoldDelay parameter wherein said parameter corresponds with the sequence distance between a first amino acid residue with its furthest native interaction amino acid partner as is determined in the 3D fold of the target protein and wherein the sequence distance is counted in the primary sequence of said target protein b) plotting the determined FoldDelay parameters for each of the amino acids in the primary sequence in a graph and c) identifying a sequence between 3 and 15 amino acids which have at least 50% higher than average FoldDelay parameters in said graph and wherein said sequence is a CWS.

[0021] In yet another embodiment the invention provides a method to identify a co-translational weak spot (CWS) in a target protein comprising a) determining for each amino acid residue in the primary sequence of said target protein a FoldDelay parameter wherein said parameter corresponds with the sequence distance between a first amino acid residue with its furthest native interaction amino acid partner present in the primary amino acid sequence of said target protein, b) plotting the determined FoldDelay parameters for each of the amino acids in the primary sequence in a graph and c) identifying a CWS between 3 and 15 amino acids which have at least 50% higher than average FoldDelay parameters in said graph.

[0022] In yet another embodiment the invention provides a method to identify a co-translational weak spot (CWS) in a target protein comprising a) determining for each amino acid residue present in the primary sequence of said target protein the interaction with its amino acid interaction partners present in the 3- dimensional fold of said target protein, b) calculating the FoldDelay value for each amino acid residue which is the amino acid sequence distance between each amino acid residue with a C-terminal amino acid interaction partner as present in the primary sequence of said target protein, c) plotting the determined FoldDelay parameters for each of the amino acids in the primary sequence of said target protein in a graph and c) identifying a CWS between 3 and 15 amino acids which have at least 50% higher than the average FoldDelay values in said graph.

[0023] In yet another embodiment the invention provides a method to identify a co-translational weak spot (CWS) in a target protein comprising the steps of: a) determining for each amino acid of a target protein the intramolecular amino acid interaction partners as present in the 3-dimensional fold of said target protein, b) calculating the FoldDelay value for each of said amino acid residues of said target protein, c) identifying a CWS consisting of between about 3 to about 15 amino acids present in said target protein, wherein the FoldDelay values of said amino acids as calculated in step b) each are at least 50 % higher than the average FoldDelay value of the amino acids of the target protein, wherein said FoldDelay value of an amino acid residue represents the distance between said amino acid residue and the most C-terminal amino acid interaction partner as present in the 3D fold of said target protein wherein C-terminal refers to the primary sequence of said target protein, and wherein the average FoldDelay value of the target protein is defined as the sum of all FoldDelay values for each of the amino acid residues as calculated in step b), divided by the number of amino acids as present in the primary sequence of the target protein.

[0024] In yet another embodiment the invention provides a method to identify a co-translational weak spot (CWS) in a target protein comprising the steps of: a) determining for each amino acid of a target protein the intramolecular amino acid interaction partners as present in the 3-dimensional fold of said target protein, b) calculating the FoldDelay value for each of said amino acid residues of said target protein, c) identifying a CWS consisting of between 3 to 15 amino acids of said target protein, wherein the FoldDelay values of said amino acids as calculated in step b) each are at least 50 % higher than the average FoldDelay value of the amino acids of the target protein, wherein said FoldDelay value of an amino acid residue represents the distance between said amino acid residue and a C-terminal amino acid interaction partner as present in the 3D fold of said target protein wherein the distance between said amino acid and said C-terminal interaction partner is measured in the primary sequence of said target protein, and wherein the average FoldDelay value of the target protein is defined as the sum of all FoldDelay values for each of the amino acid residues as calculated in step b), divided by the number of amino acids as present in the primary sequence of the target protein.

[0025] In a specific embodiment a CWS is a linear sequence present in the target protein of about 3 to about 15 amino acids.

[0026] In another specific embodiment a CWS is a drug pocket of between about 3 to about 15 amino acids present in the target protein. A drug pocket is not necessarily a linear sequence and most often consists of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or 13 amino acids which are present in different locations in the linear sequence (or in other words these amino acids are not adjacent in the linear sequence of the target protein) of the target protein but the amino acids of the drug pocket in the folded target protein interact with each other in the 3-dimensional fold.

[0027] In another particular embodiment a CWS sequence is between 5 and 15 amino acids.

[0028] In another particular embodiment a CWS sequence is between 3 and 10 amino acids.

[0029] In another particular embodiment a CWS sequence is between 5 and 10 amino acids.

[0030] In another particular embodiment a CWS sequence is between 3 and 8 amino acids.

[0031] In another particular embodiment a CWS sequence is between 5 and 8 amino acids. In another particular embodiment the CWS sequence has at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 100%, at least 200% or at least 300% higher than the average FoldDelay values in the plotted graph.

[0032] To determine for a specific amino acid of a target protein the intramolecular amino acid interaction partners in the 3-dimensional fold of said target protein an arbitrary range of 4, or 5 or 6 or 7 or 9 or 10 angstrom is taken to identify the interacting amino acids which are in the vicinity of 4, or 5 or 6 or 7 or 8 or 9 or 10 angstroms. For example for a value of 6 angstroms, amino acid residues were considered to interact with other amino acids if they contained non-hydrogen atoms within a 6 angstrom distance.

[0033] Once one or more cotranslational weakspots are identified in a target protein screening assays can be set up to screen for molecules that bind to these CWS regions and which molecules induce protein misfolding.

[0034] Accordingly, the invention provides the use of a co-translational weak spot for the screening of molecules binding to said CWS.

[0035] In another embodiment the invention provides a method for identifying molecules binding on a c- translational weak spot derived from a target protein.

[0036] Selection of a suitable meta-stable foldon comprising a CWS

[0037] One possibility is starting with the isolation of meta-stable foldons from a target protein wherein the foldon comprises a CWS. This can for example be done in a protein fragment expression assay. In this assay subdomains of the target protein comprising a CWS (a foldon) are expressed in a cellular system. In one particular embodiment a foldon comprising a CWS can be expressed with a C-terminal fluorescent marker. If the foldon comprising the CWS is meta-stable, it will be expressed and distributed homogenously in a cell. In contrast, if the foldon comprising the CWS is unstable, it will be rapidly cleared by the cell, leading to lower expression levels or non-homogenous distribution in the cell. In the assay meta-stable foldons comprising a CWS and exposing the CWS can be isolated from the cell and used in an in vitro screening assay (see further herein).

[0038] In an alternative step a cellular expression assay can be set up in which the full-length protein is expressed with a C-terminal fluorescent marker, but with specific stalling sites introduced in its sequence. Placing the stalling sites 30-60 amino acids downstream of the identified CWS will lead to longer exposure times of the CWS outside of the ribosome - making them more accessible for small molecule binding. As a result of the stalling, these constructs will show lower expression levels compared to wild-type sequences without stalling sites. This assay setup will identify the stalling sites that still allow some level of the protein to be detected. In vitro screening assays using a metastable foldon comprising a CWS or using a CWS

[0039] The constructs that were identified as meta-stable foldons comprising a CWS (see above) are expressed in a cellular system. Cells are lysed and the foldons comprising the CWS are isolated using their affinity tag. Using a BioLayer Interferometry (BLI) or a Surface Plasmon Resonance (SPR) assay, one can measure small molecule binding to these foldons comprising a CWS.

[0040] Alternatively constructs comprising a stalling site (see above) are expressed in a cellular system. Cells expressing these constructs can be used to screen libraries of small molecules. Readout of the assay are reduced protein levels or non-homogenous distribution of the target protein upon small molecule treatment. Specificity is checked by testing multiple target proteins in parallel, serving as each other's controls.

[0041] Yet another possibility is the use of immobilized CWS peptides to screen for small molecule binding. BioLayer Interferometry (BLI) or a Surface Plasmon Resonance (SPR) can be used to measure the strength of interaction of binding.

[0042] The term "compound" is used herein in the context of a "test compound" or a "drug candidate compound" described in connection with the methods of the present invention. As such, these compounds comprise organic or inorganic compounds, derived synthetically or from natural resources. The compounds include polynucleotides, lipids or hormone analogs that are characterized by low molecular weights. Other biopolymeric organic test compounds include small peptides or peptide-like molecules (peptidomimetics) comprising from about 2 to about 40 amino acids and larger polypeptides comprising from about 40 to about 500 amino acids, such as antibodies or antibody conjugates.

[0043] In yet another embodiment the invention provides an apparatus comprising control circuitry configured to perform the herein described computer-implemented methods.

[0044] Systems of the disclosure can include an intranet-based computer system that is capable of communicating with various software. A computer system includes any type of computing device or communication device. Examples of such a system can include, but are not limited to, super computers, a processor array, distributed parallel system, a desktop computer with LAN, WAN, Internet or intranet access, a laptop computer with LAN, WAN, Internet or intranet access, a smart phone, a server, a server farm, an android device (or equivalent), a tablet, smartphones, and a personal digital assistant (PDA). Further, as discussed above, such a system can have corresponding software (e.g., user software, sensor device software). The software of one system can be a part of, or operate separately but in conjunction with, the software of another system. Embodiments of the disclosure include a storage repository. The storage repository can be a persistent storage device (or set of devices) that stores software and data. Examples of a storage repository can include, but are not limited to, a hard drive, flash memory, some other form of solid-state data storage, or any suitable combination thereof. The storage repository can be located on multiple physical machines, each storing all or a portion of the database, Al platform, protocols, algorithms, or other stored data according to some example embodiments. Each storage unit or device can be physically located in the same or in a different geographic location. In embodiments, the storage repository may be stored locally, or on cloud-based serveries such as Amazon Web Services.

[0045] In one or more example embodiments, the storage repository stores one or more databases, Al Platforms, protocols, algorithms, and stored data. The protocols can include any of a number of communication protocols that are used to send, receive, or send and receive data between the processor, datastore, memory and the user. A protocol can be used for wired and / or wireless communication. Examples of a protocols can include, but are not limited to, Modbus, profibus, Ethernet, and fiberoptic.

[0046] Systems of the disclosure can include a hardware processor. The processor of the computer executes software, algorithms, and firmware in accordance with one or more example embodiments. The processor can be a central processing unit, a multi-core processing chip, SoC, a multi-chip module including multiple multi-core processing chips, or other hardware processor in one or more example embodiments. The processor is known by other names, including but not limited to a computer processor, a microprocessor, and a multi-core processor. The processor can also be an array of processors.

[0047] In one or more example embodiments, the processor executes software instructions stored in memory. Such software instructions can include generating machine learning models, executing machine learning models, performing analysis on data received from the database, and so forth. The memory includes one or more cache memories, main memory, or any other suitable type of memory. The memory can include volatile or non-volatile memory.

[0048] The processing system can be in communication with a computerized data storage system which can be stored in the storage repository. The data storage system can include a non-relational or relational data store, such as a MySQL or other relational database. Other physical and logical database types could be used. The data store may be a database server, such as Microsoft SQL Server, Oracle, IBM DB2, SQLITE, or any other database software, relational or otherwise. The data store may store the information identifying syntactical tags and any information required to operate on syntactical tags. In some embodiments, the processing system may use object-oriented programming and may store data in objects. In these embodiments, the processing system may use an object-relational mapper (ORM) to store the data objects in a relational database. The systems and methods described herein can be implemented using any number of physical data models. In one example embodiment, an RDBMS can be used. In those embodiments, tables in the RDBMS can include columns that represent coordinates. The tables can have pre-defined relationships between them. The tables can also have adjuncts associated with the coordinates.

[0049] In embodiments, the systems of the disclosure can include one or more I / O (input / output) devices allow a user to enter commands and information into the system, and also allow information to be presented to the user or other components or devices. Examples of input devices include, but are not limited to, a keyboard, a cursor control device (such as a mouse), a microphone, a touchscreen, and a scanner. Examples of output devices include, but are not limited to, a display device (e.g., a display, a monitor, or projector), speakers, outputs to a lighting network (such as a DMX card), a printer, and a network card. For example, the input devices can be used to enter data on native proteins and mutation sequences and assays. The input devices can also enter wanted functional data for a protein. The output devices can be used to output analysis data and / or engineered protein sequences resulting from Al protein design.

[0050] Various techniques are described herein in the general context of software.

[0051] Generally, software includes routines, programs, objects, components, data structures, and so forth that perform particular tasks or implement particular abstract data types. An implementation of these modules and techniques can be stored on or transmitted across some form of computer readable media. Computer readable media is any available non-transitory medium or non-transitory media that is accessible by a computing device. By way of example, and not limitation, computer readable media includes computer storage media.

[0052] The above-described embodiments of the present invention can be implemented in any of numerous ways. For example, the embodiments may be implemented using hardware, software or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers. It should be appreciated that any component or collection of components that perform the functions described above can be generically considered as one or more controllers that control the above-discussed functions. The one or more controllers can be implemented in numerous ways, such as with dedicated hardware, or with general purpose hardware (e.g., one or more processors) that is programmed using microcode or software to perform the functions recited above.

[0053] One or more processors may be interconnected by one or more networks in any suitable form, including as a local area network or a wide area network, such as an enterprise network or the Internet. Such networks may be based on any suitable technology and may operate according to any suitable protocol and may include wireless networks, wired networks, or fiber optic networks.

[0054] One or more algorithms for controlling methods or processes provided herein may be embodied as a readable storage medium (or multiple readable media) (e.g., a computer memory, one or more floppy discs, compact discs (CD), optical discs, digital video disks (DVD), magnetic tapes, flash memories, circuit configurations in Field Programmable Gate Arrays or other semiconductor devices, or other tangible storage medium) encoded with one or more programs that, when executed on one or more computers or other processors, perform methods that implement the various methods or processes described herein.

[0055] In some embodiments, a computer readable storage medium may retain information for a sufficient time to provide computer-executable instructions in a non-transitory form. Such a computer readable storage medium or media can be transportable, such that the program or programs stored thereon can be loaded onto one or more different computers or other processors to implement various aspects of the methods or processes described herein. As used herein, the term "computer-readable storage medium" encompasses only a computer-readable medium that can be considered to be a manufacture (e.g., article of manufacture) or a machine. Alternatively, or additionally, methods or processes described herein may be embodied as a computer readable medium other than a computer-readable storage medium, such as a propagating signal.

[0056] The terms "program" or "software" are used herein in a generic sense to refer to any type of code or set of executable instructions that can be employed to program a computer or other processor to implement various aspects of the methods or processes described herein. Additionally, it should be appreciated that according to one aspect of this embodiment, one or more programs that when executed perform a method or process described herein need not reside on a single computer or processor but may be distributed in a modular fashion amongst several different computers or processors to implement various procedures or operations.

[0057] Executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of the program modules may be combined or distributed as desired in various embodiments.

[0058] Also, data structures may be stored in computer-readable media in any suitable form. Non-limiting examples of data storage include structured, unstructured, localized, distributed, short-term and / or long term storage. Non-limiting examples of protocols that can be used for communicating data include proprietary and / or industry standard protocols (e.g., HTTP, HTML, XML, JSON, SQL, web services, text, spreadsheets, etc., or any combination thereof). For simplicity of illustration, data structures may be shown to have fields that are related through location in the data structure. Such relationships may likewise be achieved by assigning storage for the fields with locations in a computer-readable medium that conveys relationship between the fields. However, any suitable mechanism may be used to establish a relationship between information in fields of a data structure, including using pointers, tags, or other mechanisms that establish relationship between data elements.

[0059] Examples l.Protein translation is orders of magnitude slower than protein folding

[0060] To visualize the differences in timescales of protein folding and translation rates, we compared folding rates and estimated translation times for 133 proteins that have experimentally recorded in vitro refolding rates from denaturing conditions reported in the Protein Folding DataBase (PFDB)

[0027] (Figure 1A). The PFDB contains information on proteins from a wide array of species, and for a lot of these an accurate translation rate has never been established. Therefore, translation times were estimated by simply multiplying protein lengths with an average translation rate. We here assumed a relatively fast translation rate of 20 aas / s for all prokaryotic proteins and 5 aas / s for all eukaryotic proteins [4-6], Despite this, the distributions of translation times and folding times are clearly separated (Figure 1A). Translation times are typically on the order of seconds, whereas folding times range from microseconds to seconds, and in 126 of 133 cases (95%), the in vitro refolding time of the full-length protein is shorter than the time estimated to complete its translation (Figure 1A). Moreover, we here consider the time it takes for an entire polypeptide chain to cooperatively fold to its native conformation. Local protein conformational dynamics are generally even faster ranging from nanoseconds to microseconds

[0028] , We conclude that for many proteins folding is a co-translational process that starts as soon as the N-terminal part of the protein emerges from the ribosome tunnel and long before the full-length protein chain has been synthesized and released from the ribosome. 2.Vectorial protein translation imposes spatial and temporal constraints on folding

[0061] Protein folding studies both in vitro as well as in the cell have demonstrated that protein topology (i.e. the sequence order in which the structural elements of the tertiary fold occur in the primary sequence) is a key determinant of folding rates and efficiencies

[0029] , Protein topological complexity is often described by Contact Order (CO), a metric which calculates the average sequence distance separation of native interactions. As shown in Figure 1A, translation is relatively slow compared to folding. This means that during protein translation topology restricts folding not only spatially but also temporally: residues that are separated spatially in the primary sequence are also separated temporally as interaction partners towards the C-terminal end of the protein will simply not exist until they have been translated. Inspired by and building on CO, we here propose a metric that accounts for these temporal topological constraints, for which we have designated the term "FoldDelay" (FD). For each residue ( / ) in a protein sequence, FD measures the sequence distance (DS) from its furthest away C-terminal interactor ( / ):

[0062] FoldDelay i (aas) = S j

[0063] In other words, FD measures the number of residues that need to be synthesized before residue ( / ) can engage with all its native interaction partners. The assumption is that until that moment, its folding is delayed (hence "FoldDelay"). FD can also be expressed in time units by factoring the decoding times (tdec) of the different residues between / and j:

[0064] The FD calculation is schematically represented in Figures 1B-E. Figure IB shows the native fold for a hypothetical small globular protein. From this native structure, all interactions are mapped, resulting in a contact map (Figure 1C). While in post-translational folding all these contacts are available simultaneously (Figure 1C), in the co-translational paradigm the contact map changes over time (Figure ID and E). Residue 5, for example, interacts with residue 24 in the native structure (residues outlined in red in panels Figures 1B-E). As residue 5 emerges from the ribosome, this interaction is not available, as residue 24 has not been added to the polypeptide chain (Figure ID). Therefore, residue 5 cannot form all its native interactions until residue 24 has emerged from the ribosome and becomes physically accessible (Figure IE). Residue 5, therefore, incurs a FD of 19 aas. Importantly, residue 24 has practically no FD as all its long-range interaction partners (residues 5 and 6) have already been added to the polypeptide when it emerges from the ribosome exit tunnel. As an example, Figures 1F-H show the FD calculation for the E. coli peptidyl-prolyl isomerase B (PPI B, UniProt code P23869) protein. PPI B is an abundant cytoplasmic enzyme with a length of 164 amino acids. Its functional form is a globular shape comprised of beta sheets and alpha helices separated by several random coils (Figure IF). While PPI B has an average folding time of about 600 msits estimated translation time is eight seconds (assuming an average translation rate of 20 aas / s). Clearly, the timescales of folding and translation here are vastly different, and PPI B is likely to start folding co-translationally. Figures 1G and 1H show the contact map calculated from the PPI B native structure and the per-residue FD profile, respectively. PPI B contains a beta-sheet comprised of a strand close to the N-terminus (El), and a strand close to the C-terminal end in the primary sequence (E8). Strand El cannot fully be stabilized in its native conformation until strand E8 has been produced, which takes about 160 aas. This means that strand El sits partially exposed for about eight full seconds. On the other hand, strand E8 has a negligible FD, as its interaction partners have all been produced when it emerges from the ribosome. To explore general FD patterns across different protein topologies, we ran FD analyses on the protein domains of the SCOPe40 dataset. This dataset contains single-domain structures that have been manually classified based on their architectures and filtered so that no two domains in the set have more than 40% identical sequences [30, 31], Reflecting the vectorial nature of protein translation, FD has both spatial and temporal implications. First, the FD profiles of proteins display an N- to C-terminal gradient: N-terminal elements generally incur larger FDs than more C-terminal elements (Figure II). In addition, domain size is a big determinant of FD, as the longer a polypeptide chain, the more potential there is for long-range interactions, leading to high FDs (Figure 1J). On top of length, FD also reflects the topological specificity of the translated protein. Indeed, even when considering proteins of identical length (101 aas), proteins from different SCOPe classifications have different FD patterns (Figure 1J and K). More complex protein topologies have higher FDs and present more pronounced N- to C-terminal FD gradients, resulting in different profiles for alpha-helical or beta-sheet structured domains (Figure IK and L). Interestingly, FD profiles of large multi-domain proteins often display a sawtooth profile reflecting the domain dependence of N- to C-terminal FD gradients. On top of topology, FD is also dependent on translation rates, which can vary strongly between species. Figure IM shows the AlphaFold predicted structure of a peptidyl-prolyl cis-trans isomerase (CPR1, UniProt code P14832) from S. cerevisiae, which is homologous and structurally very similar to PPI B from E. coli. As is the case for PPI B, the N-terminal domain has a strand near its N-terminus (El) that forms contacts with a strand at the C-terminus of the domain (E8), resulting in high FD values for El (Figure IN and O). Although the FDs for the N-terminal strands in PPI B and CPR1 are similar when expressed in number of aas, the relatively slower translation rates of S. cerevisiae (estimated to be around 5 aas / s on average), means that strand El idles on the ribosome for about 32 seconds, as opposed to the 8 seconds estimated for strand El in PPI B. Therefore, differences in translation rates of different organisms can cause domains with very similar folds to incur vastly different FDs. 3. Sequence segments with high FoldDelays often consist of aggregation-prone tertiary structural elements that stabilize the native structure

[0065] Having established the FD algorithm, we next used it to explore FD patterns on a proteome-wide scale. The near-exhaustive availability of AlphaFold-predicted structures combined with the computationally inexpensive nature of our algorithm allows us to calculate FD for all residues across entire proteomes [32, 33], In addition, AlphaFold models provide a confidence measurement to assess the relative position of two residues within the predicted structure, called the Predicted Aligned Error (PAE). We used this metric to filter out interactions between residues whose relative positions with respect to each other are predicted with low confidence since these interactions most probably do not occur in the actual structure, as is the case for contacts with disordered regions or some contacts between distinct domains. We calculated the FD incurred by all residues in the E. coli and S. cerevisiae proteomes, assuming average translation rates of 20 aa / s and 5 aas / s, respectively. Interestingly, most proteins have at least one residue that has to wait for tens of seconds for the translation of all its native interacting residues (Figure 2A). Binning proteome-wide FD however, reveals that most residues have low FDs as they interact only with their neighbours (± 5 aa). While intermediate FDs are relatively rare, about 23% of residues in S. cerevisiae proteome incur FDs of more than 10 seconds (Figure 2B), while the same is true for 7% of E. coli residues. Specific secondary structures are more likely to incur FD (Figure 2C). Residues in random coils (C) are depleted from the high FD groups as they barely make any contacts. On the other hand, helical structures (G, H and I) are dominated by short-range contacts, yielding average FDs. Pi-helices (I) have higher FDs than alpha-helices (H), which makes sense given that backbone interactions in pi-helices occur at an interval of five residues, where this is four residues for alpha-helices and three for 3-turn helices (G). Finally, beta-structured elements (B, E), are enriched in residues with the highest FDs. Again, this makes sense given that contacts between beta strands are generally more long-range than those between residues in alpha helices. Looking at the sequence composition of segments with high FD, we find them to be enriched in aromatic and aliphatic residues (Figure 2D). This makes sense as these residues are often buried in the hydrophobic cores of globular proteins, where they make many contacts. Exploring this further, we find that regions of high FD are often structurally ordered - as indicated by the AlphaFold pLDDT score, which inversely correlates with disorder - (Figure 2E) and indeed constituted of buried residues - as shown by their relatively low solvent accessibility (Figure 2F). Furthermore, regions of high FD are usually stabilizing to the domain structure, as shown by their low free energy (Figure 2G). Given their propensity for beta-sheet formation and hydrophobic nature, we asked whether regions of high FD tend to be aggregation-prone. Indeed we find that the proportion of residues in aggregation- prone regions substantially increases with FD (Figure 2H), although the distribution of their aggregation propensities is quite similar (Figure 21). 4. The co-translational chaperone Ssb binds to regions with high FoldDelays

[0066] Aggregation-prone exposed regions of high hydrophobicity are the preferred binding sites of many molecular chaperones, including Hsp70s [24, 34, 35], It is proposed that Hsp70s bind to these regions to delay the folding of newly forming polypeptides until the residues required for folding emerge from the ribosome, thus preventing the formation of non-native interactions [35, 36], Given that FD reflects co- translational exposure and that regions of high FD tend to be hydrophobic, we hypothesized that FD could help explain the engagement of specific segments of the nascent chain by chaperones. To address this question, we used a dataset containing the binding footprints for the co-translational chaperone Ssb from S. cerevisiae, obtained by Dbring et al.

[0024] using selective ribosome profiling (SeRP). These binding footprints indicate the codons that are being translated by ribosomes while Ssb is bound to the emerging polypeptide chain (Figure 3A). We carried out a metagene analysis by aligning the starting site of Ssb binding footprints across the S. cerevisiae proteome and calculated the median FD value at each position. A distinct FD peak was revealed at around 50 aa towards the N-terminal side (Figure 3B). This is the exact distance that has been reported to exist between the Ssb footprint, i.e., the sequence segment protected by the ribosome at the moment of Ssb engaging the nascent chain, and the actual Ssb binding site [24, 37], This indicates that the observed FD peak is directly associated with the regions engaged by Ssb. Indeed, at these positions, we observed some of the characteristic sequence and structural properties of Ssb binding motifs [24, 25], including an enrichment in positively charged residues and p-sheet propensity, and an underrepresentation of intrinsically disordered regions. A similar FD pattern was observed using a different published dataset of Ssb binding regions

[0025] , On the other hand, a dataset of Ssb binding regions generated in the absence of RAC (RACA,

[0024] ), a cochaperone that is required for high affinity binding of Ssb to its substrates, did not show any peak around these positions. Interestingly, there is an additional FD peak between -16 and -6 aa from the start of Ssb biding footprints, which is approximately 36 residues downstream of the main Ssb binding region (Figure 3B). This peak might correspond to other Ssb binding regions, as these have been previously described to occur in proteins every 36 amino acids, on average

[0038] , Together, these results indicate that Ssb binds to regions with high FD. I ntriguingly, despite Ssb recognition motifs being extremely common within protein sequences

[0038] , SeRP data showed that many putative binding sites in vitro are actually ignored in vivo [24, 25], Thus, we investigated whether putative chaperone binding motifs with low FDs could be skipped co- transitionally by Ssb. To investigate this, we used the computational tool Limbo to predict chaperone binding sites in yeast proteins

[0039] , Although Limbo was trained to predict E. coli DnaK binding sites, these motifs have been shown to be very similar to Ssb binding regions

[0024] , In fact, Limbo regions are enriched around 50 residues upstream of Dbring et al.

[0024] Ssb footprints (figure S3F) and match with half of the identified Ssb binding regions (Figure 3C). On the other hand, only 22% of all Limbo predicted regions matched with an Ssb binding region (P-value < 0.001 by Fisher exact test), suggesting that there is a higher level of regulation beyond the amino acid sequence. Comparing Limbo regions that matched and did not match with Ssb binding regions, we saw that those that are not bound by Ssb co- translationally have lower FDs, even when analysing regions with similar relative positions in both groups to avoid biases (Figure 3D). As an example, the protein S-methyl-5'-thioadenosine phosphorylase (MTAP) has six predicted chaperone binding sites based on Limbo (Figures 3E). Out of these, only two were experimentally identified in vivo and reside in regions with high FDs. Conversely, the other four predicted binding sites are in regions with lower or even negligible FDs. It seems then that Ssb not only engages target based on amino acid composition, but also on availability, which is aptly captured by the FD metric. The experimentally determined Ssb binding sites in MTAP have a FD of about 20 seconds during which time they are available for Ssb engagement. The more C-terminal Ssb binding sites, on the other hand, have no FD as all their interacting residues have already been translated. As discussed by Dbring et al. in the original Ssb SeRP publication, the maximal lifetime of the Ssb-Nascent chain complex can be extrapolated from the width of Ssb-binding peaks

[0024] , The average width of the Ssb peaks we consider in our analysis is 6.9 aas, which corresponds to a translation time, and hence Ssb engagement time, of 1.38 seconds. I ntrigu ingly, we found that FDs of experimentally confirmed Ssb binding regions are, on average, 1.44 s higher FDs than regions from the same proteins that were sampled at random (Figure 3G). This suggests that regions that have a FD that is equal to or higher than the Ssb binding time can actually be engaged by the chaperone. To corroborate this, we asked whether Ssb binding sites with longer engagement times, i.e., wider footprints, have higher FDs. For Ssb footprints ranging in size from 5 to 11 aas, we indeed observed a strong positive correlation between the average FD at positions -53 to -35 (Ssb binding region) and the footprint size (p-value = 0.04) (Figures 3G and 3H). Moreover, the slope of this correlation roughly corresponds to the addition of one amino acid (Figure 3H). Correlations outside the Ssb binding region were weaker and not significant.

[0067] 5. Proteins with high FoldDelays are associated with co-translational misfolding and aggregation

[0068] We have shown that Ssb preferentially engages regions of high FD. To corroborate this, we used a dataset produced by Willmund et al., who mapped Ssb clients across the S. cerevisiae proteome and showed that the deletion of Ssb leads to widespread aggregation of newly synthesized polypeptides

[0023] , We used this dataset to assess whether Ssb clients indeed have higher FDs and whether proteins with high FDs are disproportionately affected by Ssb deletion. To this end, we assigned a single value to each protein by simply summing the FDs of individual residues. As expected, Ssb clients generally have higher total FDs than proteins that are not engaged by the co-translational chaperone (Figure 4A). Furthermore, Ssb clients that aggregate upon deletion of Ssb (SSBA,

[0023] ) have a higher total FD than Ssb substrates that remain soluble (Figure 4A). To examine this difference in more detail, we looked at the metagene FD profile of specific Ssb binding sites of aggregated and soluble Ssb substrates based on Dbring et al.

[0024] ribosome footprints. Ssb binding regions in proteins that aggregate in SSB! cells have, on average, a one second higher FD compared to binding regions in proteins that do not aggregate (Figure 4B). We next investigated whether proteins in the aggregated fraction upon Ssb deletion have higher intrinsic aggregation propensities. Although these proteins have a similar number of APRs per length unit (Figure 4C), we found that proteins that aggregate upon Ssb deletion have a significantly higher proportion of APRs in their Ssb binding regions (positions -53 to -35) compared to other regions in the same proteins of the same size (Figure 4D). In contrast, proteins that remain soluble have a significantly lower proportion of APRs in their Ssb binding regions (P-value < 0.0001 by Fisher exact test), similarly to other regions from the same proteins (Figure 4E). This suggests that APRs are driving the aggregation of the aggregated proteins in SSB! cells. To further corroborate these findings, we analysed a dataset produced by Jacobson et al. who identified proteins that aggregate upon treatment of yeast cells with arsenite

[0040] , a metalloid known to cause aggregation by interfering with the folding of nascent proteins

[0041] , Again, we found that arsenite stress disproportionately causes the aggregation of proteins with high total FD (Figure 4F). In eukaryotic cells, misfolded proteins are tagged through ubiquitination for degradation

[0042] , Duttler et al.

[0043] showed that a subset of cytoplasmic nascent polypeptides is often co-translationally ubiquitinated. Re-analysis of this dataset revealed that proteins that are co- translationally ubiquitinated have a significantly higher total FD compared to other abundantly translated yeast proteins (Figure 4G).

[0069] 6. FD cannot be fully compensated for through codon optimization

[0070] Our previous analyses suggest that fold-delayed regions are weak spots in co-translational folding that can jeopardize folding outcomes. We therefore asked whether genetic sequences are in any way optimized to reduce FD. In our analyses so far, we made use of flat translation rates across transcripts to estimate FoldDelays in actual time units. During translation in vivo, codons are not translated at a flat rate however. Instead, ribosome profiling studies revealed a variety in codon translation rates, even between codons encoding the same amino acid

[0044] , This means that there is a possibility reducing FD by preferentially using fast-translating, "optimal" codons in regions that span long-range interactions To test this, we attempted to find correlations between FD and several codon optimization metrics, including the Codon Adaptation Index and the more recently described %MinMax

[0045] , We were however unable to find convincing evidence of codon optimization towards reducing FD (data not shown). We next asked what the actual reduction in FD would be, given perfect codon optimization. Using the typical codon translation rate tables for S. cerevisiae determined by Dana and Tuller

[0044] and Sharma et al

[0046] (Figure 5A and B, respectively), we recalculated the FD of the longest-range interaction in each S. cerevisiae protein using either the wild-type codon, yielding the "actual" FD or the synonymous codon with the fastest typical translation rate, giving the "minimal" FD (Figure 5C and D). In effect, the minimal translation rates represent the idealized situation where every codon in between long-range interactors is optimized for speed. We then calculated the hypothetical time gained by fully optimizing sequences in between long-range interactors (Figure 5E and F). Even at very high FDs, these gains seem to be marginal. Indeed, calculating the proportional reduction of FD given perfect optimization, we find that the actual FD could only be reduced by about 20% according to the Dana & Tuller decoding timetables, and around 15% according to the tables devised by Sharma et al. (Figure 5G and H). The reason we were unable to find codon optimization towards FD may therefore simply be that there isn't much to gain. Moreover, optimizing long stretches of amino acids in between interactors is evolutionarily a tall order, given that individual point mutations will have extremely marginal FD effects, and are therefore unlikely to persist.

[0071] 7. Isolation of linear CWS regions in the human PARP protein

[0072] Figure 6 depicts the per-residue fold-delay (FD) calculation for human PARP1 (Poly [ADP-ribose] polymerase 1), a protein that mediates poly-ADP-ribosylation of other proteins and plays a key role in DNA repair.

[0073] The graph shows the amino acid positions of PARP1 on the x-axis and the fold-delay values on the y-axis. The fold delay value represents the distance in number of residues between the amino acid residue and the most C-terminal amino acid interaction partner as present in the 3D fold. In Figure 6 three regions are highlighted (with boxes and arrows 1, 2, 3) with a high fold-delay (at least 350), meaning that these amino acid stretches cannot be fully stabilized in its native conformation until the next 350 (or more) amino acids have been produced. This means those regions are partially exposed for a long time and can be seen as co-translational weak spots.

[0074] The three identified CWS regions are:

[0075] Indicated by arrow 1: Amino acid position 40 - QSPMFDG (SEQ ID NO: 1) - until amino acid position 46 (length of CWS: 7 amino acids)

[0076] Indicated by arrow 2: Amino acid position 312 - TGDVTAWTK (SEQ ID NO: 2) - until amino acid position 320 (length of CWS is 9 amino acids)

[0077] Indicated by arrow 3: Amino acid position 668 - PVQDLIKMIFDV (SEQ ID NO: 3) - until amino acid position 679 (length of CWS is 12 amino acids). 8. Isolation of a non-linear CWS region in the human PFAS protein

[0078] A non-linear CWS region was identified in the human protein phosphoribosylformylglycinamidine synthase (PFAS). The amino acid sequence of the human PFAS protein is found in the Uniprot database as 015067. The PDB structure used for calculating the amino acid interactions was the Alphafold model AF-O15067-Fl-model_v4. With the FoldDelay algorithm we identified a CWS consisting of five amino acids (VAL674, LEU705, THR917, GLU921 and PHE924) which five amino acids all interact with TRP1011 in the 3-dimensional structure, forming a three-dimensional pocket that could be targeted with a druglike molecule.

[0079] Amino acids VAL674, LEU705, THR917, GLU921 and PHE924 interact with TRP1011 (which is the most C- terminal residue in the linear amino acid sequence of PFAS with which VAL674, LEU705, THR917, GLU921 and PHE924 interact). We calculated that the FoldDelay values for these amino acids were higher than 85 residues Thus the identified CWS consisting of CWS consisting of the five amino acids (VAL674, LEU705, THR917, GLU921 and PHE924) is a novel drug pocket present in the human PFAS enzyme.

[0080] Materials and methods

[0081] Protein folding vs protein translation rates

[0082] Protein folding rates were retrieved from the Protein Folding Database

[0027] (PFD2.0). This curated dataset contains folding rates derived from experimental data. From the reported folding rates at 25 degrees C (k , we calculated average folding times (calculated as 1 / kf). For an estimation of the translation times, proteins from prokaryotic organisms were assigned translation rates of 20 aas / s, whereas proteins form eukaryotes were assigned translation rates of 5 aas / s [4-6], For an estimation of the total translation time of a protein, we simply multiplied these translation rates by the number of residues in each protein studied.

[0083] SCQPe40 analysis

[0084] We analysed FD profiles of protein domains in the SCOPe40 dataset. This dataset contains single-domain structures that have been manually classified based on their architectures and filtered so that no two domains in the set have more than 40% identical sequences [30, 31], To establish a general pattern of FD from N- to C-term within domains (Figure II), residues were assigned relative positions by dividing their position in the domain by the domain length, multiplying by 100 and rounding off to the nearest integer. For each relative position, average FDs and standard deviations were calculated. The average (or mean) FD for a domain was calculated as the sum of the FD of all residues in a domain divided by the domain length. Proteome-wide analyses

[0085] AlphaFold structures (version 4) and their corresponding predicted aligned error (PAE) matrices for the full proteomes of Escherichia coli and Saccharomyces cerevisiae (yeast) were retrieved from the AlphaFold Protein Structure Database [32, 33], Genomic sequences for both species were retrieved from NCBI Genomes FTP server. AlphaFold structures were mapped to genomic sequences using the UniProt ID mapping tool. 3,929 and 4,363 proteins were successfully matched with their corresponding codon sequences for E. coli and yeast, respectively. The energies of the structures were minimized using the FoldX "RepairPDB" command, and stability calculations for each amino acid were performed using the "SequenceDetail" command

[0055] . Protein secondary structures and absolute solvent accessibility values were obtained with DSSP based on the AlphaFold structures [56, 57], Then, the relative solvent accessibility (RSA) values were calculated by dividing the absolute solvent accessibility values by residuespecific maximal accessibility values, as extracted from Tien et al

[0058] , Disordered regions were defined using the pLDDT score provided in the AlphaFold models, as regions with low confidence scores (pLDDT < 50) have been shown to overlap largely with intrinsically disorder regions

[0059] , To exclude biases arising from intrinsically disordered proteins, proteins with more than 90% disordered residues were filtered out of the data. Aggregation prone regions were defined with the TANGO algorithm (score > 5)

[0060] at physiological conditions (pH at 7.5, temperature at 298 K, protein concentration at 1 mM, and ionic strength at 0.15 M).

[0086] FoldDelay

[0087] FoldDelay (FD) profiles were determined from protein structures for all SCOP40 domains and the E. coli and yeast proteomes based on AlphaFold models using the formulas described in the Results section. Residues were considered to interact if they contained non-hydrogen atoms within 6 A. This threshold was chosen since it is commonly used to calculate other topological parameters, such as contact order [3], For SCOP40 domains, all residue interactions were considered as the structures were solved with experimental methods. Instead, for AlphaFold predicted models, interactions between two residues whose relative position to each other is low based on the Predicted Aligned Error (PAE) metric were filtered out. Specifically, we excluded interactions with an expected position error > 6 A. We assigned a single FD value to each protein to facilitate the proteome-wide FD correlations in Figure 1 and Figure 4. The "mean FD" values correspond to the mean of the FD of individual residues in a structure. The "total FD" values reported are simply the sum of the FD of individual residues in a structure. These metrics provide a global view of the delay incurred by a polypeptide chain throughout its ribosomal production. FD optimization through decoding times

[0088] To assess whether codon optimization could alleviate FD, we recalculated it for S. cerevisiae using codonspecific decoding times as reported by Tuller et al

[0044] , as well as by Sharma et al

[0046] , the latter of which was based on data reported originally by Weismann et al

[0061] , For each protein, the maxFD was calculated as the furthest interaction in amino acids. To calculate the "actual" FD's reported in Figure 5, each codon was assigned its mean decoding time. To calculate "minimal" FD, each codon was assigned the minimal mean decoding time of the codons that encode the same amino acid as the original. Gain was calculated as the difference between the actual FD (s) and the minimal FD (s), representing by how much time FD could in theory be reduced by optimization of codons. The proportional differences were calculated as the gain (s) divided by the actual FD (s).

[0089] Ssb binding profiles

[0090] Ssb binding sites were retrieved from

[0049] , Briefly, the authors used Selective Ribosome profiling (SeRP) to determine the ribosome-protected footprint at the time of Ssb interaction with the nascent chain, extrapolating binding sites from this information. For our analysis, only footprints of 6-8 codons were analysed. To determine local FD patterns, footprints were aligned at their startcodon and median FDs across the codons before and after this region calculated. As a control, an equivalent-sized set of randomly selected regions of 6-8 amino acids from the same proteins was aligned, and local FD patterns assessed.

[0091] Limbo regions

[0092] Hsp70 binding segments were predicted using the LIMBO algorithm described in

[0062] , using default parameters.

[0093] Aggregation-Prone Region determination (APRs)

[0094] APRs were determined using the TANGO algorithm described in

[0063] using default parameters.

[0095] FD of proteins aggregating in Ssb knockout strain

[0096] We reanalyzed a dataset produced by Willmund et al.

[0023] , Through pulldowns of Ribosome Nascent Chain complexes followed by MS, the authors established the S. cerevisiae "translatome". Through Ssb pulldowns, the translatome was then stratified into a group that interacts with Ssb co-translationally ("Ssb not bound" in Figure 4A), and a group that does not. The authors further determined which proteins aggregate upon deletion of the Ssb chaperone, indicating they are dependent on Ssb for their solubility. Using this information, we divided the group of Ssb binders into a "soluble" and an "aggregated" fraction as shown in Figure 4A. FD of proteins sensitive to Arsenite stress

[0097] Ibstedt et al. report the identification of aggregated proteins in S. Cerevisiae both in physiological conditions ("Physiological" in Figure 4F), as well as upon exposure to Arsenite stress ("Arsenic" in Figure 4F)

[0041] , Aggregated fractions were separated through centrifugation and proteins in the aggregated fraction identified through LC-MS. As a background, the authors used a previously established S. cerevisiae proteome, which we copied ("MS proteome" in Figure 4F).

[0098] FD of proteins that are co-translationally ubiquitinated

[0099] Duttler et al. produced a dataset of proteins that are co-translationally ubiquitinated under physiological conditions in S. cerevisiae

[0043] , They do not report a background proteome, so we compared total FD of the co-translationally ubiquitinated proteins with the translatome reported by Willmund et al

[0023] ,

[0100] Statistics

[0101] GraphPad prism or R software were used to perform the different statistical tests. The tests used in each analysis are specified in the corresponding figure. P-values are represented as: * P-value < 0.05, ** P- value < 0.01, *** P-value < 0.001 and **** P-value < 0.0001.

[0102] Visualisations

[0103] Visualisations were performed with GraphPad prism or custom R scripts using the packages ggplot2

[0064] , Contact maps were visualized using the circlize R package

[0065] , ChimeraX was used to visualize protein structures

[0066] ,

[0104] References

[0105] [1] Anfinsen CB. Principles that Govern the Folding of Protein Chains. Science. 1973;181:223 - 30.

[0106] [2] Wagaman AS, Coburn A, Brand-Thomas I, Dash B, Jaswal SS. A comprehensive database of verified experimental data on protein folding kinetics. Protein Sci. 2014;23:1808-12.

[0107] [3] Plaxco KW, Simons KT, Baker D. Contact order, transition state placement and the refolding rates of single domain proteinsllEdited by P. E. Wright. J Mol Biol. 1998;277:985-94.

[0108] [4] Ingolia NT, Lareau LF, Weissman JS. Ribosome profiling of mouse embryonic stem cells reveals the complexity and dynamics of mammalian proteomes. Cell. 2011;147:789-802.

[0109] [5] Liang ST, Xu YC, Dennis P, Bremer H. mRNA composition and control of bacterial gene expression. J Bacteriol. 2000;182:3037-44.

[0110] [6] Chaney JL, Clark PL. Roles for Synonymous Codon Usage in Protein Biogenesis. Annu Rev Biophys. 2015;44:143-66.

[0111] [7] Ciryam P, Morimoto Rl, Vendruscolo M, Dobson CM, O'Brien EP. In vivo translation rates can substantially delay the cotranslational folding of the Escherichia coli cytosolic proteome. Proc Natl Acad Sci U S A. 2013;110:E132-40.

[0112] [8] Waudby CA, Dobson CM, Christodoulou J. Nature and Regulation of Protein Folding on the Ribosome. Trends Biochem Sci. 2019;44:914-26.

[0113] [9] Frydman J, Erdjument-Bromage H, Tempst P, Hartl FU. Nat Struct Biol. 1999;6:697-705.

[0114]

[0010] Samelson AJ, Bolin E, Costello SM, Sharma AK, O'Brien EP, Marqusee S. Kinetic and structural comparison of a protein’s cotranslational folding and refolding pathways. Science Advances. 2018;4:eaas9098.

[0115]

[0011] Zhang G, Hubalewska M, Ignatova Z. Transient ribosomal attenuation coordinates protein synthesis and co-translational folding. Nat Struct Mol Biol. 2009;16:274-80.

[0116]

[0012] Ugrinov KG, Clark PL. Cotranslational folding increases GFP folding yield. Biophys J. 2010;98:1312- 20.

[0117]

[0013] Evans MS, Sander IM, Clark PL. Cotranslational folding promotes beta-helix formation and avoids aggregation in vivo. J Mol Biol. 2008;383:683-92.

[0118]

[0014] To P, Whitehead B, Tarbox HE, Fried SD. Nonrefoldability is Pervasive Across the E. coli Proteome. J Am Chem Soc. 2021;143:11435-48.

[0119]

[0015] Huang C, Wagner-Valladolid S, Stephens AD, Jung R, Poudel C, Sinnige T, et al. Intrinsically aggregation-prone proteins form amyloid-like aggregates and contribute to tissue aging in Caenorhabditis elegans. Elife. 2019;8.

[0120]

[0016] Zhu M, Kuechler ER, Wong RWK, Calabrese G, Sitarik IM, Rana V, et al. Pulse labeling reveals the tail end of protein folding by proteome profiling. Cell Rep. 2022;40:111096.

[0017] Bitran A, Jacobs WM, Shakhnovich E. The critical role of co-translational folding: An evolutionary and biophysical perspective. Current Opinion in Systems Biology. 2024;37:100485.

[0121]

[0018] Natan E, Endoh T, Haim-Vilmovsky L, Flock T, Chaiancon G, Hopper JTS, et al. Cotranslational protein assembly imposes evolutionary constraints on homomeric proteins. Nat Struct Mol Biol. 2018;25:279- 88.

[0122]

[0019] Cassaignau AME, Wlodarski T, Chan SHS, Woodburn LF, Bukvin IV, Streit JO, et al. Interactions between nascent proteins and the ribosome surface inhibit co-translational folding. Nature Chemistry. 2021.

[0123]

[0020] Deckert A, Cassaignau AME, Wang X, Wlodarski T, Chan SHS, Waudby CA, et al. Common sequence motifs of nascent chains engage the ribosome surface and trigger factor. Proc Natl Acad Sci U S A. 2021;118.

[0124]

[0021] Deuerling E, Gamerdinger M, Kreft SG. Chaperone Interactions at the Ribosome. Cold Spring Harb Perspect Biol. 2019;ll.

[0125]

[0022] Ferbitz L, Maier T, Patzelt H, Bukau B, Deuerling E, Ban N. Trigger factor in complex with the ribosome forms a molecular cradle for nascent proteins. Nature. 2004;431:590-6.

[0126]

[0023] Willmund F, del Alamo M, Pechmann S, Chen T, Albanese V, Dammer EB, et al. The cotranslational function of ribosome-associated Hsp70 in eukaryotic protein homeostasis. Cell. 2013;152:196-209.

[0127]

[0024] Doring K, Ahmed N, Riemer T, Suresh HG, Vainshtein Y, Habich M, et al. Profiling Ssb-Nascent Chain Interactions Reveals Principles of Hsp70-Assisted Folding. Cell. 2017;170:298-311 e20.

[0128]

[0025] Stein KC, Kriel A, Frydman J. Nascent Polypeptide Domain Topology and Elongation Rate Direct the Cotranslational Hierarchy of Hsp70 and TRiC / CCT. Mol Cell. 2019;75:1117-30 e5.

[0129]

[0026] Schubert U, Anton LC, Gibbs J, Norbury CC, Yewdell JW, Bennink JR. Rapid degradation of a large fraction of newly synthesized proteins by proteasomes. Nature. 2000;404:770-4.

[0130]

[0027] Manavalan B, Kuwajima K, Lee J. PFDB: A standardized protein folding database with temperature correction. Sci Rep. 2019;9:1588.

[0131]

[0028] Xu Y, Purkayastha P, Gai F. Nanosecond Folding Dynamics of a Three-Stranded -Sheet. J Am Chem Soc. 2006;128:15836-42.

[0132]

[0029] Plaxco KW, Simons KT, Baker D. Contact order, transition state placement and the refolding rates of single domain proteinsllEdited by P. E. Wright. J Mol Biol. 1998;277:985-94.

[0133]

[0030] Fox NK, Brenner SE, Chandonia J-M. SCOPe: Structural Classification of Proteins— extended, integrating SCOP and ASTRAL data and classification of new structures. Nucleic Acids Res. 2014;42:D304- D9.

[0031] Chandonia J-M, Guan L, Lin S, Yu C, Fox Naomi K, Brenner Steven E. SCOPe: improvements to the structural classification of proteins - extended database to facilitate variant interpretation and machine learning. Nucleic Acids Res. 2022;50:D553-D9.

[0134]

[0032] Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583-9.

[0135]

[0033] Varadi M, Anyango S, Deshpande M, Nair S, Natassia C, Yordanova G, et al. AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high- accuracy models. Nucleic Acids Res. 2022;50:D439-D44.

[0136]

[0034] Sekhar A, Velyvis A, Zoltsman G, Rosenzweig R, Bouvignies G, Kay LE. Conserved conformational selection mechanism of Hsp70 chaperone-substrate interactions. Elife. 2018;7:e32764.

[0137]

[0035] Rosenzweig R, Nillegoda NB, Mayer MP, Bukau B. The Hsp70 chaperone network. Nature reviews molecular cell biology. 2019;20:665-80.

[0138]

[0036] Preissler S, Deuerling E. Ribosome-associated chaperones as key players in proteostasis. Trends BiochemSci. 2012;37:274-83.

[0139]

[0037] Zhang Y, Valentin Gese G, Conz C, Lapouge K, Kopp J, Wblfle T, et al. The ribosome-associated complex RAC serves in a relay that directs nascent chains to Ssb. Nature communications. 2020;ll:1504.

[0140]

[0038] Rudiger S, Germeroth L, Schneider-MergenerJ, Bukau B. Substrate specificity of the DnaK chaperone determined by screening cellulose-bound peptide libraries. The EMBO journal. 1997;16:1501-7.

[0141]

[0039] Van Durme J, Maurer-Stroh S, Gallardo R, Wilkinson H, Rousseau F, Schymkowitz J. Accurate prediction of DnaK-peptide binding via homology modelling and experimental data. PLoS Comput Biol. 2009;5:el000475.

[0142]

[0040] Jacobson T, Navarrete C, Sharma SK, Sideri TC, Ibstedt S, Priya S, et al. Arsenite interferes with protein folding and triggers formation of protein aggregates in yeast. Journal of cell science. 2012;125:5073-83.

[0143]

[0041] Ibstedt S, Sideri TC, Grant CM, Tamas MJ. Global analysis of protein aggregation in yeast during physiological conditions and arsenite stress. Biol Open. 2014;3:913-23.

[0144]

[0042] Klaips CL, Jayaraj GG, Hartl FU. Pathways of cellular proteostasis in aging and disease. Journal of Cell Biology. 2018;217:51-63.

[0145]

[0043] Duttler S, Pechmann S, Frydman J. Principles of Cotranslational Ubiquitination and Quality Control at the Ribosome. Molecular Cell. 2013;50:379-93.

[0146]

[0044] Dana A, Tuller T. Mean of the Typical Decoding Rates: A New Translation Efficiency Index Based on the Analysis of Ribosome Profiling Data. G3: Genes | Genomes | Genetics. 2015;5:73 LP - 80.

[0147]

[0045] Rodriguez A, Wright G, Emrich S, Clark PL. %MinMax: A versatile tool for calculating and comparing synonymous codon usage and its impact on protein folding. Protein Sci. 2018;27:356-62.

[0046] Sharma AK, Sormanni P, Ahmed N, Ciryam P, Friedrich UA, Kramer G, et al. A chemical kinetic basis for measuring translation initiation and elongation rates from ribosome profiling data. PLoS Comput Biol. 2019;15:el007070.

[0148]

[0047] Shamir M, Bar-On Y, Phillips R, Milo R. Snapshot: timescales in cell biology. Cell. 2016;164:1302-. el.

[0149]

[0048] Shmookler Reis RJ, Atluri R, Balasubramaniam M, Johnson J, Ganne A, Ayyadevara S. "Protein aggregates" contain RNA and DNA, entrapped by misfolded proteins but largely rescued by slowing translational elongation. Aging Cell. 2021;20:el3326.

[0150]

[0049] Dbring K, Ahmed N, Riemer T, Suresh HG, Vainshtein Y, Habich M, et al. Profiling Ssb-Nascent Chain Interactions Reveals Principles of Hsp70-Assisted Folding. Cell. 2017;170:298-311. e20.

[0151]

[0050] Netzer WJ, Hartl FU. Recombination of protein domains facilitated by co-translational folding in eukaryotes. Nature. 1997;388:343-9.

[0152]

[0051] Morales-Polanco F, Lee JH, Barbosa NM, Frydman J. Cotranslational Mechanisms of Protein Biogenesis and Complex Assembly in Eukaryotes. Annual Review of Biomedical Data Science. 2022;5:67- 94.

[0153]

[0052] Samatova E, Komar AA, Rodnina MV. How the ribosome shapes cotranslational protein folding. Current Opinion in Structural Biology. 2024;84:102740.

[0154]

[0053] Brar GA. Beyond the triplet code: context cues transform translation. Cell. 2016;167:1681-92.

[0155]

[0054] Wu CC-C, Zinshteyn B, Wehner KA, Green R. High-resolution ribosome profiling defines discrete ribosome elongation states and translational regulation during cellular stress. Molecular cell. 2019;73:959-70. e5.

[0156]

[0055] Schymkowitz J, Borg J, Stricher F, Nys R, Rousseau F, Serrano L. The FoldX web server: an online force field. Nucleic Acids Res. 2005;33:W382-8.

[0157]

[0056] Kabsch W, Sander C. Dictionary of protein secondary structure: pattern recognition of hydrogen- bonded and geometrical features. Biopolymers. 1983;22:2577-637.

[0158]

[0057] Joosten RP, te Beek TA, Krieger E, Hekkelman ML, Hooft RW, Schneider R, et al. A series of PDB related databases for everyday needs. Nucleic Acids Res. 2011;39:D411-9.

[0159]

[0058] Tien MZ, Meyer AG, Sydykova DK, Spielman SJ, Wilke CO. Maximum allowed solvent accessibilites of residues in proteins. PloS one. 2013;8:e80635.

[0160]

[0059] Ruff KM, Pappu RV. AlphaFold and implications for intrinsically disordered proteins. Journal of Molecular Biology. 2021;433: 167208.

[0161]

[0060] Fernandez-Escamilla AM, Rousseau F, Schymkowitz J, Serrano L. Prediction of sequence-dependent and mutational effects on the aggregation of peptides and proteins. Nat Biotechnol. 2004;22:1302-6.

[0162]

[0061] Williams CC, Jan CH, Weissman JS. Targeting and plasticity of mitochondrial proteins revealed by proximity-specific ribosome profiling. Science. 2014;346:748-51.

[0062] Van Durme J, Maurer-Stroh S, Gallardo R, Wilkinson H, Rousseau F, Schymkowitz J. Accurate prediction of DnaK-peptide binding via homology modelling and experimental data. PLoS Comp Biol. 2009;5:el000475.

[0163]

[0063] Fernandez-Escamilla A-M, Rousseau F, Schymkowitz J, Serrano L. Prediction of sequence-dependent and mutational effects on the aggregation of peptides and proteins. Nat Biotechnol. 2004;22:1302-6.

[0164]

[0064] Wickham H. Data analysis. ggplot2: Springer; 2016. p. 189-201.

[0165]

[0065] Gu Z, Gu L, Eils R, Schlesner M, Brors B. circlize Implements and enhances circular visualization in R. Bioinformatics. 2014;30:2811-2.

[0166]

[0066] Goddard TD, Huang CC, Meng EC, Pettersen EF, Couch GS, Morris JH, et al. UCSF ChimeraX: Meeting modern challenges in visualization and analysis. Protein Science. 2018;27:14-25.

Claims

Claims1. A computer-implemented method to identify a co-translational weak spot (CWS) in a target protein comprising the steps of: a) determining for each amino acid of a target protein the intramolecular amino acid interaction partners as present in the 3-dimensional fold of said target protein, b) calculating the FoldDelay value for each of said amino acid residues of said target protein, c) identifying a CWS as about 3 to about 15 amino acids present in said target protein, wherein the FoldDelay values of the amino acids in said CWS are at least 30 amino acids as measured in the primary sequence of said target protein, wherein said FoldDelay value of an amino acid residue represents the distance between said amino acid residue and the furthest stabilizing C-terminal amino acid interaction partner as defined in the 3D fold of said target protein and wherein the distance between said amino acid and said furthest stabilizing C-terminal interaction partner is measured in the primary sequence of said target protein.

2. A computer-implemented method as described in claim 1 wherein the CWS is a linear sequence of about 3 to about 15 amino acids present in the target protein.

3. A computer-implemented method as described in claim 1 wherein the CWS is a drug pocket of about 3 to about 15 amino acids present in the target protein.

4. A computer-readable storage medium which stores computer-executable instructions that, when executed by at least one processor, cause the processor to perform the method of claims 1 to 3.

5. An apparatus comprising control circuitry configured to perform the method of claims 1 to 3.

6. Use of a co-translational weak spot (CWS) to screen for compounds binding to said CWS.