Novel selection methods for the discovery of polypeptides with specific assembly and aggregation properties
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- DANMARKS TEKNISKE UNIV
- Filing Date
- 2024-06-14
- Publication Date
- 2026-04-22
AI Technical Summary
Current methods lack efficient strategies for selecting polypeptides based on specific biophysical properties such as aggregation and liquid-liquid phase separation, which are crucial for industrial and therapeutic applications, due to the vast and complex amino acid sequence space.
A method involving a display-library of polypeptide variants is used, where a molar excess of a selection polypeptide is contacted to separate variants into populations exhibiting or lacking the desired property, allowing for the recovery and identification of specific variants through mRNA display technology.
This approach enables the effective selection and mapping of amino acid residues contributing to biophysical properties, facilitating the design of customized polypeptides with desired characteristics, thereby addressing the challenges of protein aggregation and phase separation.
Smart Images

Figure IMGF000041_0001 
Figure IMGF000041_0002 
Figure IMGF000042_0001
Abstract
Description
[0001] Novel selection methods for the discovery of polypeptides with specific assembly and aggregation properties
[0002] Technical field
[0003] The present invention relates to the selection of polypeptide variants exhibiting or lacking a biophysical property of interest.
[0004] Background
[0005] The biophysical properties of a protein are of utmost importance in evaluating its industrial applicability, such as therapeutic potential. For instance, aggregation is a major problem in the pharmaceutical and biotechnological industries. Protein aggregation usually removes any functionality of a given protein (enzyme, antibody), makes it essentially useless, and in some cases toxic. Similarly, protein liquid-liquid phase separation (LLPS), in which an initially homogeneous protein solution forms highly concentrated droplets in a dilute background, may e.g. quench therapeutic targets from antibodies and render the targets inaccessible for therapeutics. For instance, protein LLPS is described in the context of neurological disease, wherein many proteins that cause diseases such as Alzheimer’s and Parkinson’s can undergo LLPS and subsequently aggregate into neurotoxic fibrils.
[0006] While the scientific community has developed several methods to search for proteins with new functions that do not yet exist in nature, methods for selecting proteins based on their biophysical properties are sorely lacking.
[0007] The key for the discovery of new types of proteins with novel biophysical properties is to screen large regions of what is known as amino acid sequence space (AASS) and then select those proteins that have an improved property. However, AASS is the sum of all possible amino acid sequences and is astronomical. Rational design is highly challenging and even advanced computational approaches to design and engineer specific biophysical properties are only feasible for simple applications.
[0008] Improved methods for screening of biophysical properties of polypeptides are thus needed. Summary
[0009] The main objective of the present disclosure is to provide a method for selecting one or more polypeptide variants exhibiting or lacking a property of interest.
[0010] Thus, in a first aspect, the present disclosure concerns a method of selecting one or more polypeptide variants exhibiting or lacking a property of interest, the method comprising the steps of: a) providing a display-library consisting of one or more polypeptide variants of a polypeptide of interest; b) contacting said display-library with a molar excess of a selection polypeptide under conditions wherein the selection polypeptide exhibits said property of interest, such as aggregation or liquid-liquid phase separation, thereby separating the display-library into at least a first population and a second population, wherein: i) the first population comprises polypeptide variants exhibiting the property of interest; and ii) the second population comprises polypeptide variants lacking the property of interest; and c) recovering said first and / or second population, thereby selecting polypeptide variants exhibiting or lacking the property of interest.
[0011] In another aspect, the present disclosure concerns a method of mapping one or more amino acid residues of one or more polypeptides or fragments thereof in relation to a property of interest, wherein said one or more polypeptide or fragments thereof are selected and recovered using the method of the present disclosure.
[0012] In an aspect, the present disclosure concerns a computer implemented method for designing a customized polypeptide sequence from a parent polypeptide sequence, said customized polypeptide sequence correlating with at least one property of interest, said method comprising the steps of: obtaining polypeptide variant information for the parent polypeptide sequence, providing a database of polypeptide variants data and of property of interest data, wherein said polypeptide variants are selected according to the method of any one of the preceding claims, said database comprising o a plurality of polypeptide variant data records, where each polypeptide variant data record consists of the amino acid sequence of each polypeptide variant, and o a plurality of property of interest data records, wherein each polypeptide variant data record is correlated with one data record, wherein the data record reflects whether the property of interest is exhibited or not by each polypeptide variant, such that the property of interest data record of each polypeptide variant is directly connected to the amino acid sequence of each polypeptide variant, preferably to the desired amino acid sequence, categorizing the polypeptide data of each polypeptide variant into categories of property of interest, said categorizing based on whether the polypeptide variant exhibits or lacks the property of interest, determining, based on the categorization, specific amino acid residues correlating with the exhibiting or lacking the property of interest, and designing the amino acid sequence of the customized polypeptide by selecting, for each residue of the customized polypeptide, an amino acid which is likely to, preferably the amino acid which is most likely to, contribute to the customized polypeptide exhibiting or lacking the property of interest, thereby obtaining the sequence of the customized polypeptide.
[0013] In another aspect, the present disclosure concerns a kit for selecting one or more polypeptide variants exhibiting or lacking a property of interest, the kit comprising: a) a display-library consisting of one or more polypeptide variants of a polypeptide of interest; b) a selection polypeptide in an amount allowing for a molar excess of said selection polypeptide to be added to the one or more polypeptide variants; and c) instructions for use.
[0014] In another aspect, the present disclosure concerns a polypeptide identified or obtained by the method of the present disclosure. In another aspect, the present disclosure concerns a polynucleotide encoding the polypeptide of the present disclosure.
[0015] In another aspect, the present disclosure concerns a vector comprising a polynucleotide of the present disclosure.
[0016] In another aspect, the present disclosure concerns a host cell comprising a vector or a polynucleotide of the present disclosure.
[0017] Description of Drawings
[0018] Figure 1
[0019] Overview of technology, a) An mRNA-display library (grey) is added to a solution of the non-mutated protein of interest undergoing LLPS (black). Each library member will partition uniquely based on their individual chemistry between the two phases, b) The two phases can be separated by centrifugation and sequencing is used to identify the library members present in the two phases.
[0020] Figure 2 mRNA display strategy for N-terminal intrinsically disordered region (I DR) of Dead-box helicase 4 (Ddx4N1). To investigate the importance of different segment of the IDR, fragments or tiles of the sequence were screened using mRNA-display. By overhang- PCR, regions needed for in vitro transcription (T7 promoter), translation (Ribosomal binding site-(RBS)), and mRNA-display (Hemagglutinin epitope tag (HA-tag), and puromycin-linker complementary sequence (Pur-lig)) were introduced. Transcribed mRNA is thus able to form a hetero-duplex with the DNA sequence of the puromycin (Pur) containing linker molecule to facilitate ssDNA / RNA ligation at the 3’end of the mRNA.
[0021] Figure 3
[0022] Average partition propensity along the sequence of Ddx4N1 calculated for different tile lengths. A) Scores for each amino acid position in the IDR are obtained by calculation of the average dilute phase preference of all fragments covering this position.
[0023] Additionally, these values are calculated separately for 4 groups each containing different lengths of tiles from the library. This shows that smaller tiles partition relatively less than longer as seen by the gradual decrease of the dilute phase preference in longer tiles (darker color). Oppositely, shorter length tiles (e.g. 14, 16, and 18 amino acids - lightest color) show more pronounced variations and thus map the sequence with higher resolution. B) Scores for each amino acid position in the I DR were obtained by calculation of the average partition free energy of all fragments covering this position. Additionally, these values were calculated separately for 4 groups each containing different lengths of tiles from the library. The results show that smaller tiles partition relatively less than longer ones as seen by the gradually more favorable partition free energy in longer tiles (darker color). On the other hand, shorter length tiles (e.g. 14, 16, and 18 amino acids - lightest color) show more pronounced variations and thus map the sequence dependence with higher resolution. Data is shown after subtracting 2.5 kJ / mol, which is the current best estimate for the contribution from the puromycin linker (see Analysis).
[0024] Figure 4
[0025] Overview of experiments, a) Dead box helicase contains two central domains which make up the helicase activity. The N- and C-terminal flanking regions are disordered and modulate the activity. The N-terminal disordered region is necessary and sufficient to form condensates, b) Design of tiling library was done with tiles of length 14-50 and spaced with two amino acids (14,16,18...). Similarly, the start of each tile was shifted in blocks of two amino acids. This resulted in -2000 different peptides, which were measured in mRNA-display. c) In a different experiment, a central 40 amino acid long fragment was mutated to investigate the effect of amino acid substitutions in this area. This library contained -30,000 sequences containing on average 8 amino acid substitutions.
[0026] Figure 5
[0027] A) Average partition propensity along the sequence of Ddx4N1. Scores for each amino acid position in the IDR are obtained by calculation the average dilute phase preference of all fragments covering this position. The standard deviation at each position is indicated by a light grey interval around this profile. This shows that local regions of the sequence have a higher propensity for partitioning, suggesting that the IDR is built from blocks of “sticky” sequences spaced out from one another. A central region used for a higher resolution mutational study is indicated in a vertical darker grey area from position 100 to 140. B) Average partition free energy along the sequence of Ddx4N1 for tiles of 14 amino acids in length. Scores for each amino acid position in the I DR were obtained by calculation the average partition free energy of all fragments covering this position. The standard deviation at each position is indicated by a light grey interval around this profile. The results show that local regions of the sequence have a higher propensity for partitioning, suggesting that the I DR is built from blocks of “sticky” sequences separated from one another by “spacer” regions. A central region used for a higher resolution mutational study is indicated in a vertical darker grey area from position 100 to 140. Data is shown after subtracting 2.5 kJ / mol, which is the current best estimate for the contribution from the puromycin linker (see Analysis).
[0028] Figure 6
[0029] A) Inferred single amino acid effects contributing additively to the partitioning of a 40 amino acid tile containing positions [100-139] in Ddx4. Original amino acids are indicated by small circles and positions with inadequate data are indicated with grey crosses. The values were fitted using L2 regularised linear regression with a lambda value of 31.6 chosen based on 5-fold cross validation. B) Similar model fit on the partition free energies of the single peptide variant data using a lambda value of 3162 instead. The model captures known interaction prone residues (Trp, Tyr, Phe, and Arg) as most strongly promoting partitioning. Additionally, the central part of the tile shows a weaker contribution of these amino acids to partitioning into the dense phase, consistent with the sequence profile of the previous tiling library.
[0030] Figure 7
[0031] Validation of mRNA display using individual fluorescently labelled peptides. Six peptides of 14 amino acids lengths from the tiling library were synthesised by solid state peptide synthesis with an additional N-terminal cysteine conjugated to an Oregon Green-488 malemide fluorophore. The partitioning of these purified peptides was measured using the in-house developed method Capflex (Stender et al, 2021) to obtain orthogonal estimates of their partition free energy. The correlation plot is shown with diagonal line of slope 1 with intercept -5.5 kJ / mol, suggesting a constant off-set potentially due to preferential partitioning of the fluorophore alone. The mRNA display data is shown after subtracting 2.5 kJ / mol, which is the current best estimate for the contribution from the puromycin linker (see Analysis). Figure 8
[0032] Proof-of-concept data showing the possibility of microfluidic diffusional separation of a library, a) Schematic overview of a co-flow device (Zhang et al 2016) showing how phase separation of sample (containing high concentration of Ddx4N1 CtoA (375 pM) and 500 mM NaCI) can be induced by the more rapid dispersion of NaCI causing the protein to phase separate. Upon formation of protein condensates / droplets, these will stay in the centre of the device and exit in the central outlet. Library members which are preferentially partitioned to the droplets (A) will then be enriched in the central outlet compared to library members excluded from the droplets (B). b) Three positions along the channel of the co-flow device show the gradual development and coarsening of Ddx4N1 condensates, as seen by bright field microscopy, c) Bright field microscopy images showing how condensates exit through the middle outlet, d) Illustration of the effect of phase separation inside the device on the distribution of Alexa-488 labelled ssDNA, which is strongly partitioned to Ddx4N1 condensates. Upon phase separation, the fluorescence is decreased in the outer outlets and enriched in the central outlet with the condensate droplets, e) Peptides 34 and 62 are the most differently partitioned peptides based on Figure 7, and fluorescently labelled versions of these peptides show indeed distinct distributions of fluorescence in the device (insets show quantifications of fluorescence across the co-flow channel immediately before the outlet including the fluorescence outside of the channel as a reference). Peptide 62 is enriched within the central channel resulting in a 2-fold increase in fluorescence of the central channel compared to outer outlet channels. This result demonstrates that the dynamic range and sensitivity of the separation is sufficient to distinguishing differentially partitioning library members from another.
[0033] Detailed description
[0034] Definitions
[0035] The term “intrinsically disordered region” refers to a region of a polypeptide that lacks a fixed or ordered three-dimensional structure. Intrinsically disordered regions range from, but are not limited to, fully unstructured to partially structured and include, but are not limited to, random coil, molten globule-like states, or flexible linkers in large multidomain polypeptides. The intrinsically disordered region may constitute the entire length of a polypeptide, wherein the polypeptide is sometimes referred to as an “intrinsically disordered protein”. The term “display-library” as used herein, refers to a physical collection containing nonidentical peptides and / or nucleic acids in which there is a direct link between the phenotype and the genotype.
[0036] The term “exhibiting” as used herein may refer to strong exhibiting of a property or a less strong exhibiting of a property. Lacking as used herein refers to polypeptide variants exhibiting the property less strongly than the variants referred to as exhibiting said property or not exhibiting the property at all. A variant lacking a property of interest compared to a reference polypeptide (e.g. the parent polypeptide) thus may still exhibit that property to some extent, but to a lower level than the reference polypeptide.
[0037] The term “mRNA library” as used herein, refers to a mRNA library which can be used in mRNA display. mRNA display is an established methodology which in brief means that the library comprises nucleic acids encoding polypeptide variants, which are transcribed (into mRNA) and translated into protein. In mRNA display, the mRNA is linked (covalently or non-covalently) to the antibiotic puromycin, this mRNA-puromycin complex is provided to a test-tube translation system, translation is stalled at a stop codon in the ribosome allowing a permanent covalent link to form between the protein sequence and the puromycin. This establishes a link between the mRNA carrying the genetic information and the protein performing some function. The skilled person knows how to generate mRNA display libraries (Josephson et a / 2014).
[0038] The term “a random artificial library” refers to a set of polynucleotide sequences that encodes a set of random peptides, and to the set of random peptides encoded by those polynucleotide sequences, as well as the fusion proteins containing those random peptides.
[0039] The term “DNA-encoded library of small molecules” refers to a technology for the synthesis and screening on collections of small molecule compounds. DNA-encoded library of small molecules involves the conjugation of chemical compounds or building blocks to short DNA fragments that serve as identification barcodes. The technique enables the mass creation and interrogation of libraries via selection.
[0040] The term “DNA-encoded library of peptides” refers to a technology for the synthesis and screening on collections of peptides. DNA-encoded library of peptides involves the conjugation of peptide compounds or building blocks to short DNA fragments that serve as identification barcodes. The technique enables the mass creation and interrogation of libraries via selection.
[0041] As used herein, "amino acid" refers to any of the naturally occurring amino acids. It also refers to those known synthetic amino acids. Unless otherwise indicated, all amino acid sequences listed in this disclosure are listed in the order from the amino terminus to the carboxyl terminus. As used herein, the abbreviations for any protective groups, amino acids and other compounds, are in accord with their common usage, recognized abbreviations, or the IUPAC-IUB Commission on Biochemical Nomenclature, unless otherwise indicated (see Biochemistry 11 : 1726 (1972)). As used herein, amino acid residues are represented by the full name thereof, by the three letter code corresponding thereto, or by the one-letter code corresponding thereto, as known in the art.
[0042] The term “unnatural amino acids” refers to all amino acids that are not naturally occurring. Such amino acids include the D-isomers of any of the 20 naturally occurring amino acids described above. Unnatural amino acids also include homoserine, ornithine, norleucine, and thyroxine. Additional unnatural amino acids are well known to one of ordinary skill in the art. An unnatural amino acid may be a D- or L-isomer. An unnatural amino acid may also be an alpha amino acid or a beta amino acid. An unnatural amino acid may also be a post-translationally modified amino acid, such as a phosphorylated serine, threonine or tyrosine, an acylated lysine, or an alkylated lysine or arginine. Many forms of post-translationally modified amino acids are known.
[0043] The term “liquid-liquid phase separation” as used herein refers to the process wherein an initially homogeneous protein solution forms highly concentrated droplets in a dilute background.
[0044] As used herein, “aggregation” refers to a phenomenon in which proteins aggregate (i.e. , accumulate and clump together) intracellularly or extracellularly (including in vitro). When proteins in pharmaceutical acceptable formulations aggregate, the effective dose of the formulation is reduced and they may cause unwanted immunological responses. Proteins may aggregate in their native states or denatured states. As used herein, the term “coacervation” refers to the separation of biomolecular solutions into two or more liquid phases. The phase(s) more concentrated in biomolecular component(s) is / are the coacervate(s) and the other phase is the equilibrium solution. Coacervation often involves interactions between charge- complementary macromolecules, such as positively charged peptides or proteins and DNA or polysaccharides. Coacervation can involve structured / folded or disordered polypeptides.
[0045] As used herein, the term “condensation” or “biomolecular condensation” refers to the formation of regions of concentrated, liquid phase in multicomponent biopolymer solutions. A solution containing many types of biomolecules can undergo condensation, which leads to the local accumulation of certain components of the solution inside condensate droplets, surrounded by a phase more dilute in the given compound. Every compound of the initial solution will display a characteristic partition coefficient between the condensate and the surrounding solution.
[0046] Polypeptide of interest
[0047] The present invention takes advantage of proteins having specific biophysical properties, e.g. their propensity to aggregate or to undergo LLPS (Figure 1). In order to determine which amino acid residues of a polypeptide of interest or parent polypeptide are involved in (i.e. responsible for, correlated with, or linked to) the given biophysical property, an mRNA display library comprising variants of the polypeptide of interest is generated. Said display library can cover the entire AASS of the polypeptide of interest. The library is then contacted with a molar excess of a selection polypeptide, which can be the polypeptide of interest, or a different polypeptide. As an example, the property of interest is LLPS, and the selection polypeptide is the polypeptide of interest, which exhibits the capability of undergoing LPPS. Upon contacting the library with a molar excess of the polypeptide of interest, some variants will undergo LLPS together with the polypeptide of interest, and others will not. This results in the formation of two phases: one containing the polypeptide of interest and the variants who can undergo LLPS, and one who contains the remaining variants, who cannot undergo LLPS. The variants from any of the phases can then be retrieved and their sequences identified. For each amino acid position, a map of residues can thus be generated, and information about the effect of different amino acids at any given position on the property of interest can be mapped. In a first aspect, the present disclosure concerns a method of selecting one or more polypeptide variants exhibiting or lacking a property of interest, the method comprising the steps of: a) providing a display-library consisting of one or more polypeptide variants of a polypeptide of interest; b) contacting said display-library with a molar excess of a selection polypeptide under conditions wherein the selection polypeptide exhibits said property of interest, such as aggregation or liquid-liquid phase separation, thereby separating the display-library into at least a first population and a second population, wherein: i) the first population comprises polypeptide variants exhibiting the property of interest; and ii) the second population comprises polypeptide variants lacking the property of interest; and c) recovering said first and / or second population, thereby selecting polypeptide variants exhibiting or lacking the property of interest.
[0048] Proteins, or polypeptides, are complex molecules that control chemical and biological processes in living organisms and may act as complicated molecular machines or simple structural materials. As described herein above, the biophysical properties of a protein are of utmost importance in evaluating its industrial applicability, such as therapeutic potential. For instance, aggregation is a major problem in the pharmaceutical and biotechnological industries.
[0049] In some embodiments, said polypeptide of interest is an antibody. In some embodiments, said polypeptide of interest is a cyclic peptide. In some embodiments, said polypeptide of interest is a therapeutic polypeptide. In some embodiments, said polypeptide of interest is an enzyme. In some embodiments, said polypeptide of interest is a polypeptide suitable for administration to a subject in need thereof. In some embodiments, said polypeptide of interest is a polypeptide suitable for administration to a human subject in need thereof. In some embodiments, said polypeptide of interest is selected from the list consisting of: lipases, carboxylases, cazymes, kinases, phosphatases, proteases, DNA and RNA polymerases and ligases.
[0050] In some embodiments, said polypeptide of interest is selected from the list consisting of: an antibody, a nanobody, a Fab, a minibody, a diabody, a tribody, a scFv, a DARPin, a de novo binder, a peptide aptamer, and an affibody.
[0051] In some embodiments, said polypeptide of interest comprises one or more intrinsically disordered regions. In some embodiments, said polypeptide of interest is an intrinsically disordered protein.
[0052] Once the polypeptide of interest is identified a display-library may be provided.
[0053] Display-library
[0054] The challenge in identifying polypeptides with specific properties of interest is to search a fraction of a virtually infinite space sufficient to have a realistic chance to find a polypeptide of interest. The most powerful methodologies that currently exist for this purpose are display technologies (a non-limiting list of examples include phage display, ribosome display, mRNA display and cDNA display). Common to these is the large number of different amino acid sequences, which are included. This is commonly done at the level of DNA, meaning that the library consists of different DNA sequences that each encode a distinct polypeptide. In this process, it is crucial to maintain a connection, in one way or another, of the DNA sequence encoding the polypeptide and the polypeptide itself. This is because efficient methods for sequencing of individual polypeptide molecules are currently lacking, whereas DNA sequencing technologies are available, making it relatively easy to retrieve the sequence information from very small quantities of DNA, down to single molecules. This retrieved DNA sequence can subsequently be translated into protein and the properties of the polypeptide are tested and evaluated, ideally in a high throughput assay or through competition emulating the process of evolution. If a given subset of the display-library is found to encode polypeptides with a property of interest, those can be identified and used for subsequent iterations of the same process, until the desired property is reached. In some embodiments, steps a-c) are performed in one or more iterations, wherein said display-library comprises or consists of the first population and / or the second population recovered from one or more previous iterations.
[0055] For example, once steps a)-c) have been performed and one of the first and second populations has been selected, this population can be used to generate a new displaylibrary. Steps a)-c) can then be performed on this new display library, from which another two populations are obtained.
[0056] In some embodiments, steps b-c) are performed in 2 iterations or more, such as 3, 4, 5, 6, 7, 8, 9, or 10 iterations or more.
[0057] The display-libraries typically take one of two forms:
[0058] - Collections of random sequences for de novo discovery.
[0059] - Variants of a pre-defined sequence (e.g. mutations of a known peptide or protein)
[0060] In some embodiments, the display-library is generated from a DNA library encoding polypeptide variants.
[0061] In some embodiments, the one or more polypeptide variants are encoded by one or more polynucleotides. In some embodiments, the one or more polynucleotides comprise a ribosome-binding site.
[0062] In some embodiments, the display-library is a RAPID display library. The RAPID display method allows the complete accomplishment of transcription, translation and linker-mRNA complex formation followed by linkage between the peptide and the linker in a single translation system by modifying mRNA display to replace the puromycin- conjugated linker by a linker molecule optimized for reconstituted cell-free translation system (EP 2 492 344 B1).
[0063] As compared with conventional mRNA display methods, the RAPID display method mainly has the following features.
[0064] (a) The 3' end of the linker molecule has a structure in which an amino acid is attached to adenosine via an ester (i.e. , aminoacylated) rather than puromycin.
[0065] (b) Aminoacylation reaction is mediated by an artificial RNA catalyst (ribozyme). (c) A reconstituted cell-free translation system is used.
[0066] (d) The fusion between the linker and an mRNA is made by complex formation based on hybridization in a translation system rather than ligation.
[0067] (e) Transcription, translation and complex formation with the linker can be performed in a single translation reaction vessel.
[0068] In some embodiments, the one or more polynucleotides are covalently linked to the corresponding one or more polypeptide variants via a linker. In some embodiments, the linker comprises or consists of puromycin or an amino acid.
[0069] In some embodiments, the linker is linked to the 3’-end of the one or more polynucleotides, optionally via a second linker, preferably a flexible linker.
[0070] In some embodiments, the one or more polynucleotides comprises a ribosome-binding site and is covalently linked to the corresponding one or more polypeptide via a linker.
[0071] In some embodiments, the linker and / or the second linker comprises or consists of a nucleic acid sequence consisting of between 16 and 40 nucleotides, such as between 17 and 39 nucleotides, such as between 18 and 38 nucleotides, such as between 19 and 37 nucleotides, such as between 20 and 36 nucleotides, such as between 21 and 35 nucleotides, such as between 22 and 34 nucleotides, such as between 23 and 33 nucleotides, such as between 24 and 32 nucleotides, such as between 25 and 31 nucleotides, such as between 26 and 30 nucleotides, such as between 27 and 29 nucleotides, such as 28 nucleotides.
[0072] In some embodiments, the second linker comprises or consists of a PEG linker.
[0073] In some embodiments, the linker and / or the second linker comprises or consists of a PEG linker and a nucleic acid sequence consisting of between 16 and 40 nucleotides, such as between 17 and 39 nucleotides, such as between 18 and 38 nucleotides, such as between 19 and 37 nucleotides, such as between 20 and 36 nucleotides, such as between 21 and 35 nucleotides, such as between 22 and 34 nucleotides, such as between 23 and 33 nucleotides, such as between 24 and 32 nucleotides, such as between 25 and 31 nucleotides, such as between 26 and 30 nucleotides, such as between 27 and 29 nucleotides, such as 28 nucleotides. In some embodiments, an organic fluorophore is linked to the one or more polynucleotides. In some embodiments, the organic fluorophore is fluorescein.
[0074] In some embodiments, the organic fluorophore is linked to position 5 of a thymine ring by a 6-carbon linker. In some embodiments, said thymine ring is within the linker and / or the second linker.
[0075] A number of display-libraries are commonly used in the field. These include mRNA display, cDNA display, phage display, yeast display, and DNA-encoded libraries of small molecules.
[0076] In some embodiments, said display-library is an mRNA library, a cDNA library, a ribosome display library, a random artificial library, a DNA-encoded library of small molecules or a DNA-encoded library of peptides.
[0077] In some embodiments, said display-library covers part or all of the sequence space of the full-length polypeptide of interest.
[0078] In some embodiments, said display-library covers part or all of the sequence space of one or more regions of interest of the polypeptide of interest.
[0079] In some embodiments, said display-library covers part or all of the sequence space of one or more amino acids of interest of the polypeptide of interest.
[0080] In some embodiments, said display-library is a mRNA library that covers part or all of the sequence space of the full-length polypeptide of interest. In some embodiments, said display-library is a mRNA library that covers part or all of the sequence space of one or more regions of interest of the polypeptide of interest. In some embodiments, said display-library is a mRNA library that covers part or all of the sequence space of one or more amino acids of interest of the polypeptide of interest. In some embodiments, the method comprises the steps of: a) providing a mRNA library consisting of one or more polypeptide variants of a polypeptide of interest; b) contacting said mRNA library with a molar excess of a selection polypeptide under conditions wherein the selection polypeptide exhibits said property of interest, such as aggregation or liquid-liquid phase separation, thereby separating the mRNA library into at least a first population and a second population, wherein: i) the first population comprises polypeptide variants exhibiting the property of interest; and ii) the second population comprises polypeptide variants lacking the property of interest; and c) recovering said first and / or second population, thereby selecting polypeptide variants exhibiting or lacking the property of interest.
[0081] In some embodiments, said display-library is a cDNA library that covers part or all of the sequence space of the full-length polypeptide of interest. In some embodiments, said display-library is a cDNA library that covers part or all of the sequence space of one or more regions of interest of the polypeptide of interest. In some embodiments, said display-library is a cDNA library that covers part or all of the sequence space of one or more amino acids of interest of the polypeptide of interest.
[0082] In some embodiments, the method comprises the steps of: a) providing a cDNA library consisting of one or more polypeptide variants of a polypeptide of interest; b) contacting said cDNA library with a molar excess of a selection polypeptide under conditions wherein the selection polypeptide exhibits said property of interest, such as aggregation or liquid-liquid phase separation, thereby separating the cDNA library into at least a first population and a second population, wherein: i) the first population comprises polypeptide variants exhibiting the property of interest; and ii) the second population comprises polypeptide variants lacking the property of interest; and c) recovering said first and / or second population, thereby selecting polypeptide variants exhibiting or lacking the property of interest.
[0083] Polypeptide variants
[0084] While the display-library allows the screening of large number of polypeptide variants the amino acid sequence space is the sum of all possible amino acid sequences and is astronomical. To illustrate, there are 1065different amino acid sequence combinations for a small protein of 50 amino acids (considering only the 20 naturally occurring amino acids). Furthermore, in some cases, only specific regions or amino acids are of interest. Thus in some cases, the searched amino acid sequence space does not cover the full-length sequence of the polypeptide of interest (Figure 2).
[0085] In some embodiments, said polypeptide variants are fragments of the polypeptide of interest.
[0086] In some embodiments, said fragments of the polypeptide of interest are at least 100 amino acids long, such as at least 200 amino acids long, such as at least 300 amino acids long, such as at least 400 amino acids long, such as at least 500 amino acids long, such as at least 600 amino acids long, such as at least 700 amino acids long, such as at least 800 amino acids long, such as at least 900 amino acids long, or such as at least 1000 amino acids long.
[0087] In some embodiments, said fragments of the polypeptide of interest are between 30 and 1000 amino acids long, such as between 30 and 900 amino acids long, between 30 and 800 amino acids long, between 30 and 700 amino acids long, between 30 and 600 amino acids long, between 30 and 500 amino acids long, between 30 and 400 amino acids long, between 30 and 300 amino acids long, between 30 and 200 amino acids long, between 30 and 100 amino acids long, between 30 and 50 amino acids long, between 50 and 1000 amino acids long, between 50 and 900 amino acids long, between 50 and 800 amino acids long, between 50 and 700 amino acids long, between 50 and 600 amino acids long, between 50 and 500 amino acids long, between 50 and 400 amino acids long, between 50 and 300 amino acids long, between 50 and 200 amino acids long, between 50 and 100 amino acids long, between 100 and 1000 amino acids long, between 100 and 900 amino acids long, between 100 and 800 amino acids long, between 100 and 700 amino acids long, between 100 and 600 amino acids long, between 100 and 500 amino acids long, between 100 and 400 amino acids long, between 100 and 300 amino acids long, between 100 and 200 amino acids long, between 200 and 1000 amino acids long, between 200 and 900 amino acids long, between 200 and 800 amino acids long, between 200 and 700 amino acids long, between 200 and 600 amino acids long, between 200 and 500 amino acids long, between 200 and 400 amino acids long, between 200 and 300 amino acids long, between 300 and 1000 amino acids long, between 300 and 900 amino acids long, between 300 and 800 amino acids long, between 300 and 700 amino acids long, between 300 and 600 amino acids long, between 300 and 500 amino acids long, between 300 and 400 amino acids long, between 400 and 1000 amino acids long, between 400 and 900 amino acids long, between 400 and 800 amino acids long, between 400 and 700 amino acids long, between 400 and 600 amino acids long, between 400 and 500 amino acids long, between 500 and 1000 amino acids long, between 500 and 900 amino acids long, between 500 and 800 amino acids long, between 500 and 700 amino acids long, between 500 and 600 amino acids long, between 600 and 1000 amino acids long, between 600 and 900 amino acids long, between 600 and 800 amino acids long, between 600 and 700 amino acids long, between 700 and 1000 amino acids long, between 700 and 900 amino acids long, between 700 and 800 amino acids long, between 800 and 1000 amino acids long, between 800 and 900 amino acids long, or such as between 900 and 1000 amino acids long.
[0088] In some embodiments, said fragments of the polypeptide of interest are up to 30 amino acids long, such as up to 25 amino acids long, up to 20 amino acids long, up to 15 amino acids long, up to 10 amino acids long, up to 5 amino acids long, up to 4 amino acids long, or such as up to 2 amino acids long.
[0089] In some embodiments, said fragments of the polypeptide of interest are between 2 and 30 amino acids long, such as between 2 and 25 amino acids long, between 2 and 20 amino acids long, between 2 and 15 amino acids long, between 2 and 10 amino acids long, between 2 and 5 amino acids long, between 2 and 4 amino acids long, between 4 and 30 amino acids long, between 4 and 25 amino acids long, between 4 and 20 amino acids long, between 4 and 15 amino acids long, between 4 and 10 amino acids long, between 4 and 5 amino acids long, between 5 and 30 amino acids long, between 5 and 25 amino acids long, between 5 and 20 amino acids long, between 5 and 15 amino acids long, between 5 and 10 amino acids long, between 10 and 30 amino acids long, between 10 and 25 amino acids long, between 10 and 20 amino acids long, between 10 and 15 amino acids long, between 15 and 30 amino acids long, between 15 and 25 amino acids long, between 15 and 20 amino acids long, between 20 and 30 amino acids long, between 20 and 25 amino acids long, between 25 and 30 amino acids long.
[0090] In some embodiments, said polypeptide variants are mutants of the polypeptide of interest.
[0091] In some embodiments, said polypeptide variants are targeted edited variants of the polypeptide of interest.
[0092] In some embodiments, said polypeptide variants comprise one or more unnatural amino acids.
[0093] In some embodiments, said display-library is a mRNA library and said polypeptide variants are fragments of the polypeptide of interest.
[0094] In some embodiments, said display-library is a mRNA library and said polypeptide variants are mutants of the polypeptide of interest.
[0095] In some embodiments, said polypeptide variants are fragments and mutants of the polypeptide of interest.
[0096] In some embodiments, said display-library is a mRNA library and said polypeptide variants are fragments and mutants of the polypeptide of interest.
[0097] Once a display-library consisting of one or more polypeptide variants of a polypeptide of interest has been provided, the library is contacted with a selection polypeptide.
[0098] In some embodiments, said display-library is a cDNA library and said polypeptide variants are fragments of the polypeptide of interest. In some embodiments, said display-library is a cDNA library and said polypeptide variants are mutants of the polypeptide of interest. In some embodiments, said display-library is a cDNA library and said polypeptide variants are fragments and mutants of the polypeptide of interest.
[0099] Selection polypeptide
[0100] The selection polypeptide constitutes a means of inducing the property of interest in the polypeptide variants exhibiting said property of interest. As described herein above, the display library is contacted with a molar excess of a selection polypeptide and a property of interest is induced. Some polypeptide variants of the display-library may exhibit the property of interest and some polypeptide variants of the display-library may lack the property of interest.
[0101] In some embodiments, said molar excess of the selection polypeptide is at least 5-fold, such as at least 10-fold, such as at least 20-fold, such as at least 50-fold, such as at least 100-fold, such as at least 200-fold, such as at least 300-fold, such as at least 400- fold, such as at least 500-fold, such as at least 600-fold, such as at least 700-fold, such as at least 800-fold, such as at least 900-fold, or such as wherein the molar excess of the selection polypeptide is at least 1000-fold.
[0102] In some embodiments, said molar excess of the selection polypeptide is between 5 and 1000-fold, such as between 5 and 900-fold, between 5 and 800-fold, between 5 and 700-fold, between 5 and 600-fold, between 5 and 500-fold, between 5 and 400-fold, between 5 and 300-fold, between 5 and 200-fold, between 5 and 100-fold, between 5 and 50-fold, between 5 and 20-fold, between 5 and 10-fold, between 10 and 1000-fold, between 10 and 900-fold, between 10 and 800-fold, between 10 and 700-fold, between 10 and 600-fold, between 10 and 500-fold, between 10 and 400-fold, between 10 and 300-fold, between 10 and 200-fold, between 10 and 100-fold, between 10 and 50-fold, between 10 and 20-fold, between 20 and 1000-fold, between 20 and 900-fold, between 20 and 800-fold, between 20 and 700-fold, between 20 and 600-fold, between 20 and 500-fold, between 20 and 400-fold, between 20 and 300-fold, between 20 and 200-fold, between 20 and 100-fold, between 20 and 50-fold, between 50 and 1000-fold, between 50 and 900-fold, between 50 and 800-fold, between 50 and 700-fold, between 50 and 600-fold, between 50 and 500-fold, between 50 and 400-fold, between 50 and 300-fold, between 50 and 200-fold, between 50 and 100-fold, between 100 and 1000-fold, between 100 and 900-fold, between 100 and 800-fold, between 100 and 700-fold, between 100 and 600-fold, between 100 and 500-fold, between 100 and 400-fold, between 100 and 300-fold, between 100 and 200-fold, between 200 and 1000-fold, between 200 and 900-fold, between 200 and 800-fold, between 200 and 700-fold, between 200 and 600-fold, between 200 and 500-fold, between 200 and 400-fold, between 200 and 300-fold, between 300 and 1000-fold, between 300 and 900-fold, between 300 and 800-fold, between 300 and 700-fold, between 300 and 600-fold, between 300 and 500-fold, between 300 and 400-fold, between 400 and 1000-fold, between 400 and 900-fold, between 400 and 800-fold, between 400 and 700-fold, between 400 and 600-fold, between 400 and 500-fold, between 500 and 1000-fold, between 500 and 900-fold, between 500 and 800-fold, between 500 and 700-fold, between 500 and 600-fold, between 600 and 1000-fold, between 600 and 900-fold, between 600 and 800-fold, between 600 and 700-fold, between 700 and 1000-fold, between 700 and 900-fold, between 700 and 800-fold, between 800 and 1000-fold, between 800 and 900-fold, or wherein said molar excess of the selection polypeptide is between 900 and 1000-fold.
[0103] In some embodiments, said mRNA library is contacted with a molar excess of the selection polypeptide, wherein the molar excess of the selection polypeptide is between 5 and 1000-fold, such as between 5 and 900-fold, between 5 and 800-fold, between 5 and 700-fold, between 5 and 600-fold, between 5 and 500-fold, between 5 and 400-fold, between 5 and 300-fold, between 5 and 200-fold, between 5 and 100- fold, between 5 and 50-fold, between 5 and 20-fold, between 5 and 10-fold, between 10 and 1000-fold, between 10 and 900-fold, between 10 and 800-fold, between 10 and 700-fold, between 10 and 600-fold, between 10 and 500-fold, between 10 and 400-fold, between 10 and 300-fold, between 10 and 200-fold, between 10 and 100-fold, between 10 and 50-fold, between 10 and 20-fold, between 20 and 1000-fold, between 20 and 900-fold, between 20 and 800-fold, between 20 and 700-fold, between 20 and 600-fold, between 20 and 500-fold, between 20 and 400-fold, between 20 and 300-fold, between 20 and 200-fold, between 20 and 100-fold, between 20 and 50-fold, between 50 and 1000-fold, between 50 and 900-fold, between 50 and 800-fold, between 50 and 700- fold, between 50 and 600-fold, between 50 and 500-fold, between 50 and 400-fold, between 50 and 300-fold, between 50 and 200-fold, between 50 and 100-fold, between 100 and 1000-fold, between 100 and 900-fold, between 100 and 800-fold, between 100 and 700-fold, between 100 and 600-fold, between 100 and 500-fold, between 100 and 400-fold, between 100 and 300-fold, between 100 and 200-fold, between 200 and 1000-fold, between 200 and 900-fold, between 200 and 800-fold, between 200 and 700-fold, between 200 and 600-fold, between 200 and 500-fold, between 200 and 400- fold, between 200 and 300-fold, between 300 and 1000-fold, between 300 and 900- fold, between 300 and 800-fold, between 300 and 700-fold, between 300 and 600-fold, between 300 and 500-fold, between 300 and 400-fold, between 400 and 1000-fold, between 400 and 900-fold, between 400 and 800-fold, between 400 and 700-fold, between 400 and 600-fold, between 400 and 500-fold, between 500 and 1000-fold, between 500 and 900-fold, between 500 and 800-fold, between 500 and 700-fold, between 500 and 600-fold, between 600 and 1000-fold, between 600 and 900-fold, between 600 and 800-fold, between 600 and 700-fold, between 700 and 1000-fold, between 700 and 900-fold, between 700 and 800-fold, between 800 and 1000-fold, between 800 and 900-fold, or wherein said molar excess of the selection polypeptide is between 900 and 1000-fold.
[0104] In some embodiments, said cDNA library is contacted with a molar excess of the selection polypeptide, wherein the molar excess of the selection polypeptide is between 5 and 1000-fold, such as between 5 and 900-fold, between 5 and 800-fold, between 5 and 700-fold, between 5 and 600-fold, between 5 and 500-fold, between 5 and 400-fold, between 5 and 300-fold, between 5 and 200-fold, between 5 and 100- fold, between 5 and 50-fold, between 5 and 20-fold, between 5 and 10-fold, between 10 and 1000-fold, between 10 and 900-fold, between 10 and 800-fold, between 10 and 700-fold, between 10 and 600-fold, between 10 and 500-fold, between 10 and 400-fold, between 10 and 300-fold, between 10 and 200-fold, between 10 and 100-fold, between 10 and 50-fold, between 10 and 20-fold, between 20 and 1000-fold, between 20 and 900-fold, between 20 and 800-fold, between 20 and 700-fold, between 20 and 600-fold, between 20 and 500-fold, between 20 and 400-fold, between 20 and 300-fold, between 20 and 200-fold, between 20 and 100-fold, between 20 and 50-fold, between 50 and 1000-fold, between 50 and 900-fold, between 50 and 800-fold, between 50 and 700- fold, between 50 and 600-fold, between 50 and 500-fold, between 50 and 400-fold, between 50 and 300-fold, between 50 and 200-fold, between 50 and 100-fold, between 100 and 1000-fold, between 100 and 900-fold, between 100 and 800-fold, between 100 and 700-fold, between 100 and 600-fold, between 100 and 500-fold, between 100 and 400-fold, between 100 and 300-fold, between 100 and 200-fold, between 200 and 1000-fold, between 200 and 900-fold, between 200 and 800-fold, between 200 and 700-fold, between 200 and 600-fold, between 200 and 500-fold, between 200 and 400- fold, between 200 and 300-fold, between 300 and 1000-fold, between 300 and 900- fold, between 300 and 800-fold, between 300 and 700-fold, between 300 and 600-fold, between 300 and 500-fold, between 300 and 400-fold, between 400 and 1000-fold, between 400 and 900-fold, between 400 and 800-fold, between 400 and 700-fold, between 400 and 600-fold, between 400 and 500-fold, between 500 and 1000-fold, between 500 and 900-fold, between 500 and 800-fold, between 500 and 700-fold, between 500 and 600-fold, between 600 and 1000-fold, between 600 and 900-fold, between 600 and 800-fold, between 600 and 700-fold, between 700 and 1000-fold, between 700 and 900-fold, between 700 and 800-fold, between 800 and 1000-fold, between 800 and 900-fold, or wherein said molar excess of the selection polypeptide is between 900 and 1000-fold.
[0105] The polypeptide of interest and the selection polypeptide may interact. A non-limiting example is wherein the polypeptide of interest and the selection polypeptide are interaction partners.
[0106] In some embodiments, the polypeptide of interest and the selection polypeptide interact.
[0107] In some embodiments, the polypeptide of interest is capable of binding to the selection polypeptide.
[0108] In some embodiments, the polypeptide of interest is the selection polypeptide.
[0109] In some embodiments, the polypeptide of interest is the selection polypeptide and the polypeptide variants are mutants of the polypeptide of interest. In some embodiments, the polypeptide of interest is the selection polypeptide and the polypeptide variants are fragments of the polypeptide of interest. In some embodiments, the polypeptide of interest is the selection polypeptide and the polypeptide variants are mutants and fragments of the polypeptide of interest.
[0110] In some embodiments, said selection polypeptide is an antibody. In some embodiments, said selection polypeptide is a cyclic peptide. In some embodiments, said selection polypeptide is a therapeutic polypeptide. In some embodiments, said selection polypeptide is an enzyme. In some embodiments, said selection polypeptide is a polypeptide suitable for administration to a subject in need thereof. In some embodiments, said selection polypeptide is a polypeptide suitable for administration to a human subject in need thereof.
[0111] In some embodiments, said enzyme is selected from the list consisting of: lipases, carboxylases, cazymes, kinases, phosphatases, proteases, DNA and RNA polymerases and ligases.
[0112] In some embodiments, said selection polypeptide is selected from the list consisting of: lipases, carboxylases, cazymes, kinases, phosphatases, proteases, DNA and RNA polymerases and ligases.
[0113] In some embodiments, said selection polypeptide is selected from the list consisting of: an antibody, a nanobody, a Fab, a minibody, a diabody, a tribody, a scFv, a DARPin, a de novo binder, a peptide aptamer, and an affibody.
[0114] The selection polypeptide may be involved in diseases.
[0115] Diseases
[0116] As described herein above, properties, including but not limited to liquid-liquid phase separation, are interesting in the context of diseases. Non-limiting examples of this include Alzheimer’s, Amyotrophic lateral sclerosis (ALS), and Parkinson’s, wherein polypeptides can undergo LLPS and subsequently aggregate into neurotoxic fibrils.
[0117] The methods of the present disclosure may be used to investigate the sequence space of polypeptides and allow the identification of amino acid sequences lacking the propensity to undergo e.g. LLPS. Without being bound by theory, the method may be used in the development of variants skewing the balance towards the lack of e.g. tendency to undergo LLPS.
[0118] In some embodiments, said polypeptide of interest is a polypeptide involved in a neurological disease or a rare disease, preferably wherein said disease is associated with protein aggregation in organs, body fluids or body tissues, such as in the brain, in the liver, and / or in the kidney. In some embodiments, said selection polypeptide is a polypeptide involved in a neurological disease or a rare disease, preferably wherein said disease is associated with protein aggregation in organs, body fluids or body tissues, such as in the brain, in the liver, and / or in the kidney.
[0119] In some embodiments, said rare disease is a systemic amyloidosis, such as lysozyme amyloidosis, transthyretin amyloidosis, dialysis-related amyloidosis, light chain amyloidosis.
[0120] In some embodiments, said neurological disease is selected from the group consisting of: amyotrophic lateral sclerosis, Alzheimer’s disease, Parkinson’s disease, Huntington’s disease, and frontotemporal lobar degeneration.
[0121] In some embodiments, said polypeptide of interest is selected from the group consisting of: alpha-synuclein (P37840), TDP-43 (Q13148) C-terminal low complexity domain, Tau (P10636), heterogeneous nuclear ribonucleoproteins (e.g. P09651) and FUS (P56959).
[0122] In some embodiments, said selection polypeptide is selected from the group consisting of: alpha-synuclein (P37840), TDP-43 (Q13148) C-terminal low complexity domain, Tau (P10636), heterogeneous nuclear ribonucleoproteins (e.g. P09651) and FUS (P56959).
[0123] In some embodiments, said polypeptide of interest is a polypeptide involved in a cancer. In some embodiments, said selection polypeptide is a polypeptide involved in a cancer. In some embodiments, said polypeptide of interest is HER2, P53, Androgen receptor, or NUP98-HOXA9. In some embodiments, said selection polypeptide is HER2, P53, Androgen receptor, or NUP98-HOXA9.
[0124] In some embodiments, the cancer is selected from the group consisting of: multiple myeloma or lymphoma, malignant melanoma, HPV induced cancers, prostate cancer, breast cancer, lung cancer, ovarian cancer, liver cancer, uterine serious carcinoma, and gastric cancer. In some embodiments, said polypeptide variants are fragments of a polypeptide of interest involved a neurological disease, such as Alzheimer's disease.
[0125] Property of interest
[0126] As described above, the present invention may be used to determine which amino acid residues of a polypeptide of interest or parent polypeptide are involved in (i.e. responsible for, correlated with, or linked to) a given biophysical property.
[0127] The property of interest may both be desired or non-desired. In a non-limiting example, the polypeptide of interest is a therapeutic polypeptide, e.g. an antibody that aggregates at high concentrations. In this case, the property of interest may be undesired, i.e. it is desirable to obtain polypeptide variants which lack the ability to aggregate at high concentrations, and the second population, i.e. the population lacking the property of interest is used for further development. In the opposing case, the property of interest is desired. In a non-limiting example, the polypeptide of interest is a small peptide targeting the selection polypeptide exhibiting e.g. LLPS. Small peptides engaging in LLPS and hence allowing the targeting of the selection polypeptide may be desired, in which case the first population, i.e. the population exhibiting the property of interest, is selected. The first population may then be used for further development.
[0128] In some embodiments, wherein said property of interest is selected from the group consisting of: liquid-liquid phase separation, coacervation, condensation, aggregation, solubility and stability.
[0129] In some embodiments, said one or more polypeptide variants are selected for their ability to bind to a target.
[0130] In some embodiments, said target is selected from the group consisting of: a polypeptide, a carbohydrate, a lipid, a small molecule, and a hormone.
[0131] Induction of the property of interest of the selection polypeptide may require specific conditions, such as a specific temperature, ionic strength, detergent concentration, time, pH, solvent concentration, presence of crowding agents, pressure, addition of a specific salt or addition of a small molecule. In some embodiments, said conditions are selected from the group consisting of: temperature, ionic strength, detergent concentration, time, pH, solvent concentration, presence of crowding agents, pressure, addition of a specific salt and addition of a small molecule.
[0132] In some embodiments, said display-library is a mRNA library and said property of interest is liquid-liquid phase separation. In some embodiments, said display-library is a mRNA library and said property of interest is aggregation.
[0133] In some embodiments, said display-library is a mRNA library, said polypeptide variants are fragments of the polypeptide of interest and said property of interest is liquid-liquid phase separation. In some embodiments, said display-library is a mRNA library, said polypeptide variants are mutants of the polypeptide of interest and said property of interest is liquid-liquid phase separation.
[0134] In some embodiments, said display-library is a mRNA library, said polypeptide variants are fragments and mutants of the polypeptide of interest and said property of interest is liquid-liquid phase separation.
[0135] In some embodiments, said display-library is a cDNA library and said property of interest is liquid-liquid phase separation. In some embodiments, said display-library is a cDNA library and said property of interest is aggregation.
[0136] In some embodiments, said display-library is a cDNA library, said polypeptide variants are fragments of the polypeptide of interest and said property of interest is liquid-liquid phase separation. In some embodiments, said display-library is a cDNA library, said polypeptide variants are mutants of the polypeptide of interest and said property of interest is liquid-liquid phase separation.
[0137] In some embodiments, said display-library is a cDNA library, said polypeptide variants are fragments and mutants of the polypeptide of interest and said property of interest is liquid-liquid phase separation. To minimize the effect of the bias that covalently attached polynucleotides exert on the property of interest, unique polynucleotide identifiers may be employed.
[0138] In some embodiments, the second population comprises variants exhibiting the property of interest less strongly compared to the first population.
[0139] Unique polynucleotide identifier
[0140] A “Unique polynucleotide identifier”, also referred to as a unique molecular identifier or a molecular barcode, is known in the art. A unique polynucleotide identifier can be a short polynucleotide sequence used to uniquely tag a molecule of interest in a displaylibrary. Unique polynucleotide identifiers are commonly used to link the phenotype and the genotype in a display-library, without the need for a linkage with the full-length coding sequence of the screened components.
[0141] In some embodiments, said polypeptide variants are linked to a unique polynucleotide identifier. In some embodiments, said polypeptide variants are independently covalently linked to a unique polynucleotide identifier.
[0142] In some embodiments, the unique polynucleotide identifier is a non-coding polynucleotide.
[0143] In some embodiments, the unique polynucleotide identifier is up to 15 nucleotides long, such as up to 14 nucleotides long, such as up to 13 nucleotides long, such as up to 12 nucleotides long, such as up to 11 nucleotides long, or such as up to 10 nucleotides long.
[0144] As described herein above, in some embodiments the display-library is generated from a DNA library encoding polypeptide variants and the one or more polypeptide variants are encoded by one or more polynucleotides. In some embodiments, the one or more polynucleotides independently comprise a unique polynucleotide identifier.
[0145] Separation
[0146] Once the property of interest has been induced for the polypeptide variants, these can be separated into separate populations. The property of interest may be induced by incubating the display-library and the selection polypeptide under conditions wherein the selection polypeptides exhibits the property of interest. Examples of conditions are described herein below; applicable conditions are otherwise known to the skilled person.
[0147] In some embodiments, separating the display-library is performed by a method comprising one or more of ultracentrifugation, ultrafiltration, chromatography, field flow fractionation, centrifugation, filtration, dialysis and / or microfluidic diffusion.
[0148] In some embodiments, the separation of the display library results in the formation of a plurality of phases, wherein the first population and the second population are comprised within different phases. In some embodiments, the plurality of phases is two phases, the first population is comprised within one of the two phases, and the second population is comprised within the other of the two phases.
[0149] In some embodiments, one of the phases is a liquid droplet, a dilute phase or a solid phase such as a precipitate.
[0150] In some embodiments, one of the phases is a solid phase, such as amyloid gels, amyloid fibrils, spherulites, amyloid liquid crystals, protein crystals, filaments, oligomers, or amorphous aggregates.
[0151] In some embodiments, the property of interest is liquid-liquid phase separation and the phases are liquid droplets and dilute phase. In some embodiment, the property of interest is liquid-liquid phase separation and separating the display-library is performed by ultracentrifugation.
[0152] Recovering
[0153] The method described in the present disclosure comprises the three steps: a) providing a display-library consisting of one or more polypeptide variants of a polypeptide of interest; b) contacting said display-library with a molar excess of a selection polypeptide under conditions wherein the selection polypeptide exhibits said property of interest, such as aggregation or liquid-liquid phase separation, thereby separating the display-library into at least a first population and a second population: and c) recovering said first and / or second population, thereby selecting polypeptide variants exhibiting or lacking the property of interest. The stochastic nature of the amino acid sequences of the display-library opens for the possibility that one of the first and second populations does not comprise any hits. This may e.g. occur if the property of interest is only found in a narrow spectrum of the sequence space which is not covered by the display-library. Thus in some embodiments, the first population comprises or is suspected of comprising polypeptide variants exhibiting the property of interest. In some embodiments, the second population comprises or is suspected of comprising polypeptide variants lacking the property of interest.
[0154] Recovering polypeptide variants from the separated populations may require affinity purification.
[0155] In some embodiments, the polypeptide variants comprise a tag such as an affinity tag.
[0156] In some embodiments, the affinity tag is an affinity tag selected from the group consisting of: human influenza hemagglutinin (HA)-tag, ubiquitin, histidine (His)-tag, FLAG-tag, Myc-Tag, glutathione S-transferase (GST)-tag, Maltose binding protein (MBP)-tag, Small Ubiquitin-like Modifier (SUMO)-tag, strep tag, strepll tag, and Protein A-tag.
[0157] In some embodiments, recovering said first and / or second population is performed by affinity purification.
[0158] In some embodiments, the affinity purification is performed using a resin comprising a suitable composition for the purification of the affinity tag, such as a non-limiting example wherein nickel nitrilotriacetic acid resin is used wherein the affinity tag is a histidine-tag.
[0159] In some embodiments, the display-library is further separated in one or more further populations exhibiting the property of interest less than the polypeptide of interest and more than the polypeptide of interest. Sequencing
[0160] The present disclosure describes a method for selecting one or more polypeptide variants exhibiting or lacking a property of interest. Specifically, a display-library is contacted with an excess of a selection polypeptide and a property of interest is induced. Some polypeptide variants of the display-library will exhibit the property of interest and some polypeptide variants of the display-library will lack the property of interest. The solution can be separated so the populations are separated (e.g. by centrifugation, microfluidics or other).
[0161] Once the populations have been separated and the populations have been recovered, the polypeptide variants of the populations may be identified by sequencing. This allows identification of what amino acid sequence substitutions modify the tendency to exhibit the property of interest.
[0162] There are a number of commercial methodologies for polynucleotide sequencing. These technologies are often referred to as "next generation sequencing," "massively parallel sequencing," or "bulk sequencing." These terms are used interchangeably to describe any sequencing method that is capable of acquiring more than one million polynucleic acid sequence tags in a single run. Typically these methods function by making highly parallelized measurements, i.e., parallelized screening of millions of DNA clones on glass slides. The methods for linking multiple polynucleic acid targets in single cells could be used in combination with any commercialized bulk sequencing method. These methods include reversible terminator chemistry, pyrosequencing using polony emulsion droplets, single molecule sequencing, and others.
[0163] In some embodiments, the method further comprises identifying said recovered polypeptide variants by sequencing the unique polypeptide identifiers.
[0164] In some embodiments, the method further comprises identifying said recovered polypeptide variants by sequencing the one or more polynucleotide encoding said one or more polypeptide variants.
[0165] In some embodiments, said sequencing is high throughput sequencing, such as second generation sequencing, next generation sequencing, nanopores or mass spectrometry. The amino acid sequence may also be identified by other means, such as mass spectrometry of the polypeptides.
[0166] In some embodiments, the amino acid sequence of the polypeptide variants are sequenced by nanopores or mass spectrometry.
[0167] Mapping of sequence space
[0168] As described herein above, the method allows identification of which amino acid residues of a polypeptide of interest or parent polypeptide are involved in (i.e. responsible for, correlated with, or linked to) the given biophysical property. The method may thus be used in generating a map of the role of individual amino acids for biophysical property of the polypeptide variant.
[0169] In some embodiments, the method further comprises mapping the amino acid residues of the recovered polypeptides in relation to the property of interest.
[0170] In an aspect, the present disclosure thus concerns a method of mapping one or more amino acid residues of one or more polypeptides or fragments thereof in relation to a property of interest, wherein said one or more polypeptide or fragments thereof are selected and recovered using the method as described herein above.
[0171] In some embodiments, the fragments are overlapping, such as tiles of the polypeptide of interest.
[0172] Software
[0173] Once the method have been used to identify the amino acid residues of a polypeptide of interest or parent polypeptide that are involved in the given biophysical property, the data can be incorporated in a software for designing a customized polypeptide sequence with a property of interest.
[0174] In some embodiments, said mapping is used for generating computational models and / or machine learning.
[0175] In an aspect, the present disclosure concerns a computer implemented method for designing a customized polypeptide sequence from a parent polypeptide sequence, said customized polypeptide sequence correlating with at least one property of interest, said method comprising the steps of: obtaining polypeptide variant information for the parent polypeptide sequence, providing a database of polypeptide variants data and of property of interest data, wherein said polypeptide variants are selected according to the method of any one of the preceding claims, said database comprising o a plurality of polypeptide variant data records, where each polypeptide variant data record consists of the amino acid sequence of each polypeptide variant, and o a plurality of property of interest data records, wherein each polypeptide variant data record is correlated with one data record, wherein the data record reflects whether the property of interest is exhibited or not by each polypeptide variant, such that the property of interest data record of each polypeptide variant is directly connected to the amino acid sequence of each polypeptide variant, preferably to the desired amino acid sequence, categorizing the polypeptide data of each polypeptide variant into categories of property of interest, said categorizing based on whether the polypeptide variant exhibits or lacks the property of interest, determining, based on the categorization, specific amino acid residues correlating with the exhibiting or lacking the property of interest, and designing the amino acid sequence of the customized polypeptide by selecting, for each residue of the customized polypeptide, an amino acid which is likely to, preferably the amino acid which is most likely to, contribute to the customized polypeptide exhibiting or lacking the property of interest, thereby obtaining the sequence of the customized polypeptide.
[0176] In some embodiments, the computer implemented method further comprises the step of obtaining sample condition information, wherein each condition information record is correlated with at least one property of interest information record, such that the property of the polypeptide is directly connected to the at least one condition.
[0177] In some embodiments, the condition is the molar excess of a selection polypeptide relative to the display-library. In some embodiments, the condition is selected from the group consisting of: temperature, ionic strength, detergent concentration, time, pH, solvent concentration, presence of crowding agents, pressure, addition of a specific salt and addition of a small molecule.
[0178] In some embodiments, the polypeptide variant, selection polypeptide and property of interest are as defined herein above.
[0179] In an aspect, the present disclosure concerns a computer implemented method for designing a customised polypeptide comprising at least one desired amino acid sequence, said method comprising the steps of: i. designing the customized polypeptide according to the method described herein, and ii. synthesising the customized polypeptide.
[0180] A kit
[0181] The method may further be distributed in the form of a kit comprising the reagents needed for the performing the steps of the method.
[0182] In another aspect, the present disclosure concerns a kit for selecting one or more polypeptide variants exhibiting or lacking a property of interest, the kit comprising: i) a display-library consisting of one or more polypeptide variants of a polypeptide of interest; ii) a selection polypeptide in an amount allowing for a molar excess of said selection polypeptide to be added to the one or more polypeptide variants; and iii) instructions for use.
[0183] In some embodiments, the polypeptide variants, the polypeptide of interest and selection polypeptide as described herein above.
[0184] In some embodiments, the one or more polypeptide variants are linked to a unique polynucleotide identifier. Polypeptide, Polynucleotide, vector and host cell
[0185] As described above, the method may be used for generating a map of residues comprising information about the effect of different amino acids at any given position on the property of interest can be mapped. This map may be used for designing a customized polypeptide sequence with a property of interest. This customized polypeptide may be encoded in a polypeptide, which may be comprised in a vector that may be comprised in a host cell.
[0186] In another aspect, the present disclosure concerns a polypeptide identified or obtained by the method as described herein above.
[0187] In some embodiments, said polypeptide is an antigen binding molecule, such as an antibody, a nanobody, a Fab, a minibody, a diabody, a tribody, a scFv, a DARPin, a de novo binder, a peptide aptamer, or an affibody.
[0188] In some embodiments, said polypeptide is a cyclic peptide
[0189] In some embodiments, said polypeptide is a therapeutic polypeptide.
[0190] In some embodiments, said polypeptide is an enzyme.
[0191] In an aspect, the present disclosure concerns a polynucleotide encoding the polypeptide as described herein above.
[0192] In an aspect, the present disclosure concerns a vector comprising a polynucleotide as described herein above.
[0193] In an aspect, the present disclosure concerns a host cell comprising a vector or a polynucleotide as described herein above.
[0194] Examples
[0195] Example 1
[0196] Materials
[0197] Puromycin linker
[0198] ( / 5Phos / CTCCCGCCCCCCG / iFluorT / CC / iSp18 / / iSp18 / / iSp18 / / iSp18 / / iSp18 / CC / 3Puro / ) (SEQ ID NO: 1) called FPur-linker was ordered from IDT (Germany). DNA primers were ordered from TAG Copenhagen (Denmark). Double distilled RNAse free water was used for all RNA work (Invitrogen™ UltraPure™ DNase / RNase-Free Distilled Water).
[0199] Methods mRNA-display
[0200] Several options exist for designing libraries, including, but not limited to, amplification from a plasmid, amplification from an oligo pool, or construction from oligo nucleotides synthesised with error referred to as doped oligo nucleotides. For the experiments described herein below, an oligo pool was used for the library
[0201] Fragment library from oligo pool
[0202] To investigate which parts of the Ddx4N1 CtoA (parent sequence is Q9NQI0-1 and exact sequence is SEQ ID NO: 4) sequence participated most strongly in the condensation, a tiling library consisting of many overlapping fragments of the protein domain was designed. The library was designed with fragments starting from 14 amino acids in length and increasing by two all the way to 50 amino acids (I = {14,16,18 ... 50}) (Figure 4). For each length, overlapping fragments were positioned to start from every other amino acid, such that for fragments of 14 amino acids length, each amino acid in Ddx4N1, sufficiently far from the N and C-terminus, was present in 7 different fragments. The library was synthesised by Twist Bioscience (USA) with 5’ and 3’ regions designed for PCR amplification. This PCR step is recommended by the supplier, but also allowed the attachment of the remaining 5’ and 3’ regions needed for mRNA-display. The library was solubilised to 2.5 ng / pL in TE buffer (10 mM Tris-HCI, 1 mM EDTA, pH=8) and 1 pL was added a total of 100 pL PCR solution (0.5 pM Fivep_flank (SEQ ID NO: 6), 0.5 pM Threep_flank_HA (SEQ ID NO: 7), 1X Phusion HF buffer, 0.2 mM dNTPs, and 0.02 U Phusion (NEB M0530, USA)). 14 cycles were run with initial denaturation for 30 s at 98°C, final extension for 60 s at 72°C, and cycling through: 10 s at 98°C, 30 s at 56°C, and 40 s at 72°C.
[0203] Transcription
[0204] Before transcription, the template DNA was cleaned using phenol-chloroform extraction. Briefly, 3 M NaCI was added to the solution to a final concentration of 300 mM. For a 100 pL solution, 100 pL (same volume) phenol-cholorform-isoamyl alcohol (PCI) at 25:24:1 ratio pH=6.5 was then added before vortexing and centrifuging at 10,000 g for 4 minutes. The top aqueous phase was then transferred to a new tube and cleaned by adding 100 pL chloroform-isoamyl alcohol (Cl) at 24:1 ratio. Again, the solution was vortexed and centrifuged to separate the two phases. Finally, the aqueous phase was then transferred to a new tube and precipitated by adding ethanol to a final 70% (v / v). The pellet was washed with fresh 25 pL of 70% ethanol, dried for a couple of minutes, and then re-solubilised in 10 pL RNAse free water, for a theoretical maximum of 5 pM (assuming complete PCR and no loss in PCI / CI extraction). Transcription was performed using the Ribomax kit (Promega, USA) in a 50 pL reaction: 10 pL 5X buffer, 15 pL 100 mM rNTPs, 2.5 pL DNA template, 0.5 pL 1 mM DDT, 0.3 pL RNAsin Plus (Promega, USA), 5 pL Enzyme mix, and 16.7 pL RNAse free water. The reaction was incubated at 37°C over night (~16 hours). To remove template DNA, 5.7 pL DNAse buffer (0.4M Tris pH 8.0, 0.1M MgSO4, 0.01 M CaCI2) and 1 pL DNAse 1 (Thermo EN0525, USA) were added and the solution was incubated for a maximum of 30 minutes at 37°C. To stop the reaction, EDTA pH=8 was added to a final concentration of 50 mM and NaCI was added to a final concentration of 300 mM. The RNA was then purified using PCI / CI as described above using a PCI solution with pH=4.5. Instead of precipitation, the aqueous phase after Cl treatment was transferred to a NAP-5 column (Cytiva, USA) equilibrated with 10 mL RNAse free water. If adding a 70 pL sample, 430 pL of RNAse free water was then added. Finally, 500 pL RNAse free water was added to elute the mRNA. 0.3 pL RNAsin was added to the tube before storing at -80°C.
[0205] Puromycin ligation
[0206] Ligation to the puromycin containing linker was done in 50 pL reaction containing 7.5 pM FPur-Linker and 5 pM mRNA. A 25 pL solution with 2x concentration of the two (15 and 10 pM, respectively) was heated to 95°C and slowly cooled to room temperature. Then, 5 pL 10x T4 RNA ligase buffer and 5 pL T4 ssRNA ligase 1 (NEB, USA) was added and the reaction was incubated for 1 hour at 37°C. 3 pL of the reaction was loaded unto a PAGE gel containing TBE (Tris-Borate-EDTA), 8 M Urea, 20% (v / v) formamide, and 8% acrylamide 19:1 to evaluate ligation. EDTA was added to the remainder of the reaction to a final concentration of 50 mM and NaCI to 300 mM, and the solution was purified using PCI / CI as described above using pH 6.5 PCI solution. The aqueous phase was the precipitated in 70% ethanol, washed with 70% ethanol, and solubilised in 7.5 pL RNAse free water. This solution contained also un-ligated mRNA, however this was removed in the HA-tag purification step. Translation
[0207] Translation of the ligated mRNA was performed using the PureExpress ARF1 ,2,3 kit (NEB, USA). A 25 pL reaction was performed using 10 pL solution A, 7.5 pL solution B, 5 pL template, and 2.5 pL RNAse free water. The solution was incubated for 1 hour at 37°C for translation to proceed. 3.25 pL 100 mM Mg(OAc)2 pH=8 were added to the solution to improve efficiency of the puromycin reaction during additional 30 min incubation at 37°C. Finally, 8.25 pL 100 mM EDTA pH=8 were added to release the ribosomes during additional 30 minutes incubation at 37°C. The translation mixture was then mixed 1:1 with 2X blocking buffer (100 mM HEPES-KOH pH 7.6, 400 mM NaCI, 0.1% (v / v) Tween-20, 0.1 g / L yeast RNA (Thermo, USA), 0.2 g / L acetylated BSA (Thermo, USA), and 5 pM GS3an_R36 (SEQ ID NO: 8)) and cooled on ice. 40 pL Pierce Anti-HA beads (Thermo, USA) was washed 3 times in ice cold wash buffer (50 mM HEPES-KOH pH 7.6, 200 mM NaCI, 0.05%(v / v) Tween-20, 1 pM GS3an_R36) using magnetic separation. The blocked translation mixture was used to re-suspend the pelleted beads and incubated 1 hour at 4°C with rotation to allow binding. The supernatant was then removed after magnetic separation and two rounds of washing was swiftly performed using ice cold washing buffer. A final wash was performed with elution buffer (20 mM NaPi, 500 mM NaCI, 0.05%(v / v) Tween-20, and trace amounts of RNAsin (Promega, USA)), before eluting the bound molecules by adding 20 pL elution buffer and incubating the beads at 80°C for 10 minutes. The solution was aliquoted and flash frozen in liquid nitrogen. Quantification was performed using qPCR or fluorescence from the fluorescein on the puromycin linker yielding a typical concentration of ~20 nM.
[0208] Partitioning experiments An aliquot of the library to be screened for partitioning was diluted 10-fold with low salt buffer (20 mM NaPi and 0.05%(v / v) Tween-20) such that this solution could induce coacervation of Ddx4N1 CtoA. In parallel, the mRNA used to create the library was diluted to the same concentration, usually ~2 nM, to be used as a control. Both of these solutions were cooled on ice and then added to two different PCR tubes, one containing 1.5 pL 600 pM Ddx4N1 in elution buffer and one with 1.5 pL elution buffer without protein, resulting in four tubes in total. All tubes were incubated for 1 minute at room temperature and then centrifuged for 10 minutes at 10,000 xg in a preequilibrated centrifuge at 20°C. After centrifugation, 1 pL of the supernatant from each tube was immediately added to the reverse transcription solution (1x SS-IV buffer, 0.25 mM dNTPs, 0.75 pM GS3an_R36, 1.5 units of SS-IV reverse transcriptase (Thermo, USA), and trace amount of RNAsin (Promega, USA)). The reaction was left to proceed at 50°C for 1 hour to ensure completion. The resulting solution was stored at -20°C.
[0209] Sequencing
[0210] Library preparation for sequencing
[0211] Attachment of Illumina adapter sequences was done through two consecutive PCR reactions. But first, the concentration of template in each of the reverse transcription reactions was estimated using qPCR, such that overcycling could be avoided. The first PCR reaction was performed in 50 pL using 0.15 pM of Rd1T7g10M_F71 (SEQ ID NO: 9) and an13Rd2_R49 (SEQ ID NO: 10), 1X HF buffer, 0.2 mM dNTPs, 0.5 pL reverse transcription solution, and 0.5 pL Phusion polymerase (NEB, USA). A varying amount of cycles (usually 10-20) were run based on the number needed to get 0.05 pM double stranded product. The PCR program included an initial denaturation for 30 s at 98°C, final extension for 60 s at 72°C, and cycling through: 10 s at 98°C, 30 s at 52°C, and 40 s at 72°C. For the second reaction, 0.15 pM of Rd2N(701)P7_R52 (SEQ ID NO: 12) and P5S(502)Rd1_F57 (SEQ ID NO: 11) were used in a similar PCR reaction using 2.5 pL of the former PCR reaction as template in a 50 pL reaction. Each sample was assigned a unique combination of multiplexing barcodes in the two primers in the regions specified by parenthesis. The samples were cleaned using QIAquick PCR clean up kit (Qiagen, Netherlands) before being sequenced using MiSeq v3 chip for 2x300 bps reads (Illumina, USA).
[0212] Processing of data
[0213] NGMerge (Gaspar, 2018) was used to merge paired end reads using default configuration. Merged reads were mapped to variants expected in the library allowing a single base pair discrepancy originating from sequencing errors, using a custom python script.
[0214] Analysis
[0215] Calculation of dilute phase preference:
[0216] The calculations were based on relative abundance in the library, and the fraction Fvof all total reads by one variant v in sample X was calculated as:
[0217] Where, and N is the total amount of variants in the library.
[0218] The dilute phase preference Dv,x, similar to the enrichment ratio in deep mutational scanning experiments, was calculated as the ratio between the abundances before and after screening. The ratio was scaled using the base-2 logarithm:
[0219] To isolate the effect originating from the peptide, the dilute phase preference calculated for the library containing pure mRNA, Dv(mRNA), was subtracted from the value obtained from the mRNA-display library, Dv(mRNA display)’.
[0220] Dv(peptide) = DvmRN Auli splay) — Dv(mRNA)
[0221] Calculation of average effect per position:
[0222] To obtain a measure of the interaction potential of each amino acid in the sequence, the dilute phase preference of all frag ments / ti les covering a given amino acid position in the sequence were averaged. By doing this procedure through the entire protein, a map of the contribution for all regions of the sequence could be obtained (Figure 3 and 5).
[0223] Calculation of partition free energy:
[0224] The inventors first calculated the concentration of each library variant, v, in both the control and in the dilute phase after phase separation, [v]_(X), by calculating the fraction of the total amount of reads F_(v,X) attributed to each variant, v. The different samples or sequencing runs are indicated by X:
[0225] The concentration of each member of the library was then calculated using the total concentration of the library c_(library,X) measured by qPCR directly during the experiment:
[0226] The fractional concentration in the dilute phase of all library members was then calculated as follows:
[0227] The partition free energy of all variants in the sample was then calculated as follows, by using the dense phase volume fraction:
[0228] Under the assumption that the partition free energy of all parts of the mRNA display construct (mRNA, linker, and peptide) are additive, the inventors could calculate the apparent effect stemming from the peptide:
[0229] This means that the apparent partition free energies of all peptides have a constant contribution from the linker. The best estimate for this effect is 2.5 kJ / mol, meaning that the linker adds an unfavourable contribution to the partition free energy. This value has been subtracted from the apparent energies for all data shown in the figures. Conclusions
[0230] The experimental design of this experiment is displayed in Figure 2 and 4. The experiment showed that fragments / tiles of the protein could be used to map regions of increased interaction potential. Several aspects of the approach were verified, including the robustness of subtracting the effect of the mRNA from the mRNA-display samples to obtain a measure of the partitioning of the displayed peptide. This was shown by the good correlation between different tile lengths when calculating the average effect for a given position (Figure 3). The effect was calculated as both the dilute phase preference (Figure 3a) and the average partition energy (Figure 3b) of each amino acid position.
[0231] Example 2
[0232] Materials
[0233] Puromycin linker ( / 5Phos / CTCCCGCCCCCCG / iFluorT / CC / iSp18 / / iSp18 / / iSp18 / / iSp18 / / iSp18 / CC / 3Puro / ) (SEQ ID NO: 1) called FPur-linker was ordered from IDT (Germany). DNA primers were ordered from TAG Copenhagen (Denmark). Double distilled RNAse free water was used for all RNA work (Invitrogen™ UltraPure™ DNase / RNase-Free Distilled Water).
[0234] Methods mRNA-display
[0235] Several options exist for designing libraries, including, but not limited to, amplification from a plasmid, amplification from an oligo pool, or construction from oligo nucleotides synthesised with error referred to as doped oligo nucleotides. For the experiments described herein below, a set of doped oligo nucleotides were used for the library, which were ordered from IDT (Germany).
[0236] Mutation library from doped oligo nucleotides
[0237] In each vial for synthesis, a small amount of each of the other three nucleotides are added. This results in a highly diverse synthesised product, with almost all molecules being unique, given a sufficiently high mutagenesis rate. Two such oligos were designed, Pep6_lib_fw (SEQ ID NO: 2) and Pep6_lib_rev (SEQ ID NO: 3), to generate a randomly mutated library of a central region in Ddx4N1 positions (SEQ ID NO: 5) [100:140). The two oligos were designed with constant regions to anneal to Fivep_flank (SEQ ID NO: 6) and Threep_flank_HA (SEQ ID NO: 7) respectively for PCR amplification. Additionally, the two oligos overlapped with 15 base pairs at their 3’-ends allowing them to anneal and be extended to a double stranded product using the Klenow fragment polymerase (NEB M0210, USA). The extension reaction was performed in 100 pL with 1 pM of each oligo in NEBuffer 2 (50 mM NaCI, 10 mM Tris-HCI, 10 mM MgCh, 1 mM DTT, pH 7.9). The oligos were annealed in the PCR solution, but without dNTPs and polymerase by heating the solution to 94°C for 2 minutes and then ramping down the temperature at 0.2 °C / s down to 50°C. The solution was kept for 2 minutes, and then ramped down to 37°C at the same speed.
[0238] Then 2.5 pL 10 mM dNTPs and 1 pL Klenow fragment was added and the reaction was incubated at 37°C for 30 minutes. To limit the amount of variants in the library, the 10 nM double stranded product from the Klenow extension was diluted 10-fold (5 pL in 45 pL 0.05% triton X-100) 6 times to a theoretical maximum amount of molecules in the solution in 1 pL of -600,000.
[0239] 1 pL was added a total of 100 pL PCR solution (0.5 pM Fivep_flank (SEQ ID NO: 6), 0.5 pM Threep_flank_HA (SEQ ID NO: 7), 1X Phusion GC buffer, 0.2 mM dNTPs, and 0.02 U Phusion (NEB M0530, USA)). 33 cycles were run with initial denaturation for 30 s at 98°C, final extension for 60 s at 72°C, and cycling through: 10 s at 98°C, 30 s at 56°C, and 40 s at 72°C.
[0240] Transcription
[0241] Before transcription, the template DNA was cleaned using phenol-chloroform extraction. Briefly, 3 M NaCI was added to the solution to a final concentration of 300 mM. For a 100 pL solution, 100 pL (same volume) phenol-cholorform-isoamyl alcohol (PCI) at 25:24:1 ratio pH=6.5 was then added before vortexing and centrifuging at 10,000 g for 4 minutes. The top aqueous phase was then transferred to a new tube and cleaned by adding 100 pL chloroform-isoamyl alcohol (Cl) at 24:1 ratio. Again, the solution was vortexed and centrifuged to separate the two phases. Finally, the aqueous phase was then transferred to a new tube and precipitated by adding ethanol to a final 70% (v / v). The pellet was washed with fresh 25 pL of 70% ethanol, dried for a couple of minutes, and then re-solubilised in 10 pL RNAse free water, for a theoretical maximum of 5 pM (assuming complete PCR and no loss in PCI / CI extraction). Transcription was performed using the Ribomax kit (Promega, USA) in a 50 pL reaction: 10 pL 5X buffer, 15 pL 100 mM rNTPs, 2.5 pL DNA template, 0.5 pL 1 mM DDT, 0.3 pL RNAsin Plus (Promega, USA), 5 pL Enzyme mix, and 16.7 pL RNAse free water. The reaction was incubated at 37°C over night (~16 hours). To remove template DNA, 5.7 pL DNAse buffer (0.4M Tris pH 8.0, 0.1M MgSO4, 0.01 M CaCI2) and 1 pL DNAse 1 (Thermo EN0525, USA) were added, and the solution was incubated for a maximum of 30 minutes at 37°C. To stop the reaction, EDTA pH=8 was added to a final concentration of 50 mM and NaCI was added to a final concentration of 300 mM. The RNA was then purified using PCI / CI as described above using a PCI solution with pH=4.5. Instead of precipitation, the aqueous phase after Cl treatment was transferred to a NAP-5 column (Cytiva, USA) equilibrated with 10 mL RNAse free water. If adding a 70 pL sample, 430 pL of RNAse free water was then added. Finally, 500 pL RNAse free water was added to elute the mRNA. 0.3 pL RNAsin was added to the tube before storing at -80°C.
[0242] Puromycin ligation
[0243] Ligation to the puromycin containing linker was done in 50 pL reaction containing 7.5 pM FPur-Linker and 5 pM mRNA. A 25 pL solution with 2x concentration of the two (15 and 10 pM, respectively) was heated to 95°C and slowly cooled to room temperature. Then, 5 pL 10x T4 RNA ligase buffer and 5 pL T4 ssRNA ligase 1 (NEB, USA) was added and the reaction was incubated for 1 hour at 37°C. 3 pL of the reaction was loaded unto a PAGE gel containing TBE (Tris-Borate-EDTA), 8 M Urea, 20% (v / v) formamide, and 8% acrylamide 19:1 to evaluate ligation. EDTA was added to the remainder of the reaction to a final concentration of 50 mM and NaCI to 300 mM, and the solution was purified using PCI / CI as described above using pH 6.5 PCI solution. The aqueous phase was the precipitated in 70% ethanol, washed with 70% ethanol, and solubilised in 7.5 pL RNAse free water. This solution contained also un-ligated mRNA, however this was removed in the HA-tag purification step.
[0244] Translation
[0245] Translation of the ligated mRNA was performed using the PureExpress ARF1 ,2,3 kit (NEB, USA). A 25 pL reaction was performed using 10 pL solution A, 7.5 pL solution B, 5 pL template, and 2.5 pL RNAse free water. The solution was incubated for 1 hour at 37°C for translation to proceed. 3.25 pL 100 mM Mg(OAc)2pH=8 were added to the solution to improve efficiency of the puromycin reaction during additional 30 min incubation at 37°C. Finally, 8.25 pL 100 mM EDTA pH=8 were added to release the ribosomes during additional 30 minutes incubation at 37°C. The translation mixture was then mixed 1:1 with 2X blocking buffer (100 mM HEPES-KOH pH 7.6, 400 mM NaCI, 0.1% (v / v) Tween-20, 0.1 g / L yeast RNA (Thermo, USA), 0.2 g / L acetylated BSA (Thermo, USA), and 5 pM GS3an_R36 (SEQ ID NO: 8)) and cooled on ice. 40 pL Pierce Anti-HA beads (Thermo, USA) was washed 3 times in ice cold wash buffer (50 mM HEPES-KOH pH 7.6, 200 mM NaCI, 0.05%(v / v) Tween-20, 1 pM GS3an_R36) using magnetic separation. The blocked translation mixture was used to re-suspend the pelleted beads and incubated 1 hour at 4°C with rotation to allow binding. The supernatant was then removed after magnetic separation and two rounds of washing was swiftly performed using ice cold washing buffer. A final wash was performed with elution buffer (20 mM NaPi, 500 mM NaCI, 0.05%(v / v) Tween-20, and trace amounts of RNAsin (Promega, USA)), before eluting the bound molecules by adding 20 pL elution buffer and incubating the beads at 80°C for 10 minutes. The solution was aliquoted and flash frozen in liquid nitrogen. Quantification was performed using qPCR or fluorescence from the fluorescein on the puromycin linker yielding a typical concentration of ~20 nM.
[0246] Partitioning experiments An aliquot of the library to be screened for partitioning was diluted 10-fold with low salt buffer (20 mM NaPi and 0.05%(v / v) Tween-20) such that this solution could induce coacervation of Ddx4N1 CtoA. In parallel, the mRNA used to create the library was diluted to the same concentration, usually ~2 nM, to be used as a control. Both of these solutions were cooled on ice and then added to two different PCR tubes, one containing 1.5 pL 600 pM Ddx4N1 in elution buffer and one with 1.5 pL elution buffer without protein, resulting in four tubes in total. All tubes were incubated for 1 minute at room temperature and then centrifuged for 10 minutes at 10,000 xg in a preequilibrated centrifuge at 20°C. After centrifugation, 1 pL of the supernatant from each tube was immediately added to the reverse transcription solution (1x SS-IV buffer, 0.25 mM dNTPs, 0.75 pM GS3an_R36, 1.5 units of SS-IV reverse transcriptase (Thermo, USA), and trace amount of RNAsin (Promega, USA)). The reaction was left to proceed at 50°C for 1 hour to ensure completion. The resulting solution was stored at -20°C. Sequencing
[0247] Library preparation for sequencing
[0248] Attachment of Illumina adapter sequences was done through two consecutive PCR reactions. But first, the concentration of template in each of the reverse transcription reactions was estimated using qPCR, such that overcycling could be avoided. The first PCR reaction was performed in 50 pL using 0.15 pM of Rd1T7g10M_F71 (SEQ ID NO: 9) and an13Rd2_R49 (SEQ ID NO: 10), 1X HF buffer, 0.2 mM dNTPs, 0.5 pL reverse transcription solution, and 0.5 pL Phusion polymerase (NEB, USA). A varying amount of cycles (usually 10-20) were run based on the number needed to get 0.05 pM double stranded product. The PCR program included an initial denaturation for 30 s at 98°C, final extension for 60 s at 72°C, and cycling through: 10 s at 98°C, 30 s at 52°C, and 40 s at 72°C. Phasing version of the two primers were used, such that 0-8 additional random nucleotides were inserted in the primers to off-set the read which created increased diversity during sequencing. For the second reaction, 0.15 pM of Rd2N(701)P7_R52 (SEQ ID NO: 12) and P5S(502)Rd1_F57 (SEQ ID NO: 11) were used in a similar PCR reaction using 2.5 pL of the former PCR reaction as template in a 50 pL reaction. Each sample was assigned a unique combination of multiplexing barcodes in the two primers as indicated by the parts of the names and sequences in the parenthesis. The samples were cleaned using Mag-Bind® TotalPure NGS (Omega Bio-tek, USA) and sequenced using a SP Flow-cell on NovaSeq (Illumina, USA).
[0249] Processing of data
[0250] NGMerge (Gaspar, 2018) was used to merge paired end reads using default configuration. The reads were cut to only contain the mutated area and unique sequences were then counted using custom code written in C++. As the variation in the library is in nature similar to sequencing errors, the only way to remove sequencing noise is to require a minimum number of reads. This cut-off was established by requiring that the sequence was present in all four samples and the cut-off should maximise the difference between the standard error of synonymous and non- synonymous variants. A cut-off of minimum 50 counts per sequence was chosen, and additionally it was required that between the sample with and without LLPS, a total of 500 reads must have been observed. Analysis and model fitting
[0251] Calculation of dilute phase preference:
[0252] The calculations were based on relative abundance in the library, and the fraction Fvof all total reads by one variant v in sample X was calculated as:
[0253] Where, and / V is the total amount of variants in the library.
[0254] The dilute phase preference Dv,x, similar to the enrichment ratio in deep mutational scanning experiments, was calculated as the ratio between the abundances before and after screening. The ratio was scaled using the base-2 logarithm:
[0255] To isolate the effect originating from the peptide, the dilute phase preference calculated for the library containing pure mRNA, Dv(mRNA), was subtracted from the value obtained from the mRNA-display library, Dv(mRNA display):
[0256] Dv(peptide) — Dv(mRNA_di splay
[0257] Calculation of partition free energy:
[0258] The inventors first calculated the concentration of each library variant, v, in both the control and in the dilute phase after phase separation, [v]_(X), by calculating the fraction of the total amount of reads F_(v,X) attributed to each variant, v. The different samples or sequencing runs are indicated by X: readsVx
[0259] Y readsi x
[0260] The concentration of each member of the library was then calculated using the total concentration of the library c_(library,X) measured by qPCR directly during the experiment:
[0261] The fractional concentration in the dilute phase of all library members was then calculated as follows:
[0262] The partition free energy of all variants in the sample was then calculated as follows, by using the dense phase volume fraction:
[0263] Under the assumption that the partition free energy of all parts of the mRNA display construct (mRNA, linker, and peptide) are additive, the inventors could calculate the apparent effect stemming from the peptide:
[0264] This means that the apparent partition free energies of all peptides have a constant contribution from the linker. The best estimate for this effect is 2.5 kJ / mol, meaning that the linker adds an unfavourable contribution to the partition free energy. This value has been subtracted from the apparent energies for all data shown in the figures.
[0265] Linear model of amino acid effects
[0266] Before any model fitting, the data was split in 9:1 for adjusting the hyper-parameter (training set) and testing, respectively. For the mutation library, the dilute phase preference of the peptide was modelled with Tikhonov regularization using the identity matrix, which is also known as Ridge or L2 regularization. This was done in Python using the package scikit-learn and the regularisation parameter, alpha, was chosen as 31.6 based on 5-fold cross validation. Pearson correlation for prediction of the training set and test set was 0.47 and 0.42, respectively. For analysis using the partition free energy instead, lambda was chosen as 3162 and weighted pearson correlation for the training and test sets was 0.48 and 0.46, respectively.
[0267] Conclusions
[0268] The experimental design of this experiment is displayed in Figure 4. The experiment showed that effects from individual amino acids could be obtained by inferring these from a linear model. This approach therefore reports on the average effect on partitioning into the WT Ddx4N1 condensates of an amino acid substitution at each position in the sequence indicated in the shaded area of (Figure 5a). This approach further reports on the partition free energy along the WT Ddx4N1 condensates of an amino acid substitution at each position in the sequence indicated in the shaded area of (Figure 5b). In particular, aromatic amino acids were very important for the peptide to partition into the condensates, although not all regions of this particular sequence showed strong effects from the aromatic amino acids (Figure 6a). These results may similarly be presented as partition free energy (Figure 6b). This experiment therefore contains additional information which can be extracted using machine learning models to predict and design peptides with high propensity to form liquid coacervates.
[0269] Example 3
[0270] Materials and methods
[0271] Six peptides of 14 amino acids lengths from the tiling library were synthesised by solid state peptide synthesis with an additional N-terminal cysteine conjugated to an Oregon Green-488 malemide fluorophore. The partitioning of these purified peptides was measured using the in-house developed method Capflex (Stender et al, 2021) to obtain orthogonal estimates of their partition free energy.
[0272] Conclusion
[0273] As shown in Figure 7, purified peptides behave similarly to what is extracted from the mRNA display experiment. Example 4
[0274] Materials and methods
[0275] Synthesis of six peptides from the tiling library with an additional N-terminal cysteine conjugated to an Oregon Green-488 malemide fluorophore was ordered from Schafer- N (Copenhagen, Denmark). Peptides were dissolved in water and diluted to 100 nM in 20 mM sodium phosphate (NaPi) buffer pH = 6.5 containing 0.05% Tween-20. 16 pL of this solution was used to induce phase separation of 4 pL 300 pM Ddx4N1 CtoA dissolved in a buffer containing 500 mM NaCI and 20 mM NaPi pH 6.5 by reducing the NaCI concentration to 100 mM. Immediately after mixing, Capflex measurements were performed on the sample by flowing it through a 75 pm diameter and 1 m length Fida 1 capillary using 400 mbar positive pressure for 300s at 20°C. The dilute phase signal was recorded as the 10% lowest quantile of the signal after reaching the detector to filter out signal stemming from the condensed phase droplets. A control sample lacking protein was recorded afterwards, where 16 pL of the same solution containing 100 nM peptide was mixed with 4 pL 500 mM NaCI buffer without protein. Partition free energies where calculated from the fraction of peptide concentration remaining in the dilute phase, by assuming a dense phase volume fraction of 0.2% (based on measured dilute phase protein concentration of 30 pM and a dense phase concentration of ~15 mM as reported by Brady et al, 2017) using the equation:
[0276] The amino acid sequence of the peptides (without N-terminal cysteine):
[0277] Peptide 05 (SEQ ID NO: 13): INPHMSSYVPIFEK
[0278] Peptide 34 (SEQ ID NO: 14): RDAGEANKRDNTST
[0279] Peptide 54 (SEQ ID NO: 15): GFWRESSNDAEDNP
[0280] Peptide 62 (SEQ ID NO: 16): NRGFSKRGGYRDGN
[0281] Peptide 78 (SEQ ID NO: 17): ARGGFGLGSPNNDL
[0282] Peptide 98 (SEQ ID NO: 18): DTSQSESGSGSERG
[0283] Microfluidic separation
[0284] The microfluidic chip used for these experiments had a “co-flow” path length of 150 mm, width of 150 pm and hight of 50 pm. The resistances of the inlet paths were designed so that in the “co-flow” path, the ratio of buffer-sample-buffer was 2-1-2, when the positive driving pressure for both inlets was kept constant at 200 mbar. Buffer condition of the solution applied to the outer inlet was 20 mM phosphate buffer, pH 6.5. Buffer condition for the solution applied to the inner inlet was 20 mM phosphate buffer, 500 NaCI, pH 6.5, with or without 370 pM DdxN1 CtoA, containing different indicators: ssDNA-Alexa488 (SEQ ID NO: 19) (TTTTTCCTAGAGAGTAGAGCCTGCTTCGTGG) and two different peptide fragments of Ddx4N1 CtoA labelled with OG488 (Peptide 34 and Peptide 62) at 100 nM concentration. Fluorescent microscopy images were recorded using a Carl Zeiss™ Axio Vert.AI microscope equipped with a 10x objective. The microscope featured a LED light source (CoolLED, pE-300), with images captured at 100% light intensity and a 500 ms exposure time in fluorescence channels with a filtering cube of 475 / 525 nm. Imaged (FIJI) software was used for image processing and fluorescence signal quantification.
[0285] Conclusion
[0286] Figure 8 shows proof-of-concept data on how microfluidic diffusional separation can be used as a more general platform for selection and screening of libraries.
[0287] Sequence overview
[0288] Sequences and primers
[0289] SEQ ID NO: 1: Puromycin linker
[0290] CTCCCGCCCCCCGTCC
[0291] SEQ ID NO: 2: Pep6_lib_fw.
[0292] TTAACTTTAAGAAGGAGATATACATATGcgtttcgaagacggcgatagcagcggtttttggcgtgag agcagcaacgatgcggaagacaa
[0293] SEQ ID NO: 3: Pep6Jib_rev.
[0294] GGATAGCTACCGCTACCgctgttgttgccatcacgataaccgccacgtttgctgaaaccacggttacgggtcg ggttgtcttccgcatcg
[0295] SEQ ID NO: 4: Ddx4N1 CtoA
[0296] GMGDEDWEAEINPHMSSYVPIFEKDRYSGENGDNFNRTPASSSEMDDGPSRRDHF MKSGFASGRNFGNRDAGEANKRDNTSTMGGFGVGKSFGNRGFSNSRFEDGDSSG FWRESSNDAEDNPTRNRGFSKRGGYRDGNNSEASGPYRRGGRGSFRGARGGFGL GSPNNDLDPDEAMQRTGGLFGSRRPVLSGTGNGDTSQSRSGSGSERGGYKGLNEE
[0297] VITGSGKNSWKSEAEGGES
[0298] SEQ ID NO: 5: Ddx4N1 CtoA [100:140)
[0299] MRFEDGDSSGFWRESSNDAEDNPTRNRGFSKRGGYRDGNNS
[0300] SEQ ID NO: 6: Fivep_flank
[0301] TAATACGACTCACTATAGGGTTAACTTTAAGAAGGAGATATACATATG
[0302] SEQ ID NO: 7: Threep_flank_HA
[0303] TTTCCGCCCCCCGTCCTAGCTGCCGCTGCCGCTGCCTGCATAATCCGGAACATCA
[0304] TACGGATAGCTACCGCTACC
[0305] SEQ ID NO: 8: GS3an_R36
[0306] TTTCCGCCCCCCGTCCTAGCTGCCGCTGCCGCTGCC
[0307] SEQ ID NO: 9: Rd1T7g10M_F71
[0308] ACACTCTTTCCCTACACGACGCTCTTCCGATCT(N)_(O-
[0309] 8)TAATACGACTCACTATAGGGTTAACTTTAAGAAGGAGA
[0310] SEQ ID NO: 10: an13Rd2_R49
[0311] GACTGGAGTTCAGACGTGTGCTCTTCCGATCT(N)_(0-8)TTTCCGCCCCCCGTCCT
[0312] SEQ ID NO: 11 : P5S(502)Rd1_F57
[0313] AATGATACGGCGACCACCGAGATCTACAC(CTCTCTAT)ACACTCTTTCCCTACACG AC
[0314] SEQ ID NO: 12: Rd2N(701)P7_R52
[0315] CAAGCAGAAGACGGCATACGAGAT(TCGCCTTA)GTGACTGGAGTTCAGACGTG
[0316] SEQ ID NO: 13: Peptide 05.
[0317] INPHMSSYVPIFEK
[0318] SEQ ID NO: 14: Peptide 34.
[0319] RDAGEANKRDNTST SEQ ID NO: 15: Peptide 54.
[0320] GFWRESSNDAEDNP
[0321] SEQ ID NO: 16: Peptide 62.
[0322] NRGFSKRGGYRDGN
[0323] SEQ ID NO: 17: Peptide 78.
[0324] ARGGFGLGSPNNDL
[0325] SEQ ID NO: 18: Peptide 98.
[0326] DTSQSESGSGSERG
[0327] SEQ ID NO: 19: ssDNA-Alexa488.
[0328] TTTTTCCTAGAGAGTAGAGCCTGCTTCGTGG
[0329] References
[0330] Gaspar, John M. (Dec. 2018). “NGmerge: merging paired-end reads via novel empirically-derived models of sequencing errors”, en. In: BMC Bioinformatics 19.1, p. 536. issn: 1471-2105. doi: 10.1186 / s12859-018-2579-2.
[0331] Josephson K, Ricardo A, Szostak JW. mRNA display: from basic principles to macrocycle drug discovery. Drug Discov Today. 2014 Apr; 19(4):388-99.
[0332] EP2492344B1 RAPID DISPLAY METHOD IN TRANSLATIONAL SYNTHESIS OF PEPTIDE. PEPTIDREAM INC [JP],
[0333] Stender, E.G.P., Ray, S., Norrild, R.K., Larsen, J. A., Petersen, D., Farzadfard, A., Galvagnion, C., Jensen, H., Buell, A.K. Nature Communications 2021, 12, 7289.
[0334] Y. Zhang, A. K. Buell, T. Muller, E. De Genst, J. Benesch, C. M. Dobson, T. P. J. Knowles, ChemBioChem 2016, 17, 1920.
[0335] J.P. Brady, P.J. Farber, A. Sekhar, Y. Lin, R. Huang, A. Bah, T.J. Nott, H.S. Chan, A. J. Baldwin, J.D. Forman-Kay, and L.E. Kay. PNAS 2017, E8194-E8293. Items
[0336] 1. A method of selecting one or more polypeptide variants exhibiting or lacking a property of interest, the method comprising the steps of: a) providing a display-library consisting of one or more polypeptide variants of a polypeptide of interest; b) contacting said display-library with a molar excess of a selection polypeptide under conditions wherein the selection polypeptide exhibits said property of interest, such as aggregation or liquid-liquid phase separation, thereby separating the display-library into at least a first population and a second population, wherein: i) the first population comprises polypeptide variants exhibiting the property of interest; and ii) the second population comprises polypeptide variants lacking the property of interest; and c) recovering said first and / or second population, thereby selecting polypeptide variants exhibiting or lacking the property of interest.
[0337] 2. The method according to item 1 , wherein steps a-c) are performed in one or more iterations, wherein said display-library comprises or consists of the first population and / or the second population recovered from one or more previous iterations.
[0338] 3. The method according to item 2, wherein steps b-c) are performed in 2 iterations or more, such as 3, 4, 5, 6, 7, 8, 9, or 10 iterations or more.
[0339] 4. The method according to any one of the preceding items, wherein said polypeptide variants are linked to a unique polynucleotide identifier.
[0340] 5. The method according to any one of the preceding items, wherein said polypeptide variants are independently covalently linked to a unique polynucleotide identifier. The method according to any one of the preceding items, wherein said molar excess of the selection polypeptide is at least 5-fold, such as at least 10-fold, such as at least 20-fold, such as at least 50-fold, such as at least 100-fold, such as at least 200-fold, such as at least 300-fold, such as at least 400-fold, such as at least 500-fold, such as at least 600-fold, such as at least 700-fold, such as at least 800-fold, such as at least 900-fold, or such as wherein the molar excess of the selection polypeptide is at least 1000-fold. The method according to any one of the preceding items, wherein said molar excess of the selection polypeptide is between 5 and 1000-fold, such as between 5 and 900-fold, between 5 and 800-fold, between 5 and 700-fold, between 5 and 600-fold, between 5 and 500-fold, between 5 and 400-fold, between 5 and 300-fold, between 5 and 200-fold, between 5 and 100-fold, between 5 and 50-fold, between 5 and 20-fold, between 5 and 10-fold, between
[0341] 10 and 1000-fold, between 10 and 900-fold, between 10 and 800-fold, between 10 and 700-fold, between 10 and 600-fold, between 10 and 500-fold, between 10 and 400-fold, between 10 and 300-fold, between 10 and 200-fold, between 10 and 100-fold, between 10 and 50-fold, between 10 and 20-fold, between 20 and 1000-fold, between 20 and 900-fold, between 20 and 800-fold, between 20 and 700-fold, between 20 and 600-fold, between 20 and 500-fold, between 20 and 400-fold, between 20 and 300-fold, between 20 and 200-fold, between 20 and 100-fold, between 20 and 50-fold, between 50 and 1000-fold, between 50 and 900-fold, between 50 and 800-fold, between 50 and 700-fold, between 50 and 600-fold, between 50 and 500-fold, between 50 and 400-fold, between 50 and 300-fold, between 50 and 200-fold, between 50 and 100-fold, between 100 and 1000-fold, between 100 and 900-fold, between 100 and 800-fold, between 100 and 700-fold, between 100 and 600-fold, between 100 and 500-fold, between 100 and 400-fold, between 100 and 300-fold, between 100 and 200- fold, between 200 and 1000-fold, between 200 and 900-fold, between 200 and 800-fold, between 200 and 700-fold, between 200 and 600-fold, between 200 and 500-fold, between 200 and 400-fold, between 200 and 300-fold, between 300 and 1000-fold, between 300 and 900-fold, between 300 and 800-fold, between 300 and 700-fold, between 300 and 600-fold, between 300 and 500- fold, between 300 and 400-fold, between 400 and 1000-fold, between 400 and 900-fold, between 400 and 800-fold, between 400 and 700-fold, between 400 and 600-fold, between 400 and 500-fold, between 500 and 1000-fold, between 500 and 900-fold, between 500 and 800-fold, between 500 and 700-fold, between 500 and 600-fold, between 600 and 1000-fold, between 600 and 900- fold, between 600 and 800-fold, between 600 and 700-fold, between 700 and 1000-fold, between 700 and 900-fold, between 700 and 800-fold, between 800 and 1000-fold, between 800 and 900-fold, or wherein said molar excess of the selection polypeptide is between 900 and 1000-fold.
[0342] 8. The method according to any one of the preceding items, wherein said polypeptide variants are fragments of the polypeptide of interest.
[0343] 9. The method according to any one of the preceding items, wherein said polypeptide variants are mutants of the polypeptide of interest.
[0344] 10. The method according to any one of the preceding items, wherein said polypeptide variants are targeted edited variants of the polypeptide of interest.
[0345] 11. The method according to any one of the preceding items, wherein said polypeptide variants comprise one or more unnatural amino acids.
[0346] 12. The method according to item 8, wherein said fragments of the polypeptide of interest are at least 100 amino acids long, such as at least 200 amino acids long, such as at least 300 amino acids long, such as at least 400 amino acids long, such as at least 500 amino acids long, such as at least 600 amino acids long, such as at least 700 amino acids long, such as at least 800 amino acids long, such as at least 900 amino acids long, or such as at least 1000 amino acids long.
[0347] 13. The method according to item 8, wherein said fragments of the polypeptide of interest are between 30 and 1000 amino acids long, such as between 30 and 900 amino acids long, between 30 and 800 amino acids long, between 30 and
[0348] 700 amino acids long, between 30 and 600 amino acids long, between 30 and
[0349] 500 amino acids long, between 30 and 400 amino acids long, between 30 and
[0350] 300 amino acids long, between 30 and 200 amino acids long, between 30 and
[0351] 100 amino acids long, between 30 and 50 amino acids long, between 50 and 1000 amino acids long, between 50 and 900 amino acids long, between 50 and 800 amino acids long, between 50 and 700 amino acids long, between 50 and
[0352] 600 amino acids long, between 50 and 500 amino acids long, between 50 and
[0353] 400 amino acids long, between 50 and 300 amino acids long, between 50 and
[0354] 200 amino acids long, between 50 and 100 amino acids long, between 100 and
[0355] 1000 amino acids long, between 100 and 900 amino acids long, between 100 and 800 amino acids long, between 100 and 700 amino acids long, between 100 and 600 amino acids long, between 100 and 500 amino acids long, between 100 and 400 amino acids long, between 100 and 300 amino acids long, between 100 and 200 amino acids long, between 200 and 1000 amino acids long, between 200 and 900 amino acids long, between 200 and 800 amino acids long, between 200 and 700 amino acids long, between 200 and 600 amino acids long, between 200 and 500 amino acids long, between 200 and 400 amino acids long, between 200 and 300 amino acids long, between 300 and 1000 amino acids long, between 300 and 900 amino acids long, between 300 and 800 amino acids long, between 300 and 700 amino acids long, between 300 and 600 amino acids long, between 300 and 500 amino acids long, between 300 and 400 amino acids long, between 400 and 1000 amino acids long, between 400 and 900 amino acids long, between 400 and 800 amino acids long, between 400 and 700 amino acids long, between 400 and 600 amino acids long, between 400 and 500 amino acids long, between 500 and 1000 amino acids long, between 500 and 900 amino acids long, between 500 and 800 amino acids long, between 500 and 700 amino acids long, between 500 and 600 amino acids long, between 600 and 1000 amino acids long, between 600 and 900 amino acids long, between 600 and 800 amino acids long, between 600 and 700 amino acids long, between 700 and 1000 amino acids long, between 700 and 900 amino acids long, between 700 and 800 amino acids long, between 800 and 1000 amino acids long, between 800 and 900 amino acids long, or such as between 900 and 1000 amino acids long. The method according to item 8, wherein said fragments of the polypeptide of interest are up to 30 amino acids long, such as up to 25 amino acids long, up to 20 amino acids long, up to 15 amino acids long, up to 10 amino acids long, up to 5 amino acids long, up to 4 amino acids long, or such as up to 2 amino acids long.
[0356] 15. The method according to item 8, wherein said fragments of the polypeptide of interest are between 2 and 30 amino acids long, such as between 2 and 25 amino acids long, between 2 and 20 amino acids long, between 2 and 15 amino acids long, between 2 and 10 amino acids long, between 2 and 5 amino acids long, between 2 and 4 amino acids long, between 4 and 30 amino acids long, between 4 and 25 amino acids long, between 4 and 20 amino acids long, between 4 and 15 amino acids long, between 4 and 10 amino acids long, between 4 and 5 amino acids long, between 5 and 30 amino acids long, between 5 and 25 amino acids long, between 5 and 20 amino acids long, between 5 and 15 amino acids long, between 5 and 10 amino acids long, between 10 and 30 amino acids long, between 10 and 25 amino acids long, between 10 and 20 amino acids long, between 10 and 15 amino acids long, between 15 and 30 amino acids long, between 15 and 25 amino acids long, between 15 and 20 amino acids long, between 20 and 30 amino acids long, between 20 and 25 amino acids long, between 25 and 30 amino acids long
[0357] 16. The method according to any one of the preceding items, wherein said displaylibrary is an mRNA library, a cDNA library, a ribosome display library, a random artificial library, a DNA-encoded library of small molecules or a DNA-encoded library of peptides.
[0358] 17. The method according to any one of the preceding items, wherein said displaylibrary covers part or all of the sequence space of the full-length polypeptide of interest.
[0359] 18. The method according to any one of the preceding items, wherein said displaylibrary covers part or all of the sequence space of one or more regions of interest of the polypeptide of interest.
[0360] 19. The method according to any one of the preceding items, wherein said displaylibrary covers part or all of the sequence space of one or more amino acids of interest of the polypeptide of interest. 20. The method according to any one of the preceding items, wherein the displaylibrary is generated from a DNA library encoding polypeptide variants.
[0361] 21. The method according to any one of the preceding items, wherein the one or more polypeptide variants are encoded by one or more polynucleotides.
[0362] 22. The method according to item 21 , wherein the one or more polynucleotides are covalently linked to the corresponding one or more polypeptide variants via a linker.
[0363] 23. The method according to any one of items 21 to 22, wherein the one or more polynucleotides comprise a ribosome-binding site.
[0364] 24. The method according to any one of items 21 to 23, wherein the linker comprises or consists of puromycin or an amino acid.
[0365] 25. The method according to item 24, wherein the linker is linked to the 3’-end of the one or more polynucleotides, optionally via a second linker, preferably a flexible linker.
[0366] 26. The method according to item 25, wherein the linker and / or the second linker comprises or consists of a nucleic acid sequence consisting of between 16 and 40 nucleotides, such as between 17 and 39 nucleotides, such as between 18 and 38 nucleotides, such as between 19 and 37 nucleotides, such as between 20 and 36 nucleotides, such as between 21 and 35 nucleotides, such as between 22 and 34 nucleotides, such as between 23 and 33 nucleotides, such as between 24 and 32 nucleotides, such as between 25 and 31 nucleotides, such as between 26 and 30 nucleotides, such as between 27 and 29 nucleotides, such as 28 nucleotides.
[0367] 27. The method according to any one of items 25 to 26, wherein the second linker comprises or consists of a PEG linker. 28. The method according to any one of items 21 to 27, wherein an organic fluorophore is linked to the one or more polynucleotides.
[0368] 29. The method according to item 28, wherein the organic fluorophore is fluorescein.
[0369] 30. The method according to any one of the items 28 to 29, the organic fluorophore is linked to position 5 of a thymine ring by a 6-carbon linker.
[0370] 31. The method according to any one of the items 28 to 30, said thymine ring is within the linker and / or the second linker.
[0371] 32. The method according to any one of the items 4 to 31, wherein the unique polynucleotide identifier is a non-coding polynucleotide.
[0372] 33. The method according to any one of the items 4 to 32, wherein the unique polynucleotide identifier is up to 15 nucleotides long, such as up to 14 nucleotides long, such as up to 13 nucleotides long, such as up to 12 nucleotides long, such as up to 11 nucleotides long, or such as up to 10 nucleotides long.
[0373] 34. The method according to any one of the preceding items, wherein said polypeptide of interest is an antibody.
[0374] 35. The method according to any one of items 1 to 33, wherein said polypeptide of interest comprises one or more intrinsically disordered regions.
[0375] 36. The method according to any one of items 1 to 33, wherein said polypeptide of interest is an intrinsically disordered protein.
[0376] 37. The method according to any one of items 1 to 33, wherein said polypeptide of interest is a cyclic peptide.
[0377] 38. The method according to any one of items 1 to 33, wherein said polypeptide of interest is a therapeutic polypeptide. 39. The method according to any one of items 1 to 33, wherein said polypeptide of interest is an enzyme.
[0378] 40. The method according to any one of items 1 to 33, wherein said polypeptide of interest is a polypeptide suitable for administration to a subject in need thereof.
[0379] 41. The method according to any one of items 1 to 33, wherein said polypeptide of interest is a polypeptide suitable for administration to a human subject in need thereof.
[0380] 42. The method according to any one of items 1 to 33, wherein said polypeptide of interest is a polypeptide involved in a neurological disease or a rare disease, preferably wherein said disease is associated with protein aggregation in organs, body fluids or body tissues, such as in the brain, in the liver, and / or in the kidney.
[0381] 43. The method according to item 42, wherein said rare disease is a systemic amyloidosis, such as lysozyme amyloidosis, transthyretin amyloidosis, dialysis- related amyloidosis, light chain amyloidosis.
[0382] 44. The method according to item 42, wherein said neurological disease is selected from the group consisting of: amyotrophic lateral sclerosis, Alzheimer’s disease, Parkinson’s disease, Huntington’s disease, and frontotemporal lobar degeneration.
[0383] 45. The method according to any one of the items 1 to 33, wherein said polypeptide of interest is selected from the group consisting of: alpha-synuclein (P37840), TDP-43 (Q13148) C-terminal low complexity domain, Tau (P10636), heterogeneous nuclear ribonucleoproteins (e.g. P09651) and FUS (P56959).
[0384] 46. The method according to any one of the items 1 to 16, wherein said polypeptide of interest is a polypeptide involved in a cancer. 47. The method according to item 46, wherein said polypeptide of interest is HER2, P53, Androgen receptor, or NUP98-HOXA9.
[0385] 48. The method according to item 46, wherein the cancer is selected from the group consisting of: multiple myeloma or lymphoma, malignant melanoma, HPV induced cancers, prostate cancer, breast cancer, lung cancer, ovarian cancer, liver cancer, uterine serious carcinoma, and gastric cancer.
[0386] 49. The method according to any one of the preceding items, wherein said conditions are selected from the group consisting of: temperature, ionic strength, detergent concentration, time, pH, solvent concentration, presence of crowding agents, pressure, addition of a specific salt and addition of a small molecule.
[0387] 50. The method according to any one of the preceding items, wherein said property of interest is selected from the group consisting of: liquid-liquid phase separation, coacervation, condensation, aggregation, solubility and stability.
[0388] 51. The method according to any one of the preceding items, wherein said one or more polypeptide variants are selected for their ability to bind to a target.
[0389] 52. The method according to item 51 , wherein said target is selected from the group consisting of: a polypeptide, a carbohydrate, a lipid, a small molecule, and a hormone.
[0390] 53. The method according to any one of the preceding items, wherein separating the display-library is performed by a method comprising one or more of ultracentrifugation, ultrafiltration, chromatography, field flow fractionation, centrifugation, filtration, dialysis and / or microfluidic diffusion.
[0391] 54. The method according to any one of the preceding items, wherein the separation of the display library results in the formation of a plurality of phases, wherein the first population and the second population are comprised within different phases. 55. The method according to item 54, wherein one of the phases is a liquid droplet, a dilute phase or a solid phase such as a precipitate.
[0392] 56. The method according to item 54, wherein one of the phases is a solid phase, such as amyloid gels, amyloid fibrils, spherulites, amyloid liquid crystals, protein crystals, filaments, oligomers, or amorphous aggregates.
[0393] 57. The method according to any one of items 4 to 56, wherein the method further comprises identifying said recovered polypeptide variants by sequencing the unique polypeptide identifiers.
[0394] 58. The method according to any one of items 21 to 56, wherein the method further comprises identifying said recovered polypeptide variants by sequencing the one or more polynucleotide encoding said one or more polypeptide variants.
[0395] 59. The method according to item 58, wherein said sequencing is high throughput sequencing, such as second generation sequencing, next generation sequencing, nanopores or mass spectrometry.
[0396] 60. The method according to any one of the preceding items, wherein the amino acid sequence of the polypeptide variants are sequenced by nanopores or mass spectrometry.
[0397] 61. The method according to any one of the preceding items, wherein the method further comprises mapping the amino acid residues of the recovered polypeptides in relation to the property of interest.
[0398] 62. The method according to any one of the preceding items, wherein the polypeptide variants comprise a tag such as an affinity tag.
[0399] 63. The method according to item 62, wherein the affinity tag is an affinity tag selected from the group consisting of: human influenza hemagglutinin (HA)-tag, ubiquitin, histidine (His)-tag, FLAG-tag, Myc-Tag, glutathione S-transferase (GST)-tag, Maltose binding protein (MBP)-tag, Small Ubiquitin-like Modifier (SUMO)-tag, strep tag, strepll tag, and Protein A-tag. 64. The method according to any one of the preceding items, wherein recovering said first and / or second population is performed by affinity purification.
[0400] 65. The method according to any one of the preceding items, wherein the displaylibrary is further separated in one or more further populations exhibiting the property of interest less than the polypeptide of interest and more than the polypeptide of interest.
[0401] 66. The method according to any one of the preceding items, wherein the polypeptide of interest and the selection polypeptide interact.
[0402] 67. The method according to any one of the preceding items, wherein the polypeptide of interest is the selection polypeptide.
[0403] 68. The method according to any one of the preceding items, wherein the polypeptide of interest is capable of binding to the selection polypeptide.
[0404] 69. A method of mapping one or more amino acid residues of one or more polypeptides or fragments thereof in relation to a property of interest, wherein said one or more polypeptide or fragments thereof are selected and recovered using the method according to any one of items 1 to 61.
[0405] 70. The method according to item 69, wherein the fragments are overlapping, such as tiles of the polypeptide of interest.
[0406] 71. A computer implemented method for designing a customized polypeptide sequence from a parent polypeptide sequence, said customized polypeptide sequence correlating with at least one property of interest, said method comprising the steps of: obtaining polypeptide variant information for the parent polypeptide sequence, providing a database of polypeptide variants data and of property of interest data, wherein said polypeptide variants are selected according to the method of any one of the preceding claims, said database comprising o a plurality of polypeptide variant data records, where each polypeptide variant data record consists of the amino acid sequence of each polypeptide variant, and o a plurality of property of interest data records, wherein each polypeptide variant data record is correlated with one data record, wherein the data record reflects whether the property of interest is exhibited or not by each polypeptide variant, such that the property of interest data record of each polypeptide variant is directly connected to the amino acid sequence of each polypeptide variant, preferably to the desired amino acid sequence, categorizing the polypeptide data of each polypeptide variant into categories of property of interest, said categorizing based on whether the polypeptide variant exhibits or lacks the property of interest, determining, based on the categorization, specific amino acid residues correlating with the exhibiting or lacking the property of interest, and designing the amino acid sequence of the customized polypeptide by selecting, for each residue of the customized polypeptide, an amino acid which is likely to, preferably the amino acid which is most likely to, contribute to the customized polypeptide exhibiting or lacking the property of interest, thereby obtaining the sequence of the customized polypeptide. The computer implemented method according to item 71 , wherein the computer implemented method further comprises the step of obtaining sample condition information, wherein each condition information record is correlated with at least one property of interest information record, such that the property of the polypeptide is directly connected to the at least one condition. The computer implemented method according to item 72, wherein the condition is the molar excess of a selection polypeptide relative to the display-library. The computer implemented method according to any one of items 72 to 73, wherein the condition is selected from the group consisting of: temperature, ionic strength, detergent concentration, time, pH, solvent concentration, presence of crowding agents, pressure, addition of a specific salt and addition of a small molecule. 75. The computer implemented method according to any one of items 71 to 74, wherein the polypeptide variant, selection polypeptide and property of interest are as defined in any one of items 1 to 70.
[0407] 76. A computer implemented method for designing a customised polypeptide comprising at least one desired amino acid sequence, said method comprising the steps of: i. designing the customized polypeptide according to the method of any one of claims 69 to 73, and ii. synthesising the customized polypeptide.
[0408] 77. A kit for selecting one or more polypeptide variants exhibiting or lacking a property of interest, the kit comprising: i) a display-library consisting of one or more polypeptide variants of a polypeptide of interest; ii) a selection polypeptide in an amount allowing for a molar excess of said selection polypeptide to be added to the one or more polypeptide variants; and iii) instructions for use.
[0409] 78. The kit according to item 77, wherein the polypeptide variants, the polypeptide of interest and selection polypeptide are as defined in any one of items 1 to 66.
[0410] 79. The kit according to any one of items 77 to 78, wherein the one or more polypeptide variants are linked to a unique polynucleotide identifier.
[0411] 80. A polypeptide identified or obtained by the method according to any one of items 1 to 66.
[0412] 81. The polypeptide of item 80, wherein said polypeptide is an antigen binding molecule, such as an antibody, a nanobody, a Fab, a minibody, a diabody, a tribody, a scFv, a DARPin, a de novo binder, a peptide aptamer, or an affibody.
[0413] 82. The polypeptide of item 80, wherein said polypeptide is a cyclic peptide. 83. The polypeptide of item 80, wherein said polypeptide is a therapeutic polypeptide. 84. The polypeptide of item 80, wherein said polypeptide is an enzyme.
[0414] 85. A polynucleotide encoding the polypeptide according to any one of items 80 to 84. 86. A vector comprising a polynucleotide according to item 85.
[0415] 87. A host cell comprising a vector or a polynucleotide according to any one of items 85 or 86.
Claims
Claims1 . A method of selecting one or more polypeptide variants exhibiting or lacking a property of interest, the method comprising the steps of: a) providing a display-library consisting of one or more polypeptide variants of a polypeptide of interest; b) contacting said display-library with a molar excess of a selection polypeptide under conditions wherein the selection polypeptide exhibits said property of interest, wherein said property of interest is coacervation, condensation, aggregation or liquid-liquid phase separation, thereby separating the display-library into at least a first population and a second population, wherein: i) the first population comprises polypeptide variants exhibiting the property of interest; and ii) the second population comprises polypeptide variants lacking the property of interest; and c) recovering said first and / or second population, thereby selecting polypeptide variants exhibiting or lacking the property of interest, wherein said polypeptide variants are linked to a unique polynucleotide identifier, such as wherein said polypeptide variants are independently covalently linked to a unique polynucleotide identifier and wherein said display-library is an mRNA library, a cDNA library, a DNA-encoded library of small molecules or a DNA-encoded library of peptides and optionally, wherein said molar excess of the selection polypeptide is at least 5-fold.
2. The method according to claim 1 , wherein said molar excess of the selection polypeptide is between 5 and 1000-fold, such as between 5 and 900-fold, between 5 and 800-fold, between 5 and 700-fold, between 5 and 600-fold, between 5 and 500-fold, between 5 and 400-fold, between 5 and 300-fold, between 5 and 200-fold, between 5 and 100-fold, between 5 and 50-fold, between 5 and 20-fold, between 5 and 10-fold, between 10 and 1000-fold, between 10 and 900-fold, between 10 and 800-fold, between 10 and 700-fold, between 10 and 600-fold, between 10 and 500-fold, between 10 and 400-fold, between 10 and 300-fold, between 10 and 200-fold, between 10 and 100-fold, between 10 and 50-fold, between 10 and 20-fold, between 20 and 1000-fold,between 20 and 900-fold, between 20 and 800-fold, between 20 and 700-fold, between 20 and 600-fold, between 20 and 500-fold, between 20 and 400-fold, between 20 and 300-fold, between 20 and 200-fold, between 20 and 100-fold, between 20 and 50-fold, between 50 and 1000-fold, between 50 and 900-fold, between 50 and 800-fold, between 50 and 700-fold, between 50 and 600-fold, between 50 and 500-fold, between 50 and 400-fold, between 50 and 300-fold, between 50 and 200-fold, between 50 and 100-fold, between 100 and 1000- fold, between 100 and 900-fold, between 100 and 800-fold, between 100 and 700-fold, between 100 and 600-fold, between 100 and 500-fold, between 100 and 400-fold, between 100 and 300-fold, between 100 and 200-fold, between 200 and 1000-fold, between 200 and 900-fold, between 200 and 800-fold, between 200 and 700-fold, between 200 and 600-fold, between 200 and 500- fold, between 200 and 400-fold, between 200 and 300-fold, between 300 and 1000-fold, between 300 and 900-fold, between 300 and 800-fold, between 300 and 700-fold, between 300 and 600-fold, between 300 and 500-fold, between 300 and 400-fold, between 400 and 1000-fold, between 400 and 900-fold, between 400 and 800-fold, between 400 and 700-fold, between 400 and 600- fold, between 400 and 500-fold, between 500 and 1000-fold, between 500 and 900-fold, between 500 and 800-fold, between 500 and 700-fold, between 500 and 600-fold, between 600 and 1000-fold, between 600 and 900-fold, between 600 and 800-fold, between 600 and 700-fold, between 700 and 1000-fold, between 700 and 900-fold, between 700 and 800-fold, between 800 and 1000- fold, between 800 and 900-fold, or wherein said molar excess of the selection polypeptide is between 900 and 1000-fold.
3. The method according to any one of the preceding claims, wherein said displaylibrary covers part or all of the sequence space of the full-length polypeptide of interest.
4. The method according to any one of the preceding claims, wherein said polypeptide of interest is selected from the group consisting of: an antibody, a cyclic peptide, a therapeutic polypeptide, an enzyme, a polypeptide suitable for administration to a subject in need thereof, and a polypeptide suitable for administration to a human subject in need thereof, and / or wherein said polypeptide of interest comprises one or more intrinsically disordered regions,such as wherein said polypeptide of interest is an intrinsically disordered protein.
5. The method according to any one of the preceding claims, wherein said polypeptide of interest is a polypeptide involved in a neurological disease or a rare disease, preferably wherein said disease is associated with protein aggregation in organs, body fluids or body tissues, such as in the brain, in the liver, and / or in the kidney.
6. The method according to any one of the preceding claims, wherein the polypeptide of interest and the selection polypeptide interact or the polypeptide of interest is the selection polypeptide.
7. A method of mapping one or more amino acid residues of one or more polypeptides or fragments thereof in relation to a property of interest selected from coacervation, condensation, aggregation and liquid-liquid phase separation, wherein said one or more polypeptide or fragments thereof are selected and recovered using the method according to any one of claims 1 to 6.
8. A computer implemented method for designing a customized polypeptide sequence from a parent polypeptide sequence, said customized polypeptide sequence correlating with at least one property of interest selected from coacervation, condensation, aggregation and liquid-liquid phase separation, said method comprising the steps of: obtaining polypeptide variant information for the parent polypeptide sequence, providing a database of polypeptide variants data and of property of interest data, wherein said polypeptide variants are selected according to the method of any one of the preceding claims, said database comprising o a plurality of polypeptide variant data records, where each polypeptide variant data record consists of the amino acid sequence of each polypeptide variant, and o a plurality of property of interest data records,wherein each polypeptide variant data record is correlated with one data record, wherein the data record reflects whether the property of interest is exhibited or not by each polypeptide variant, such that the property of interest data record of each polypeptide variant is directly connected to the amino acid sequence of each polypeptide variant, preferably to the desired amino acid sequence, categorizing the polypeptide data of each polypeptide variant into categories of property of interest, said categorizing based on whether the polypeptide variant exhibits or lacks the property of interest, determining, based on the categorization, specific amino acid residues correlating with the exhibiting or lacking the property of interest, and designing the amino acid sequence of the customized polypeptide by selecting, for each residue of the customized polypeptide, an amino acid which is likely to, preferably the amino acid which is most likely to, contribute to the customized polypeptide exhibiting or lacking the property of interest, thereby obtaining the sequence of the customized polypeptide, wherein the polypeptide variant, selection polypeptide and property of interest are as defined in any one of claims 1 to 6, optionally wherein the computer implemented method further comprises the step of obtaining sample condition information, wherein each condition information record is correlated with at least one property of interest information record, such that the property of the polypeptide is directly connected to the at least one condition, and / or optionally wherein the condition is selected from the group consisting of: the molar excess of a selection polypeptide relative to the display-library, temperature, ionic strength, detergent concentration, time, pH, solvent concentration, presence of crowding agents, pressure, addition of a specific salt and addition of a small molecule.
9. A computer implemented method for designing a customised polypeptide comprising at least one desired amino acid sequence, said method comprising the steps of: i. designing the customized polypeptide according to the method of claim 8, and ii. synthesising the customized polypeptide.
10. A kit for selecting one or more polypeptide variants exhibiting or lacking a property of interest selected from coacervation, condensation, aggregation and liquid-liquid phase separation, the kit comprising: i) a display-library consisting of one or more polypeptide variants of a polypeptide of interest; ii) a selection polypeptide in an amount allowing for a molar excess of said selection polypeptide to be added to the one or more polypeptide variants; and iii) instructions for use, wherein said display-library is an mRNA library, a cDNA library, a DNA-encoded library of small molecules or a DNA- encoded library of peptides and wherein the one or more polypeptide variants are linked to a unique polynucleotide identifier, preferably wherein the polypeptide variants, the polypeptide of interest and selection polypeptide are as defined in any one of claims 1 to 6.
11. A polypeptide identified or obtained by the method according to any one of claims 1 to 6.
12. A polynucleotide encoding the polypeptide according to claim 11.
13. A vector comprising a polynucleotide according to claim 12.
14. A host cell comprising a vector or a polynucleotide according to any one of claims 12 or 13.