Methods for screening peptides using multiple libraries

By barcoding peptide-nucleic acid complexes and screening mixed libraries, the problem of low peptide screening efficiency in mixed libraries was solved, enabling efficient screening of peptides that bind to target molecules, expanding the library size, and identifying the source of peptides.

CN121986174APending Publication Date: 2026-05-05CHUGAI PHARMA CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHUGAI PHARMA CO LTD
Filing Date
2024-10-10
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

When using multiple cell-free translation systems to generate peptide-nucleic acid complex libraries, it is difficult to effectively screen out peptides that bind to target molecules. Furthermore, the association between genotype and phenotype fails when mixing libraries, making it difficult to identify which library the enriched peptides are derived from. In addition, bias exists, making it difficult to enrich low-enriched peptides.

Method used

By barcoding peptide-nucleic acid complexes, multiple independent nucleic acid display libraries were prepared, mixed, and then contacted with target molecules. The nucleic acid portion that bound to the target molecule was amplified using barcoded primers, and the library from which the peptide was derived was identified.

Benefits of technology

This technology enables efficient screening of peptides that bind to target molecules in mixed libraries, expands library size, and allows identification of which library each peptide is derived from based on barcode sequences, avoiding the neglect of low-enriched peptides and improving screening efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present invention provides a method for screening for candidate peptides capable of binding to a target molecule, the method comprising the steps of: (1) preparing a plurality of nucleic acid display libraries comprising barcoded peptide-(nucleic acid) complexes wherein: the barcoded peptide-(nucleic acid) complexes comprise nucleic acid moieties and peptide moieties, the nucleic acid portion comprises a barcode sequence and a nucleic acid sequence encoding the peptide, and the plurality of nucleic acid display libraries are generated by respective translations using a cell-free translation system; (2) a step of mixing the plurality of nucleic acid display libraries to prepare a mixed nucleic acid display library; (3) a step of bringing the mixed nucleic acid display library into contact with a target molecule; and (4) a step of amplifying a nucleic acid corresponding to the nucleic acid portion in the barcoded peptide-(nucleic acid) complex, which has been bound to the target molecule, using a barcode primer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for screening peptides using multiple libraries. Background Technology

[0002] mRNA display technology using a recombinant cell-free translation system of Escherichia coli (E. coli) has been used to synthesize peptide libraries containing non-natural amino acids and to screen for peptides with high binding capacity (Patent Document 1).

[0003] Library diversity is crucial for effectively identifying useful peptides during screening. Therefore, increasing the diversity of the building blocks (amino acids) of the peptides contained in the library is important. However, when peptides are synthesized via translation, the types of amino acids that can be incorporated into a codon table are limited, and thus the number of available amino acids is also limited. Therefore, the creation of highly diverse libraries using different cell-free translation systems is being investigated (Non-Patent Literature 1 and 2).

[0004] [List of Citations]

[0005] [Patent Literature]

[0006] [Patent Document 1] International Publication No. WO 2013 / 100132

[0007] [Non-patent literature]

[0008] [Non-Patent Literature 1] DEHacker et al. ACS Chem. Biol., 2017, 12, 795-804.

[0009] [Non-Patent Literature 2] MEBrousseau et al. Cell Chem. Biol., 2022, 29, 249-258. Summary of the Invention

[0010] [Technical Issues]

[0011] Theoretically, when multiple cell-free translation systems are used to generate peptide-nucleic acid complex libraries in which phenotypes and genotypes are associated, the diversity of peptides as phenotypes can be increased overall. However, in this case, since genotypes are common in the libraries, it is necessary to process the libraries independently for each cell-free translation system, and the time and effort required for panning increases with the number of libraries, which is problematic. Therefore, it is conceivable to use mixed libraries and panning to more easily select candidate peptides that can bind to target molecules. However, when mixing and panning against target molecule pairs, the association between genotype (nucleic acid) and phenotype (peptide) fails, and it is difficult to identify which library the enriched peptide (specifically, the peptide as a phenotype corresponding to the enriched nucleic acid) is derived from. Furthermore, when mixing and panning libraries, if there is a bias, such as the enrichment of peptides that readily bind to target molecules in a particular library, peptides derived from that library are easily enriched, while peptides derived from other libraries are relatively difficult to enrich (leading to low enrichment), and therefore it may be difficult to obtain hit peptides with a variety of structures.

[0012] This invention was made with such considerations in mind, and the object of this invention is to provide an effective method for screening multiple peptide-nucleic acid complex libraries.

[0013] [Solution to the problem]

[0014] To address the aforementioned problems, the inventors have conducted in-depth research and have thus discovered a method that can identify which library a peptide binding to a target molecule originates from by assigning a nucleic acid barcode (barcoding) to the peptide-nucleic acid complex contained in the library. This method can even identify which library a peptide binding to a target molecule originates from when multiple libraries are mixed, and has ultimately led to the completion of this invention.

[0015] The present invention provides the following [A1] to [A67], [B1] to [B3], [C1] to [C5] or [D1] to [D9].

[0016] [A1] A method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0017] (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety containing a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system;

[0018] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0019] (3) Contact the mixed nucleic acid display library with the target molecule; and

[0020] (4) Amplify the nucleic acid corresponding to the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers.

[0021] [A2] According to the method of [A1], wherein step (1) includes independently translating multiple nucleic acid libraries using a cell-free translation system to prepare multiple peptide-nucleic acid complexes.

[0022] [A3] According to the method of [A1] or [A2], wherein step (1) includes translating a nucleic acid encoding a peptide to prepare a peptide and linking the peptide to the nucleic acid to prepare a peptide-nucleic acid complex.

[0023] [A4] The method according to any one of [A1] to [A3], wherein step (1) includes barcoding the peptide-nucleic acid complex to prepare a barcoded peptide-nucleic acid complex.

[0024] [A5] The method according to any one of [A1] to [A4], wherein the barcoding is performed by reverse transcription of the nucleic acid portion of the peptide-nucleic acid complex.

[0025] [A6] The method according to any one of [A1] to [A5], wherein the barcoding is performed by reverse transcription of the nucleic acid portion of the peptide-nucleic acid complex using a first primer containing a barcode sequence.

[0026] [A7] The method according to any one of [A1] to [A6], wherein the plurality of nucleic acid display libraries have different barcode sequences for each nucleic acid display library.

[0027] [A8] The method according to any one of [A1] to [A7], wherein the barcoded peptide-nucleic acid complex contained in the plurality of nucleic acid display libraries has a different barcode sequence for each nucleic acid display library.

[0028] [A9] The method according to any one of [A1] to [A8], wherein the plurality of nucleic acid display libraries are nucleic acid display libraries that are independently generated by translating each other using different cell-free translation systems.

[0029] [A10] The method according to any one of [A1] to [A9], wherein the plurality of nucleic acid display libraries are nucleic acid display libraries produced by translation based on each other’s different genetic code tables.

[0030] [A11] The method according to any one of [A1] to [A10], wherein in the plurality of nucleic acid display libraries, at least one codon has a different correspondence with an amino acid.

[0031] [A11.1] The method according to any one of [A1] to [A10], wherein in the plurality of nucleic acid display libraries, the correspondence between two or more codons and amino acids is different.

[0032] [A11.2] The method according to any one of [A1] to [A10], wherein in the plurality of nucleic acid display libraries, the correspondence between three or more codons and amino acids is different.

[0033] [A11.3] The method according to any one of [A1] to [A10], wherein in the plurality of nucleic acid display libraries, four or more codons have different correspondences with amino acids.

[0034] [A11.4] The method according to any one of [A1] to [A10], wherein in the plurality of nucleic acid display libraries, five or more codons have different correspondences with amino acids.

[0035] [A11.5] The method according to any one of [A1] to [A10], wherein in the plurality of nucleic acid display libraries, the correspondence between 10 or more codons and amino acids is different.

[0036] [A11.6] The method according to any one of [A1] to [A10], wherein in the plurality of nucleic acid display libraries, the correspondence between 15 or more codons and amino acids is different.

[0037] [A11.7] The method according to any one of [A1] to [A10], wherein in the plurality of nucleic acid display libraries, the correspondence between one or more and 48 or fewer codons and amino acids is different.

[0038] [A11.8] The method according to any one of [A1] to [A10], wherein in the plurality of nucleic acid display libraries, the correspondence between one or more and 30 or fewer codons and amino acids is different.

[0039] [A11.9] The method according to any one of [A1] to [A10], wherein in the plurality of nucleic acid display libraries, the correspondence between two or more and 30 or fewer codons and amino acids is different.

[0040] [A12] The method according to any one of [A1] to [A11.9], wherein in step (2), two or more nucleic acid display libraries are mixed.

[0041] [A12.01] The method according to any one of [A1] to [A11.9], wherein in step (2), three or more nucleic acid display libraries are mixed.

[0042] [A12.02] The method according to any one of [A1] to [A11.9], wherein in step (2), four or more nucleic acid display libraries are mixed.

[0043] [A12.03] The method according to any one of [A1] to [A11.9], wherein in step (2), two or more and 400 or fewer nucleic acid display libraries are mixed.

[0044] [A12.04] The method according to any one of [A1] to [A11.9], wherein in step (2), two or more and 200 or fewer nucleic acid display libraries are mixed.

[0045] [A12.05] The method according to any one of [A1] to [A11.9], wherein in step (2), two or more and 100 or fewer nucleic acid display libraries are mixed.

[0046] [A12.06] The method according to any one of [A1] to [A11.9], wherein in step (2), two or more and 50 or fewer nucleic acid display libraries are mixed.

[0047] [A12.07] The method according to any one of [A1] to [A11.9], wherein in step (2), two or more and 15 or fewer nucleic acid display libraries are mixed.

[0048] [A12.08] The method according to any one of [A1] to [A11.9], wherein in step (2), two or more and ten or fewer nucleic acid display libraries are mixed.

[0049] [A12.09] The method according to any one of [A1] to [A11.9], wherein in step (2), three or more and 50 or fewer nucleic acid display libraries are mixed.

[0050] [A12.10] The method according to any one of [A1] to [A11.9], wherein in step (2), four or more and 50 or fewer nucleic acid display libraries are mixed.

[0051] [A13] The method according to any one of [A1] to [A12.1], wherein the barcoded peptide-nucleic acid complex is barcoded such that a library derived from the barcoded peptide-nucleic acid complex can be identified.

[0052] [A14] According to the method of [A13], wherein the “library from which the barcoded peptide-nucleic acid complex is derived” is any one of the plurality of nucleic acid display libraries.

[0053] [A15] According to the method of [A13] or [A14], wherein the barcode sequence is different for each nucleic acid display library, so that nucleic acid display libraries containing the barcode sequence can be identified.

[0054] [A16] The method according to any one of [A1] to [A15], wherein the peptide-nucleic acid complex comprises a nucleic acid portion and a peptide portion, and the nucleic acid portion comprises a nucleic acid sequence encoding the peptide portion.

[0055] [A17] The method according to any one of [A1] to [A16], wherein the nucleic acid sequence encoding the peptide portion contains a random sequence.

[0056] [A18] The method according to [A17], wherein the random sequence contains at least one codon specifying a non-natural amino acid.

[0057] [A19] The method according to [A17] or [A18], wherein the random sequence contains at least two triplets of one or more types of triplets selected from a variety of types.

[0058] [A19.1] The method according to [A17] or [A18], wherein the random sequence contains 2 to 1000 triplets of one or more types of triplets selected from a variety of types.

[0059] [A19.2] The method according to [A17] or [A18], wherein the random sequence contains 2 to 500 triplets of one or more types of triplets selected from a variety of types.

[0060] [A19.3] The method according to [A17] or [A18], wherein the random sequence contains 2 to 100 triplets of one or more types of triplets selected from a variety of types.

[0061] [A19.4] The method according to [A17] or [A18], wherein the random sequence contains 2 to 20 triplets of one or more types of triplets selected from a variety of types.

[0062] [A19.5] The method according to [A17] or [A18], wherein the random sequence contains 3 to 18 triplets of one or more types of triplets selected from a variety of types.

[0063] [A19.6] The method according to [A17] or [A18], wherein the random sequence contains 4 to 18 triplets of one or more types of triplets selected from a variety of types of triplets.

[0064] [A19.7] The method according to [A17] or [A18], wherein the random sequence contains 4 to 15 triplets of one or more types of triplets selected from a variety of types.

[0065] [A19.8] The method according to [A17] or [A18], wherein the random sequence contains 5 to 15 triplets of one or more types of triplets selected from a variety of types.

[0066] [A19.9] The method according to [A17] or [A18], wherein the random sequence contains 5 to 12 triplets of one or more types of triplets selected from a variety of types.

[0067] [A20] The method according to any one of [A19] to [A19.9], wherein the multiple types of triplets are 64 or fewer types of triplets.

[0068] [A21] The method according to any one of [A19] to [A19.9], wherein the one or more types of triplets are randomly selected from the plurality of types of triplets.

[0069] [A22] The method according to any one of [A1] to [A21], wherein the peptide-nucleic acid complex is a cyclic peptide-nucleic acid complex.

[0070] [A23] The method according to any one of [A1] to [A22], wherein the peptide-nucleic acid complex is selected from the group consisting of peptide-DNA complex and peptide-mRNA complex.

[0071] [A24] The method according to any one of [A1] to [A23], wherein the peptide-nucleic acid complex is a peptide-mRNA complex.

[0072] [A25] The method according to any one of [A1] to [A24], wherein the peptide-nucleic acid complex is a cyclic peptide-mRNA complex.

[0073] [A26] The method according to any one of [A1] to [A25] further comprises reverse transcription of the nucleic acid portion of the peptide-nucleic acid complex to obtain the barcoded peptide-nucleic acid complex.

[0074] [A27] The method according to any one of [A1] to [A26] includes reverse transcription using a first primer containing a barcode sequence to obtain the barcoded peptide-nucleic acid complex.

[0075] [A28] The method according to any one of [A1] to [A27] includes reverse transcription of the nucleic acid portion of the peptide-nucleic acid complex using a first primer containing a barcode sequence to obtain the barcoded peptide-nucleic acid complex.

[0076] [A29] The method according to any one of [A1] to [A28], wherein the barcoded peptide-nucleic acid complex is a barcoded cyclic peptide-nucleic acid complex.

[0077] [A30] The method according to any one of [A1] to [A29], wherein the nucleic acid portion of the barcoded peptide-nucleic acid complex contains a barcode sequence and a nucleic acid sequence encoding the peptide.

[0078] [A31] The method according to any one of [A1] to [A30], wherein the nucleic acid portion of the barcoded peptide-nucleic acid complex is selected from the group consisting of DNA / mRNA duplexes and DNA / DNA duplexes.

[0079] [A32] The method according to any one of [A1] to [A31], wherein the nucleic acid portion of the barcoded peptide-nucleic acid complex is a DNA / mRNA duplex.

[0080] [A33] The method according to [A32], wherein the DNA / mRNA duplex is a cDNA / mRNA duplex.

[0081] [A34] The method according to [A6] or [A27], wherein the first primer contains an identification barcode sequence.

[0082] [A34.1] The method according to [A6] or [A27], wherein the first primer contains a common barcode sequence.

[0083] [A34.2] The method according to [A6] or [A27], wherein the first primer contains an identification barcode sequence and a common barcode sequence.

[0084] [A34.3] The method according to [A6], [A27], [A34] or [A34.2], wherein the first primer contains an identification barcode sequence having 15 to 30 bases.

[0085] [A34.4] The method according to [A6], [A27], [A34.1] or [A34.2], wherein the first primer contains a common barcode sequence having 15 to 30 bases.

[0086] [A34.5] The method according to [A6] or [A27], wherein the first primer contains a barcode sequence having 25 to 150 bases.

[0087] [A35] The method according to any one of [A1] to [A34.5], wherein the cell-free translation system is a reconstructed cell-free translation system.

[0088] [A36] The method according to any one of [A1] to [A35], wherein the cell-free translation system contains amino acid combinations that are different from each other in the plurality of nucleic acid display libraries.

[0089] [A37] The method according to any one of [A1] to [A36], wherein the nucleic acid sequence of the nucleic acid contained in the nucleic acid library contains a random sequence.

[0090] [A38] The method according to [A37], wherein the random sequence contains at least two triplets of one or more types of triplets selected from a variety of types.

[0091] [A38.1] The method according to [A37], wherein the random sequence contains 2 to 1000 triplets of one or more types of triplets selected from a variety of types.

[0092] [A38.2] The method according to [A37], wherein the random sequence contains 2 to 500 triplets of one or more types of triplets selected from a variety of types.

[0093] [A38.3] The method according to [A37], wherein the random sequence contains 2 to 100 triplets of one or more types of triplets selected from a variety of types.

[0094] [A38.4] The method according to [A37], wherein the random sequence contains 2 to 20 triplets of one or more types of triplets selected from a variety of types.

[0095] [A38.5] The method according to [A37], wherein the random sequence contains 3 to 18 triplets of one or more types of triplets selected from a variety of types of triplets.

[0096] [A38.6] The method according to [A37], wherein the random sequence contains one or more types of triplets selected from a variety of types of triplets, of 4 to 18 triplets.

[0097] [A38.7] The method according to [A37], wherein the random sequence contains 4 to 15 triplets of one or more types of triplets selected from a variety of types of triplets.

[0098] [A38.8] The method according to [A37], wherein the random sequence contains 5 to 15 triplets of one or more types of triplets selected from a variety of types.

[0099] [A38.9] The method according to [A37], wherein the random sequence contains 5 to 12 triplets of one or more types of triplets selected from a variety of types.

[0100] [A39] The method according to any one of [A38] to [A38.9], wherein the multiple types of triplets are 64 or fewer types of triplets.

[0101] [A40] The method according to any one of [A38] to [A39], wherein the one or more types of triplets are randomly selected from the plurality of types of triplets.

[0102] [A41] The method according to any one of [A1] to [A40], wherein the nucleic acid library is an mRNA library.

[0103] [A41-1] The method according to any one of [A1] to [A40], wherein the nucleic acid library contains nucleic acids with different numbers of repeats.

[0104] [A41-2] The method according to any one of [A1] to [A40], wherein the nucleic acid library contains nucleic acids having the same number of repeats.

[0105] [A42] The method according to any one of [A1] to [A41], wherein the peptide portion of the peptide-nucleic acid complex contains a predetermined amino acid, and the nucleic acid portion contains a triplet encoding the amino acid.

[0106] [A43] The method according to any one of [A1] to [A42] includes, between steps (3) and (4), a step (3A) of eluting the nucleic acid portion from the barcoded peptide-nucleic acid complex bound to the target molecule.

[0107] [A44] According to the method of [A43], the elution in step (3A) is carried out by enzymatic treatment, heating, application of chemical stimuli or light irradiation.

[0108] [A45] The method according to [A44], wherein the enzyme is a TEV protease.

[0109] [A46] According to the method of [A44], the light to be irradiated is light having a wavelength of 300 nm or greater and 500 nm or less.

[0110] [A47] The method according to [A44], wherein the chemical irritant is an acid or a base.

[0111] [A48] The method according to any one of [A1] to [A47] further includes, between steps (3) and (4), a step (3B) of identifying the amino acid sequence of the peptide portion of the barcoded peptide-nucleic acid complex that binds to the target molecule.

[0112] [A49] The method according to [A48], wherein the amino acid sequence is identified based on the nucleic acid sequence of the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule.

[0113] [A50] The method according to [A48] or [A49], wherein the amino acid sequence is identified based on the nucleic acid sequence and barcode sequence of the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule.

[0114] [A51] The method according to any one of [A1] to [A50], wherein in step (4), the nucleic acid corresponding to the nucleic acid portion of the barcoded peptide-nucleic acid complex having the same barcode sequence as the target molecule is selectively amplified.

[0115] [A52] The method according to any one of [A1] to [A51], wherein the barcode primer is a second primer containing a barcode sequence.

[0116] [A53] The method according to any one of [A1] to [A52], wherein the barcode primer is a second primer composed of the barcode sequence.

[0117] [A54] The method according to [A53], wherein the barcode sequence has 15 to 30 bases.

[0118] [A55] The method according to any one of [A1] to [A54], wherein in step (4), the nucleic acid of the barcoded peptide-nucleic acid complex having the same barcode sequence as the barcode sequence of the barcode primer is selectively amplified.

[0119] [A56] The method according to any one of [A1] to [A55], wherein in step (4), the nucleic acid is amplified by PCR.

[0120] [A57] The method according to any one of [A1] to [A56], wherein the amplified nucleic acid has the same bases as the DNA contained in the barcoded peptide-nucleic acid complex.

[0121] [A58] The method according to any one of [A1] to [A57], wherein steps (1) to (4) constitute a loop and the loop is repeated multiple times.

[0122] [A59] The method according to any one of [A1] to [A58] further includes step (5) after step (4) or before starting the next cycle, further amplifying the nucleic acid amplified in step (4) with primers that do not contain the barcode sequence.

[0123] [A60] The method according to any one of [A1] to [A58] includes, after step (4) or before starting the next cycle, a step (5) of further amplifying the nucleic acid amplified in step (4) with primers that do not contain the barcode sequence, wherein the nucleic acid that does not contain the barcode sequence is amplified.

[0124] [A61] The method according to any one of [A1] to [A60], wherein the primer without barcode sequence is a primer annealed to the 3' region of the nucleic acid library.

[0125] [A62] The method according to any one of [A1] to [A61], wherein the nucleic acid display library has 10 4 Or a more diverse library.

[0126] [A63] The method according to any one of [A1] to [A60], wherein the nucleic acid display library has 10 10 Or a more diverse library.

[0127] [A64] The method according to any one of [A1] to [A63], wherein the nucleic acid display library contains 10 4 A library of one or more types of peptide-nucleic acid complexes.

[0128] [A65] The method according to any one of [A1] to [A64], wherein the nucleic acid display library contains 10 10 A library of one or more types of peptide-nucleic acid complexes.

[0129] [A66] The method according to any one of [A1] to [A65], wherein the nucleic acid display library is an mRNA display library.

[0130] [A67] The method according to any one of [A1] to [A66], wherein the target molecule is a protein.

[0131] [B1] A method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0132] (1) Multiple nucleic acid libraries were independently translated using a cell-free translation system to prepare peptide-nucleic acid complexes containing both nucleic acid and peptide moieties.

[0133] (1-1) The peptide-nucleic acid complex was barcoded to prepare multiple nucleic acid display libraries containing the barcoded peptide-nucleic acid complex;

[0134] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0135] (3) Contact the mixed nucleic acid display library with the target molecule; and

[0136] (4) Amplify the nucleic acid contained in the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers.

[0137] [B2] According to the method of [B1], wherein the barcoded peptide-nucleic acid complex contains a barcoded nucleic acid sequence (barcoded sequence) in the nucleic acid portion.

[0138] [B3] The method according to [B1] or [B2] includes any one of the features in [A2] to [A67].

[0139] "A feature of any one of [A2] to [A67]" means the matter specifying the invention described in each of [A2] to [A67] above, and the terms used in conjunction with the descriptions in [B1] and [B2] are interchangeable and can be understood.

[0140] [C1] A method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0141] (1) Prepare peptide-nucleic acid complexes by independently translating multiple nucleic acid libraries using a single cell-free translation system.

[0142] (1-1) The peptide-nucleic acid complex was barcoded to prepare multiple nucleic acid display libraries containing the barcoded peptide-nucleic acid complex;

[0143] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0144] (3) Contact the mixed nucleic acid display library with the target molecule; and

[0145] (4) Amplify the nucleic acid contained in the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers.

[0146] [C2] According to the method of [C1], wherein the nucleic acid sequence of the nucleic acid contained in the nucleic acid library contains a random sequence, and the random sequence contains at least two triplets of one or more types of triplets selected from a variety of types of triplets.

[0147] [C3] The method according to [C1] or [C2], wherein the nucleic acid library has a different number of replicates from other nucleic acid libraries.

[0148] [C4] The method according to any one of [C1] to [C3], wherein the peptide portion of the peptide-nucleic acid complex contains a predetermined amino acid, and the nucleic acid portion contains a triplet encoding the amino acid.

[0149] [C5] The method according to any one of [C1] to [C4], wherein the plurality of nucleic acid display libraries are nucleic acid display libraries produced by translation based on the same genetic code table.

[0150] [C6] The method according to any one of [C1] to [C5] includes any one of the features of [A1] to [A67].

[0151] [D1] A method for producing a nucleic acid display library, the method comprising the steps (a) and (b):

[0152] (a) Design nucleic acid display libraries containing peptide-nucleic acid complexes using the results of principal component analysis with descriptors as indicators;

[0153] (b) Using a cell-free translation system to translate nucleic acid libraries to prepare peptide-nucleic acid complexes.

[0154] [D2] According to the method described in [D1], it further includes step (c) after step (b) barcoding the peptide-nucleic acid complex to prepare a nucleic acid display library containing the barcoded peptide-nucleic acid complex.

[0155] [D3] The method according to [D1] or [D2] further includes the step (d) of preparing nucleic acid display libraries to provide multiple nucleic acid display libraries.

[0156] [D4] The method according to any one of [D1] to [D3], wherein the plurality of nucleic acid display libraries are libraries in which the peptide-nucleic acid complex has almost no structural overlap.

[0157] [D5] The method according to any one of [D1] to [D4], wherein the descriptor is a descriptor comprising the following items: 'NumValenceElectrons', 'PEOE_VSA3', 'fr_Al_OH', 'fr_Al_OH_noTert', 'fr_Ar_N', 'fr_C_O_noCOO', 'fr_C_S', 'fr_HOCCN', 'fr_NH1', 'fr_Ndealkylation2', 'fr_Nhpyrrole', 'fr_SH', 'fr_alkyl_halide', 'fr_allylic_oxid' ,'fr_aryl_methyl','fr_azide','fr_azo','fr_bicyclic','fr_diazo','fr_dihydropyridine','fr_epoxide','fr_ether','fr _furan', 'fr_halogen', 'fr_hdrzine', 'fr_hdrzone', 'fr_imidazole', 'fr_imide', 'fr_isocyan', 'fr_isothiocyan', 'fr_keto ne', 'fr_ketone_Topliss', 'fr_lactam', 'fr_lactone', 'fr_methoxy', 'fr_morpholine', 'fr_nitrile', 'fr_nitro', 'fr_nitro _arom', 'fr_nitro_arom_nonortho', 'fr_nitroso', 'fr_oxazole', 'fr_phenol', 'fr_phenol_noOrthoHbond', 'fr_phos_acid', ' fr_phos_ester', 'fr_piperdine', 'fr_piperzine', 'fr_priamide', 'fr_prisulfonamd', 'fr_pyridine', 'fr_quatN', 'fr_sulfi de', 'fr_sulfonamd', 'fr_sulfone', 'fr_term_acetylene', 'fr_tetrazole', 'fr_thiazole', 'fr_thiocyan', and 'fr_thiophene'.

[0158] [D6] The method according to any one of [D1] to [D5], wherein the nucleic acid display library has 10 4Or a more diverse library.

[0159] [D7] The method according to any one of [D1] to [D5], wherein the nucleic acid display library contains 10 4 A library of one or more types of peptide-nucleic acid complexes.

[0160] [D8] The method according to any one of [D1] to [D7], wherein the nucleic acid display library is an mRNA display library.

[0161] [D9] A nucleic acid display library obtained according to any one of [D1] to [D8].

[0162] [D10] The method according to any one of [A1] to [A67], wherein the nucleic acid display library is the library according to [D9].

[0163] "A feature of any one of [A1] to [A67]" refers to the matter specifying the invention described in each of [A1] to [A67] above, and the terms used in conjunction with the descriptions in [C1] to [C4] are interchangeable with those used in the above numbering. In the above numbering, unless otherwise stated, a number referenced in a dependent entry includes the branch number of that number. For example, a reference to [A1] in a dependent entry indicates that it includes not only [A1] but also its branch number, such as [A1.1]. This also applies to other numbering.

[0164] [Beneficial effects of the invention]

[0165] The screening method according to the present invention enables the simultaneous screening of multiple libraries and more effectively identifies peptides that bind to target molecules. Furthermore, the screening method according to the present invention not only expands the library size by using a mixed nucleic acid display library composed of multiple mixed libraries, but also allows for easy identification of which library each peptide is derived from based on the barcode sequence binding to each peptide. Moreover, peptides capable of binding to target molecules can be enriched according to each library by selectively amplifying the nucleic acid portion encoding the peptide using primers corresponding to each barcode sequence, making it difficult to ignore low-enriched peptides and effectively identifying peptides capable of binding to target molecules. Attached Figure Description

[0166] [Figure 1] Figure 1 is a graph of the values ​​of each principal component of the random sequence in each code table (random 1: PURE1, random 2: PURE2, random 3: PURE3, random 4: PURE4).

[0167] [Figure 2] Figure 2 is a graph showing the values ​​of each principal component of the enriched sequence in each codon table.

[0168] [Figure 3] Figure 3 shows the enriched sequences (shown in black) and all random sequences (shown in gray), where each shape corresponds to the following cell-free translation system: PURE1, PURE2, A graph of the values ​​of each principal component (PURE3, PURE4). Detailed Implementation

[0169] An embodiment of the invention is described in detail below.

[0170] As used in this article, the “amino acids” and “amino acid analogs” that make up a peptide can be referred to as “amino acid residues” and “amino acid analog residues”, respectively.

[0171] As used herein, "amino acid" refers to α-, β-, or γ-amino acids, and is not limited to naturally occurring amino acids and may also be non-naturally occurring amino acids. Naturally occurring amino acids in this application refer to the 20 amino acids contained in proteins, specifically Gly, Ala, Ser, Thr, Val, Leu, Ile, Phe, Tyr, Trp, His, Glu, Asp, Gln, Asn, Cys, Met, Lys, Arg, and Pro. In the case of α-amino acids, they may be L-amino acids or D-amino acids, or α,α-dialkyl amino acids. The side chains of amino acids are not particularly limited, and, except for hydrogen atoms, the side chains are freely selected from alkyl groups, alkenyl groups, alkynyl groups, aryl groups, heteroaryl groups, aralkyl groups, cycloalkyl groups, etc. Substituents may be added to each side chain, and these substituents are freely selected from, for example, any functional group, including N, O, S, B, Si, and P atoms (i.e., optionally substituted alkyl, alkenyl, alkynyl, aryl, heteroaryl, aralkyl, cycloalkyl, etc.).

[0172] Examples of substituents include halogen-derived substituents such as fluorine (-F), chlorine (-Cl), bromine (-Br), and iodine (-I). Further examples include those substituted with one or more of these, such as alkyl groups, cycloalkyl groups, alkenyl groups, alkynyl groups, aryl groups, heteroaryl groups, or aralkyl groups, which may have halogens as substituents.

[0173] Examples of substituents used to form ethers as substituents derived from the O atom include alkoxy groups (-OR), and the alkoxy group is selected from alkylalkoxy groups, cycloalkylalkoxy groups, alkenylalkoxy groups, alkynylalkoxy groups, arylalkoxy groups, heteroarylalkoxy groups, aralkylalkoxy groups, etc. Examples of substituents used to form alcohol moieties include hydroxyl groups (-OH). Examples of substituents used to form carbonyl groups include carbonyl groups (-C(=O)-R), and the carbonyl group is selected from hydrocarbon groups (-C=OH, aldehydes are obtained as compounds), alkylcarbonyl groups (ketones are obtained as compounds), cycloalkylcarbonyl groups, alkenylcarbonyl groups, alkynylcarbonyl groups, arylcarbonyl groups, heteroarylcarbonyl groups, aralkylcarbonyl groups, etc. Examples of substituents used to form carboxylic acids (-CO2H) include carboxyl groups. Examples of substituents used to form ester groups include oxycarbonyl groups (-OC(=O)-R) and carbonylalkoxy groups (-C(=O)-OR). The carbonyl alkoxy group is selected from alkyl oxycarbonyl groups, cycloalkyl oxycarbonyl groups, alkenyl oxycarbonyl groups, alkynyl oxycarbonyl groups, aryl oxycarbonyl groups, heteroaryl oxycarbonyl groups, aralkyl oxycarbonyl groups, etc., and the oxycarbonyl group is selected from alkyl carbonyl oxy groups, cycloalkyl carbonyl oxy groups, alkenyl carbonyl oxy groups, alkynyl carbonyl oxy groups, aryl carbonyl oxy groups, heteroaryl carbonyl oxy groups, aralkyl carbonyl oxy groups, etc.

[0174] Examples of substituents used to form thioesters include thiol carbonyl groups (-SC(=O)-R) and carbonyl alkyl thiol groups (-C(=O)-SR), selected from thiol alkyl carbonyl groups, thiol cycloalkyl carbonyl groups, thiol alkenyl carbonyl groups, thiol alkynyl carbonyl groups, thiol aryl carbonyl groups, thiol heteroaryl carbonyl groups, thiol aralkyl carbonyl groups, etc., or carbonyl alkyl thiol groups, carbonyl cycloalkyl thiol groups, carbonyl alkenyl thiol groups, carbonyl alkynyl thiol groups, carbonyl aryl thiol groups, carbonyl heteroaryl thiol groups, carbonyl aralkyl thiol groups, etc.

[0175] Examples of substituents used to form amide groups include aminoalkyl carbonyl groups (-NH-C(=O)-R), aminocycloalkyl carbonyl groups, aminoenyl carbonyl groups, aminoynyl carbonyl groups, aminocycloalkyl carbonyl groups, aminoaryl carbonyl groups, aminoheteroaryl carbonyl groups, aminoaralkyl carbonyl groups, etc., or carbonylalkylamino groups (-C(=O)-NHR), carbonylcycloalkylamino groups, carbonylenylamino groups, carbonylynylamino groups, carbonylarylamino groups, carbonylheteroarylamino groups, carbonylaralkylamino groups, etc. Further examples include compounds in which the H atom bonded to the N atom is replaced by an alkyl group, cycloalkyl group, alkenyl group, ynyl group, aryl group, heteroaryl group, or aralkyl group.

[0176] Examples of substituents used to form urethane groups include aminoalkylurethane groups (-NH-C(=O)-OR), aminocycloalkylurethane groups, aminoalkenylurethane groups, aminoynylurethane groups, aminocycloalkylurethane groups, aminoarylurethane groups, aminoheteroarylurethane groups, and aminoaralkylurethane groups. Further examples include compounds in which the H atom bonded to the N atom is replaced by an alkyl group, cycloalkyl group, alkenyl group, ynyl group, aryl group, heteroaryl group, or aralkyl group.

[0177] Examples of substituents used to form the sulfonamide group include aminoalkylsulfonyl groups (-NH-SO2-R), aminocycloalkylsulfonyl groups, aminoenylsulfonyl groups, aminoynylsulfonyl groups, aminocycloalkylsulfonyl groups, aminoarylsulfonyl groups, aminoheteroarylsulfonyl groups, aminoaralkylsulfonyl groups, etc., or sulfonylalkylamino groups (-SO2-NHR), sulfonylcycloalkylamino groups, sulfonylenylamino groups, sulfonylynylamino groups, sulfonylarylamino groups, sulfonylheteroarylamino groups, sulfonylaralkylamino groups, etc. Further examples include compounds in which the H atom bonded to the N atom is replaced by an alkyl group, cycloalkyl group, alkenyl group, ynyl group, aryl group, heteroaryl group, or aralkyl group.

[0178] Examples of substituents used to form the sulfonamide group include aminoalkylaminosulfonyl groups (-NH-SO2-NHR), aminocycloalkylaminosulfonyl groups, aminoalkenylaminosulfonyl groups, aminoynylaminosulfonyl groups, aminocycloalkylaminosulfonyl groups, aminoarylaminosulfonyl groups, aminoheteroarylaminosulfonyl groups, and aminoaralkylaminosulfonyl groups. Further examples include compounds in which the H atom bonded to the N atom is replaced by a substituent selected from alkyl groups, cycloalkyl groups, alkenyl groups, ynyl groups, aryl groups, heteroaryl groups, and aralkyl groups, which may consist of any two identical or different substituents, or may form a cyclic structure.

[0179] Examples of substituents used to form thiocarboxylic acids include thiocarboxylic acid groups (-C(=O)-SH), and examples of functional groups used to form keto acids include keto acid groups (-C(=O)-CO2H).

[0180] Examples of substituents forming thiol groups derived from the S atom include thiol groups (-SH) and the formation of alkyl thiols, cycloalkyl thiols, alkenyl thiols, alkynyl thiols, aryl thiols, heteroaryl thiols, and aralkyl groups. Substituents used to form thioethers (-SR) are selected from alkyl thiol groups, cycloalkyl thiol groups, alkenyl thiol groups, alkynyl thiol groups, aryl thiol groups, heteroaryl thiol groups, aralkyl thiol groups, etc. Substituents used to form sulfoxide groups (-S(=O)-R) are selected from alkyl sulfoxide groups, cycloalkyl sulfoxide groups, alkenyl sulfoxide groups, alkynyl sulfoxide groups, aryl sulfoxide groups, heteroaryl sulfoxide groups, aralkyl sulfoxide groups, etc. Substituents used to form sulfone groups (-SO2-R) are selected from alkyl sulfone groups, cycloalkyl sulfone groups, alkenyl sulfone groups, alkynyl sulfone groups, aryl sulfone groups, heteroaryl sulfone groups, aralkyl sulfone groups, etc. Examples of substituents used to form sulfonic acids include the sulfonic acid group (-SO3H).

[0181] Examples of substituents derived from the N atom include azide groups (-N3); nitrile groups (-CN); amino groups (-NH2) as substituents forming primary amines; alkylamino groups, cycloalkylamino groups, alkenylamino groups, alkynylamino groups, arylamino groups, heteroarylamino groups, aralkylamino groups, etc., as substituents forming secondary amines (-NH-R); substituents selected from alkyl groups, cycloalkyl groups, alkenyl groups, alkynyl groups, aryl groups, heteroaryl groups, aralkyl groups, etc., as substituents forming tertiary amines (-NR(R')), such as alkyl (aralkyl)amino groups, which can consist of any two identical or different substituents, or can form a cyclic structure; amidine groups (-C(=NH)-NH2), or where N... The three substituents on the atom are replaced by any three identical or different groups of alkyl, cycloalkyl, alkenyl, alkynyl, aryl, heteroaryl, and aralkyl groups as substituents forming an amidine group (-C(=NR)-NR'R'') (such as alkyl(aralkyl)(aryl)amidinium); and guanidine groups (-NH-C(=NH)-NH2) or groups in which R, R', R'', R''' are selected from alkyl, cycloalkyl, alkenyl, alkynyl, aryl, heteroaryl, and aralkyl groups (which may consist of any four identical or different substituents or may form a cyclic structure) as substituents forming a guanidine group (-NR-C(=NR'')-NR'R'').

[0182] Examples of substituents used to form a urea group include an aminocarbamoyl group (-NR-C(=O)-NR'R''). Examples of R, R', and R'' include substituents selected from hydrogen atoms, alkyl groups, cycloalkyl groups, alkenyl groups, alkynyl groups, aryl groups, heteroaryl groups, and aralkyl groups, which may consist of any three identical or different substituents, or may form a cyclic structure.

[0183] Examples of functional groups derived from the B atom include alkylboranes (-BR(R')) and alkoxyboranes (-B(OR)(OR')). Examples of the two substituents include substituents selected from alkyl groups, cycloalkyl groups, alkenyl groups, alkynyl groups, aryl groups, heteroaryl groups, and aralkyl groups, which can consist of any two identical or different substituents, or can form a cyclic structure.

[0184] In this way, one or more functional groups containing O, N, S, B, P, Si atoms, and halogen atoms (typically used in low molecular weight compounds, such as halogen groups) can be added. That is, one or more additional substituents can be added to alkyl, cycloalkyl, alkenyl, alkynyl, aryl, heteroaryl, or aralkyl groups, which are shown as substituents. The condition of satisfying all these functional groups is defined as free choice of substituents. As with α-amino acids, any conformation is acceptable for β- and γ-amino acids, and the choice of side chains is also the same as in α-amino acids, without particular restrictions.

[0185] The skeletal amino group of an amino acid can be unsubstituted (NH2 group) or substituted (i.e., NHR group: R represents an optionally substituted alkyl group, alkenyl group, alkynyl group, aryl group, heteroaryl group, aralkyl group, or cycloalkyl group. Alternatively, the carbon atom bonded to the N atom and the carbon atom on the side chain extending from the carbonyl α position can combine to form a ring, such as proline. Substituents can be freely chosen, and examples include halogen groups, ether groups, and hydroxyl groups). Such amino acids having substituted skeletal amino groups are referred to herein as “N-substituted amino acids”. As used herein, “N-substituted amino acids” are preferably N-alkyl amino acids, and can be N-C1-C6 alkyl amino acids or N-C1-C4 alkyl amino acids. More preferably, N-substituted amino acids can be N-methyl amino acids or N-ethyl amino acids, with N-methyl amino acids being particularly preferred.

[0186] As used herein, “amino acid analog” refers to α-hydroxycarboxylic acid. α-Hydroxycarboxylic acids can have various substituents as side chains at the α-position of the hydroxyl group, as in amino acids. When an α-hydroxycarboxylic acid has the aforementioned side chain, the carbon atom attached to the side chain can be an asymmetric center. The three-dimensional structure of an α-hydroxycarboxylic acid can correspond to the L- or D-form of an amino acid. The choice of side chain is not particularly limited and is freely selected from, for example, alkyl groups, alkenyl groups, alkynyl groups, aryl groups, heteroaryl groups, aralkyl groups, cycloalkyl groups, etc., optionally substituted. The number of substituents is not limited to one and can be two or more. For example, the substituent has an S atom and can further have functional groups such as amino groups or halogen groups.

[0187] The "amino acids" and "amino acid analogs" that constitute a peptide include all combinations of isotopes corresponding to each atom. An isotope of an "amino acid" and "amino acid analog" is a form in which at least one atom is replaced by an atom with the same atomic number (number of protons) but a different mass number (total number of protons and neutrons). Examples of atoms contained in the "amino acids" or "amino acid analogs" that constitute a peptide include hydrogen atoms, carbon atoms, nitrogen atoms, oxygen atoms, phosphorus atoms, sulfur atoms, fluorine atoms, chlorine atoms, etc., and their isotopes include... 2 H, 3 H; 13 C 14 C; 15 N; 17 O、 18 O; 31 P, 32 P; 35 S; 18 F; 36 Cl et al.

[0188] As used herein, "peptide" is not limited, as long as two or more natural and / or non-natural amino acids are linked by amide bonds and may have ester bonds in a portion of the backbone, such as phenolic peptides. The number of amino acid residues contained in a peptide is preferably 5 to 30 residues, more preferably 7 to 20 residues. A peptide can be a linear peptide, a cyclic peptide, or a peptide with structures in which these are attached.

[0189] As used herein, "nucleic acid" includes deoxyribonucleic acid (DNA), ribonucleic acid (RNA), nucleotide derivatives with artificial bases, and peptide nucleic acids (PNA), and may be any mixture thereof, as long as the genetic information of interest is preserved. That is, the nucleic acids of this invention also include DNA-RNA hybrid nucleotides or chimeric nucleic acids, wherein different nucleic acids such as DNA and RNA are linked in a single strand.

[0190] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0191] (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety containing a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system;

[0192] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0193] (3) Contact the mixed nucleic acid display library with the target molecule; and

[0194] (4) Amplify the nucleic acid corresponding to the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers.

[0195] Step (1)

[0196] Step (1) is the step of preparing multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes.

[0197] As used in this article, a "nucleic acid display library" refers to a library in which peptides are associated with the RNA or DNA encoding those peptides. Peptides that bind to a desired target can be enriched by contacting the library with the desired target and washing away molecules that are not bound to the target (panning method). The sequence of the protein that binds to the target can be identified by analyzing the genetic information associated with the peptides selected through such a process. For example, methods utilizing the antibiotic puromycin (an analogue of aminoacyl-tRNA) for nonspecific binding to proteins in the ribosome-mediated translational elongation of mRNA have been reported as mRNA display (ProcNatl Acad Sci USA. 1997; 94: 12297-302. RNA-peptide fusions for the in vitro selection of peptides and proteins. Roberts RW, Szostak JW.) or in vitro viruses (FEBS Lett. 1997;414:405-8. In vitro virus: bonding of mRNA bearing puromycin at the 3'-terminal end to the C-terminal end of its encoded protein on the ribosome in vitro. Nemoto N, Miyamoto-Sato E, Husimi Y, Yanagawa H.).

[0198] The peptide-nucleic acid complex comprises a nucleic acid moiety and a peptide moiety, wherein the nucleic acid moiety contains a nucleic acid sequence encoding the nucleic acid of the peptide contained in the peptide moiety. The peptide moiety of the peptide-nucleic acid complex contains a predetermined amino acid, and the nucleic acid moiety contains a triplet encoding that amino acid.

[0199] As used herein, a "triad" refers to a sequence of three consecutive nucleobases in a nucleic acid sequence, usually written in order starting from the 5' end. Examples of triads when the nucleic acid is DNA include TTT, TTG, CTT, CTG, ATT, ATG, GTT, GTG, TCT, TCG, CCG, ACT, GCT, TAC, CAT, CAG, AAC, GAT, GAG, TGC, TGG, CGT, CGG, AGT, AGG, and GGT. These correspond to codons. When translated using a cell-free translation system, identical triads are translated into identical amino acids, resulting in peptides produced by translating mRNA containing multiple identical triads that will contain multiple identical amino acids.

[0200] The nucleic acid sequence encoding the peptide contained in the peptide moiety contains a random sequence. The random sequence in this paper is a sequence distinct for each nucleic acid sequence (a non-common sequence) and composed of arbitrarily selected amino acids and / or amino acid analogs. The random sequence may include at least one triplet (codon) specifying a non-natural amino acid. The random sequence preferably contains at least one, 2 to 1000, 2 to 750, 2 to 500, 2 to 250, 2 to 100, 2 to 50, 2 to 20, 2 to 18, 2 to 15, 2 to 12, at least 3, 3 to 1000, 3 to 750, 3 to 500, 3 to 250, 3 to 100, 3 to 50, 3 to 18, 3 to 15, 3 to 12, at least 4, 4 to 1000, 4 to 750, 4 to 500, 4 to 250, 4 to 100, 4 to 50, 4 to 18, preferably including 4 to 15, 4 to 12, at least 5, 5 1 to 1000, 5 to 750, 5 to 500, 5 to 250, 5 to 100, 5 to 50, 5 to 18, 5 to 15, or 5 to 12 triplets. As used herein, the number of triplets is referred to as the “number of repetitions.” The number of types of triplets can be 64 or less, 60 or less, 55 or less, 50 or less, 45 or less, 40 or less, 35 or less, 30 or less, 25 or less, 20 or less, 15 or less, 10 or less, or 5 or less, and preferably 64 or less, more preferably 50 or less, and particularly preferably 40 or less. Each triplet can be the same as or different from each other and can be randomly selected. Preferably, the peptide portion of the peptide-nucleic acid complex contains a predetermined amino acid, and the nucleic acid portion contains a triplet encoding that amino acid.

[0201] The peptide-nucleic acid complex can bind to the nucleic acid at the C-terminus of the peptide moiety. A linker may be present between the nucleic acid moiety and the peptide moiety of the peptide-nucleic acid complex. The peptide-nucleic acid complex preferably contains a linker formed from puromycin or a derivative thereof, a polymer of RNA, DNA, or hexaethylene glycol (spc18: 18-O-dimethoxytriphenylmethylhexaethylene glycol, 1-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide). Specifically, for example, the 3' end of the nucleic acid moiety and the C-terminus of the peptide moiety are linked via a linker containing puromycin, etc. The peptide-nucleic acid complex is preferably a peptide-DNA complex or a peptide-mRNA complex, more preferably a peptide-mRNA complex, and particularly preferably a cyclic peptide-mRNA complex. A schematic diagram illustrating an example of a peptide-nucleic acid complex is shown below.

[0202] [Formula 1]

[0203]

[0204] When using a cell-free translation system, it is preferable to include a nucleic acid sequence that includes a spacer (a linear portion contained in the peptide and linked to the nucleic acid) downstream of the target nucleic acid. Examples of spacer sequences include, but are not limited to, sequences containing glycine or serine.

[0205] Peptide-nucleic acid complexes can be produced, for example, by the method described in WO 2013 / 100132. Specifically, peptide-nucleic acid complexes can be produced by a method comprising the following steps:

[0206] A) Acylate dinucleotide pdCpA or pCpA with the desired amino acid or amino acid analog to provide aminoacyl pdCpA or aminoacyl pCpA;

[0207] B) Provide tRNAs lacking CA at the 3' end;

[0208] C) Link the aminoacyl pdCpA or aminoacyl pCpA from step A) to the CA-deficient tRNA from step B) to provide an initiator tRNA acylated with the desired amino acid or an amino acid analog.

[0209] D) Provides a cell-free translation system containing the initiator tRNA of step C) and free from at least one of methionine, methionyl tRNA synthetase (MetRS), methionine translation initiator tRNA, formyl donor, and methionyl tRNA transferase;

[0210] E) Provide template DNA in which the cysteine ​​codon UGU or UGC follows the translation initiation ATG downstream of the promoter, and a peptide sequence containing the anticodon of the tRNA corresponding to that in step C) is located further downstream.

[0211] F) Provide mRNA from the template DNA in step E).

[0212] G) Attach the adapter to the 3' end of the mRNA from step F), and

[0213] H) The mRNA that binds to the adapter from step G) is added to the cell-free translation system of step D) and translated to provide a pre-cyclized peptide-mRNA complex.

[0214] When the peptide is a cyclic peptide, a step of further cyclizing the pre-cyclized peptide-mRNA complex in step H) can be provided. In step H), a desulfurization reaction can be performed if necessary. Furthermore, after step G), a step of synthesizing cDNA using primers annealed to the 3' region of the mRNA can be included.

[0215] In steps A through C, RNA can be synthesized by preparing template DNA encoding the desired tRNA sequence, placing a T7, T3, or SP6 promoter upstream, and transcribing it using an RNA polymerase suitable for the promoter (such as T7 RNA polymerase and T3 and SP6 RNA polymerase). Alternatively, the target tRNA can be extracted and purified from cells and extracted using a probe with a complementary sequence to the tRNA sequence. Cells transformed with an expression vector of the target tRNA can then be used as the source. The RNA of the target sequence can also be synthesized chemically. For example, aminoacyl tRNA can be obtained by combining tRNA in which the CA is removed from the CCA sequence obtained in this manner at the 3' end with an aminoacylated pdCpA prepared separately using an RNA ligase (pdCpA method). Alternatively, full-length tRNA can be prepared and aminoacylated using flexizyme, a ribozyme for preparing active esters of various non-natural amino acids carried on tRNA. Acylated tRNA can also be obtained by using the methods described below.

[0216] In steps E through H, DNA with the desired nucleotide sequence positioned downstream of a promoter (such as the T7 promoter) is first chemically synthesized and used as a template to obtain double-stranded DNA via primer extension. The obtained double-stranded DNA is then transcribed into mRNA using an RNA polymerase (such as T7 RNA polymerase). A linker containing puromycin is attached to the 3' end of the transcribed mRNA, and the mRNA is translated into a protein using a known cell-free translation system (such as the PURE system). Puromycin is then incorporated into the protein, and the mRNA and its encoded protein are linked via puromycin. In this way, a peptide-mRNA complex in which the mRNA and its encoded product (peptide) are associated can be constructed.

[0217] Transfer RNA (tRNA) is a 73-93 base chain with a molecular weight of 25,000 to 30,000, containing a 3' CCA sequence. As an amino acid-linked tRNA, linked by esterification at the C-terminus of an amino acid at its 3' end, it forms a triplet with peptide elongation factor (EF-Tu) and GTP, and is transported to the ribosome. During the ribosome-mediated translation of the mRNA's nucleic acid sequence into an amino acid sequence, tRNA participates in codon identification by forming base pairs between the anticodon in the tRNA sequence and the mRNA codon. Intracellularly synthesized tRNA contains covalently modified bases, which may affect the conformation of tRNA and the formation of anticodon base pairs, and may contribute to codon identification via tRNA. Although tRNA synthesized by general in vitro transcription consists of the so-called nucleobases adenine, uracil, guanine, and cytosine, tRNA prepared intracellularly or chemically synthesized may contain modified bases, such as methylated bases, sulfur-containing derivatives, deamination derivatives, and adenosine derivatives containing isopentenyl and threonine, and may also contain deoxy bases when using methods such as pdCpA.

[0218] According to the general genetic code table (also known as the "codon table"), methionine is translated as the translation initiation amino acid. However, the method of translating the N-terminus as the desired amino acid can be used by translating with an acyl tRNA attached to the desired amino acid or an amino acid analog. It is known that the N-terminal introduction of non-natural amino acids allows for a higher amount of amino acid extension than extension, and amino acids or amino acid analogs with structures significantly different from those of natural amino acids can be used (see non-patent literature: J Am Chem Soc. 2009 Apr 15;131(14):5040-1. Translationinitiation with initiator tRNA charged with exotic peptides. Goto Y, Suga H.). For example, in this invention, the method for introducing a desired amino acid or amino acid analog other than methionine into the N-terminus includes the following steps. Acylated tRNA with the desired amino acid or amino acid analogue is added as translation initiation tRNA to a translation system lacking methionine, a formyl donor, or methionyltransferase and a translation initiation codon (e.g., ATG), encoding the amino acid or amino acid analogue, and translated to construct a pre-cyclized peptide with the amino acid or amino acid analogue at the end. The N-terminus can be diversified when using various combinations of translation initiation tRNA where the anticodon is not CAU with codons corresponding to the anticodons as the translation initiation tRNA and start codon. In other words, peptides or peptide libraries can be created prior to cyclization by aminoacylated multiple types of translation initiation tRNAs, respectively, with the desired amino acid, amino acid analogue, or N-terminal carboxylic acid analogue to different anticodons, and translating mRNAs or mRNA libraries with codons corresponding to those as start codons, where the N-terminal residues are not limited to one type. Specifically, peptides or peptide libraries can be created prior to cyclization using methods described, for example, in Mayer C et al., *Anticodon sequence mutants of Escherichia coli initiator tRNA: effects of overproduction of aminoacyl-tRNA synthetases, methionyl-tRNA formyltransferase, and initiation factor 2 onactivity in initiation*. *Biochemistry*. 2003, 42, 4787-99.(E. coli with a mutation that has an anticodon other than CAU in its initiating tRNA translates from amino acids other than f-Met and expresses a protein containing the same codon in the middle).

[0219] Peptides can be translated by adding mRNA to the PURE system, which contains a mixture of protein factors required for E. coli translation (methionyl-tRNA, formyltransferase, EF-G, RF1, RF2, RF3, RRF, IF1, IF2, IF3, EF-Tu, EF-Ts, ARS (selected as needed from AlaRS, ArgRS, AsnRS, AspRS, CysRS, GlnRS, GluRS, GlyRS, HisRS, IleRS, LeuRS, LysRS, MetRS, PheRS, ProRS, SerRS, ThrRS, TrpRS, TyrRS, ValRS), ribosomes, amino acids, creatine kinase, myokinase, inorganic pyrophosphatase, nucleoside diphosphate kinase, E. coli-derived tRNA, creatine phosphate, potassium glutamate, HEPES-KOH pH 7.6, magnesium acetate, spermidine, dithiothreitol, GTP, ATP, CTP, UTP, etc.). When T7 RNA is added... When polymerase is used, transcription and translation from template DNA containing a T7 promoter can also be performed in combination. In this case, peptides containing non-natural amino acid groups can be synthesized by adding a set of desired aminoacyl-tRNAs or a set of ARS-allowed non-natural amino acids (e.g., F-Tyr) to the system (Kawakami T et al. Ribosomal synthesis of polypeptoids and peptoid-peptide hybrids. J Am Chem Soc. 2008, 130, 16861-3., Kawakami T et al. Diverse backbone-cyclized peptides via codonre programming. Nat Chem Biol. 2009, 5, 888-90.).Alternatively, ribosomes and EF-Tu mutants can be used to increase the efficiency of introducing non-natural amino acids through translation (Dedkova LM et al., Construction of modified ribosomes for incorporation of D-amino acids into proteins. Biochemistry 2006, 45, 15541-51.; Doi Y et al., Elongation factor Tumutants expand amino acid tolerance of protein biosynthesis system. J Am Chem Soc. 2007, 129, 14458-62; Park HS et al., Expanding the genetic code of Escherichia coli with phosphoserine. Science 2011, 333, 1151-4.).

[0220] Cell-free translation systems are mixtures of tRNA, amino acids, ATP, and other translation-related factors extracted from the cells of the material. Cell-free translation systems do not contain living cells.Examples of cells used in the study include *Escherichia coli* (MethodsEnzymol. 1983;101:674-90. Prokaryotic coupled transcription-translation. Chen HZ, Zubay G.), yeast (J. Biol. Chem. 1979 254: 3965-3969. The preparation and characterization of a cell-free system from *Saccharomyces cerevisiae* that translates natural messenger ribonucleic acid. E Gasior, F Herrera, I Sadnik, CS McLaughlin, and K Moldave), wheat germ (Methods Enzymol. 1983;96:38-50. Cell-free translation of messenger RNA in a wheat germ system. Erickson AH, Blobel G.), and rabbit reticulocytes (Methods Enzymol. 1983;96:50-74. Preparation and use of nuclease-treated rabbit reticulocytes). lysates for the translation of eukaryotic messenger RNA. Jackson RJ, Hunt T.), HeLa cells (Methods Enzymol.1996;275:35-57. Assays for poliovirus polymerase, 3D(Pol), and authentic RNAreplication in HeLa S10 extracts.Barton DJ, Morasco BJ, Flanegan JB.) and insect cells (Comp Biochem Physiol B. 1989;93:803-6. Cell-free translation in lysates from Spodoptera frugiperda (Lepidoptera:Noctuidae) cells. Swerdel MR, FallonAM.).RNA polymerases, such as T7 RNA polymerase, can be added to the cells of the material to assemble from the transcription and translation of DNA. Meanwhile, the PURE system is a reconstructed cell-free translation system in which the protein factors, energy regeneration system enzymes, and ribosomes required for translation in *E. coli* are extracted and purified, and then mixed with tRNA, amino acids, ATP, GTP, etc. The PURE system not only has a low impurity content but also allows for the easy preparation of systems free of protein factors or amino acids that need to be eliminated, as it is a reconstructed system ((i) Nat Biotechnol. 2001; 19: 751-5. Cell-free translation reconstituted with purified components. Shimizu Y, Inoue A, Tomari Y, Suzuki T, Yokogawa T, Nishikawa K, Ueda T. (ii) Methods Mol Biol.2010; 607: 11-21. PURE technology. Shimizu Y, Ueda T.). Reconstructed cell-free translation systems can be created using appropriate known methods.

[0221] Various factors contained in cell-free translation systems, such as ribosomes and tRNA, can be purified from *E. coli* and yeast cells using methods well known to those skilled in the art. As tRNA and aminoacyl-tRNA synthases (also referred to herein as "ARS"), not only naturally occurring tRNAs can be used, but also artificial tRNAs and artificial aminoacyl-tRNA synthases that recognize non-natural amino acids. Peptides in which site-specific introduction of non-natural amino acids can be synthesized using artificial tRNAs and artificial aminoacyl-tRNA synthases can be used.

[0222] In the process of introducing non-natural amino acids into peptides through translation, orthogonal and efficient incorporation of tRNA into the ribosome is required for aminoacylation ((i) Biochemistry 2003; 42: 9598-608. Adaptation of anorthogonal archaeal leucyl-tRNA and synthetase pair for four-base, amber, and opal suppression. Anderson JC, Schultz PG., (ii) Chem Biol. 2003; 10: 1077-84. Using a solid-phase ribozyme aminoacylation system to reprogram the genetic code. Murakami H, Kourouklis D, Suga H.). Examples of methods for aminoacylation of tRNA may include the following.

[0223] Within the cell, an aminoacyl-tRNA synthetase is independently provided for each amino acid to acylate tRNA. Therefore, methods can be used to utilize certain aminoacyl-tRNA synthetases that allow non-natural amino acids such as N-Me-His, as well as methods for preparing and utilizing modified aminoacyl-tRNA synthetases that allow non-natural amino acids ((i) Proc. Natl. Acad. Sci. US A. 2002;99:9715-20. An engineered Escherichiacoli tyrosyl-tRNA synthetase for site-specific incorporation of an unnatural amino acid into proteins in eukaryotic translation and its application in awheat germ cell-free system. Kiga D, Sakamoto K, Kodama K, Kigawa T, MatsudaT, Yabuki T, Shirouzu M, Harada Y, Nakayama H, Takio K, Hasegawa Y, Endo Y, Hirao I, Yokoyama S.; (ii) Science 2003;301:964-7. An expanded eukaryotic genetic code. Chin JW, Cropp TA, Anderson JC, Mukherji M, Zhang Z, Schultz PG. Chin, JW.; (iii) Proc. Natl. Acad. Sci. US A. 2006;103:4356-61. Enzymatic aminoacylation of tRNA with unnatural amino acids. Hartman MC, Josephson K, Szostak JW. Alternatively, aminoacylation of tRNA followed by chemical modification of amino acids in the test tube can be used (J AmChem Soc. 2008; 130: 6131-6. Ribosomal synthesis of N-methyl peptides. Subtelny AO, Hartman MC, Szostak JW.).Aminoacyl-tRNA can also be obtained by binding tRNA in which CA is removed from the CCA sequence at the 3' end to an aminoacylated pdCpA prepared separately using RNA ligase (Biochemistry 1984; 23: 1468-73. T4 RNA ligase mediated preparation of novel "chemically misacylated" tRNAPheS. Heckler TG, Chang LH, Zama Y, NakaT, Chorghade MS, Hecht SM.). Aminoacylation can also be performed using a flexizyme, a ribozyme for preparing active esters of various non-natural amino acids carried on tRNA (J Am Chem Soc. 2002; 124: 6834-5. Aminoacyl-tRNA synthesis by a resin-immobilized ribozyme. Murakami H, Bonzagni NJ, Suga H.). Alternatively, tRNA and amino acid thioesters can be sonicated in cationic micelles (Chem Commun (Camb). 2005; (34): 4321-3. Simple and quickchemical aminoacylation of tRNA in cationic micellar solution underultrasonic agitation. Hashimoto N, Ninomiya K, Endo T, Sisido M.). Aminoacylation can also be achieved by adding an amino acid thioester that binds to PNA, which is complementary to the 3' end of the tRNA, to the tRNA (J Am Chem Soc. 2004; 126: 15984-9. In situ chemical aminoacylation with amino acid thioesters linked to a peptide nucleic acid. Ninomiya K, Minohata T, Nishimura M, Sisido M.).

[0224] While many methods have been reported using stop codons as codons to introduce unnatural amino acids, the PURE system described above can be used to construct synthetic systems that do not include natural amino acids and ARS, allowing the introduction of unnatural amino acids to replace excluded natural amino acids for codons encoding amino acids (J. Am. Chem. Soc. 2005;127:11727-35. Ribosomal synthesis of unnatural peptides. Josephson K, Hartman MC, Szostak JW.). Furthermore, by decoding the degeneracy of codons, unnatural amino acids can be added without excluding natural amino acids (Kwon I, et al. Breaking the degeneracy of the genetic code. J Am Chem Soc. 2003, 125, 7512-3.). Peptides containing N-methyl amino acids can be synthesized in ribosomes using cell-free translation systems such as the PURE system.

[0225] Besides mRNA display, known examples of display libraries utilizing cell-free translation systems include cDNA display in which a complex of peptide and puromycin is bound to cDNA encoding the peptide (Nucleic Acids Res. 2009;37(16):e108. cDNA display: a novel screening method for functional disulfide-rich peptides by solid-phase synthesis and stabilization of mRNA-protein fusions. Yamaguchi J, Naimuddin M, Biyani M, Sasaki T, Machida M, Kubo T, Funatsu T, Husimi Y, Nemoto N.), and ribosomal libraries utilizing relatively stable complexes of ribosomes and translation products during mRNA translation (Proc Natl Acad Sci US A. 1994;91:9022-6. An in vitro polysome display system for identifying ligands from very large peptide libraries. Mattheakis LC, Bhatt RR, Dower WJ.), including the covalent display of phage endonuclease P2A forming a covalent bond with DNA (Nucleic Acids Res. 2005;33:e10. Covalent antibody display--an in vitro antibody-DNA library selection system. Reiersen H,Lobersli I, Loset GA, Hvattum E, Simonsen B, Stacy JE, McGregor D, Fitzgerald K, Welschof M, Brekke OH, Marvik OJ.) and the CIS display utilizing the binding of the replication initiation protein RepA of microbial plasmids to the replication origin ori (Proc Natl Acad Sci US A. 2004;101:2806-10. CIS display: In vitro selection of peptides from libraries of protein-DNA complexes).Odegrip R, Coomber D, Eldridge B, Hederer R, Kuhlman PA, Ullman C, FitzGerald K, McGregor D.). In vitro compartmentalization is also known, in which the transcription-translation system is encapsulated in an oil-in-water emulsion or liposome for each DNA molecule constituting the DNA library, and the translation reaction takes place (NatBiotechnol. 1998;16:652-6. Man-made cell-like compartments for molecular evolution. Tawfik DS, Griffiths AD.). Display libraries can be created using appropriate known methods.

[0226] As used herein, a "barcode sequence" (also known as a barcode) refers to a nucleic acid sequence that can be detected and identified. Barcode sequences can have, for example, 10 to 100 nucleotides and can be incorporated into various nucleic acids. Therefore, in a library, nucleic acids incorporated with barcode sequences can be identified or grouped using barcodes.

[0227] As used herein, "barcode" refers to the presence of a barcode sequence. When a peptide-nucleic acid complex is barcoded, the barcode sequence can be identified by using a correspondence between a specific library and the barcode sequence, thereby identifying libraries containing peptide hits during screening. Once the library is identified, it becomes easier to identify nucleic acid sequences encoding peptides that readily bind to the target. The barcode sequence can contain an identifying barcode sequence and / or a common barcode sequence, and the barcode sequence can be a combination of both types of identifying barcode sequences and common barcode sequences. The identifying barcode sequence is different for each nucleic acid display library, allowing the identification of nucleic acid display libraries that derive from barcoded peptide-nucleic acid complexes containing the barcode sequence. The common barcode sequence is a barcode sequence common to any nucleic acid display library, and if the entire mixture of nucleic acid display libraries is to be amplified by PCR, the entire nucleic acid can be amplified by using primers annealed to the common barcode sequence.

[0228] The method used for barcoding is not particularly limited, but a barcoded peptide-nucleic acid complex can be obtained by reverse transcription of a nucleic acid sequence using a primer containing a barcoded sequence (also called a first barcoded primer or first primer) and reverse transcriptase to prepare a double strand of the nucleotide chain of the nucleic acid portion of the peptide-nucleic acid complex and a nucleotide chain extended with the primer (complementary DNA (cDNA)). The barcoded sequence is attached to the 5' end of the cDNA. The resulting double strand is a DNA / mRNA double strand or a DNA / DNA double strand, preferably a cDNA / mRNA double strand. Because the nucleic acid portion of the peptide-nucleic acid complex is double-stranded, the stability of the peptide-nucleic acid complex is improved, and the peptide-nucleic acid complex is easily and stably handled under various conditions such as amplification, replication, transcription, reverse transcription, and purification. Other methods for barcoding include introducing barcodes during the preparation of a nucleic acid (DNA) library or during ligation with puromycin (ligation). For example, for a DNA library, PCR is performed using primers containing a barcoded sequence, followed by transcription to prepare mRNA containing the barcoded sequence. The mRNA is then linked to a peptide using a puromycin adapter, as shown in Figure 1 of Douthwaite JA et al., "Ribosome Display and Related Technologies: Methods and Protocols" (published by Humana, 2011, pp. 113-135), or the mRNA containing the barcoded sequence is introduced into the nucleic acid portion of the puromycin adapter and linked to the peptide. The barcoded peptide-nucleic acid complex can then be obtained by translating the mRNA. The nucleic acid is DNA, a DNA / mRNA duplex, or a DNA / DNA duplex, preferably cDNA, a cDNA / mRNA duplex, or a cDNA / cDNA duplex.

[0229] The first barcode primer is preferably a primer containing an identification barcode sequence and a common barcode sequence, and the common barcode sequence is preferably located at the 5' end of the identification barcode sequence.

[0230] The number of bases in the first barcode primer is not limited, as long as it can function as a primer, and can be, for example, 15 to 150, 15 to 140, 15 to 130, 15 to 120, 15 to 110, 15 to 100, 15 to 90, 15 to 80, 15 to 70, 15 to 60, 15 to 50, 15 to 40, 15 to 30, 25 to 150, 25 to 140, 25 to 130, 25 to 120, 25 to 110, 25 to 100, 25 to 90, 25 to 80, 25 to 70, 25 to 60, 25 to 50, 25 to 40, or 25 to 30. The number of bases in the first barcode primer is preferably 25 to 150, more preferably 25 to 100, and particularly preferably 25 to 60.

[0231] The number of bases in the identification barcode sequence is not limited, as long as it can perform the function of a barcode sequence, and can be, for example, 15 to 30, 15 to 27, 15 to 25, 15 to 22, 15 to 20, 15 to 17, 17 to 30, 17 to 27, 17 to 25, 17 to 22, 17 to 20, 20 to 30, 20 to 27, 20 to 25, 20 to 22, 22 to 30, 22 to 27, 22 to 25, or 25 to 30. The number of bases in the identification barcode sequence is preferably 15 to 30, more preferably 17 to 25, and particularly preferably 17 to 22.

[0232] The number of bases in the common barcode sequence is not limited, as long as it can perform the function of a barcode sequence, and can be, for example, 15 to 30, 15 to 27, 15 to 25, 15 to 22, 15 to 20, 15 to 17, 17 to 30, 17 to 27, 17 to 25, 17 to 22, 17 to 20, 20 to 30, 20 to 27, 20 to 25, 20 to 22, 22 to 30, 22 to 27, 22 to 25, or 25 to 30. The number of bases in the common barcode sequence is preferably 15 to 30, more preferably 17 to 25, and particularly preferably 17 to 22.

[0233] In this embodiment, a cell-free translation system is used to independently translate template mRNA to prepare multiple corresponding nucleic acid display libraries. Even when using the same template mRNA, different peptides can be prepared depending on the cell-free translation system used during translation. The cell-free translation system used to prepare multiple nucleic acid display libraries contains amino acid combinations that differ from each other in the multiple nucleic acid display libraries. In this embodiment, when using different template mRNAs (e.g., mRNAs with different numbers of random sequence repeats), different peptides can be prepared even when using the same cell-free translation system used during translation.

[0234] In one aspect, multiple nucleic acid display libraries are produced for their respective translations based on different genetic code tables. In another aspect, multiple nucleic acid display libraries are produced for their respective translations based on the same genetic code table. The nucleic acid sequence of the template mRNA and the amino acid sequence of the peptide as the translation product depend on the codon table in the cell-free translation system used. In one aspect, in multiple nucleic acid display libraries, the codon-amino acid correspondences of at least one, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, or 15 or more are different from each other. In one aspect, in multiple nucleic acid display libraries, the codon-amino acid correspondences of 48 or fewer, 45 or fewer, 40 or fewer, 35 or fewer, 30 or fewer, 25 or fewer, 20 or fewer, or 15 or fewer are different from each other. In multiple nucleic acid display libraries, the codon-amino acid correspondences differ between each other, ranging from 1 or more to 48 or fewer, between 1 or more to 45 or fewer, between 1 or more to 40 or fewer, between 1 or more to 35 or fewer, between 1 or more to 30 or fewer, between 1 or more to 25 or fewer, between 1 or more to 20 or fewer, between 2 or more to 48 or fewer, between 2 or more to 40 or fewer, between 2 or more to 35 or fewer, between 2 or more to 30 or fewer, between 2 or more to 25 or fewer, or between 2 or more to 20 or fewer. Preferably, in the multiple nucleic acid display libraries, there are 1 or more, 48 or fewer, or 1 or more and 48 or fewer; more preferably, there are 1 or more, 40 or fewer, or 1 or more and 40 or fewer; particularly preferably, there are 2 or more, 30 or fewer, or 2 or more and 30 or fewer codons with different correspondences with each other.

[0235] The number of nucleic acid display libraries provided in step (1) is at least 2, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, 120 or more, 130 or more, 140 or more, 150 or more, 160 or more, 170 or more, 180 or more, 190 or more, 200 or more, 210 or more, 220 or more, 230 or more, 240 or more, 250 or more, 260 or more, 270 or more, 280 or more, 290 or more, 300 or more, 310 or more, 320 or more, 330 or more, 340 or more, or 350. Or more. The number of nucleic acid display libraries provided in step (1) is 400 or less, 390 or less, 380 or less, 370 or less, 360 or less, 350 or less, 340 or less, 330 or less, 320 or less, 310 or less, 300 or less, 290 or less, 280 or less, 270 or less, 260 or less, 250 or less, 240 or less, 230 or less, 220 or less, 210 or less, 200 or less, 190 or less, 180 or less, 170 or less, 160 or less, 150 or less, 140 or less, 130 or less, 120 or less, 110 or less, 100 or less, 90 or less, 80 or less, 70 Or less, 60 or less, 50 or less, 40 or less, 30 or less, 20 or less, or 10 or less.The number of nucleic acid display libraries provided in step (1) is 2 or more and 400 or less, 2 or more and 200 or less, 2 or more and 100 or less, 2 or more and 50 or less, 2 or more and 30 or less, 2 or more and 15 or less, 2 or more and 10 or less, 3 or more and 400 or less, 3 or more and 200 or less, 3 or more and 100 or less, 3 or more and 50 or less, 3 or more and 30 or less, 3 or more and 15 or less, 3 or more and 10 or less, 4 or more and 400 or less, 4 or more and 200 or less, 4 or more and 100 or less, 4 or more and 50 or less, 4 or more and 30 or less, 4 or more and 15 or less, or 4 or more and 10 or less. Preferably, the number of nucleic acid display libraries provided in step (1) is 2 or more, 400 or less, or 2 or more and 400 or less; more preferably 2 or more, 100 or less, or 2 or more and 100 or less; particularly preferably 2 or more, 50 or less, or 2 or more and 50 or less; most preferably 2 or more, 10 or less, or 2 or more and 10 or less.

[0236] The nucleic acid display library has 10 3 or more, 10 4 or more, 10 5 or more, 10 6 or more, 10 7 or more, 10 8 or more, 10 9 or more, or 10 10 or more, 10 11 or more, or 10 12 Libraries with greater or more diversity. This diversity is not limited to measured values ​​and is theoretical. As used herein, “diversity” means the types of peptide-nucleic acid complexes contained in a nucleic acid display library, and theoretically, the upper limit depends on the type of random sequence in the template DNA used.

[0237] When determining the "theoretical (total) number of variations" of peptide-nucleic acid complexes contained in the nucleic acid display library of this disclosure, those that cannot actually be manufactured as peptide-nucleic acid complexes are not included in the theoretical (total) number. For example, in the following cases (1) and (2), the corresponding amino acids are not translated and synthesized, and therefore peptide-nucleic acid complexes containing such amino acids are not included in the theoretical (total) number: (1) when the amino acid is added as material for the peptide to the translation solution, but does not contain a nucleic acid encoding the same amino acid as a template; (2) when the translation solution used contains a nucleic acid base sequence that serves as a template for the peptide, but does not contain the corresponding amino acid of the peptide. In addition, if there are by-reactants or unreacted compounds in the process of producing cyclic peptide compounds, those are not included in the theoretical (total) number.

[0238] For example, if a peptide with 11 residues is synthesized from a base sequence with randomly arranged codons corresponding to a single amino acid using a codon table of one set of codons, then the theoretical number of variants is 20 when the translation initiation amino acid is fixed to methionine. 10 This is the upper limit.

[0239] The nucleic acid display library contains 10 3 One or more, 10 4 One or more, 10 5 One or more, 10 6 One or more, 10 7 One or more, 10 8 One or more, 10 9 One or more, 10 10 One or more, 10 11 One or more, or 10 12 A library of one or more peptide-nucleic acid complexes.

[0240] Step (2)

[0241] Step (2) is the step of mixing the multiple nucleic acid display libraries obtained in step (1) to prepare a mixed nucleic acid display library.

[0242] In step (2), at least 2, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, 120 or more, 130 or more, 140 or more, 150 or more, 160 or more, 170 or more, 180 or more, 190 or more, 200 or more, 210 or more, 220 or more, 230 or more, 240 or more, 250 or more, 260 or more, 270 or more, 280 or more, 290 or more, 300 or more, 310 or more, 320 or more, 330 or more, 340 or more, or 350 Mixing one or more nucleic acid display libraries to prepare a mixed nucleic acid display library. In step (2), 400 or less, 390 or less, 380 or less, 370 or less, 360 or less, 350 or less, 340 or less, 330 or less, 320 or less, 310 or less, 300 or less, 290 or less, 280 or less, 270 or less, 260 or less, 250 or less, 240 or less, 230 or less, 220 or less, 210 or less, 200 or less, 190 or less, 180 or less, 170 or less, 160 or less, 150 or less, 140 or less, 130 or less, 120 or less, 110 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 Mixing 1 or fewer, 50 or fewer, 40 or fewer, 30 or fewer, 20 or fewer, or 10 or fewer nucleic acid display libraries to prepare a mixed nucleic acid display library.In step (2), two or more and 400 or fewer, two or more and 200 or fewer, two or more and 100 or fewer, two or more and 50 or fewer, two or more and 20 or fewer, two or more and 15 or fewer, two or more and 10 or fewer, three or more and 400 or fewer, three or more and 200 or fewer, three or more and 100 or fewer, three or more and 50 or fewer, three or more and 20 or fewer, three or more and 15 or fewer, three or more and 10 or fewer, four or more and 400 or fewer, four or more and 200 or fewer, four or more and 100 or fewer, four or more and 50 or fewer, four or more and 20 or fewer, four or more and 15 or fewer, or four or more and 10 or fewer nucleic acid display libraries are mixed to prepare a mixed nucleic acid display library. In step (2), preferably two or more, 400 or fewer, or two or more and 400 or fewer; more preferably two or more, 100 or fewer, or two or more and 100 or fewer; particularly preferably two or more and 10 or fewer nucleic acid display libraries are mixed to prepare a mixed nucleic acid display library.

[0243] In step (2), more diverse libraries can be prepared from a single template RNA by mixing multiple nucleic acid display libraries to prepare mixed nucleic acid display libraries. For example, when using a single template RNA and 10 cell-free translation systems to prepare 10 types of nucleic acid display libraries (assuming a diversity of 10 for each nucleic acid display library), 3 Furthermore, when the resulting libraries were mixed to prepare a mixed nucleic acid display library, the diversity of the mixed nucleic acid display library was 10. 4 In other words, when performing a single screening operation using a mixed nucleic acid display library, the same results can be obtained as when screening each of the 10 nucleic acid display libraries individually, and therefore screening can be performed very efficiently. Furthermore, by selecting different barcode sequences for each library (or the cell-free translation system used in its preparation), the barcode sequences can be used as markers to identify which library (or the cell-free translation system used in its preparation) the peptide-nucleic acid complex is derived from. For example, a mixed nucleic acid display library contains peptide-nucleic acid complexes with different identification barcode sequences and a common barcode sequence, and each identification barcode sequence corresponds to a nucleic acid display library from which the peptide-nucleic acid complex is derived, as shown below.

[0244] [Equation 2]

[0245]

[0246] Step (3)

[0247] Step (3) involves contacting the mixed nucleic acid display library obtained in step (2) with the target molecule. When the library is contacted with the target molecule, the peptide portion of some barcoded peptide-nucleic acid complexes contained in the library binds to the target molecule. The target molecule can be immobilized in a manner well known to those skilled in the art. The peptide-nucleic acid complexes that bind to the target molecule (peptide-nucleic acid complexes with high affinity for the target molecule) can be selected (panning method) by contacting the mixed nucleic acid display library with the target molecule and washing away peptide-nucleic acid complexes that do not bind to the target molecule. Step (3) can be repeated multiple times, for example, two to five times. By repeating step (3), the accuracy of selection can be improved.

[0248] From the thus selected barcoded peptide-nucleic acid complex, the nucleic acid contained in the nucleic acid moiety (e.g., cDNA) is eluted. The nucleic acid contains a barcoded sequence derived from the first barcoded primer. The nucleic acid (cDNA) corresponding to the nucleic acid moiety can be amplified from the eluted nucleic acid by PCR. The obtained nucleic acid can be amplified by PCR and the nucleic acid sequence analyzed to identify the sequence of the peptide that binds to the target molecule. The method according to this embodiment may further include steps (3A) and / or (3B) between steps (3) and (4).

[0249] Step (3A)

[0250] Step (3A) is performed between steps (3) and (4) and is the step of eluting nucleic acids (e.g., cDNA) from a barcoded peptide-nucleic acid complex bound to the target molecule. The barcoded peptide-nucleic acid complex can be eluted by treatment with a specific enzyme, heating, application of a chemical stimulus, or irradiation with specific light. Examples of enzymes that can be used for elution include TEV proteases. When irradiating to elute nucleic acids, light with a wavelength of 300 nm or greater and 500 nm or less is preferred. Examples of chemical stimuli include treatment with an acid or base. The eluted nucleic acid contains a barcoded sequence derived from the first barcoded primer at its 5' end. The eluted nucleic acid can be converted into a double-stranded nucleic acid (e.g., double-stranded cDNA) by primer extension using a forward primer.

[0251] Step (3B)

[0252] Step (3B) is performed between steps (3) and (4) (e.g., after step (3A)) and is a step for identifying the amino acid sequence of the peptide moiety of the barcoded peptide-nucleic acid complex that binds to the target molecule. The amino acid sequence of the peptide moiety is identified based on the nucleic acid sequence and / or barcode sequence of the nucleic acid moiety of the barcoded peptide-nucleic acid complex that binds to the target molecule. Furthermore, the barcode sequence of the nucleic acid (e.g., cDNA) eluted in step (3A) can be used as a marker to identify the library of the derived peptide-nucleic acid complex (or the cell-free translation system used in its preparation). For example, when two types of template mRNA are translated using two types of cell-free translation systems to prepare a total of four nucleic acid display libraries, the chemical structure of the peptide moiety can be theoretically understood from the combination by using primers containing sequences complementary to the barcode sequences corresponding to the cell-free translation systems used for reverse transcription and identifying the nucleic acid sequence and barcode sequence of the nucleic acid moiety of the peptide-nucleic acid complex.

[0253] Based on the identified amino acid sequence information, peptides can be synthesized by any method. These peptides can be used to assess binding and inhibitory activity against target molecules, or to obtain the desired peptide through cell or animal evaluation.

[0254] In this disclosure, the target molecules used in the screening methods are not particularly limited, and examples include proteins, peptides, nucleic acids, sugars, lipids, etc., but target proteins are preferred. In one aspect, the screening methods of this disclosure can obtain peptides that inhibit protein-protein interactions (PPIs). Specifically, it is believed that the surfaces of both proteins involved in the PPI contain sites called hotspots, which enhance the binding force of the PPI. Examples of PPIs include hydrophobic interactions, electrostatic interactions, van der Waals forces, hydrogen bonds, photocrosslinking, and salt bridges. Such PPIs can induce protein phosphorylation, conformational changes or localization, the formation of complexes between proteins (e.g., multi-subunitization), etc., and can participate in protein modification, transport, folding changes, signal transduction, etc. While not intended to be bound by a particular theory, it is believed that the side chains of the amino acids constituting the peptides contained in the library of this disclosure can inhibit PPIs by entering these sites. In other words, according to the library and screening methods of this disclosure, it is believed that peptides that inhibit PPIs can be obtained regardless of the type of target molecule.

[0255] In a non-limiting aspect, the target molecules used in the screening methods of this disclosure are immobilized on a support for use. The support is not particularly limited if the target molecules can be immobilized, and examples include beads or resins. Target molecules can be immobilized to the support by known methods.

[0256] The target molecules of the screening methods in this disclosure are not particularly limited. Because the libraries in this disclosure include peptide compounds capable of specifically binding to a wide variety of target molecules, in one respect, it can be used to develop drugs for previously challenging targets in drug discovery.

[0257] Step (4)

[0258] Step (4) is to amplify the nucleic acid or its cDNA corresponding to the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule using primers containing a sequence complementary to the barcoded sequence. The amplified nucleic acid has the same nucleic acid sequence as the nucleic acid contained in the barcoded peptide-nucleic acid complex.

[0259] In step (4), the nucleic acid corresponding to the barcoded peptide-nucleic acid complex bound to the target molecule is amplified by PCR. First, the nucleic acid corresponding to the nucleic acid moiety is amplified by PCR using a forward primer (Fw primer). The Fw primer has a sequence complementary to the nucleic acid sequence in the 5' region of the nucleic acid moiety. When the obtained nucleic acid is amplified by PCR using a primer containing a sequence complementary to the common barcoded sequence (common barcoded primer), cDNA can be amplified regardless of differences in the nucleic acid display library of the derived cDNA. This series of reactions can be achieved by simultaneously adding the Fw primer and the common barcoded primer and repeating the amplification by PCR.

[0260] [Formula 3]

[0261]

[0262] On the other hand, when nucleic acids corresponding to the nucleic acid moiety are amplified by PCR using primers containing sequences complementary to the identification barcode sequence (also known as second barcode primers or second primers), cDNA can be selectively amplified for each nucleic acid display library of the derived cDNA, as shown below. Generally, in panning methods, one can see less enriched and more enriched nucleic acids. When enriched by panning, less enriched nucleic acids become relatively few, so when the less enriched nucleic acid is a peptide-nucleic acid complex that binds to the target molecule, it may be ignored. By selectively amplifying less enriched nucleic acids based on the identification barcode sequence of the barcoded peptide-nucleic acid complex that binds to the target molecule using PCR, the identification barcode sequence of the barcoded peptide-nucleic acid complex that binds to the target molecule can be amplified, thereby avoiding the loss of promising peptides.

[0263] [Formula 4]

[0264]

[0265] The second barcode primer is preferably a primer containing a sequence complementary to the identification barcode sequence, and may be a primer composed of sequences complementary to the identification barcode sequence. The sequence complementary to the identification barcode sequence contained in the second barcode primer is selected such that it corresponds to the nucleic acid display library to be selectively amplified. The second barcode primer can have 10 to 150, 10 to 140, 10 to 130, 10 to 120, 10 to 110, 10 to 100, 10 to 90, 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30, 10 to 20, 15 to 150, 15 to 140, 15 to 130, 15 to 120, 15 to 110, 15 to 100, 15 to 90, 15 to 80, 15 to 70, 15 to 60, 15 to 50, 15 to 40, 15 to 30, 25 to 150 Primers may have 1, 25 to 140, 25 to 130, 25 to 120, 25 to 110, 25 to 100, 25 to 90, 25 to 80, 25 to 70, 25 to 60, 25 to 50, 25 to 40, or 25 to 30 bases. The second barcode primer may be a primer having preferably 10 to 150, more preferably 10 to 50, and particularly preferably 15 to 30 bases.

[0266] In this embodiment, preferably, steps (1) to (4) constitute one cycle (also referred to as "one round"), and this cycle is repeated multiple times. The number of repetitions can be, for example, 2 or more cycles, 3 or more cycles, 4 or more cycles, 5 or more cycles, 6 or more cycles, 7 or more cycles, 8 or more cycles, 9 or more cycles, or 10 or more cycles. By repeating this cycle, nucleic acids of peptide-nucleic acid complexes that bind to the target can be enriched.

[0267] Step (5)

[0268] The method according to this embodiment may also include a step (5) following step (4). For example, step (5) may be performed after step (4) without starting the next cycle. Step (5) is a step of further amplifying the nucleic acid amplified in step (4) using forward primers (Fw primers) and reverse primers (Rv primers). The nucleic acid amplified in step (5) does not contain the identification barcode sequence. By using Fw primers and Rv primers for amplification via PCR, the identification barcode sequence can be removed, and only the nucleic acid corresponding to the peptide portion can be amplified.

[0269] The Fw primer has a sequence complementary to the nucleic acid sequence in the 5' region of the nucleic acid moiety, and the Rv primer has a sequence complementary to the nucleic acid sequence in the 3' region of the nucleic acid moiety. That is, the Fw and Rv primers are primers capable of annealing to the 5' and 3' regions of the nucleic acid moiety of the peptide-nucleic acid complex, respectively, and do not contain sequences complementary to the barcode sequence. Through step (5), the nucleic acid corresponding to the peptide moiety can be obtained without identifying the barcode sequence, and the next round (a series of steps (1) to (4)) can be performed. If the nucleic acid sequence corresponding to the peptide moiety can be identified, the desired peptide can be synthesized by translating the nucleic acid sequence using a cell-free translation system corresponding to the type of barcode sequence identified.

[0270] [Formula 5]

[0271]

[0272] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0273] (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety containing a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system;

[0274] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0275] (3) Contact the mixed nucleic acid display library with the target molecule;

[0276] (4) Amplify the nucleic acid corresponding to the nucleic acid moiety of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers; and

[0277] (6) Repeat steps (1) to (4) multiple times (multiple cycles).

[0278] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0279] (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety containing a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system;

[0280] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0281] (3) Contact the mixed nucleic acid display library with the target molecule;

[0282] (4) Amplify the nucleic acid corresponding to the nucleic acid moiety of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers; and

[0283] (5) Remove the identification barcode sequence contained in the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule.

[0284] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0285] (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety containing a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system;

[0286] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0287] (3) Contact the mixed nucleic acid display library with the target molecule;

[0288] (4) Amplify the nucleic acid corresponding to the nucleic acid portion of the barcoded peptide-nucleic acid complex bound to the target molecule using barcoded primers;

[0289] (5) Removing the identification barcode sequence from the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule; and

[0290] (6) Repeat steps (1) to (4) multiple times (multiple cycles).

[0291] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0292] (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety containing a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system;

[0293] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0294] (3) Contact the mixed nucleic acid display library with the target molecule;

[0295] (3A) Elution of nucleic acids from barcoded peptide-nucleic acid complexes bound to target molecules; and

[0296] (4) Amplify the nucleic acid corresponding to the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers.

[0297] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0298] (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety containing a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system;

[0299] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0300] (3) Contact the mixed nucleic acid display library with the target molecule;

[0301] (3A) Elution of nucleic acids from barcoded peptide-nucleic acid complexes bound to target molecules;

[0302] (4) Amplify the nucleic acid corresponding to the nucleic acid moiety of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers; and

[0303] (6) Repeat steps (1) to (4) multiple times (multiple cycles).

[0304] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0305] (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety containing a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system;

[0306] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0307] (3) Contact the mixed nucleic acid display library with the target molecule;

[0308] (3A) Elution of nucleic acids from barcoded peptide-nucleic acid complexes bound to target molecules;

[0309] (4) Amplify the nucleic acid corresponding to the nucleic acid moiety of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers; and

[0310] (5) Remove the identification barcode sequence contained in the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule.

[0311] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0312] (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety containing a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system;

[0313] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0314] (3) Contact the mixed nucleic acid display library with the target molecule;

[0315] (3A) Elution of nucleic acids from barcoded peptide-nucleic acid complexes bound to target molecules;

[0316] (4) Amplify the nucleic acid corresponding to the nucleic acid portion of the barcoded peptide-nucleic acid complex bound to the target molecule using barcoded primers;

[0317] (5) Removing the identification barcode sequence from the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule; and

[0318] (6) Repeat steps (1) to (4) multiple times (multiple cycles).

[0319] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0320] (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety containing a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system;

[0321] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0322] (3) Contact the mixed nucleic acid display library with the target molecule;

[0323] (3A) Elution of nucleic acids from barcoded peptide-nucleic acid complexes bound to target molecules;

[0324] (3B) Identifying the amino acid sequence of the peptide moiety of a barcoded peptide-nucleic acid complex that binds to a target molecule; and

[0325] (4) Amplify the nucleic acid corresponding to the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers.

[0326] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0327] (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety containing a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system;

[0328] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0329] (3) Contact the mixed nucleic acid display library with the target molecule;

[0330] (3A) Elution of nucleic acids from barcoded peptide-nucleic acid complexes bound to target molecules;

[0331] (3B) Identify the amino acid sequence of the peptide portion of a barcoded peptide-nucleic acid complex that binds to a target molecule;

[0332] (4) Amplify the nucleic acid corresponding to the nucleic acid moiety of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers; and

[0333] (6) Repeat steps (1) to (4) multiple times (multiple cycles).

[0334] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0335] (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety containing a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system;

[0336] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0337] (3) Contact the mixed nucleic acid display library with the target molecule;

[0338] (3A) Elution of nucleic acids from barcoded peptide-nucleic acid complexes bound to target molecules;

[0339] (3B) Identify the amino acid sequence of the peptide portion of a barcoded peptide-nucleic acid complex that binds to a target molecule;

[0340] (4) Amplify the nucleic acid corresponding to the nucleic acid moiety of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers; and

[0341] (5) Remove the identification barcode sequence contained in the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule.

[0342] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0343] (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety containing a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system;

[0344] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0345] (3) Contact the mixed nucleic acid display library with the target molecule;

[0346] (3A) Elution of nucleic acids from barcoded peptide-nucleic acid complexes bound to target molecules;

[0347] (3B) Identify the amino acid sequence of the peptide portion of a barcoded peptide-nucleic acid complex that binds to a target molecule;

[0348] (4) Amplify the nucleic acid corresponding to the nucleic acid portion of the barcoded peptide-nucleic acid complex bound to the target molecule using barcoded primers;

[0349] (5) Removing the identification barcode sequence from the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule; and

[0350] (6) Repeat steps (1) to (4) multiple times (multiple cycles).

[0351] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0352] (1) Multiple nucleic acid libraries were independently translated using a cell-free translation system to prepare peptide-nucleic acid complexes containing both nucleic acid and peptide moieties;

[0353] (1-1) The peptide-nucleic acid complex was barcoded to prepare multiple nucleic acid display libraries containing the barcoded peptide-nucleic acid complex;

[0354] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0355] (3) Contact the mixed nucleic acid display library with the target molecule; and

[0356] (4) Amplify the nucleic acid contained in the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers.

[0357] The definition of each term and the description of each step in this embodiment can be found in the description in the above embodiments.

[0358] In this article, a "nucleic acid library" is a collection of nucleic acids that serve as templates for preparing nucleic acid display libraries. The nucleic acids contained in a nucleic acid library can be DNA, mRNA, cDNA, or any combination thereof. A preferred nucleic acid library is an mRNA library.

[0359] In this embodiment, a peptide-nucleic acid complex containing both a peptide and a nucleic acid moiety can be prepared by translating multiple nucleic acid libraries using a cell-free translation system. The nucleic acid moiety of the resulting peptide-nucleic acid complex can be converted into a double-stranded form with cDNA by reverse transcription using primers. Alternatively, it can be converted into a barcoded peptide-nucleic acid complex by reverse transcription using primers containing a barcoded sequence. By combining multiple cell-free translation systems with multiple nucleic acid libraries during the preparation of the peptide-nucleic acid complex, display libraries of peptide-nucleic acid complexes with multiple modes can be prepared. Reverse transcription using primers containing a specific barcoded sequence associated with the cell-free translation system used (a first barcoded primer) allows for identification based on the barcoded sequence. Even when multiple nucleic acid display libraries are mixed to prepare a mixed nucleic acid display library, the library (or cell-free translation system) from which the barcoded peptide-nucleic acid complex is derived can be identified based on the barcoded sequence. In one aspect, the nucleic acid library has a different number of repeats than other nucleic acid libraries. In another aspect, the nucleic acid library may contain at least two or more nucleic acids with different numbers of repeats in a single nucleic acid library. In one respect, a nucleic acid library can be a library composed of nucleic acids having the same number of repeats in a single nucleic acid library.

[0360] The nucleic acid sequences contained in the nucleic acid library include random sequences. The random sequences in this paper are sequences distinct for each nucleic acid sequence (non-common sequences) and composed of arbitrarily selected amino acids and / or amino acid analogs. The random sequences may include at least one triplet (codon) specifying a non-natural amino acid. The random sequence preferably contains at least one, 2 to 1000, 2 to 750, 2 to 500, 2 to 250, 2 to 100, 2 to 50, 2 to 20, 2 to 18, 2 to 15, 2 to 12, at least 3, 3 to 1000, 3 to 750, 3 to 500, 3 to 250, 3 to 100, 3 to 50, 3 to 18, 3 to 15, 3 to 12, at least 4, 4 to 1000, 4 to 750, 4 to 500, 4 to 250, 4 to 100, 4 to 50, 4 to 18, preferably including 4 to 15, 4 to 12, at least 5, 5 1 to 1000, 5 to 750, 5 to 500, 5 to 250, 5 to 100, 5 to 50, 5 to 18, 5 to 15, or 5 to 12 triplets. The number of different types of triplets is preferably 64 or less, 60 or less, 55 or less, 50 or less, 45 or less, 40 or less, 35 or less, 30 or less, 25 or less, 20 or less, 15 or less, 10 or less, or 5 or less. Each triplet may be the same as or different from the others and may be randomly selected. Preferably, the peptide moiety of the peptide-nucleic acid complex contains a predetermined amino acid, and the nucleic acid moiety contains a triplet encoding that amino acid.

[0361] One embodiment of the present invention is a method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps:

[0362] (1) Prepare peptide-nucleic acid complexes by independently translating multiple nucleic acid libraries using a single cell-free translation system.

[0363] (1-1) The peptide-nucleic acid complex was barcoded to prepare multiple nucleic acid display libraries containing the barcoded peptide-nucleic acid complex;

[0364] (2) Mix the multiple nucleic acid display libraries to prepare a mixed nucleic acid display library;

[0365] (3) Contact the mixed nucleic acid display library with the target molecule; and

[0366] (4) Amplify the nucleic acid contained in the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers.

[0367] The definition of each term and the description of each step in this embodiment can be found in the description in the above embodiments.

[0368] One embodiment of the present invention is a method for producing nucleic acid display libraries, the method comprising the following steps:

[0369] (a) Design nucleic acid display libraries containing peptide-nucleic acid complexes using the results of principal component analysis with descriptors as indicators;

[0370] (b) Using a cell-free translation system to translate nucleic acid libraries to prepare peptide-nucleic acid complexes.

[0371] In this embodiment, descriptors are used as indicators to design nucleic acid display libraries containing peptide-nucleic acid complexes. Highly diverse libraries can be designed by performing principal component analysis on multiple descriptors and mapping the first and second principal components to a two-dimensional vector space (plotting the values ​​of the first principal component on the X-axis and the values ​​of the second principal component on the Y-axis), as shown in Figures 1 to 3. Compound information from the open-source library RDKit (https: / / www.rdkit.org / ) can be used to calculate descriptors. For example, descriptors can be tailored from those provided on the RDKit website (http: / / rdkit.org / docs / source / rdkit.Chem.html). Instances of descriptors include, but are not limited to, 'NumValenceElectrons' (number of valence electrons), 'PEOE_VSA3' (MOE Charge VSA descriptor 3 (-0.25≤ x < -0.20)), 'fr_Al_OH' (number of aliphatic hydroxyl groups), 'fr_Al_OH_noTert' (number of aliphatic hydroxyl groups, excluding tertiary-OH), 'fr_Ar_N' (number of aromatic nitrogen groups), and 'fr_C_O_noCOO' (carbonyl O). The number of 'fr_C_S' (excluding COOH), 'fr_HOCCN' (the number of C(OH)CCN-C tertiary-alkyl or C(OH)CCN rings), 'fr_NH1' (the number of secondary amines), 'fr_Ndealkylation2' (the number of tertiary alicyclic amines (without heteroatoms, non-quinine-bridged N)), 'fr_Nhpyrrole' (the number of H-pyrrole nitrogen), 'fr_SH' (the number of thiol groups), 'fr_alkyl_halide' (the number of alkyl halides), 'fr_allylic_oxid' (the number of allyl oxidation sites).Excluding steroid dienes), 'fr_aryl_methyl' (number of aryl methyl sites for hydroxylation), 'fr_azide' (number of azide groups), 'fr_azo' (number of azo groups), 'fr_bicyclic' (bicyclic), 'fr_diazo' (number of diazo groups), 'fr_dihydropyridine' (number of dihydropyridine), 'fr_epoxide' (number of epoxide rings), 'fr_ether' (number of ether oxides (including phenoxy groups)), and 'fr_furan' (number of furan rings). ), 'fr_halogen' (number of halogens), 'fr_hdzrine' (number of hydrazine groups), 'fr_hdrzone' (number of hydrazone groups), 'fr_imidazole' (number of imidazole rings), 'fr_imide' (number of imide groups), 'fr_isocyan' (number of isocyanates), 'fr_isothiocyan' (number of isothiocyanates), 'fr_ketone' (number of ketones), 'fr_ketone_Topliss' (number of ketones, excluding α,β-unsaturated diene ketones Cα). The following are the numbers of β-lactams, 'fr_lactam' (number of β-lactams), 'fr_lactone' (number of cyclic esters / lactones), 'fr_methoxy' (number of methoxy groups -OCH3), 'fr_morpholine' (number of morpholine rings), 'fr_nitrile' (number of nitriles), 'fr_nitro' (number of nitro groups), 'fr_nitro_arom' (number of nitrobenzene ring substituents), 'fr_nitro_arom_nonortho' (number of non-ortho-nitrobenzene ring substituents), 'fr_nitroso' (number of nitroso groups, excluding NO2), 'fr_oxazole' (number of oxazole rings), 'fr_phenol' (number of phenols), 'fr_phenol_noOrthoHbond' (number of phenol OH groups).Excluding intramolecular H bond substituents, 'fr_phos_acid' (number of phosphate groups), 'fr_phos_ester' (number of phosphate ester groups), 'fr_piperdine' (number of piperidine rings), 'fr_piperzine' (number of piperazine rings), 'fr_priamide' (number of primary amides), 'fr_prisulfonamd' (number of primary sulfonamides), 'fr_pyridine' (number of pyridine rings), 'fr_quatN' (number of quaternary nitrogen groups), 'fr_sulfide' (number of sulfide groups), 'fr_sulfonamd' (number of sulfonamides), 'fr_sulfone' (number of sulfone groups), 'fr_term_acetylenes' (number of terminal acetylenes), 'fr_tetrazole' (number of tetrazolium rings), 'fr_thiazole' (number of thiazole rings), 'fr_thiocyan' (number of thiocyanates), and 'fr_thiophene' (number of thiophenes). This method allows for the design of libraries with minimal structural overlap of peptide-nucleic acid complexes found in multiple nucleic acid display libraries.

[0372] The definition of each term and the description of each step in this embodiment can be found in the description in the above embodiments.

[0373] [Example]

[0374] Example 1: Using Split & Pool technology to perform selection from multiple PURE systems

[0375] The peptide display library translated by the PURE system (hereinafter also referred to as PURE), a recombinant cell-free protein synthesis system derived from multiple prokaryotes, was used to perform mRNA display panning to verify whether various peptide sequence groups that bind to target proteins can be obtained by this method.

[0376] In the first round of selection, peptide display libraries translated by four types of PURE were independently selected, and the selected libraries were amplified by PCR reaction.

[0377] In the second round and thereafter, the four types of selection libraries obtained in the previous round were translated using each PURE with four types of reverse transcription primers corresponding to each PURE used to prepare peptide displays. Then, mixed libraries using these libraries were selected, and the selection libraries corresponding to each PURE were then independently amplified by PCR using the four types of primers, and this process was repeated.

[0378] Example 1-1: Preparation of target proteins for panning

[0379] Glutathione S-transferase (GST), expressed and purified using *E. coli*, was used as the target protein for panning. GST-HisTEV-Bio was prepared from GST-HisTEVvi according to the method described in WO 2022 / 138892, wherein a His tag, a TEV protease cleavage tag, and a biotinylate recognition tag were added to the C-terminus.

[0380] Example 1-2: Preparation of peptide display libraries

[0381] Example 1-2-1: Synthesis of acylated tRNA for translation

[0382] The PUREs used in this selection are defined as PURE1, PURE2, PURE3, and PURE4. The acylated tRNAs used for these were prepared according to the methods described in WO 2018 / 143145 and WO 2018 / 225864.

[0383] In PURE 1, a mixture of elongation factor amino acidified tRNAs was prepared using 19 amino acids, consisting of Asp(SMe), Phe(3-Cl), Hph(3-Cl), MeAla(3-Pyr), MeGly, MeHnl(7-F2), MeHph, MeSer(nPr), MeSer(tBuOH), Nle, Pro(4-pip-4-F2), Pic(2), Ser(3-F-5-Me-Pyr), Ser(NtBu-Aca), Ser(Ph-2-Cl), Ser(iPen), cisPro(4-pip-4-F2), D-MeSer, and nBuGly.

[0384] In PURE2, a mixture of elongation factor amino acidified tRNAs was prepared using 23 amino acids, including Asp(SMe), MeAla, MeGly, MeHnl(7-F2), MeSer(nPr), MeSer(tBuOH), MeVal, Nle, Pro(4-pip-4-F2), Pic(2), Ser(NtBu-Aca), and D-MeSer.

[0385] In PURE3, a mixture of elongation factor amino acidified tRNAs was prepared using 21 amino acids, including Asp(SMe), MeAla(3-Pyr), MeGly, MeHnl(7-F2), MeSer(nPr), MeSer(tBuOH), Nle, Pic(2), cisPro(4-pip-4-F2), and nBuGly.

[0386] In PURE4, a mixture of elongation factor amino acidified tRNAs was prepared using 23 amino acids, including Asp(SMe), MeAla, MeGly, MeHnl(7-F2), MeSer(nPr), MeSer(tBuOH), MeVal, Nle, Pro(4-pip-4-F2), Pic(2), cisPro(4-pip-4-F2), and nBuGly.

[0387] The final concentration of each acylated tRNA in the translation solution was 10 μM to 20 μM. Pnaz-protected pCpA amino acids were extracted with phenol and then processed without deprotection. The initiator aminoacylated tRNA was the same compound as Acbz-MeCys(StBu)-tRNAfMetCAU described in WO 2018 / 225864, and it was added to the translation solution to a final concentration of 25 μM.

[0388] The relationship between the abbreviations of amino acids and their structures presented in this article is shown in Table 1.

[0389] [Table 1]

[0390]

[0391] Example 1-2-2: Randomized double-stranded DNA library encoding peptide library

[0392] The DNA library was constructed according to the method described in WO 2013 / 100132. In the prepared DNA library, 26 triplets of TTT, TTG, CTT, CTG, ATT, ATG, GTT, GTG, TCT, TCG, CCG, ACT, GCT, TAC, CAT, CAG, AAC, GAT, GAG, TGC, TGG, CGT, CGG, AGT, AGG, and GGT appeared randomly 8 or 9 times (SEQ ID NO: 1 and 2). These were used to prepare mRNA-puromycin adapter ligation products for use. In the sequences below, each "N" refers to any one of adenine, cytosine, guanine, and thymine, and there are 8 or 9 triplets represented by "NNN".

[0393] DNA library (SEQ ID NO: 1)

[0394] GTAATACGACTCACTATAGGGTTAACTTTAATAAGGAGATATAAATATG(NNN)8TAGCCGACCGGCACCGGCACCGGCAAAAAAA

[0395] DNA library (SEQ ID NO: 2)

[0396] GTAATACGACTCACTATAGGGTTAACTTTAATAAGGAGATATAAATATG(NNN)9TAGCCGACCGGCACCGGCACCGGCAAAAAAA

[0397] Example 1-2-3: Translation and Cycling of Peptide Display Libraries

[0398] The translation system used is the PURE system, a recombinant cell-free protein synthesis system derived from prokaryotes. Specifically, in the four types of PURE (PURE1 to 4), the translation reaction mixture was supplemented with 1 mM MTGTP, 1 mM ATP, 20 mM creatine phosphate, 50 mM HEPES-KOH pH 7.6, 100 mM potassium acetate, 2–4 mM magnesium acetate, 2 mM spermidine, 1 mM dithiothreitol, 1 mg / mL *E. coli* MRE 600 (RNase-negative)-derived tRNA (manufactured by F. Hoffmann-La Roche, Ltd.) (with some tRNA removed by the method described in *Nucleic Acids Research*, 2010, Vol. 38, No. 6 e89), 4 μg / mL creatine kinase, 6.72 U / mL kinase, 2 units / mL inorganic pyrophosphatase, 1.1 μg / mL nucleoside diphosphate kinase, 0.26 μM EF-G, 2.7 μM IF1, and 0.4 μM... IF2, 1.5 μM IF3, 40 μM EF-Tu, 59 μM EF-Ts, 1.2 μM ribosomes, 1 μM GlyRS, 0.04-0.4 μM IleRS, 0.16 μM ProRS, 0.09 μM ThrRS, 0.11 μM LysRS, 3 μM in vitro transcribed E. coli tRNA Ala1B, 250 μM glycine, 10-100 μM isoleucine, 250 μM proline, 250 μM threonine, 250 μM lysine, elongation factor aminoacylated tRNA mixture, 25 μM initiator aminoacylated tRNA (WO 2020 / 138336 A1), 0.5 μM penicillin G amidase (PGA), 1 μM EF-P-Lys, 0.4 units / μL RNasin (R) ribonuclease inhibitor (Promega) Corporation) and the mRNA-purinemycin-linked construct prepared from the above double-stranded DNA library according to WO 2013 / 100132 to 1 μM.

[0399] In PURE1, in addition to the above, 0.5 μM variant PheRS (WO 2016 / 148044), 1.37 μM AlaRS, 1 μM variant ValRS (WO 2016 / 148044), 1 μM variant SerRS (WO 2016 / 148044), 5 mM N-methylphenylalanine, 5 mM N-methylvaline, 5 mM N-methylserine, and 2.5 mM N-methylalanine were added. In PURE3, in addition to the above, 2.73 μM AlaRS, 1 μM variant ValRS (WO 2016 / 148044), 5 mM N-methylvaline, and 5 mM N-methylalanine were added. After mixing the above factors, the mixture was allowed to stand at 37°C for 1 hour for translation.

[0400] Example 1-3: Selection

[0401] According to WO 2013 / 100132, five rounds of selection were conducted using the translated presentation library.

[0402] Example 1-3-1: Demonstrating reverse transcription of a library

[0403] In the first round, 3 μM Rv primer 1-1 (SEQ ID NO: 3), 50 mM TrisHCl (pH 8.3), 75 mM KCl, 3 mM MgCl2, 0.5 mM dGTP, 0.5 mM dATP, 0.5 mM dCTP, 0.5 mM dTTP, and 8 U / μL M-MLV reverse transcriptase (H-) (Promega Corporation, M368B) were added to a total of four types of 100 μL display library solutions. These solutions were cyclized, desulfurized, and purified, and then brought to a final volume of 125 μL with nuclease-free water. The solutions were incubated at 42°C for 60 min for reverse transcription. A buffer of 1 x TBS and 2 mg / mL BSA (invitrogen) was added to this reverse transcription solution to prepare a 0.6 μM display library.

[0404] In the second round and thereafter, RT primer 1 (SEQ ID NO: 4), RT primer 2 (SEQ ID NO: 5), RT primer 3 (SEQ ID NO: 6), and RT primer 4 (SEQ ID NO: 7) were used as reverse transcription primers to translate, cyclize, desulfurize, and purify display libraries from PURE1, PURE2, PURE3, and PURE4 translation products, respectively. Specifically, 3 μM reverse transcription primers corresponding to the PURE type, 50 mM TrisHCl (pH 8.3), 75 mM KCl, 3 mM MgCl2, 0.5 mM dGTP, 0.5 mM dATP, 0.5 mM dCTP, 0.5 mM dTTP, and 8 U / μL M-MLV reverse transcriptase (H-) (Promega Corporation, M368B) were added to 24 μL of display library solution for translation, cyclization, desulfurization, and purification. The solution was then brought to a final volume of 30 μL with nuclease-free water and incubated at 42°C for 60 minutes for reverse transcription. A 0.25 μM display library was prepared by adding 1 x TBS and 2 mg / ml BSA (Invitrogen) buffer to this reverse transcription solution. Subsequently, display libraries generated by translating each of the four types of PURE were mixed to form a single display library solution.

[0405] Rv primer 1-1 (SEQ ID NO: 3)

[0406] TTTTTTTgccggtgccggtgccggtCGG

[0407] RT primer 1 (SEQ ID NO: 4)

[0408] AAACACGTGGCAAACATTCCAAGAATTACTGACCCCTCGGTTTTTTTgccggtgccggtg

[0409] RT primer 2 (SEQ ID NO: 5)

[0410] AAACACGTGGCAAACATTCCAAAGCACTCTTAGGCCTCTGTTTTTTTgccggtgccggtg

[0411] RT primer 3 (SEQ ID NO: 6)

[0412] AAACACGTGGCAAACATTCCAAGGACTGCATACCAGGTTGTTTTTTTgccggtgccggtg

[0413] RT primer 4 (SEQ ID NO: 7)

[0414] AAACACGTGGCAAACATTCCAAGGCCCAGAAGGATACAACTTTTTTTgccggtgccggtg

[0415] Example 1-3-2: Using four types of PURE libraries for selection

[0416] In the first round, GST-HisTev-Bio was immobilized on streptavidin-coated magnetic beads to a concentration of 0.4 μM and then added to four different types of display libraries, respectively, and reacted at 4 °C for 1 h. After the reaction, the beads were magnetically collected, the supernatant was removed, and the beads were washed several times with tTBS (manufactured by NACALAI TESQUE, INC.). The beads were then treated with a TEV protease that recognizes and cleaves the TEV protease recognition sequence to elute the nucleic acids.

[0417] Specifically, 1 x TEV buffer and 10 mM DTT, 0.1 U / μL AcTEV protease (ThermoFisher Scientific, Inc., #12575015) were added to the washed beads as TEV elution buffer, and the reaction was carried out. After the reaction, the supernatant was recovered, and PCR was performed using library recognition primers (Fw primer (SEQ ID NO: 8), Rv primers 1-2 (SEQ ID NO: 9)) to obtain the input library for the next round.

[0418] In addition, qPCR was performed using a portion of the TEV elution product, and the library recovery rate was evaluated for each round. For qPCR primers, the primers described above (Fw primer (SEQ ID NO: 8), Rv primers 1-3 (SEQ ID NO: 10)) were used. For the qPCR reaction solution, 1 x Ex Taq buffer, 0.2 mM dNTP, 0.5 μM Fw primer, 0.5 μM Rv primers 1-2 or 1-3, 200,000-fold diluted SYBR Green I (manufactured by Lonza KK, catalog number 50513), and 0.02 U / μL Ex Taq polymerase (manufactured by Takara Bio Inc.) were mixed, and PCR was performed by heating at 95°C for 2 min, followed by 40 cycles of 95°C for 10 s, 57°C for 20 s, and 72°C for 30 s.

[0419] Fw primer (SEQ ID NO: 8)

[0420] GTAATACGACTCACTATAGGGTTAACTTTAATAAGGAG

[0421] Rv primers 1-2 (SEQ ID NO: 9)

[0422] TTTTTTTgccggtgccggtgccggtCGGCTA

[0423] Rv primers 1-3 (SEQ ID NO: 10)

[0424] TTTGTGGTgccggtgccggtgccggtCGG

[0425] In rounds 2 through 4, GST-HisTev-Bio was immobilized on streptavidin-coated magnetic beads to a concentration of 0.4 μM and then added to a mixture of four display libraries. The mixture was incubated at 4 °C for 1 hour. After the reaction, the beads were magnetically collected, the supernatant was removed, and the beads were washed several times with tTBS (NACALAI TESQUE, INC.) as in round 1. The beads were then treated with TEV protease, which recognizes and cleaves the TEV protease recognition sequence, to elute the nucleic acids. After the reaction, the supernatant was recovered, and PCR was performed using library recognition primers (Fw primer (SEQ ID NO: 8), Rv primer 2 (SEQ ID NO: 11)).

[0426] Rv primer 2 (SEQ ID NO: 11)

[0427] AAACACGTGGCAAACATTCC

[0428] The PCR products were PCR-treated with Fw primer (SEQ ID NO: 8) and Rv primer 3 (SEQ ID NO: 12), Rv primer 4 (SEQ ID NO: 13), Rv primer 5 (SEQ ID NO: 14) or Rv primer 6 (SEQ ID NO: 15) to obtain four types of PCR products derived from each PURE.

[0429] Rv primer 3 (SEQ ID NO: 12)

[0430] AAGAATTACTGACCCCTCGG

[0431] Rv primer 4 (SEQ ID NO: 13)

[0432] AAAGCACTCTTAGGCCTCTG

[0433] Rv primer 5 (SEQ ID NO: 14)

[0434] AAGGACTGCATACCAGGTTG

[0435] Rv primer 6 (SEQ ID NO: 15)

[0436] AAGGCCCAGAAGGATACAAC

[0437] Subsequently, four types of PCR products were independently PCR-treated using primers (Fw primer (SEQ ID NO: 8), Rv primers 1-2 (SEQ ID NO: 9)) to obtain input libraries for the next round. PURE1, PURE2, PURE3, and PURE4 were used for those derived from Rv primer 3, Rv primer 4, Rv primer 5, and Rv primer 6, respectively, for the next round of translation.

[0438] In addition, qPCR was performed using a portion of the TEV elution product, and the recovery rate of each library was evaluated in each round. For qPCR primers, the primers described above were used (Fw primer (SEQ ID NO: 8) and Rv primer 3 (SEQ ID NO: 12), Rv primer 4 (SEQ ID NO: 13), Rv primer 5 (SEQ ID NO: 14) or Rv primer 6 (SEQ ID NO: 15)). For the qPCR reaction solution, 1 x Ex Taq buffer, 0.2 mM dNTP, 0.5 μM Fw primer, 0.5 μM Rv primer, 200,000-fold diluted SYBR Green I (manufactured by Lonza KK, catalog number 50513) and 0.02 U / μL Ex Taq polymerase (manufactured by Takara Bio Inc.) were mixed, and PCR was performed by heating at 95°C for 2 min and then repeating the cycle at 95°C for 10 s, 57°C for 20 s and 72°C for 30 s for 40 cycles.

[0439] To plot calibration curves to estimate library recovery, the input library was diluted to 1E+8 / μL, 1E+6 / μL, and 1E+4 / μL, and qPCR was performed under the same conditions.

[0440] In the second and subsequent rounds, the input libraries were divided into two groups: one for selection in the presence of the target protein and the other for selection in the absence of the target protein. The nucleic acid recovery rates of both groups were evaluated based on qPCR results. The recovery rate (%) for each round of selection was calculated using the following expression. The results are shown in Table 2.

[0441] Library recovery rate (%) = Output nucleic acid quantity (molecules) / Input nucleic acid quantity (molecules) × 100

[0442] [Table 2]

[0443]

[0444] Therefore, as shown in Table 2, there is a difference in recovery rates between the selection in the presence of the target protein and the selection in the absence of the target protein in the latter half of the round. Thus, molecules recovered only in the presence of the target protein are likely present in the library of each round and are enriched with each repeated round.

[0445] Example 1-3-3 Enrichment Sequence Analysis

[0446] The base sequences of the DNA pool selected in each round of screening were analyzed, and enrichment sequence analysis was performed.

[0447] Sequences whose 11th residue of the peptide generated by translation is Asp (SMe) in the DNA pool selected in the presence of the target protein and whose occurrence frequency is greater than 0.05% in at least one round, and whose occurrence frequency when the target is added is 10 times or more than that in the pool divided in the previous round and when the target is not added, are extracted as sequences enriched only in the presence of the target protein.

[0448] In this analysis, principal component analysis (PCA) was used to evaluate the characteristics of the enriched sequences. In PCA, 60 descriptors were calculated for each sequence (as shown below), and after standardizing each value, they were mapped to a vector space with a horizontal axis representing the first principal component and a vertical axis representing the second principal component. Compound information from the open-source library RDKit (https: / / www.rdkit.org / ) was used to calculate the descriptors.

[0449] Names of 60 descriptors

[0450] ‘NumValenceElectrons’, ‘PEOE_VSA3’, ‘fr_Al_OH’, ‘fr_Al_OH_noTert’, ‘fr_Ar_N’, ‘fr_C_O_noCOO’, ‘fr_C_S’, ‘fr_HOCCN’, ‘fr_NH1’, ‘fr_Ndealkylation2’, ‘fr_Nhpyrrole’, ‘fr_SH’, ‘fr_alkyl_halide’, ‘fr_allylic_oxid’, ‘fr_aryl_methyl’, ‘fr_azide’, ‘fr_azo’, ‘fr_bicyclic’, ‘fr_diazo’, ‘fr_dihydropyridine’, ‘fr_epoxide’, ‘fr_ether’, ‘fr_furan’, ‘fr_halogen’, ‘fr_hdrzine’, ‘fr_hdrzone’, ‘fr_imidazole’, ‘fr_imide’, ‘fr_isocyan’, ‘fr_isothiocyan’, ‘fr_ketone’, ‘fr_ketone_Topliss’, ‘fr_lactam’, ‘fr_lactone’, ‘fr_methoxy’, ‘fr_morpholine’, ‘fr_nitrile’, ‘fr_nitro’, ‘fr_nitro_arom’, ‘fr_nitro_arom_nonortho’, ‘fr_nitroso’, ‘fr_oxazole’, ‘fr_phenol’, ‘fr_phenol_noOrthoHbond’, ‘fr_phos_acid’, ‘fr_phos_ester’, ‘fr_piperdine’, ‘fr_piperzine’, ‘fr_priamide’, ‘fr_prisulfonamd’, ‘fr_pyridine’, ‘fr_quatN’, ‘fr_sulfide’, ‘fr_sulfonamd’, ‘fr_sulfone’, ‘fr_term_acetylene’, ‘fr_tetrazole’, ‘fr_thiazole’, ‘fr_thiocyan’ and ‘fr_thiophene’

[0451] As a control group for sequences enriched only in the presence of the target protein, 1000 sequences of 11-residue peptide random sequences capable of being synthesized from each combination of the four codon tables PURE1, PURE2, PURE3, and PURE4 were generated. Figure 1 is a graph of the values ​​of each principal component of the random sequences in each codon table (random 1: PURE1, random 2: PURE2, random 3: PURE3, random 4: PURE4). Figure 2 is a graph of the values ​​of each principal component of the enriched sequences in each codon table. Figure 3 shows the enriched sequences (shown in black) and all random sequences (shown in gray, where each shape corresponds to the following cell-free translation system: PURE1, PURE2, A graph of the values ​​of each principal component (PURE3, PURE4).

[0452] Since each numerical vector in Figures 1 to 3 is a vector representing the characteristics of an actual amino acid sequence, amino acid sequences that are similar to each other are represented as similar vectors.

[0453] In Figure 1, most values ​​for each sequence are close across the four types of random sequences, but some values ​​are unique for sequences generated from a single codon table. In other words, it can be confirmed that while random sequences generated from the four types of codon tables are generally close to each other's corresponding amino acid sequences, there are also amino acid sequences generated from only one codon table.

[0454] In Figures 2 and 3, the values ​​of random sequences and each sequence in the enriched sequences are close, but some values ​​in the enriched sequences are close to those distributed only in specific codon tables. In other words, the enriched sequences are close to the amino acid sequences generated by the four types of codon tables and also reflect the characteristics of sequences generated by the four types of codon tables. It can be confirmed that some peptides have different characteristics in the sequence, which is the result of using multiple PUREs. This confirms that using the Split & Pool technique with multiple PURE systems to screen peptides enriches various amino acid sequences that cannot be obtained from a single codon table.

[0455] Example 2: Using Split & Pool Techniques to Select from Multiple Document Libraries

[0456] The method was validated using mRNA display panning to determine whether various peptide sequence sets binding to the target protein could be obtained using peptide display libraries with random sequences of varying lengths. Furthermore, without using this method, mRNA display panning was performed using libraries obtained by pre-mixing random sequences of varying lengths, and the results were used as a basis for comparison.

[0457] When using the Split & Pool technique for selection, translation was performed using four types of libraries of different lengths and one type of PURE, and peptide display libraries were prepared using four types of reverse transcription primers corresponding to each library length. Then, selection was performed using mixed libraries of these libraries, and the selected libraries corresponding to each PURE were then independently amplified by PCR using the four types of primers, and this process was repeated.

[0458] In a selection process conducted under comparative conditions, four types of libraries were mixed and translated using one type of PURE primer, and peptide display libraries were prepared using one type of reverse transcription primer. Selection and PCR were then performed, and these procedures were repeated.

[0459] Example 2-1: Preparation of target proteins for panning

[0460] Glutathione S-transferase (GST), expressed and purified using *E. coli*, was used as the target protein for panning. GST-HisTEV-Bio was prepared from GST-HisTEVvi according to the method described in WO 2022 / 138892, wherein a His tag, a TEV protease cleavage tag, and a biotinylate recognition tag were added to the C-terminus.

[0461] Example 2-2: Preparation of peptide display libraries

[0462] Example 2-2-1: Synthesis of acylated tRNA for translation

[0463] Acylated tRNAs for these purposes were prepared according to the methods described in WO 2018 / 143145 and WO 2018 / 225864. A mixture of elongation factor amino acid-modified tRNAs was prepared using 19 amino acids, including Asp(SMe), Phe(3-Cl), Hph(3-Cl), MeAla(3-Pyr), MeGly, MeHnl(7-F2), MeHph, MeSer(nPr), MeSer(tBuOH), Nle, Pro(4-pip-4-F2), Pic(2), Ser(3-F-5-Me-Pyr), Ser(NtBu-Aca), Ser(Ph-2-Cl), Ser(iPen), cisPro(4-pip-4-F2), D-MeSer, and nBuGly. The final concentration of each acylated tRNA in the translation solution was from 10 μM to 20 μM. The pCpA amino acids protected by Pnaz were extracted with phenol and then processed without deprotection. The initiator aminoacylated tRNA was the same compound as Acbz-MeCys(StBu)-tRNAfMetCAU described in WO 2018 / 225864, and it was added to the translation solution to a final concentration of 25 μM.

[0464] Example 2-2-2: Randomized double-stranded DNA library encoding peptide library

[0465] Prepare synthetic DNA oligonucleotides as template DNA, wherein triplets of NNB are randomly repeated 5, 7, 9, or 11 times (IDT, Inc.) (SEQ ID NO: 16, 17, 18, 19). In the following sequences, "N" refers to any one of the four DNA bases A, T, G, and C, and each "N" may be the same or different from each other. "B" refers to any one of the three DNA bases T, G, and C, and each "B" may be the same or different from each other. Four double-stranded DNA libraries (SEQ ID NO: 21-24) were prepared by mixing each of the following: 0.01 μM synthetic oligonucleotide and 1 μM synthetic oligonucleotide 1 (SEQ ID NO: 20), Rv primers 1-2 (SEQ ID NO: 9), 1 x PrimeSTAR buffer, 0.2 mM dNTP, and 0.025 U / μL PrimeSTAR HS DNA polymerase (TakaraBio Inc.), heating at 95°C for 2 minutes, and then repeating the cycle for 10 cycles of 95°C for 10 seconds, 57°C for 20 seconds, and 72°C for 30 seconds. mRNA-puromycin adapter ligation products were prepared using these methods according to WO 2013 / 100132 for use.

[0466] 7-Residue DNA Oligonucleotide (SEQ ID NO: 16)

[0467] GAGATATAAATATG(NNB)5TAGCCGACCGGCACCGGC

[0468] 9-Residue DNA Oligonucleotide (SEQ ID NO: 17)

[0469] GAGATATAAATATG(NNB)7TAGCCGACCGGCACCGGC

[0470] 11-residue DNA library (SEQ ID NO: 18)

[0471] GAGATATAAATATG(NNB)9TAGCCGACCGGCACCGGC

[0472] 13-residue DNA library (SEQ ID NO: 19)

[0473] GAGATATAAATATG(NNB) 11 TAGCCGACCGGCACCGGC

[0474] Synthetic DNA oligonucleotide 1 (SEQ ID NO: 20)

[0475] GTAATACGACTCACTATAGGGTTAACTTTAATAAGGAGATATAAATATG

[0476] 7-residue DNA library (SEQ ID NO: 21)

[0477] GTAATACGACTCACTATAGGGTTAACTTTAATAAGGAGATATAAATATG(NNB)5TAGCCGACCGGCACCGGCACCGGCAAAAAAA

[0478] 9-residue DNA library (SEQ ID NO: 22)

[0479] GTAATACGACTCACTATAGGGTTAACTTTAATAAGGAGATATAAATATG(NNB)7TAGCCGACCGGCACCGGCACCGGCAAAAAAA

[0480] 11-residue DNA library (SEQ ID NO: 23)

[0481] GTAATACGACTCACTATAGGGTTAACTTTAATAAGGAGATATAAATATG(NNB)9TAGCCGACCGGCACCGGCACCGGCAAAAAAA

[0482] 13-residue DNA library (SEQ ID NO: 24)

[0483] GTAATACGACTCACTATAGGGTTAACTTTAATAAGGAGATATAAATATG(NNB) 11 TAGCCGACCGGCACCGGCACCGGCAAAAAAA

[0484] Example 2-2-3: Translation and Cycling of Peptide Display Libraries

[0485] The translation system used is the PURE system, a recombinant cell-free protein synthesis system derived from prokaryotes. Specifically, the following were added to the translation reaction mixture: 1 mM GTP, 1 mM ATP, 20 mM creatine phosphate, 50 mM MEPES-KOH pH 7.6, 100 mM potassium acetate, 4 mM magnesium acetate, 2 mM spermidine, 1 mM dithiothreitol, 1 mg / mL *E. coli* MRE 600 (RNase-negative)-derived tRNA (F. Hoffmann-La Roche, Ltd.) (with some tRNA removed by the method described in *Nucleic Acids Research*, 2010, Vol. 38, No. 6 e89), 4 μg / mL creatine kinase, 6.72 U / mL kinase, 2 units / mL inorganic pyrophosphatase, 1.1 μg / mL nucleoside diphosphate kinase, 0.26 μM EF-G, 2.7 μM IF1, 0.4 μM IF2, 1.5 μM IF3, and 40 μM... EF-Tu, 59 μM EF-Ts, 1.2 μM ribosomes, 1 μM GlyRS, 0.04 μM IleRS, 0.16 μM ProRS, 0.09 μM ThrRS, 0.11 μM LysRS, 0.5 μM variant PheRS (WO 2016 / 148044), 1.37 μM AlaRS, 1 μM variant ValRS (WO 2016 / 148044), 1 μM variant SerRS (WO 2016 / 148044), 3 μM in vitro transcribed E. coli tRNA Ala1B, 250 μM glycine, 10 μM isoleucine, 250 μM proline, 250 μM threonine, 250 μM lysine, 5 mM N-methylphenylalanine, 5 mM N-methylvaline, 5 mM N-methylserine, 2.5 mM The mixture of N-methylalanine, elongation factor aminoacylated tRNA, 25 μM initiator aminoacylated tRNA (WO 2020 / 138336 A1), 0.5 μM penicillin G amidase (PGA), 1 μM EF-P-Lys, 0.4 units / μL RNasin (R) ribonuclease inhibitor (Promega Corporation), and mRNA-purinemycin-ligated construct prepared from the above double-stranded DNA library according to WO 2013 / 100132 was added to 1 μM, and the mixture was then incubated at 37°C for 1 hour for translation.

[0486] Example 2-3: Selection

[0487] According to WO 2013 / 100132, five rounds of selection were conducted using the translated presentation library.

[0488] Example 2-3-1: Demonstrating reverse transcription of a library

[0489] During the screening using the Split & Pool technique, RT primer 5 (SEQ ID NO: 25), RT primer 2 (SEQ ID NO: 5), RT primer 6 (SEQ ID NO: 26), and RT primer 7 (SEQ ID NO: 27) were used as reverse transcription primers for translation, cyclization, desulfurization, and purification of display libraries from initial libraries containing 7, 9, 11, and 13 NNB residues, respectively. Specifically, 3 μM reverse transcription primers of type PURE, 50 mM TrisHCl (pH 8.3), 75 mM KCl, 3 mM MgCl2, 0.5 mM dGTP, 0.5 mM dATP, 0.5 mM dCTP, 0.5 mM dTTP, and 8 U / μL M-MLV reverse transcriptase (H-) (Promega Corporation, catalog number M368B) were added to a 24 L display library solution for translation, cyclization, desulfurization, and purification. The solution was then brought to a final volume of 30 L with nuclease-free water and incubated at 42 °C for 60 min for reverse transcription. A 0.25 μM display library was prepared by adding 1 x TBS and 2 mg / mL BSA (Invitrogen) buffer to this reverse transcription solution. Subsequently, in round 1, the four types of display libraries were mixed in a ratio of 7 residues: 9 residues: 11 residues: 13 residues = 1: 1: 10: 10 to form a display library solution. In the second round and thereafter, they were mixed in equal amounts for use.

[0490] RT primer 5 (SEQ ID NO: 25)

[0491] AAACACGTGGCAAACATTCCAAACCGGAGCCATACAGTACTTTTTTTgccggtgccggtg

[0492] RT primer 6 (SEQ ID NO: 26)

[0493] AAACACGTGGCAAACATTCCAACGATGATGCTCACTCTCGTTTTTTgccggtgccggtg

[0494] RT primer 7 (SEQ ID NO: 27)

[0495] AAACACGTGGCAAACATTCCAAGACGATCCGAGCCATTACTTTTTTTgccggtgccggtg

[0496] Under comparative conditions, a mixture of 7 residues: 9 residues: 11 residues: 13 residues in a ratio of 1:1:10:10 was prepared as the initial library. 3 μM Rv primer 1-1 (SEQ ID NO: 3), 50 mM TrisHCl (pH 8.3), 75 mM KCl, 3 mM MgCl2, 0.5 mM dGTP, 0.5 mM dATP, 0.5 mM dCTP, 0.5 mM dTTP, and 8 U / μL M-MLV reverse transcriptase (H-) (Promega Corporation, catalog number M368B) were added to a 24 μL display library solution that had been translated, cyclized, desulfurized, and purified from the initial library. The solution was then brought to a final volume of 30 μL with nuclease-free water and incubated at 42°C for 60 minutes for reverse transcription. Add 1 x TBS and 2 mg / mL BSA (invitrogen) buffer to the reverse transcription solution to prepare a 0.25 μM display library.

[0497] Example 2-3-2 Selection

[0498] In rounds 1 through 4, GST-HisTev-Bio was immobilized on streptavidin-coated magnetic beads at a concentration of 0.4 μM and then added to four different types of display libraries, respectively, and reacted at 4°C for 1 hour. After the reaction, the beads were magnetically collected, the supernatant was removed, and the beads were washed several times with tTBS (manufactured by NACALAI TESQUE, INC.). The beads were then treated with a TEV protease that recognizes and cleaves the TEV protease recognition sequence to elute the nucleic acids.

[0499] Specifically, 1 x TEV buffer and 10 mM DTT, 0.1 U / μL AcTEV protease (manufactured by ThermoFisher Scientific, Inc., #12575015) were added to the washed beads as TEV elution buffer, and the reaction was carried out. After the reaction, the supernatant was recovered.

[0500] When using the Split & Pool technique for screening, PCR was performed using library recognition primers (Fw primer (SEQ ID NO: 8) and Rv primer 2 (SEQ ID NO: 11)).

[0501] The PCR products were subjected to PCR using Fw primer (SEQ ID NO: 8) and Rv primer 7 (SEQ ID NO: 28), Rv primer 4 (SEQ ID NO: 13), Rv primer 8 (SEQ ID NO: 29), or Rv primer 9 (SEQ ID NO: 30) to obtain four types of PCR products derived from each library.

[0502] Rv primer 7 (SEQ ID NO: 28)

[0503] AAACCGGAGCCATACAGTAC

[0504] Rv primer 8 (SEQ ID NO: 29)

[0505] AACGATGATGCTCACTCTCG

[0506] Rv primer 9 (SEQ ID NO: 30)

[0507] AAGACGATCCGAGCCATTAC

[0508] Subsequently, PCR was performed independently on the four types of PCR products using primers (Fw primer (SEQ ID NO: 8), Rv primers 1-2 (SEQ ID NO: 9)) to obtain input libraries for the next round. For those derived from the initial DNA libraries of 7, 9, 11, and 13 residues, the subsequent operations were repeated as for those derived from Rv primer 7, Rv primer 4, Rv primer 8, and Rv primer 9, respectively.

[0509] In addition, qPCR was performed using a portion of the TEV elution product, and the recovery rate of each library was evaluated in each round. For qPCR primers, the primers described above were used (Fw primer (SEQ ID NO: 8) and Rv primer 7 (SEQ ID NO: 28), Rv primer 4 (SEQ ID NO: 13), Rv primer 8 (SEQ ID NO: 29) or Rv primer 9 (SEQ ID NO: 30)). For the qPCR reaction solution, 1 x Ex Taq buffer, 0.2 mM dNTP, 0.5 μM Fw primer, 0.5 μM Rv primer, 200,000-fold diluted SYBR Green I (manufactured by Lonza KK, catalog number 50513) and 0.02 U / μL Ex Taq polymerase (manufactured by Takara Bio Inc.) were mixed, and PCR was performed by heating at 95°C for 2 min and then repeating the cycle at 95°C for 10 s, 57°C for 20 s and 72°C for 30 s for 40 cycles.

[0510] When panning with mixed libraries, PCR is performed using library recognition primers (Fw primer (SEQ ID NO: 8), Rv primers 1-2 (SEQ ID NO: 9)) to obtain the input library for the next round.

[0511] In addition, qPCR was performed using a portion of the TEV elution product, and the library recovery rate was evaluated for each round. For qPCR primers, the primers described above (Fw primer (SEQ ID NO: 8), Rv primer 1-1 (SEQ ID NO: 3)) were used. For the qPCR reaction solution, 1 x ExTaq buffer, 0.2 mM dNTP, 0.5 μM Fw primer, 0.5 μM Rv primer, 200,000-fold diluted SYBR Green I (manufactured by Lonza KK, catalog number 50513), and 0.02 U / μL ExTaq polymerase (manufactured by Takara Bio Inc.) were mixed, and PCR was performed by heating at 95°C for 2 min, followed by 40 cycles of 95°C for 10 sec, 57°C for 20 sec, and 72°C for 30 sec.

[0512] To plot calibration curves to estimate library recovery, the input library was diluted to 1E+8 / μL, 1E+6 / μL, and 1E+4 / μL, and qPCR was performed under the same conditions.

[0513] In the second and subsequent rounds, the input libraries were divided into two groups: one for selection in the presence of the target protein and the other for selection in the absence of the target protein. The nucleic acid recovery rates of both groups were evaluated based on qPCR results. The recovery rate (%) for each round of selection was calculated using the following expression. The results are shown in Table 3.

[0514] Library recovery rate (%) = (Number of nucleic acids output from the library (number of molecules) / Number of nucleic acids input to the library (number of molecules)) × 100

[0515] [Table 3]

[0516]

[0517] Therefore, as shown in Table 3, there is a difference in recovery rates between the selection in the presence of the target protein and the selection in the absence of the target protein in the latter half of the round. Thus, molecules recovered only in the presence of the target protein are likely present in the library of each round and are enriched with each repeated round.

[0518] Example 2-3-3 Enrichment Sequence Analysis

[0519] Analyze the base sequences of the DNA pools selected in each round. Sequences differing by only one amino acid residue are grouped into the same cluster. The maximum NGS read count of each sequence is divided by the maximum NGS read count of the sequence with the maximum read count in the same cluster. The result is a value greater than 0.01. Sequences with a frequency greater than 0.05% and a frequency 10 times or more when the target was added in at least one round than when the pool was divided in the previous round and the target was not added during the selection process are extracted as sequences enriched only in the presence of the target protein, excluding sequences that appear due to NGS or PCR errors.

[0520] For each selection result, the number of extracted sequences is shown in Table 4. In Table 4, “AA” represents amino acid residues, and “7AA”, “9AA”, “11AA”, and “13AA” represent sequences that appeared in the 7-residue initial library, the 9-residue initial library, the 11-residue initial library, and the 13-residue initial library, respectively.

[0521] [Table 4]

[0522]

[0523] Therefore, as shown in Table 4, sequences derived from the 11-residue initial library were significantly enriched. However, in the mixed library screening, sequences derived from other initial libraries were scarce. In contrast, in the Split & Pool screening, even sequences derived from the 7-, 9-, and 13-residue initial libraries, which were difficult to enrich using mixed libraries, were found in large quantities. Therefore, Split & Pool screening can be used to find various candidate binding sequences.

Claims

1. A method for screening candidate peptides capable of binding to a target molecule, the method comprising the following steps: (1) Prepare multiple nucleic acid display libraries containing barcoded peptide-nucleic acid complexes, wherein the barcoded peptide-nucleic acid complexes contain a nucleic acid moiety and a peptide moiety, the nucleic acid moiety contains a barcoded sequence and a nucleic acid sequence encoding the peptide, and each of the multiple nucleic acid display libraries is an independently generated nucleic acid display library by translation using a cell-free translation system. (2) Mix the plurality of nucleic acid display libraries to prepare a mixed nucleic acid display library; (3) Contact the mixed nucleic acid display library with the target molecule; as well as (4) Amplify the nucleic acid corresponding to the nucleic acid portion of the barcoded peptide-nucleic acid complex that binds to the target molecule using barcoded primers.

2. The method of claim 1, wherein step (1) comprises translating a nucleic acid encoding a peptide to prepare a peptide and linking the peptide to the nucleic acid to prepare a peptide-nucleic acid complex.

3. The method according to claim 1 or 2, wherein step (1) comprises barcoding the peptide-nucleic acid complex to prepare a barcoded peptide-nucleic acid complex.

4. The method according to any one of claims 1 to 3, wherein the barcoding is a reverse transcription of the nucleic acid portion of the peptide-nucleic acid complex.

5. The method according to any one of claims 1 to 4, wherein the barcoding is performed by reverse transcription of the nucleic acid portion of the peptide-nucleic acid complex using a first primer containing a barcode sequence.

6. The method according to any one of claims 1 to 5, wherein the plurality of nucleic acid display libraries have different barcode sequences for each nucleic acid display library.

7. The method according to any one of claims 1 to 6, wherein each of the plurality of nucleic acid display libraries is an independently generated nucleic acid display library produced by translating each other using different cell-free translation systems.

8. The method according to any one of claims 1 to 7, wherein in step (2), two or more nucleic acid display libraries are mixed.

9. The method according to any one of claims 1 to 8, wherein the method comprises, between step (3) and step (4), a step (3A) of eluting the nucleic acid portion from the barcoded peptide-nucleic acid complex bound to the target molecule.

10. The method according to any one of claims 1 to 9, wherein the method further comprises, between step (3) and step (4), a step (3B) of identifying the amino acid sequence of the peptide portion of the barcoded peptide-nucleic acid complex that binds to the target molecule.

11. The method according to any one of claims 1 to 10, wherein in step (4), the nucleic acid of a barcoded peptide-nucleic acid complex having the same barcode sequence as the barcode sequence of the barcode primer is selectively amplified.

12. The method according to any one of claims 1 to 11, wherein the barcode primer is a second primer containing a barcode sequence.

13. The method according to any one of claims 1 to 12, wherein steps (1) to (4) constitute a loop, and the loop is repeated multiple times.

14. The method according to any one of claims 1 to 13, wherein the method further comprises, after step (4) or before starting the next cycle, a step (5) of further amplifying the nucleic acid amplified in step (4) with primers that do not contain a barcode sequence.

15. The method according to any one of claims 1 to 14, wherein the nucleic acid display library has 10 4 Or a more diverse library.

Citation Information

Patent Citations

  • Peptide-compound cyclization method

    WO2013100132A1

  • MODIFIED AMINOACYL-tRNA SYNTHETASE AND USE THEREOF

    WO2016148044A1

  • Method for synthesizing peptides in cell-free translation system

    WO2018143145A1

  • Cyclic peptide compound having high membrane permeability, and library containing same

    WO2018225864A1

  • Mutated trna for codon expansion

    WO2020138336A1