Compositions and Methods of Use of Target Binding Moieties

By designing the target binding part of polynucleotide barcoded, the problem of low library size and sensitivity of target binding molecules in the prior art is solved, and target binding with high library size and high sensitivity is achieved, which is suitable for antibody screening and disease diagnosis.

CN112236450BActive Publication Date: 2025-07-22GUANGZHOU CHENGYUAN BIOIMMUNOLOGY TECHNOLOGY CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN201980035598.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-06-19
Filing Date
2019-03-25
Publication Date
2025-07-22
Estimated Expiration
2039-03-25

AI Technical Summary

Technical Problem

The prior art faces the problems of small size and low sensitivity when preparing target-binding molecular libraries, especially when preparing peptide arrays, it is difficult to efficiently screen molecules bound to antibodies.

Method used

High library size and high sensitivity target binding by designing a polynucleotide barcoded target binding moiety, including spacer-separated first and second peptide sequences, enables it to bind simultaneously to the antigen-binding domain of the antibody, forming target binding units on soluble or solid support.

Benefits of technology

A target-binding molecular library with a large library size and high sensitivity can be realized, which can effectively screen peptide sequences that bind to antibodies and improve binding affinity, and is suitable for antibody profile analysis and disease diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112236450B_ABST
    Figure CN112236450B_ABST
Patent Text Reader

Abstract

The present disclosure provides compositions and methods for identifying binding elements (e.g., peptides, peptoids, or proteins) that can be bound by an immune receptor (e.g., an antibody). The binding element can be provided in a target binding unit that includes two binding elements separated by a spacer such that the two binding elements simultaneously bind to a single molecule that includes an antigen-binding domain of an antibody. The present disclosure provides various strategies for constructing the spacer. The identified binding elements can also be used to fabricate arrays that can be used to profile antibodies obtained from a blood sample.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference

[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 648,218, filed Mar. 26, 2018, and U.S. Provisional Patent Application No. 62 / 686,858, filed Jun. 19, 2018, each of which is incorporated herein by reference in its entirety. BACKGROUND OF THE INVENTION

[0003] Methods for interrogating antibody libraries have been developed. These methods fall into two broad categories: (1) sequencing the coding regions of antibodies of interest, and (2) using large molecule (e.g., peptide) libraries to examine which of these molecules can bind to the antibodies of interest. These two categories are referred to herein as "sequencing methods" and "binding methods," respectively.

[0004] With the help of NextGen Sequencing, sequencing methods can be relatively easily performed, but little is known about the molecules to which the antibodies of interest can bind.

[0005] Binding methods can include protein and peptide arrays. Cloning and expression of different regions of a target gene and purification of recombinant truncated proteins can be conventional methods for mapping antigenic epitopes. However, this process can be time-consuming, and some recombinant proteins may be difficult to purify. Due to low binding affinity, printing synthetic peptides on a solid support to prepare peptide arrays may have low sensitivity and may have a small library size. Thus, the success of binding methods may be partially limited because it may be challenging to prepare libraries of molecules (e.g., peptides) that are large enough and / or have high sensitivity. SUMMARY OF THE INVENTION

[0006] It is recognized herein that there is a need to generate libraries of target-binding molecules (e.g., peptides that bind to antibodies) that have a large library size and / or high sensitivity. According to one aspect of the present disclosure, there is provided a composition comprising a polynucleotide-barcode-tagged target-binding moiety, wherein the polynucleotide-barcode-tagged target-binding moiety comprises (a) a nucleic acid sequence that is linked via a linker to (b) a target-binding unit, the target-binding unit comprising (i) a first peptide sequence comprising a first binding region and (ii) a second peptide sequence comprising a second binding region; wherein the first binding region and the second binding region are separated by a spacer and are spaced apart such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule comprising the antigen-binding domain of an antibody; wherein the nucleic acid sequence encodes the first peptide sequence and / or the second peptide sequence; and wherein the composition is soluble.

[0007] According to another aspect of the present disclosure, there is provided herein a composition comprising a plurality of polynucleotide-barcode-target-binding moieties, each polynucleotide-barcode-target-binding moiety of the plurality of polynucleotide-barcode-target-binding moieties comprising (a) a nucleic acid sequence linked via a linker to (b) a target-binding unit, the target-binding unit comprising (i) a first peptide sequence comprising a first binding region, and (ii) a second peptide sequence comprising a second binding region; wherein the first binding region and the second binding region are separated by a spacer and are spaced apart such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule comprising an antigen-binding domain of an antibody; wherein the nucleic acid sequence of each polynucleotide-barcode-target-binding moiety of the plurality of polynucleotide-barcode-target-binding moieties is unique; and wherein the composition is soluble.

[0008] In some embodiments, the single molecule comprises a first antigen-binding domain and a second antigen-binding domain; wherein the first binding region and the second binding region are spaced apart such that the first binding region binds to the first antigen-binding domain and the second binding region binds to the second antigen-binding domain. In some embodiments, the first binding region and the second binding region have the same sequence. In some embodiments, the first binding region and the second binding region have the same structure recognized by the single molecule. In some embodiments, the nucleic acid sequence of each polynucleotide barcoded target-binding moiety among the plurality of polynucleotide barcoded target-binding moieties comprises a unique barcode sequence. In some embodiments, the nucleic acid sequence further comprises a barcode. In some embodiments, the nucleic acid sequence is single-stranded. In some embodiments, the nucleic acid sequence is double-stranded. In some embodiments, the nucleic acid sequence is deoxyribonucleic acid (DNA). In some embodiments, the nucleic acid sequence is ribonucleic acid (RNA). In some embodiments, the nucleic acid hybridizes to a primer. In some embodiments, the polynucleotide barcoded target-binding moiety comprises a single target-binding unit. In some embodiments, the polynucleotide barcoded target-binding moiety comprises two or more target-binding units. In some embodiments, the composition comprises a plurality of polynucleotide barcoded target-binding moieties. In some embodiments, each polynucleotide barcoded target-binding moiety among the plurality of polynucleotide barcoded target-binding moieties comprises a single target-binding unit. In some embodiments, at least one polynucleotide barcoded target-binding moiety among the plurality of polynucleotide barcoded target-binding moieties comprises a single target-binding unit. In some embodiments, two or more polynucleotide barcoded target-binding moieties among the plurality of polynucleotide barcoded target-binding moieties comprise a single target-binding unit. In some embodiments, each polynucleotide barcoded target-binding moiety among the plurality of polynucleotide barcoded target-binding moieties comprises two or more target-binding units. In some embodiments, at least one polynucleotide barcoded target-binding moiety among the plurality of polynucleotide barcoded target-binding moieties comprises two or more target-binding units. In some embodiments, two or more polynucleotide barcoded target-binding moieties among the plurality of polynucleotide barcoded target-binding moieties comprise two or more target-binding units.In some embodiments, the plurality of polynucleotide barcoded target binding moieties includes at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1x10. 6 or more polynucleotide barcoded target binding moieties. In some embodiments, the composition includes two or more target binding units. In some embodiments, the composition includes from 2 to 1000 target binding units. In some embodiments, the composition includes three or more target binding units. In some embodiments, the spacer is a polymer. In some embodiments, the spacer is polyethylene glycol. In some embodiments, the spacer includes an identical amino acid sequence. In some embodiments, the spacer includes a folded polypeptide, secondary structure, and / or tertiary structure. In some embodiments, the spacer includes a coiled coil structure or a β-sheet structure. In some embodiments, the spacer includes two or more separate peptide chains, wherein at least one of the two or more separate peptide chains includes an α-helix or a β-strand. In some embodiments, the spacer includes a single peptide chain folded into at least two α-helices or at least two β-strands. In some embodiments, the single peptide chain includes four α-helices or four β-strands. In some embodiments, the spacer includes a first peptide chain, a second peptide chain, and a third peptide chain, wherein the second peptide chain interacts with a first portion of the first peptide chain, and wherein the third peptide chain interacts with a second portion of the first peptide chain. In some embodiments, the first peptide chain, the second peptide chain, and / or the third peptide chain fold into an α-helix. In some embodiments, the second peptide chain and the third chain are attached to opposite ends of a double-stranded polynucleotide. In some embodiments, the spacer includes an oligonucleotide. In some embodiments, the oligonucleotide is a double-stranded oligonucleotide. In some embodiments, the oligonucleotide includes 20 to 40 nucleotides or base pairs.

[0009] In some embodiments, the first binding region and the second binding region include the same epitope. In some embodiments, the sequence of the first peptide sequence is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99% or 100% identical to the sequence of the second peptide sequence. In some embodiments, the nucleic acid sequence is a double-stranded DNA-RNA hybrid. In some embodiments, the linker includes puromycin or a derivative thereof. In some embodiments, the target binding portion of the polynucleotide barcode is a linear polymer chain or a branched polymer chain. In some embodiments, the first peptide sequence, the second peptide sequence and the spacer are adjacent in a single polypeptide chain. In some embodiments, the first peptide sequence and the second peptide sequence are connected to the spacer by a non-peptide bond. In some embodiments, the antigen binding domain is scFv, Fab or F(ab)2. In some embodiments, the length of the first peptide sequence and the second peptide sequence is at least 5 amino acid residues. In some embodiments, the polynucleotide barcoded target binding moiety or each polynucleotide barcoded target binding moiety of the plurality of polynucleotide barcoded target binding moieties is within a container. In some embodiments, each polynucleotide barcoded target binding moiety of the plurality of polynucleotide barcoded target binding moieties is within a different container of a plurality of containers. In some embodiments, the container is a droplet. In some embodiments, the droplet is a water-in-oil droplet. In some embodiments, the antibody is a biomarker.

[0010] According to another aspect of the present disclosure, a solid support is provided herein, comprising a plurality of discrete regions, wherein each of the plurality of discrete regions comprises a target binding portion attached thereto via a linker, wherein the target binding portion comprises (a) a first peptide sequence comprising a first binding region, and (b) a second peptide sequence comprising a second binding region; wherein the first binding region and the second binding region are separated by a spacer and are spaced a distance apart such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule comprising an antigen binding domain of an antibody. In some embodiments, the plurality of discrete regions comprises at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 10 3 , at least about 10 4 , at least about 10 5 or at least about 10 6Discrete regions. In some embodiments, the target binding moiety comprises a first reactive group. In some embodiments, the solid support comprises a second reactive group attached thereto. In some embodiments, the linker is generated by reacting the first reactive group with the second reactive group. In some embodiments, the spacer comprises a polymer chain. In some embodiments, the polymer chain comprises a polynucleotide, a polypeptide, or polyethylene glycol. In some embodiments, the polynucleotide is double-stranded DNA, double-stranded RNA, or a double-stranded DNA-RNA hybrid. In some embodiments, the polypeptide comprises a folded polypeptide, a secondary structure, and / or a tertiary structure. In some embodiments, the spacer comprises a coiled-coil structure or a β-sheet structure. In some embodiments, the spacer comprises two or more separate peptide chains, wherein at least one of the two or more separate peptide chains comprises an α-helix or a β-strand. In some embodiments, the spacer comprises a single peptide chain folded into at least two α-helices or at least two β-strands. In some embodiments, the single peptide chain comprises four α-helices or four β-strands. In some embodiments, the spacer comprises a first peptide chain, a second peptide chain, and a third peptide chain, wherein the second peptide chain interacts with a first portion of the first peptide chain, and wherein the third peptide chain interacts with a second portion of the first peptide chain. In some embodiments, the first peptide chain, the second peptide chain, and / or the third peptide chain fold into α-helices. In some embodiments, the second peptide chain and the third chain are attached to opposite ends of a double-stranded polynucleotide. In some embodiments, the single molecule comprises a first antigen-binding domain and a second antigen-binding domain; wherein the first binding region and the second binding region are spaced apart such that the first binding region binds to the first antigen-binding domain and the second binding region binds to the second antigen-binding domain. In some embodiments, the first binding region and the second binding region have the same sequence. In some embodiments, the first binding region and the second binding region have the same structure recognized by the single molecule. In some embodiments, the antigen-binding domain is an scFv, a Fab, or an F(ab)2. In some embodiments, the sequence of the first peptide sequence is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% or 100% identical to the sequence of the second peptide sequence. In some embodiments, the target binding moiety comprises two or more peptide sequences, each of the two or more peptide sequences comprising a binding region.

[0011] The present invention also provides methods for preparing target-binding moieties by RNA display. The present invention also provides methods for preparing target-binding moieties by in vitro compartmentalization. The present invention also provides methods for using target-binding moieties. The present invention also provides methods for profiling antibody mixtures using target-binding moieties. The present invention also provides methods for diagnosing diseases using target-binding moieties. In some embodiments, the disease is cancer or an autoimmune disease.

[0012] According to another aspect of the present disclosure, there is provided a method comprising translating an RNA sequence of an RNA, wherein the RNA sequence encodes a peptide sequence, wherein the RNA is linked at its 3' end to a peptide receptor; linking the peptide receptor to amino acid residues of a translated peptide comprising the peptide sequence, thereby forming a nucleic acid-peptide fusion molecule, wherein the nucleic acid-peptide fusion molecule comprises a polynucleotide barcoded target binding portion, which comprises: (a) a first peptide sequence comprising a first binding region, and (b) a second peptide sequence comprising a second binding region; wherein the first binding region and the second binding region are separated by a spacer and are spaced apart such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule comprising an antigen-binding domain of an antibody. In some embodiments, the method further comprises providing a DNA, wherein the DNA encodes the RNA. In some embodiments, the method further comprises transcribing the DNA. In some embodiments, the method further comprises reverse transcribing the RNA. In some embodiments, the DNA molecule and / or the RNA sequence encodes the first peptide sequence, the second peptide sequence, and the spacer. In some embodiments, the nucleic acid-peptide fusion molecule comprises a plurality of nucleic acid-peptide fusion molecules. In some embodiments, each nucleic acid-peptide fusion molecule of the plurality of nucleic acid-peptide fusion molecules comprises a unique nucleic acid sequence. In some embodiments, the first binding region and the second binding region have the same sequence. In some embodiments, the first binding region and the second binding region have the same structure recognized by the single molecule. In some embodiments, the spacer of each nucleic acid-peptide fusion molecule is the same or comprises the same amino acid sequence. In some embodiments, the spacer comprises a folded polypeptide, secondary structure, and / or tertiary structure. In some embodiments, the spacer comprises a coiled coil structure or a β-sheet structure. In some embodiments, the sequence of the first peptide sequence is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, or 100% identical to the sequence of the second peptide sequence. In some embodiments, the antigen-binding domain is a scFv, Fab, or F(ab)2. In some embodiments, the first peptide sequence and the second peptide sequence are 4 to 30 amino acids in length, 5 to 20 amino acids in length, 6 to 10 amino acids in length, 8 to 11 amino acids in length, or 10 to 20 amino acids in length. In some embodiments, the translation comprises in vitro translation.

[0013] In another aspect of the present disclosure, provided herein is a method comprising expressing a first peptide sequence encoded by a nucleic acid and a second peptide sequence encoded by a nucleic acid in each of a plurality of containers, wherein each of the plurality of containers comprises a scaffold, the scaffold comprising a first attachment site and a second attachment site separated by a spacer region; and binding the first peptide sequence to the first attachment site and binding the second peptide sequence to the second attachment site, wherein the first peptide sequence bound to the first attachment site and the second peptide sequence bound to the second attachment site are spaced apart such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule comprising an antigen-binding domain of an antibody, thereby forming a plurality of target-binding moieties. In some embodiments, the nucleic acid is linked to the scaffold via a linker. In some embodiments, the 5' end or 3' end of the nucleic acid is linked to the scaffold via the linker. In some embodiments, the nucleic acid is a single nucleic acid molecule. In some embodiments, the nucleic acid is linked to the scaffold before or after generating the plurality of containers. In some embodiments, the nucleic acid molecule is double-stranded or single-stranded. In some embodiments, the nucleic acid molecule is DNA, RNA, or a combination thereof. In some embodiments, the expression comprises transcription and / or translation. In some embodiments, the first peptide sequence and the second peptide sequence comprise the same sequence. In some embodiments, the sequence of the first peptide sequence is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% or 100% identical to the sequence of the second peptide sequence. In some embodiments, the method further comprises pooling the plurality of polynucleotide-barcode-target-binding moieties of the containers. In some embodiments, the containers are droplets. In some embodiments, the droplets are water-in-oil droplets. In some embodiments, it further comprises barcoding the target-binding moieties among the plurality of target-binding moieties. In some embodiments, barcoding comprises attaching a barcode to the target-binding moiety. In some embodiments, the scaffold is attached to a barcoded polynucleotide before expression. In some embodiments, the scaffold is attached to the nucleic acid encoding the first peptide sequence and / or the nucleic acid encoding the second peptide sequence. In some embodiments, the scaffold is attached to the nucleic acid encoding the first peptide sequence and / or the nucleic acid encoding the second peptide sequence before expression.

[0014] According to another aspect of the present disclosure, provided herein is a method comprising: contacting a mixture of antibodies with a population of target binding moieties, wherein each target binding moiety in the population comprises a target binding unit having a first binding region and a second binding region, wherein the first binding region and the second binding region are separated by a spacer and are spaced apart such that the first binding region and the second binding region simultaneously bind to a single antibody molecule in the mixture. In some embodiments, the target binding unit further comprises a first peptide and / or peptidomimetic sequence having the first binding region and a second peptide and / or peptidomimetic sequence having the second binding region. In some embodiments, the first binding region and the second binding region have the same sequence. In some embodiments, the first binding region and the second binding region have the same structure recognized by the single molecule. In some embodiments, the population of target binding moieties is provided on a solid support. In some embodiments, the population of target binding moieties is immobilized on the solid support. In some embodiments, the solid support has at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 10 3 、at least about 10 4 、at least about 10 5 or at least about 10 6discrete regions. In some embodiments, each of the discrete regions has a different target-binding moiety from the population immobilized thereon. In some embodiments, the different target-binding moieties from the population include the same binding region. In some embodiments, the different target-binding moieties from the population include the same peptide and / or peptidomimetic sequence. In some embodiments, each of the discrete regions has two or more copies of the different target-binding moiety. In some embodiments, the method further comprises removing unbound antibodies from the mixture. In some embodiments, the method further comprises quantifying the amount of antibody bound at each of the discrete regions. In some embodiments, the quantification comprises detecting a fluorescence signal, an electrochemical signal, a chemiluminescence signal, a chromogenic signal, or a combination thereof. In some embodiments, the method further comprises obtaining a blood sample from a subject. In some embodiments, the method further comprises preparing a serum sample from the blood sample, wherein the serum sample comprises the antibody mixture. In some embodiments, the subject comprises a diseased subject and / or a healthy subject. In some embodiments, the method further comprises obtaining a first serum sample from the diseased subject and a second serum sample from the healthy subject, wherein the quantification comprises quantifying the amount of antibody bound at each of the discrete regions in the first serum sample on a first solid support and in the second serum sample on a second solid support. In some embodiments, the method further comprises comparing the amount of antibody bound at each of the discrete regions on the first solid support and the amount of antibody bound at each of the discrete regions on the second solid support. In some embodiments, the method further comprises selecting a set of peptides of the discrete regions that differ in the amount of antibody bound on the first solid support and the second solid support. In some embodiments, the population of target-binding moieties is provided in solution. In some embodiments, each of the population of target-binding moieties is linked to a coding moiety. In some embodiments, the coding moiety is a polynucleotide. In some embodiments, the polynucleotide is deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a combination thereof. In some embodiments, the polynucleotide is single-stranded DNA, single-stranded RNA, double-stranded DNA, double-stranded RNA, or a double-stranded DNA-RNA hybrid. In some embodiments, the coding moiety comprises a nucleic acid sequence that identifies the sequence of the target-binding moiety. In some embodiments, the nucleic acid sequence encodes the peptide sequence of the two or more target-binding moieties. In some embodiments, the population of target-binding moieties comprises at least about 100, at least about 10 3 、at least about 10 4 、at least about 10 5 or at least about 10 6Different types of target-binding moieties. In some embodiments, each type of target-binding moiety from the population comprises the same peptide and / or peptidomimetic sequence. In some embodiments, the coding moiety is unique for each type of target-binding moiety of the population. In some embodiments, the method further comprises capturing, or enriching, or isolating the antibody-binding fraction of the population of target-binding moieties. In some embodiments, the method further comprises amplifying the coding moiety of the antibody-binding fraction of the target-binding moieties. In some embodiments, the method further comprises quantifying the coding moiety or copies of the coding moiety of the antibody-binding fraction. In some embodiments, quantifying comprises sequencing the coding moiety or copies of the coding moiety of the antibody-binding fraction of the target-binding moieties. In some embodiments, the method further comprises obtaining a blood sample from a subject. In some embodiments, the method further comprises preparing a serum sample from the blood sample, wherein the serum sample comprises the antibody mixture. In some embodiments, the subject comprises a diseased subject and / or a healthy subject. In some embodiments, the method further comprises obtaining a first serum sample from the diseased subject and a second serum sample from the healthy subject, wherein the quantifying comprises quantifying the amount of the antibody-binding fraction of the target-binding moieties in the first serum sample and the amount of the antibody-binding fraction of the target-binding moieties in the second serum sample. In some embodiments, the method further comprises comparing the amount of the antibody-binding fraction of the target-binding moieties in the first serum sample and the amount of the antibody-binding fraction in the second serum sample. In some embodiments, the method further comprises selecting a set of peptides, wherein each peptide in the set has a difference in the amount of the antibody-binding fraction of the target-binding moieties in the first serum sample and the second serum sample. In some embodiments, each peptide in the set is identified in large amounts in the first serum sample but not in the second serum sample, or each peptide in the set is identified in large amounts in the second serum sample but not in the first serum sample. In some embodiments, the method further comprises preparing the population of target-binding moieties each linked to the coding moiety by RNA display. In some embodiments, each of the population of target-binding moieties further comprises a puromycin moiety or a variant thereof. In some embodiments, the sequence of the first peptide and / or peptidomimetic sequence is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% or 100% identical to the sequence of the second peptide and / or peptidomimetic sequence. In some embodiments, the first peptide sequence, the second peptide sequence, and the spacer are linked by peptide bonds. In some embodiments, the first peptide and / or peptidomimetic sequence, the second peptide and / or peptidomimetic sequence, and the spacer are linked by non-peptide bonds.In some embodiments, the target binding portion comprises two or more target binding units. In some embodiments, the first peptide and / or peptidomimetic sequence, or the second peptide and / or peptidomimetic sequence has a length of at least 5 residues. In some embodiments, the spacer comprises a polymer. In some embodiments, the spacer comprises a pre-designed amino acid sequence. In some embodiments, the spacer is a polypeptide, polynucleotide or polyethylene glycol. In some embodiments, the spacer comprises a folded polypeptide, secondary structure and / or tertiary structure. In some embodiments, the folded polypeptide comprises a coiled-coil structure or a β-sheet. In some embodiments, the coiled-coil structure is formed by two separate peptide chains, wherein each of the two separate peptide chains is folded into an α-helix. In some embodiments, the coiled-coil structure is formed by a single peptide chain, the single peptide chain comprising at least two regions folded into α-helices. In some embodiments, the single peptide chain comprises four regions folded into α-helices. In some embodiments, the coiled-coil structure is formed by a first peptide chain, a second peptide chain and a third peptide chain, wherein the second peptide chain interacts with a first portion of the first peptide chain, and the third peptide chain interacts with a second portion of the first peptide chain. In some embodiments, the second peptide chain and the third peptide chain are attached to opposite ends of a double-stranded polynucleotide. In some embodiments, the spacer comprises double-stranded deoxyribonucleic acid. In some embodiments, the antibody mixture is a mixture of monoclonal antibodies, a mixture of polyclonal antibodies or a combination thereof. In some embodiments, the antibody mixture comprises a biomarker.

[0015] According to another aspect of the present disclosure, provided herein is a method for selecting a set of peptides, comprising: (a) providing two or more copies of an array comprising a first array and a second array; (b) obtaining a first antibody mixture from a diseased subject and a second antibody mixture from a healthy subject; (c) contacting the first mixture with the first array and contacting the second mixture with the second array, wherein each array has at least 10 4 discrete regions, wherein each of the discrete regions has a unique type of target binding portion having a first target binding region and a second target binding region, wherein the first binding region and the second binding region are separated by a spacer and spaced apart such that the first binding region and the second binding region simultaneously bind to a single antibody molecule in the mixture; (d) removing unbound antibody fractions on the first array and the second array; (e) quantifying the amount of bound antibody on each of the discrete regions of the first array and the second array; and (f) identifying peptides in the second array that are not bound by the antibodies on the first array.

[0016] According to another aspect of the present disclosure, there is provided herein a method for selecting a set of peptides, comprising: (a) providing a first solution and a second solution; (b) obtaining a first antibody mixture from a diseased subject and a second antibody mixture from a healthy subject; (c) contacting the first mixture with the first solution and the second mixture with the second solution, wherein each of the first solution and the second solution comprises a plurality of polynucleotide barcoded target binding moieties, wherein each of the plurality of polynucleotide barcoded target binding moieties comprises a nucleic acid sequence linked via a linker to a target binding unit, the target binding unit comprising a first peptide sequence containing a first binding region and a second peptide sequence containing a second binding region; wherein the first binding region and the second binding region are separated by a spacer and spaced apart such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule comprising the antigen-binding domain of an antibody; (d) capturing the antibody-bound components in the population in the first solution and the second solution; (e) sequencing the captured coding portions from the first solution and the second solution; and (f) selecting the set of peptides.

[0017] According to another aspect of the present disclosure, provided herein is a method for profiling an antibody mixture from a subject, comprising: contacting the antibody mixture with an array having at least 10 discrete regions, wherein each of the discrete regions has a unique type of target-binding moiety, and wherein each of the unique types of target-binding moieties comprises a first binding region and a second binding region, wherein the first binding region and the second binding region are separated by a spacer and are spaced apart such that the first binding region and the second binding region simultaneously bind to a single antibody molecule in the mixture. In some embodiments, the method further comprises removing the unbound fraction of the antibodies in the mixture. In some embodiments, the method further comprises detecting the bound fraction of the antibodies in the mixture on the array, wherein a signal is observed at each of the discrete regions to which an antibody has bound, thereby generating a signal pattern on the array. In some embodiments, the method further comprises identifying a disease of the subject. In some embodiments, the disease is an autoimmune disease, cancer, or an infectious disease. In some embodiments, each of the unique types of target-binding moieties further comprises a first peptide and / or peptidomimetic sequence having the first binding region and a second peptide and / or peptidomimetic sequence having the second binding region. In some embodiments, the first binding region and the second binding region have the same sequence. In some embodiments, the first binding region and the second binding region have the same structure recognized by the single molecule. In some embodiments, the sequence of the first peptide and / or peptidomimetic sequence is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% or 100% identical to the sequence of the second peptide and / or peptidomimetic sequence. In some embodiments, the first peptide sequence, the second peptide sequence, and the spacer are linked by peptide bonds. In some embodiments, the first peptide and / or peptidomimetic sequence, the second peptide and / or peptidomimetic sequence, and the spacer are linked by non-peptide bonds. In some embodiments, the target-binding moiety comprises two or more binding regions. In some embodiments, the first peptide and / or peptidomimetic sequence, or the second peptide and / or peptidomimetic sequence, has a length of at least 5 residues. In some embodiments, the spacer comprises a polymer. In some embodiments, the spacer comprises the same amino acid sequence. In some embodiments, the spacer is a polypeptide, polynucleotide, or polyethylene glycol. In some embodiments, the spacer comprises a folded polypeptide, secondary structure, and / or tertiary structure. In some embodiments, the folded polypeptide comprises a coiled-coil structure or a β-sheet. In some embodiments, the coiled-coil structure is formed by two separate peptide chains, wherein each of the two separate peptide chains is folded into an α-helix. In some embodiments, the coiled-coil structure is formed by a single peptide chain, the single peptide chain forming at least two regions folded into α-helices.In some embodiments, the single peptide chain comprises four regions folded into an α-helix. In some embodiments, the coiled-coil structure is formed by a first peptide chain, a second peptide chain, and a third peptide chain, wherein the second peptide chain interacts with a first portion of the first peptide chain, and the third peptide chain interacts with a second portion of the first peptide chain. In some embodiments, the second peptide chain and the third peptide chain are attached to opposite ends of a double-stranded polynucleotide. In some embodiments, the spacer region comprises double-stranded deoxyribonucleic acid. In some embodiments, the antibody mixture is a mixture of monoclonal antibodies, a mixture of polyclonal antibodies, or a combination thereof. In some embodiments, the antibody mixture comprises a biomarker.

[0018] In another aspect of the present disclosure, provided herein are a plurality of nucleic acid molecules, wherein each nucleic acid molecule of the plurality of nucleic acid molecules encodes a polypeptide target-binding portion, the peptide target-binding portion comprising: (i) a first peptide sequence comprising a first binding region, and (ii) a second peptide sequence comprising a second binding region; wherein the first binding region and the second binding region are (i) separated by a spacer region, and (ii) spaced apart such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule; wherein the plurality of nucleic acid molecules comprises at least about 10 6 at least about 10 7 at least about 10 8 at least about 10 9 at least about 10 10 at least about 10 11 at least about 10 12 at least about 10 13 at least about 10 14 at least about 10 15 or at least about 10 unique sequences. In some embodiments, the single molecule comprises an antigen-binding domain. In some embodiments, the plurality of nucleic acid molecules are a plurality of double-stranded DNA molecules. In some embodiments, the plurality of nucleic acid molecules are a plurality of single-stranded RNA molecules. In some embodiments, each nucleic acid molecule of the plurality of nucleic acid molecules is a circular molecule. Also provided herein is a method for preparing a library of polynucleotide-peptide fusion molecules using the library of nucleic acid molecules described herein. The method comprises: providing a plurality of nucleic acid molecules and expressing the plurality of nucleic acid molecules in an RNA display assay to produce a plurality of peptides, wherein each nucleic acid molecule is linked by a linker to a peptide expressed from the nucleic acid molecule. In some embodiments, the linker comprises puromycin or a derivative thereof.

[0019] Incorporated by reference

[0020] All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the invention will be obtained from the following detailed description taken in conjunction with the accompanying drawings which illustrate exemplary embodiments of the principles of the invention.

[0022] In the figures:

[0023] Figure 1 Two exemplary structures of the target-binding portion described herein or the target-binding portion linked to the coding portion are shown.

[0024] Figure 2 An exemplary structure of the target-binding unit described herein is shown.

[0025] Figure 3 An exemplary structure of the spacer described herein is shown.

[0026] Figure 4 An example of an array having different discrete regions and an example structure of a scaffold fixed to the discrete regions of the array are shown.

[0027] Figure 5A An exemplary embodiment of generating a polynucleotide-barcode target-binding portion using in vitro compartmentalization is shown.

[0028] Figure 5B An exemplary embodiment of generating a polynucleotide-barcode target-binding portion using in vitro compartmentalization is shown.

[0029] Figure 5C An exemplary embodiment of generating a polynucleotide-barcode target-binding portion using in vitro compartmentalization is shown.

[0030] Figure 6A An example structure for generating a polynucleotide-barcode binding element by split-and-pool synthesis is shown.

[0031] Figure 6B An example structure for generating a polynucleotide-barcode target-binding portion by split-and-pool synthesis is shown.

[0032] Figure 7A An example protocol for generating a library of target-binding portions is shown.

[0033] Figure 7B An example protocol for generating a library of target-binding portions is shown.

[0034] Figure 7C

[0034] Figure 7C illustrate an exemplary scheme for generating a library of target-binding moieties.

[0035] Figure 7D

[0035] Figure 7D illustrate an exemplary scheme for generating a library of target-binding moieties.

[0036] Figure 7E

[0036] Figure 7E illustrate an exemplary scheme for generating a library of target-binding moieties.

[0037] Figure 8A

[0037] Figure 8A illustrate an exemplary scheme for generating a library of target-binding moieties.

[0038] Figure 8B

[0038] Figure 8B illustrate a denaturing polyacrylamide gel image showing the presence of products of the desired size (e.g., D1, D4, D7, D12, and D14) generated during different steps of the exemplary scheme shown in Figure 8A Figure 8A .

[0039] Figure 9

[0039] Figure 9 illustrate a denaturing polyacrylamide gel image showing the presence of products of the desired size generated by further processing the DNA product (536 bp) of D14 in different steps of RNA display to prepare an RNA-peptide fusion molecule.

[0040] Figure 10A

[0040] Figure 10A illustrate a schematic diagram of using RNA display to generate a polynucleotide-peptide fusion molecule.

[0041] Figure 10B

[0041] Figure 10B illustrate an example of a polynucleotide-barcode-target-binding moiety.

[0042] Figure 10C

[0042] Figure 10C illustrate an example of a polynucleotide-barcode-target-binding moiety with a flexible spacer that can be further manipulated to generate a rigid spacer. Detailed Description

[0043] In this disclosure, the use of the singular includes the plural unless specifically stated otherwise. Additionally, unless otherwise indicated, the use of "or" means "and / or". Similarly, "comprising", "including", and "containing" do not constitute a limitation.

[0044] Overview

[0045] Determining the specific antigen-binding region or epitope of an antibody can be beneficial not only for antibody-based test platforms, but also for their therapeutic applications, vaccine development, protein interaction studies, autoimmune diseases, etc. Methods for interrogating an antibody library can include (1) sequencing the coding region of an antibody of interest, (2) using a large library of molecules (e.g., peptides, peptoids, or proteins, which are collectively referred to herein as potential immunoreceptor-binding molecules or PIRMs) to examine which of these molecules can be bound by a target of interest (e.g., the antibody of interest). To overcome problems associated with methods known in the art, the compositions and methods provided herein can be used to generate PIRM libraries having a large library size and / or high sensitivity. For example, the methods provided herein can generate PIRM libraries that include at least about 100, at least about 10 3 、at least about 10 4 、at least about 10 5 、at least about 10 6 、at least about 10 7 、at least about 10 8 、at least about 10 9 、at least about 10 10 、at least about 10 11 、at least about 10 12 、at least about 10 13 、at least about 10 14 、at least about 10 15 、at least about 10 16 、at least about 10 17 、at least about 10 18 、at least about 10 19 or at least about 10 20 different species (e.g., unique sequences).

[0046] The present disclosure provides compositions and methods for screening / identifying binding elements that bind to an analyte. For example, the present disclosure provides compositions and methods for screening / identifying peptide or peptoid sequences that bind to an antibody. The compositions and methods disclosed herein offer several advantages over traditional peptide arrays, including high library size and high sensitivity. The identified peptide sequences can then be used to fabricate arrays having an addressable library of peptide sequences for antibody profiling in a given sample. In various embodiments disclosed herein, two peptides are provided in a pair as a target-binding unit, wherein the two peptides are spaced a certain distance apart such that they can bind to a single molecule simultaneously.

[0047] In one aspect, the present disclosure provides a composition comprising a target-binding unit, wherein the target-binding unit comprises two binding elements separated by a spacer, and wherein the two binding elements simultaneously bind to a single molecule having an antigen-binding domain of an antibody. In various embodiments, the binding element is a peptide sequence or a peptidomimetic sequence. The two peptides of the target-binding unit are separated by a spacer such that they are positioned to simultaneously bind to two antigen-binding domains (e.g., Fab) of an antibody. The strategy provided herein utilizes the two antigen-binding domains of an antibody and can increase the binding affinity by at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 6-fold, at least 7-fold, at least 8-fold, at least 9-fold, at least 10-fold, at least 11-fold, at least 12-fold, at least 13-fold, at least 14-fold, at least 15-fold, at least 16-fold, at least 17-fold, at least 18-fold, at least 19-fold, or at least 20-fold. In some embodiments, the two peptides of the target-binding unit have the same sequence. The present disclosure also provides a composition comprising a target-binding moiety, wherein the target-binding moiety comprises one or more target-binding units. One or more target-binding units in each target-binding moiety may be linked. The target-binding moiety comprises two or more binding elements (e.g., peptide sequences or peptidomimetic sequences). In some embodiments, the target-binding moiety comprises a first peptide sequence and a second peptide sequence, wherein the first peptide sequence and the second peptide sequence are separated by a spacer such that they can simultaneously bind to a single antibody molecule. In some embodiments, the two or more peptide sequences of the target-binding moiety have the same sequence.

[0048] In another aspect, the present disclosure provides methods of using a target-binding unit or a target-binding moiety to screen for and identify peptide or peptidomimetic sequences that differ in binding to antibodies in patient samples and antibodies in healthy samples. For example, in a peptide screening and identification method, a plurality of target-binding units will be provided, wherein each target-binding unit comprises two peptides having the same sequence or binding region (e.g., epitope). In some cases, a composition comprising the target-binding unit is provided in solution. In such cases, the target-binding unit or the target-binding moiety may be further linked to a coding moiety. In some applications, the coding moiety serves as a barcode corresponding to or identifying the peptide or peptidomimetic sequence of the target-binding unit. In some other applications, the coding moiety comprises a sequence that can both encode the peptide of the target-binding unit and serve as a barcode for identifying the peptide sequence. The coding moiety is capable of identifying the peptide sequence by high-throughput sequencing. In some cases, a composition comprising the target-binding unit or the target-binding moiety is provided on a solid surface. In such cases, the target-binding unit or the target-binding moiety having a unique binding element (e.g., a peptide sequence) will be immobilized on discrete regions of the solid surface, separated from another target-binding unit having a different binding element. Thus, when the target-binding unit is provided on a solid surface, a coding moiety may not be required.

[0049] On the other hand, the present disclosure provides methods for profiling antibodies in a sample from a subject using an identified peptide or peptidomimetic sequence, which can be used to infer whether a subject has a disease. Since the peptide or peptidomimetic is identified by the screening / identification methods disclosed herein, it is known that the peptide or peptidomimetic is capable of distinguishing patient samples from healthy samples and that the peptide or peptidomimetic is also associated with the disease associated with the patient sample. The identified peptide can be provided on a solid surface (e.g., an array) and can generate a characteristic pattern on the solid surface when profiling the antibodies in the sample. In some cases, the characteristic pattern can be similar to the patient sample, indicating that the subject may have a disease. In some other cases, the characteristic pattern can be similar to the healthy sample, indicating that the subject is healthy. As disclosed herein, the peptide sequence on the solid surface is provided as a target binding unit, wherein each target binding unit comprises two peptides having the same sequence or epitope, and the two peptides are separated by a spacer region.

[0050] Definition

[0051] The term “about” or “approximately” means, with respect to a particular value, within an acceptable error range as determined by a person of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to the practice in the art, “about” can mean within 1 or more standard deviations. Alternatively, “about” can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Or, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5-fold, more preferably within 2-fold of a value. Where a particular value is described in the present application and claims, unless otherwise stated, the term “about” should be assumed to mean within an acceptable error range of that particular value.

[0052] The terms "polynucleotide", "nucleic acid" and "oligonucleotide" are used interchangeably. They can refer to polymeric forms of nucleotides of any length, whether deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides can include one or more nucleotides selected from adenosine (A), cytosine (C), guanine (G), thymine (T) and uracil (U) or variants thereof. The polynucleotides provided herein can be double-stranded or single-stranded. Nucleotides generally include a nucleoside and at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more phosphate (PO3) groups. Nucleotides can include a nucleobase, a pentose sugar (ribose or deoxyribose) and one or more phosphate groups. Polynucleotides can have any three-dimensional structure and can perform any known or unknown function. The following are non-limiting examples of polynucleotides: coding or non-coding regions of genes or gene fragments, loci defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), circular RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes and primers. Polynucleotides can include one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If there are modifications to the nucleotide structure, they can be made before or after polymer assembly. The sequence of nucleotides can be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, such as by conjugation with a labeling component. In some cases, the polynucleotides provided herein are coding portions that can be used to indicate the identity of another part. In some other cases, the polynucleotides provided herein are spacer regions that can be used to link two entities.

[0053] The polynucleotide can include one or more nucleotide variants, including non-standard nucleotides, unnatural nucleotides, nucleotide analogs, and / or modified nucleotides. In some cases, the nucleotide can include a modification in its phosphate moiety, including modifications to the triphosphate moiety. Non-limiting examples of such modifications include a longer length of the phosphate chain (e.g., a phosphate chain having 4, 5, 6, 7, 8, 9, 10, or more phosphate moieties) and modifications of the thiol moiety (e.g., α-thiotriphosphate and β-thiotriphosphate). The nucleic acid molecule can be modified at the base moiety (e.g., one or more atoms that are typically available to form hydrogen bonds with complementary nucleotides and / or one or more atoms that are typically not capable of forming hydrogen bonds with complementary nucleotides), the sugar moiety, or the phosphate backbone. The nucleic acid molecule can contain amine-modified groups such as aminoallyl 1-dUTP (aa-dUTP) and aminohexyl acrylamide-dCTP (aha-dCTP) to permit covalent attachment of amine-reactive moieties such as N-hydroxysuccinimide esters (NHS). The polynucleotide can be modified at one or more positions to enhance stability introduced during chemical synthesis or subsequent enzymatic modification or polymerase replication. These modifications include, but are not limited to, introducing one or more alkylated nucleic acids, locked nucleic acids (LNA), peptide nucleic acids (PNA), phosphonates, phosphorothioates, etc. into the oligomer. Examples of modified nucleotides include, but are not limited to, 2,6-diaminopurine, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouracil, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylqueosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-methyladenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosylqueosine, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid (v), wybutoxosine, pseudouracil, queosine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, methyl uracil-5-oxyacetate, uracil-5-oxyacetic acid (v), 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, 2,6-diaminopurine, etc. Substitutes for the standard DNA base pairs or RNA base pairs in the oligonucleotides of the present disclosure can provide a higher density (bit / mm 3) Higher security (resistance to accidental or purposeful synthesis of natural toxins), easier discrimination in photo-programmable polymerases, or lower secondary structure. Such alternative base pairs compatible with natural and mutant polymerases for de novo synthesis and / or amplified synthesis are described in Betz K, Malyshev DA, Lavergne T, Welte W, Diederichs K, Dwyer TJ, Ordoukhanian P, Romesberg FE, Marx A. Nat Chem Biol. 2012 Jul;8(7):612-4, which is incorporated herein by reference for all purposes.

[0054] As used herein, the term "polypeptide" refers to two or more amino acids joined together by peptide bonds (or amide bonds), also referred to as "peptide". A peptide bond, also known as an amide bond, is a covalent chemical bond that links two consecutive amino acid monomers along a peptide or protein chain. In the context of this specification, it should be understood that amino acids can be L - optical isomers or D - optical isomers. Amino acids can be non - naturally encoded amino acids. Peptides can include one or more non - naturally encoded amino acids. A "non - naturally encoded amino acid" refers to an amino acid that is not one of the common amino acids or pyrrolysine or selenocysteine. Other terms that can be used synonymously with the term "non - naturally encoded amino acid" are "non - natural amino acid", "non - naturally occurring amino acid", and various forms with and without hyphens. The term "non - naturally encoded amino acid" also includes, but is not limited to, amino acids that are produced by modification (e.g., post - translational modification) of naturally encoded amino acids (including, but not limited to, the 20 common amino acids or pyrrolysine and selenocysteine), but are not themselves naturally incorporated into the growing polypeptide chain by the translation complex. Examples of such non - naturally occurring amino acids include, but are not limited to, N - acetylglucosamine - L - serine, N - acetylglucosamine - L - threonine, and O - phosphoryl tyrosine. The length of a peptide is two or more amino acid monomers, and typically can exceed 20 amino acid monomers. Polypeptides can be linear and unstructured or folded into a three - dimensional structure. Stereoisomers of the twenty conventional amino acids, non - natural amino acids (such as α,α - disubstituted amino acids, N - alkyl amino acids, lactic acid), and other non - conventional amino acids (e.g., D - amino acids) can also be suitable components of the polypeptides of the present disclosure. Examples of non - conventional amino acids include: 4 - hydroxyproline, γ - carboxyglutamic acid, ε - N,N,N - trimethyllysine, ε - N - acetyllysine, O - phosphorylserine, N - acetylserine, N - formylmethionine, 3 - methylhistidine, 5 - hydroxylysine, σ - N - methylarginine, and other similar amino acids and imino acids (e.g., 4 - hydroxyproline). In the polypeptide notation used herein, according to standard usage and convention, the left - hand direction is the amino - terminal direction, and the right - hand direction is the carboxyl - terminal direction. Structured polypeptides can be proteins. As used herein, "protein" refers to a long polymer of amino acid residues linked by peptide bonds, which can consist of one or more polypeptide chains. More specifically, the term "protein" refers to a molecule composed of one or more amino acids in a specific order; for example, the order is determined by the base sequence of nucleotides in the gene encoding the protein. Proteins are essential for the structure, function, and regulation of human cells, tissues, and organs, and each protein has a unique function. Examples are hormones, enzymes, antibodies, and any fragments thereof. A protein can be part of a protein, such as a domain, sub - domain, or motif of a protein.The protein can be a variant (or mutant) of a protein, in which one or more amino acid residues are inserted into, deleted from, and / or substituted for the naturally occurring (or at least known) amino acid sequence of the protein. The protein can be a modified protein. Non-limiting examples of protein modification include phosphorylation, acetylation, glycosylation, amidation, hydroxylation, methylation, alkylation, acylation, ubiquitination, pyrrolidone carboxylic acid, and sulfation. The protein or its variant can be naturally occurring or recombinant.

[0055] As used herein, the term "peptoid", also known as poly-N-substituted glycine, is a class of peptidomimetics in which the side chains are attached to the nitrogen atom of the peptide backbone rather than to the α-carbon atom (as they are in amino acids). In peptoids, the side chains are attached to the nitrogen of the peptide backbone rather than to the α-carbon as in peptides. Notably, peptoids lack the amide hydrogens that contribute to many secondary structure elements in peptides and proteins. Peptoids can be used to mimic protein / peptide products to aid in the discovery of protease-stable small molecule drugs. In the synthesis of peptoids, each residue can be installed in two steps by a method called the submonomer approach: acylation and substitution. In the acylation step, a haloacetic acid, typically bromoacetic acid activated by diisopropylcarbodiimide, reacts with the amine of the previous residue. In the substitution step (a classic SN2 reaction), the amine displaces the halide to form an N-substituted glycine residue. The submonomer approach allows any commercially available or synthetically accessible amine to be used in combinatorial chemistry. Peptoids are generally resistant to proteolysis and are thus advantageous for therapeutic applications where proteolysis is a major concern. Since the secondary structure in peptoids generally does not involve hydrogen bonding, they are not typically denatured by solvents, temperature, or chemical denaturants such as urea. Due to the flexibility of the backbone methylene groups and the lack of stabilizing hydrogen bond interactions along the backbone, peptoid oligomers may be conformationally unstable. However, by choosing appropriate side chains, it is possible to form specific steric or electronic interactions that favor the formation of stable secondary structures such as helices. In particular, peptoids with C-α branched side chains are known to adopt a structure similar to that of polyproline I helices. Different strategies can be used to predict and characterize peptoid secondary structures, with the ultimate goal of developing fully folded peptoid protein structures. Cis / trans amide bond isomerization can lead to conformational heterogeneity, which may not allow the formation of homogeneous peptoid foldamers. However, trans-inducing agents such as N-aryl side chains that favor polyproline II-type helices, as well as strong cis-inducing agents such as large naphthylethyl and tert-butyl side chains, have been discovered. In various embodiments provided herein, the peptide sequence of the target binding unit can be replaced with a peptoid sequence.

[0056] The term "coiled coil" or "coiled coil structure" refers to a structural motif in which multiple α-helices coil together like strands of a multi-stranded rope (e.g., dimers and trimers are examples of common types). For example, leucine zippers are coiled coil structures.

[0057] The terms "functional group", "active moiety", "activating group", "leaving group", "reaction site", "reacting group", "chemical reaction group", and "chemical reaction moiety" refer to distinct, definable portions or units within a molecule. The term is used herein to denote the portion of a molecule that performs some function or activity or reacts with other molecules.

[0058] The term "linkage" or "connector" refers to a group or bond that is typically formed as a result of a chemical reaction.

[0059] The terms "spacer" or "spaced apart" are used interchangeably herein. When describing two peptide / peptidomimetic sequences that are separated or spaced apart by a spacer region, it is meant that the two peptides are not directly connected to each other, but rather are connected via the spacer region.

[0060] As used herein, the term "sequence" and its grammatical equivalents can refer to a polypeptide sequence or a polynucleotide sequence. A polynucleotide or nucleotide sequence can be DNA or RNA; it can be linear, circular, or branched; and it can be single-stranded or double-stranded. The sequence can be mutated. The sequence can have any length, e.g., between 2 and 1,000,000 or more amino acids or nucleotides (or any integer value between or above the two), e.g., between about 100 and about 10,000 nucleotides or between about 200 and about 500 amino acids or nucleotides.

[0061] "Solid support", "support", "solid-phase support", "matrix", and other grammatical equivalents herein refer to any material that can be modified to contain discrete individual sites suitable for the attachment or association of molecules and that can be adapted for at least one detection method. They can be a material or a group of materials having one or more rigid or semi-rigid surfaces. The number of possible matrices can be very large. Possible matrices include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylic, polystyrene, and copolymers of styrene with other materials, polypropylene, polyethylene, polybutene, polyurethane, polytetrafluoroethylene, etc.), polysaccharides, nylon or nitrocellulose, resins, silica or silica-based materials (including silicon and modified silicon), carbon, metals, inorganic glass, plastics, fiber optic bundles, and various other polymers. In some cases, the matrix allows for optical detection and may not fluoresce itself. The matrix can be flat (planar), although other configurations of the matrix can also be used; for example, the matrix can use a three-dimensional structure, such as by embedding the target-binding moiety into a plastic porous block that allows the sample to access the target-binding moiety. In some embodiments, the target-binding moiety can be placed on the inner surface of a tube for flow-through sample analysis to minimize the sample volume. The matrix can include fiber optic bundles, as well as flat planar matrices such as glass, polystyrene, and other plastics and acrylics. In some embodiments, the solid support or matrix can be a microtiter plate. In some embodiments, at least one surface of the solid-phase support can be substantially flat, although in some embodiments, physically separating regions for different molecules or reactions with features such as holes, raised areas, pins, etched grooves, etc. can be useful. In some embodiments, the solid support can take the form of beads, resins, gels, microspheres, or other geometric configurations.

[0062] "Addressable" with respect to a target-binding unit or target-binding moiety means that the peptide / peptidomimetic sequence, or other physical or chemical feature, of the target-binding unit or target-binding moiety can be determined from its address; for example, there is a one-to-one correspondence between the sequence or other property of the target-binding unit or target-binding moiety and its spatial position on the solid-phase support or the features of the solid-phase support to which it is attached.

[0063] "Array" or "microarray" refers to a solid support having a planar surface that can carry an array of nucleic acids or peptides, wherein each spatially defined region or site of the array comprises a copy of an oligonucleotide or peptide immobilized on that spatially defined region or site, and the regions or sites do not overlap with other regions or sites of the array; i.e., the regions or sites are spatially discrete. The spatially defined binding sites can additionally be "addressable" in that, for example, prior to use, their location and the identity of the oligonucleotide or peptide immobilized thereon are known or pre-determined. A microarray can comprise at least one planar solid support, such as a glass microscope slide. In some embodiments, the oligonucleotide or peptide is covalently attached to the solid support. In some embodiments, the oligonucleotide can be attached to the solid support via the 5'-end or 3'-end. In some embodiments, the peptide can be indirectly attached via an oligonucleotide or via a non-natural amino acid incorporated into the peptide. For reviews of microarray technology, see: Schena, Editor, Microarrays: A Practical Approach ((lRL Press, Oxford, 2000); Southern, Current Opin. Chem. Biol., 2:404-410 (1998); Nature Genetics Supplement, 21:1-60 (1999). "Random microarray" refers to a microarray in which the spatially discrete regions of oligonucleotides or peptides may not be spatially addressable. For example, at least initially, the identity of the attached oligonucleotide or peptide may not be discernible from its location. In some aspects, a random microarray is a planar array of microbeads, wherein each microbead is attached to a single species of oligonucleotide or peptide. Microbead arrays can be formed in a variety of ways, e.g., Brenner et al., Nature Biotechnology, 18:630-634 (2000); Tulley et al., U.S. Patent No. 6,133,043; Stuelpnagel et al., U.S. Patent No. 6,396,995; Chee et al., U.S. Patent No. 6,544,732; etc. Similarly, after formation, the microbeads or their oligonucleotides / peptides in a random array can be identified in a variety of ways, including by optical labeling, e.g., fluorescence dye ratios or quantum dots, shape, sequence analysis, etc.

[0064] The term "label" refers to a composition capable of generating a detectable signal that indicates the presence of a target polynucleotide in an assay sample. Suitable labels include radioisotopes, nucleotide chromophores, enzymes, substrates, fluorescent molecules, chemiluminescent moieties, magnetic particles, bioluminescent moieties, etc. Thus, a label can be any composition detectable by spectroscopic, photochemical, biochemical, immunochemical, electrical, optical, or chemical means.

[0065] "Sample" means an amount of material from a biological, environmental, medical, or patient source in which detection or measurement of an analyte is sought. Samples can include specimens or cultures (e.g., microbial cultures). Samples can include biological and environmental samples. Samples can include specimens of synthetic origin. Biological samples can include materials taken from patients or healthy subjects, including but not limited to, blood, saliva, spinal fluid, pleural fluid, milk, lymph fluid, sputum, semen, aspirates, plasma, serum, external portions of skin, respiratory, intestinal, and urogenital tracts, tears, cells (including but not limited to blood cells), tumors, organs, and samples of in vitro cell culture components. Biological samples can be obtained from all sorts of domestic and wild animals, including but not limited to, ungulates, bears, fish, rodents, etc. Environmental samples can include environmental materials such as surface substances, soil, water, and industrial samples, and samples obtained from food and dairy processing instruments, devices, equipment, utensils, disposable and non-disposable items. These examples should not be construed as limiting the types of samples applicable to the present disclosure.

[0066] As used herein, the term "container" refers to a compartment (e.g., a microfluidic channel, a well, or a droplet) in which biochemical reactions (e.g., target protein and antibody binding, nucleic acid hybridization, and primer extension) can occur. The terms "container" and "compartment" can be used interchangeably. The volume of the compartment can be as large as 1 mL or as small as 1 picoliter. In some embodiments, the median size of the compartments in a plurality of compartments is from about 1 picoliter to about 10 picoliters, about 10 picoliters to about 100 picoliters, about 100 picoliters to about 1 nanoliter, about 1 nanoliter to about 10 nanoliters, about 10 nanoliters to about 100 nanoliters, about 100 nanoliters to about 1 microliter, about 1 microliter to about 10 microliters, about 10 microliters to about 100 microliters, or about 100 microliters to about 1000 microliters. The volume of water contained in the compartment can be less than or approximately equal to the volume of the compartment. In some embodiments, the median volume of water contained in the compartment is 1 microliter or less.

[0067] "Droplet" means a compartment surrounded by a liquid rather than a solid. Droplets can be water-in-oil; water-in-oil-in-water, or water-in-a lipid layer (liposome). In some embodiments, droplets can have a uniform size or a non-uniform size. In some embodiments, the median diameter of the droplets in a plurality of droplets can range from about 0.001 μm to about 1 mm. In some embodiments, the median volume of the droplets in a plurality of droplets can range from about 0.01 nanoliter to about 1 microliter.

[0068] The term "epitope" refers to any protein or peptide capable of specific binding to an immunoglobulin or antibody. Epitope influencing factors can include the chemically reactive surface groups of a molecule, such as amino acids or sugar side chains, and can have three-dimensional structural features as well as charge characteristics. An antibody can specifically bind an antigen when the dissociation constant is ≤1 μM or ≤100 nM or ≤10 nM. As used herein, an epitope can also be used to refer to a peptidomimetic capable of specific binding to an immunoglobulin or antibody.

[0069] As used herein, the term "partitioning" can be a verb or a noun. When used as a verb (e.g., "partition into" or "perform partitioning"), the term generally refers to fractionating (e.g., subdividing) a species or sample (e.g., a polynucleotide sample) into a container that can be used to isolate one fraction (or subdivision) from another. Such a container is denoted by the noun "partition". Partitioning can be performed, for example, using microfluidic techniques, dilution, dispensing, vortexing, etc. A partition can be, for example, a well, a micro-well, a hole, a droplet (e.g., a droplet in an emulsion), the continuous phase of an emulsion, a test tube, a spot, a capsule, a bead, the surface of a bead in a dilution solution, or any other suitable container that isolates a small portion of a sample from another fraction. Partitioning can also include another partition.

[0070] The percent (%) sequence identity relative to a reference polypeptide sequence (or nucleic acid sequence) is defined as the percentage of amino acid residues (or nucleotides in the case of a nucleic acid sequence) in the candidate sequence that are identical to the amino acid residues (or nucleotides) in the reference polypeptide sequence (or nucleic acid sequence) after aligning the sequences and introducing gaps as needed to achieve the maximum percent sequence identity, without considering any conservative substitutions as part of the sequence identity. For purposes of determining the percent amino acid sequence identity, the alignment can be achieved in various ways within the skill in the art, e.g., using publicly available computer software such as BLAST, BLAST-2, ALIGN, or Megalign (DNASTAR) software. One of skill in the art can determine the appropriate parameters for aligning the sequences, including any algorithms needed to achieve maximal alignment over the full length of the sequences being compared. However, for purposes herein, the % amino acid sequence identity values are generated using the sequence comparison computer program ALIGN-2. The ALIGN-2 sequence comparison computer program is licensed from Genentech, Inc., and the source code has been filed with the U.S. Copyright Office (Washington, D.C. 20559) together with the user documentation, and is registered in U.S. Copyright Registration No. TXU510087. The ALIGN-2 program is publicly available from Genentech, Inc., South San Francisco, California, and can also be compiled from the source code. The ALIGN-2 program should be compiled for use on the UNIX operating system (including Digital UNIX V4.0D). All sequence comparison parameters are set by the ALIGN-2 program and are not varied.

[0071] In the case of amino acid sequence comparison using ALIGN-2, the % amino acid sequence identity of a given amino acid sequence A with respect to a given amino acid sequence B (which can also be stated as a given amino acid sequence A having or including a certain % amino acid sequence identity with a given amino acid sequence B) is calculated as follows: the fraction X / Y multiplied by 100, where X is the number of amino acid residues scored as identical matches by the sequence alignment program ALIGN-2 in the alignment of A and B by that program, and where Y is the total number of amino acid residues in B. It should be understood that if the length of amino acid sequence A is not equal to the length of amino acid sequence B, the % amino acid sequence identity of A with B will not be equal to the % amino acid sequence identity of B with A. Unless otherwise specifically stated, all % amino acid sequence identity values used herein are obtained using the ALIGN-2 computer program as described in the previous paragraph.

[0072] Analyte

[0073] The present disclosure provides compositions and methods for identifying binding elements (e.g., peptide sequences or peptidomimetic sequences) that can bind to an analyte from a sample. As used herein, an analyte can refer to a target or a target molecule. The binding element can be a peptide sequence or a peptidomimetic sequence. The analyte can be a polypeptide. The analyte can be a protein. The analyte can be an endogenous protein or an artificial protein. The analyte can be a recombinant protein. The analyte described herein can be obtained or isolated from a subject. In various embodiments, the methods provided herein include contacting the analyte with a target binding moiety or a target binding unit described herein.

[0074] The analyte can be an antibody or an immunoglobulin. In some embodiments, the analyte is an antibody fragment. In some embodiments, the antibody or its fragment described in the compositions and methods includes two antigen-binding (Fab) fragments. In some embodiments, the two Fab fragments bind the same antigen. In some embodiments, the antibody is a bispecific antibody. In some embodiments, the antibody is a bispecific antibody, wherein the bispecific antibody includes two Fab fragments that bind different antigens. In some embodiments, the antibody or its fragment includes non-conventional amino acids. In some embodiments, the antibody is a monoclonal antibody, a polyclonal antibody, or a combination thereof. In some embodiments, the analyte is a mixture of antibodies. In some embodiments, the analyte is a mixture of monoclonal antibodies. In some embodiments, the analyte is a mixture of polyclonal antibodies. In some embodiments, the mixture of antibodies includes monoclonal antibodies and polyclonal antibodies. In some embodiments, a blood sample contains a mixture of antibodies. In some embodiments, a serum sample contains a mixture of antibodies.

[0075] In some embodiments, the analyte described herein can be a molecule that binds to the antigen-binding domain of an antibody. In some embodiments, the analyte described herein can be a molecule that includes the antigen-binding domain of an antibody. The antigen-binding domain can be an scFv, a Fab, or an F(ab′)2. Other examples of antigen-binding domains include, in particular, Fab, Fab′, F(ab′)2, Fv, dAb, and complementarity-determining region (CDR) fragments, single-chain antibodies (scFv), single-domain antibodies, chimeric antibodies, diabodies, and polypeptides containing at least a portion of an immunoglobulin that is sufficient to effect specific antigen binding with the polypeptide. For the purposes described herein, linear antibodies are also included. As used herein, the term "diabody" refers to a small antibody fragment having two antigen-binding sites, the fragment including a heavy-chain variable domain (V H -V L ) linked to a light-chain variable domain (V L ) in the same polypeptide chain (V H) By using linkers that are too short for two domains on the same chain to pair, these domains are forced to pair with complementary domains on the other chain, creating two antigen-binding sites.

[0076] The analytes described herein can be obtained or isolated from a subject. In some applications, the analyte is obtained or isolated from a blood sample of the subject. In some applications, the analyte is obtained or isolated from a serum sample of the subject. In some applications, the antibody is obtained or isolated from a blood sample of the subject. In some applications, the antibody is obtained or isolated from a serum sample of the subject.

[0077] In some embodiments, the methods provided herein include contacting an analyte with a target-binding moiety or target-binding unit. For example, the method can include contacting an antibody with a target-binding moiety or target-binding unit. In some embodiments, the antibody is an antibody mixture. Contacting the antibody mixture with one or more target-binding moieties can be used to identify the target-binding moieties that bind to the antibodies in the antibody mixture. The target-binding moieties that bind to the analyte can be identified in solution or on a surface.

[0078] In some embodiments, the analyte can be a biomarker. In some embodiments, the amount of analyte from a patient sample is different from the amount of analyte from a healthy sample. In some embodiments, quantifying the amount of analyte from patient samples and healthy samples can be used for disease diagnosis and prognosis. An array comprising the target-binding moieties or target-binding units described herein can be used to detect or quantify analytes from patient samples and healthy samples.

[0079] Antibody

[0080] In some embodiments, the analyte is an immunoglobulin, antibody, or an antigen-binding domain thereof. A complete immunoglobulin or antibody typically consists of four polypeptides: two identical copies of a heavy (H) chain polypeptide and two identical copies of a light (L) chain polypeptide. In mammals, antibodies are divided into five isotypes: IgG, IgM, IgA, IgD, and IgE. Isotypes vary in their biological properties, functional locations, and ability to handle different antigens. The type of heavy chain present defines the class of the antibody. Mammalian Ig heavy chains are of five types, designated by Greek letters: α, δ, ε, γ, and μ. These chains are found in IgA, IgD, IgE, IgG, and IgM antibodies, respectively. Heavy chains vary in size and composition; α and γ contain approximately 450 amino acids, while μ and ε contain approximately 550 amino acids. Each heavy chain can contain an N-terminal variable (V H ) region and three C-terminal constant (C H 1, C H 2, and C H 3) regions, and each light chain can contain an N-terminal variable (V L) region and a C-terminal constant (C L ) region. Immunoglobulin light chains can be assigned to one of two different types based on the amino acid sequence of their constant domain, kappa (κ) or lambda (λ). In a typical immunoglobulin, each light chain can be linked to a heavy chain by a disulfide bond, and the two heavy chains can be linked to each other by a disulfide bond. In some embodiments, the provided heavy chain, light chain, and / or antibody agent have a structure that includes one or more disulfide bonds. In some embodiments, one or more disulfide bonds are or include disulfide bonds at the expected positions of IgG4 immunoglobulins. The variable region of the light chain can be aligned with the variable region of the heavy chain, and the light chain constant region can be aligned with the first constant region of the heavy chain. The remaining constant regions of the heavy chain can be aligned with each other.

[0081] The variable regions of each pair of light and heavy chains can form the antigen-binding site of the antibody. V H and V LThe domains can have the same general structure, with each domain including four framework (FW or FR) regions that are connected by three complementarity-determining regions (CDRs). As used herein, the term "framework region" can refer to a relatively conserved amino acid sequence within a variable domain that lies between the hypervariable or complementarity-determining regions (CDRs). In a typical immunoglobulin, there can be four framework regions in each variable region, designated FR1, FR2, FR3, and FR4, respectively. The framework regions form beta-sheets that provide the structural framework of the variable region (see, e.g., C.A. Janeway et al. (eds.), Immunobiology, 5th ed., Garland Publishing, New York, NY (2001)). In a typical immunoglobulin, there can be three complementarity-determining regions (CDRs) in each variable domain, designated CDR1, CDR2, and CDR3, respectively. The CDRs form the "hypervariable regions" of the antibody and are responsible for antigen binding. The CDRs form loops that connect the beta-sheet structures formed by the framework regions and, in some cases, constitute part of the beta-sheet structures formed by the framework regions. Exemplary hypervariable loops occur at amino acid residues 26-32 (L1), 50-52 (L2), 91-96 (L3), 26-32 (H1), 53-55 (H2), and 96-101 (H3) (Chothia and Lesk, J. Mol. Biol., 196:901-917 (1987)). Exemplary CDRs (CDR-L1, CDR-L2, CDR-L3, CDR-H1, CDR-H2, and CDR-H3) occur at amino acid residues 24-34 of L1, 50-56 of L2, 89-97 of L3, 31-35B of H1, 50-65 of H2, and 95-102 of H3 (Kabat et al., Sequences of Proteins of Immunological Interest, 5th ed. (1991)). In addition to V HExcept for CDR1 in it, CDRs generally include amino acid residues that form hypervariable loops. CDRs also include "specificity-determining residues" or "SDRs", which are residues that contact the antigen. SDRs are included within CDR regions called abbreviated CDRs or a-CDRs. Exemplary a-CDRs (a-CDR-L1, a-CDR-L2, a-CDR-L3, a-CDR-H1, a-CDR-H2, and a-CDR-H3) occur at amino acid residues 31-34 of L1, 50-55 of L2, 89-96 of L3, 31-35B of H1, 50-58 of H2, and 95-102 of H3 (see, e.g., Fransson, Front. Biosci., 13:1619-1633 (2008)). Unless otherwise specified, residues in the variable domain are numbered herein according to that described by Kabat et al. supra. The variable region is the domain of the antibody heavy or light chain that is involved in the binding of the antibody to the antigen (see, e.g., Kindt et al., Kuby Immunology, 6th ed., W.H. Freeman and Co., p. 91 (2007)). A single V H or V L domain may be sufficient to confer antigen-binding specificity. In addition, V H or V L domains can be isolated from an antibody that binds an antigen to screen a library of complementary V L or V H domains, respectively (see, e.g., Portolano et al., J. Immunol., 150:880-887 (1993); Clarkson et al., Nature 352:624-628 (1991)).

[0082] Antibodies can include antigen-binding fragments (Fab) and fragment crystallizable regions (Fc). The Fc region can interact with cell surface receptors, which can allow the antibody to activate the immune system. In IgG, IgA, and IgD antibody isotypes, the Fc region consists of two identical protein fragments that are derived from the second and third constant domains of the two heavy chains of the antibody, respectively; the Fc regions of IgM and IgE contain three heavy chain constant domains (C H domains 2-4) in each polypeptide chain. The Fc region of IgG bears a highly conserved N-glycosylation site. Glycosylation of the Fc fragment may be crucial for Fc receptor-mediated activities. The N-glycans attached to this site are mainly complex-type core fucosylated biantennary structures. Examples of antibody fragments include, but are not limited to, (1) Fab fragments, which are composed of V L 、V H 、C L and C H(1) Monovalent fragments consisting of 1 domain, (2) F(ab')2 fragments, which are divalent fragments comprising two Fab fragments linked by disulfide bonds in the hinge region, (3) Fv fragments, which consist of the V L and V H domains, (4) Fab' fragments, which are generated by disrupting the disulfide bonds of F(ab')2 fragments under mild reducing conditions, (5) disulfide-stabilized Fv fragments (dsFv), and (6) single-domain antibodies (sdAb), which are single variable regions (V H or V L ) of an antibody that specifically bind an antigen.

[0083] Although the constant regions of the light and heavy chains may not directly participate in the binding of an antibody to an antigen, the constant regions can affect the orientation of the variable regions. The constant regions can also exhibit various effector functions, such as participating in antibody-dependent complement-mediated lysis or antibody-dependent cytotoxicity by interacting with effector molecules and cells.

[0084] Antibodies can also include chimeric antibodies, humanized antibodies, and recombinant antibodies, human antibodies produced from transgenic non-human animals, and antibodies selected from libraries using enrichment techniques available to those skilled in the art.

[0085] Antibodies can be proteins found in the blood or other body fluids of vertebrates and are used by the immune system to identify and neutralize foreign substances such as bacteria and viruses. Antibodies can include monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies and polyreactive antibodies), and antibody fragments. Thus, antibodies can include, but are not limited to, any specific binding member, immunoglobulin class, and / or isotype (e.g., IgG1, IgG2, IgG3, IgG4, IgM, IgA, IgD, IgE, and IgM); and biologically relevant fragments thereof or specific binding members thereof, including but not limited to Fab, F(ab')2, Fv, and scFv (single-chain or related entities). Antibody fragments are produced by recombinant DNA technology or by enzymatic or chemical cleavage of intact antibodies. Antibodies other than "bispecific" or "bifunctional" antibodies are understood to have identical binding sites at each of their binding sites. Monoclonal antibodies can be obtained from a population of antibodies that are substantially isotypic, i.e., each antibody constituting the population is identical except for possibly minor naturally occurring mutations that may be present. Polyclonal antibodies can be preparations that include different antibodies directed against different determinants (epitopes).

[0086] Antibody biomarkers

[0087] Antibodies can be used as biomarkers for disease diagnosis or prognosis. In some embodiments, the antibody can be a biomarker for cancer diagnosis or prognosis. In some embodiments, the antibody can be a biomarker for infectious disease diagnosis or prognosis. In some embodiments, the antibody can be a biomarker for autoimmune disease diagnosis or prognosis.

[0088] Circulating antibodies produced by a patient's own immune system upon exposure to cancer proteins can be biomarkers for early detection of cancer. In some embodiments, such circulating antibodies are autoantibodies. One advantage of autoantibodies as biomarkers is that they can be produced in large amounts despite the relatively small amount of the corresponding antigen. Due to limited proteolysis and clearance, autoantibodies are also expected to have a persistent concentration and a long half-life. The immune system constantly monitors the invasion of microorganisms and foreign molecules in the human body. A tightly regulated network of antibodies, T lymphocytes, antigen-presenting cells, cytokines, and microenvironmental signals ensures the development of appropriate targeted immune responses against infections. B lymphocytes recognize foreign extracellular and surface antigens and produce responses by secreting antibodies. To elicit a sustained antibody response, B cells can utilize additional signals from T helper cells, which present the relevant antigen in the form of peptide fragments 15-25 amino acids in length complexed with class II major histocompatibility complex (MHC). Antigens can also stimulate CD8+ T lymphocytes. These cells are activated by certain intracellular and membrane proteins that are processed through the endogenous processing pathway and presented as peptides 8-12 amino acids in length complexed with class I MHC. These two systems are highly coordinated, and in most cases, a high-affinity immunoglobulin G (IgG) antibody response may require both B and T lymphocytes to recognize the antigen. In the early stages of the development of the immune system, it is estimated that more than half of the newly generated B cell receptors are capable of binding to self-antigens. However, most autoreactive B cells are eliminated during B cell maturation, preventing mature B cells from reacting with self-molecules. This selection provides the basis for the development of self-tolerance, the ability of the immune system to recognize and ignore the body's own cells and tissues. Sometimes this mechanism fails due to overexpression, mutation, altered protein half-life, misfolding, abnormal degradation of self-proteins, or changes in post-translational modifications of proteins (e.g., glycosylation and phosphorylation), and the immune system reacts with self-antigens. Autoantibodies have long been recognized in autoimmune diseases, including systemic lupus erythematosus, myasthenia gravis, and rheumatoid arthritis. In some diseases, autoantibodies play an important role in their pathogenesis (e.g., myasthenia gravis). For example, in rheumatoid arthritis, testing for anti-IgG antibodies (also known as rheumatoid factors) is useful, with a sensitivity of approximately 80%. Autoantibodies against various antigens can be detected in cancer patients. The antigen is mainly present in cancer cells and rarely in healthy cells.

[0089] Examples of tumor - associated antibodies that can be used as biomarkers include, but are not limited to, anti - IL6, anti - IL8, anti - CA - 125, anti - c - myc, anti - p53, anti - CEA, anti - CA 15 - 3, anti - MUC - 1, anti - survivin, anti - bHCG, anti - osteopontin, anti - PDGF, anti - Her2 / neu, anti - Akt1, and anti - cytokeratin 19 antibodies. In some embodiments, the presence of abnormal levels of two or more antibodies in a sample from a patient's blood indicates the presence of cancer in the patient. Arrays can also be provided to quantify the levels of antibody biomarkers in a patient's blood. In some embodiments, a method of predicting the onset of cancer includes determining the change in concentration over time of two or more antibodies in a sample from a patient's blood.

[0090] Target binding unit

[0091] Compositions and methods are provided herein that include a target - binding unit. The target - binding unit can be a structural arrangement of two binding elements (e.g., peptide or peptidomimetic sequences), each binding element including a target - binding region. For example, the target - binding unit can be a structural arrangement of two peptide sequences. As used herein, a binding element can refer to a potential immunoreceptor - binding molecule, or PIRM. In various embodiments, the peptide sequences can be replaced with peptidomimetic sequences. The target - binding unit can be the minimal unit desired to bind to an analyte molecule (e.g., an antibody or a fragment thereof).

[0092] According to one aspect of the present disclosure, the target binding unit includes a first peptide sequence and a second peptide sequence, wherein the first peptide sequence and the second peptide are separated by a spacer and spaced apart such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule including the antigen-binding domain of an antibody. The first peptide sequence may have the same sequence as the second peptide sequence. The first peptide sequence may have a sequence different from that of the second peptide sequence. For example, the sequence of the first peptide sequence is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99% or 100% identical to the sequence of the second peptide sequence. In some embodiments, the first peptide sequence includes a first binding region, and the second peptide sequence includes a second binding region. As used herein, "binding region" and "target binding region" may be used interchangeably and refer to the region where the binding element interacts with the analyte, e.g., an epitope. In some embodiments, the first binding region and the second binding region are separated by a spacer and spaced apart such that the first binding region and the second binding region simultaneously bind to a single molecule including the antigen-binding domain of an antibody. The binding region may be a part of a peptide sequence that physically interacts with or attaches to the antigen-binding domain of a molecule. For example, the binding region may be an epitope that directly interacts with or attaches to the Fab of an antibody. In some embodiments, the first binding region and the second binding region have the same sequence. In some embodiments, the first binding region and the second binding region have the same epitope. In some embodiments, the first binding region and the second binding region have the same structure recognized by a single molecule.

[0093] According to another aspect of the present disclosure, the target binding unit includes a first class of peptide sequences and a second class of peptide sequences, wherein the first class of peptide sequences and the second class of peptide sequences simultaneously bind to a single molecule including the antigen-binding domain of an antibody.

[0094] In some embodiments, the first peptide sequence, the second peptide sequence, and the spacer are adjacent and included in a single peptide chain. For example, the C-terminus of the first peptide sequence and the N-terminus of the spacer, and the C-terminus of the spacer and the N-terminus of the second peptide sequence are connected by peptide bonds.

[0095] In some embodiments, the first peptide sequence, the second peptide sequence, and the spacer are not included in a single peptide chain. For example, the spacer may serve as a scaffold, and the first peptide sequence and the second peptide sequence are respectively attached to opposite ends of the spacer. In some embodiments, the first peptide sequence, the second peptide sequence, and the spacer are connected by non-peptide bonds.

[0096] The spacer can be any type of polymer and can be used to separate the first peptide sequence and the second peptide sequence such that the first peptide sequence and the second peptide sequence bind simultaneously to a single molecule of the analyte. The spacer can include one or more strands of the polymer. In some embodiments, the spacer includes one or more strands of the same type of polymer. In some embodiments, the spacer includes one or more strands of different types of polymers. In some embodiments, the polymer can be a polypeptide, a polynucleotide, or a polyalkylene glycol. Polyalkylene glycols can include polyethylene glycol (PEG) and polypropylene glycol (PPG). Other examples of polymers include, but are not limited to, cellulose acetate (including cellulose diacetate), ethylene vinyl alcohol copolymer, hydrogels (e.g., acrylic resins), polyacrylonitrile, polyvinyl acetate, cellulose acetate butyrate, nitrocellulose, copolymers of urethane / carbonate, copolymers of styrene / maleic acid, polylactide, polyglycolide, polycaprolactone, polyanhydrides, polyamides, polyurethanes, polyesteramides, polyorthoesters, polydioxanes, polyacetals, polyketals, polycarbonates, polyorthocarbonates, polyphosphazenes, polyhydroxybutyrate, polyhydroxyvalerate, polyalkylene oxalates, polyalkylene succinates, poly(malic acid), poly(amino acids), polyvinylpyrrolidone, polyethylene glycol, polyhydroxycellulose, chitin, chitosan, and copolymers, terpolymers, fibrin, gelatin, collagen, and any combination thereof.

[0097] In some cases, two peptide or peptidomimetic sequences are linked to the spacer by a linker. In some cases, the linker is formed by a first reactive group on the peptide or peptidomimetic sequence and a second reactive group on the spacer. For example, the spacer can include two reactive groups that can be linked to two peptide or peptidomimetic sequences.

[0098] The spacer can be an unstructured linker. In some embodiments, the spacer folds into a structure. The structure can be a double helix structure of a polynucleotide, or any secondary or tertiary structure of a polypeptide or polynucleotide. In some embodiments, the spacer includes a pre-designed amino acid sequence. For example, the pre-designed amino acids can include multiple sets of glycine and serine repeats, such as (Gly4Ser). n, where n is a positive integer equal to or greater than 1. In some embodiments, the spacer comprises a folded polypeptide. In some embodiments, the spacer comprises a secondary structure and / or a tertiary structure. In some embodiments, the spacer comprises a coiled-coil structure or a β-sheet structure. In some embodiments, the spacer comprises two or more separate peptide chains, wherein at least one of the two or more separate peptide chains comprises an α-helix or a β-strand. In some embodiments, the spacer comprises two separate peptide chains, wherein each of the two peptide chains comprises an α-helix, and the two α-helices form a coiled-coil structure. In some embodiments, the spacer comprises two separate peptide chains, wherein each of the two peptide chains comprises a β-strand, and the two β-strands form a β-sheet. In some embodiments, the spacer comprises a single peptide chain folded into at least two α-helices or at least two β-strands. In some embodiments, the spacer comprises a single peptide chain folded into a coiled-coil structure or a β-sheet structure. In some embodiments, the spacer comprises a single peptide chain, and the single peptide chain is folded into four α-helices or four β-strands.

[0099] In some embodiments, the spacer comprises a coiled-coil structure formed by three separate peptide chains. For example, the spacer comprises a first peptide chain, a second peptide chain, and a third peptide chain, wherein the second peptide chain interacts with a first portion of the first peptide chain, and wherein the third peptide chain interacts with a second portion of the first peptide chain. In this case, the first peptide chain, the second peptide chain, and the third peptide chain fold into α-helices. In some embodiments, the second peptide chain and the third peptide chain are attached to opposite ends of a double-stranded polynucleotide. In some embodiments, the spacer comprises an oligonucleotide. In some embodiments, the spacer comprises an oligonucleotide, wherein the oligonucleotide is a double-stranded oligonucleotide.

[0100] The length of the spacer region can be designed to space two binding elements a certain distance apart such that the two binding elements simultaneously bind to a single molecule having an antigen-binding domain of an antibody. In some embodiments, the spacer region is a polynucleotide. In such cases, the length of the spacer region can be from 20 to 40 nucleotides, from 30 to 50 nucleotides, from 40 to 60 nucleotides, from 50 to 70 nucleotides, from 60 to 80 nucleotides, from 70 to 90 nucleotides, from 80 to 100 nucleotides, from 90 to 110 nucleotides, from 100 to 120 nucleotides, from 110 to 130 nucleotides, from 120 to 140 nucleotides, or 130 to 150 nucleotides. In some cases, the length of the spacer region can be at least 10 nucleotides, at least 15 nucleotides, at least 20 nucleotides, at least 25 nucleotides, at least 30 nucleotides, at least 35 nucleotides, at least 40 nucleotides, at least 45 nucleotides, at least 50 nucleotides, at least 55 nucleotides, at least 60 nucleotides, at least 65 nucleotides, at least 70 nucleotides, at least 75 nucleotides, at least 80 nucleotides, at least 85 nucleotides, at least 90 nucleotides, at least 95 nucleotides, at least 100 nucleotides, or more. The polynucleotide can be double-stranded or single-stranded. In the case of a double-stranded polynucleotide, the length in units of "nucleotides" can mean "base pairs". For example, a length of 20 nucleotides can mean a length of 20 base pairs. In some embodiments, the spacer region is a polypeptide. In such cases, the length of the spacer region can be from 20 to 40 amino acid residues, from 30 to 50 amino acid residues, from 40 to 60 amino acid residues, from 50 to 70 amino acid residues, from 60 to 80 amino acid residues, from 70 to 90 amino acid residues, from 80 to 100 amino acid residues, from 90 to 110 amino acid residues, from 100 to 120 amino acid residues, from 110 to 130 amino acid residues, from 120 to 140 amino acid residues, from 130 to 150 amino acid residues, from 140 to 160 amino acid residues, from 150 to 170 amino acid residues, from 160 to 180 amino acid residues, from 170 to 190 amino acid residues, or from 180 to 200 amino acid residues. In some cases, the length of the spacer region can be at least 20 amino acid residues, at least 30 amino acid residues, at least 40 amino acid residues, at least 50 amino acid residues, at least 60 amino acid residues, at least 70 amino acid residues, at least 80 amino acid residues, at least 90 amino acid residues, at least 100 residues, at least 120 amino acid residues, at least 150 amino acid residues, or more.

[0101] The target-binding unit may include two binding elements. In some embodiments, the binding element is a peptide sequence or a peptidomimetic sequence. In some embodiments, the binding element may be a peptide sequence having a length of from 4 to 30 amino acids, from 5 to 20 amino acids, from 6 to 10 amino acids, from 8 to 11 amino acids, or from 10 to 20 amino acids. In some embodiments, the binding element may be a peptide sequence having a length of from 5 to 15 amino acids, from 6 to 16 amino acids, from 7 to 17 amino acids, from 8 to 18 amino acids, from 9 to 19 amino acids, from 10 to 20 amino acids, from 11 to 21 amino acids, from 12 to 22 amino acids, from 13 to 23 amino acids, from 14 to 24 amino acids, from 15 to 25 amino acids, from 16 to 26 amino acids, from 17 to 27 amino acids, from 18 to 28 amino acids, from 19 to 29 amino acids, or from 20 to 30 amino acids. In some embodiments, the binding element may be a peptide sequence having a length of at least 5 amino acids, at least 10 amino acids, at least 15 amino acids, at least 20 amino acids, at least 25 amino acids, at least 30 amino acids, at least 35 amino acids, at least 40 amino acids, at least 45 amino acids, at least 50 amino acids, or longer. In some embodiments, the binding element is a peptide sequence having a folded structure. For example, the binding element may be a protein or a portion thereof. In such a case, the binding element may have a length of at least 50 amino acids, at least 60 amino acids, at least 70 amino acids, at least 80 amino acids, at least 90 amino acids, at least 100 amino acids, at least 150 amino acids, at least 200 amino acids, at least 250 amino acids, at least 300 amino acids, at least 400 amino acids, at least 500 amino acids, or longer.

[0102] In some embodiments, the binding element is a peptoid sequence. In some embodiments, the binding element can be a peptoid sequence having a length of from 4 to 30 residues, from 5 to 20 residues, from 6 to 10 residues, from 8 to 11 residues, or from 10 to 20 residues. In some embodiments, the binding element can be a peptoid sequence having a length of from 5 to 15 residues, from 6 to 16 residues, from 7 to 17 residues, from 8 to 18 residues, from 9 to 19 residues, from 10 to 20 residues, from 11 to 21 residues, from 12 to 22 residues, from 13 to 23 residues, from 14 to 24 residues, from 15 to 25 residues, from 16 to 26 residues, from 17 to 27 residues, from 18 to 28 residues, from 19 to 29 residues, or from 20 to 30 residues. In some embodiments, the binding element can be a peptoid sequence having a length of at least 5 residues, at least 10 residues, at least 15 residues, at least 20 residues, at least 25 residues, at least 30 residues, at least 35 residues, at least 40 residues, at least 45 residues, at least 50 residues, or longer.

[0103] Target binding portion

[0104] The target binding moiety refers to a construct comprising one or more target binding units linked together. In some cases, the target binding moiety comprises only one target binding unit, and in such cases, the target binding unit and the target binding moiety are equivalent.

[0105] Compositions and methods are provided herein that include a target binding moiety, wherein the target binding moiety comprises one or more target binding units. The target binding unit can comprise two binding elements (e.g., peptides or peptoids) separated by a spacer region.

[0106] Compositions and methods are also provided herein that include a polynucleotide-barcode-tagged target binding moiety, wherein the polynucleotide-barcode-tagged target binding moiety comprises one or more target binding units. In some cases, the polynucleotide can be used as a barcode that identifies or corresponds to the identity of the target binding moiety, e.g., the peptide or peptoid sequence of the target binding moiety. In some cases, the polynucleotide can encode (e.g., can be transcribed and / or translated into) the peptide sequence of the target binding moiety. The polynucleotide can be detected by any available method in the art to determine the sequence of the polynucleotide and thereby determine the sequence of the corresponding peptide sequence. For example, the polynucleotide can be detected by sequencing.

[0107] In some embodiments, provided herein is a polynucleotide barcoded target-binding moiety comprising a nucleic acid sequence linked via a linker to a target-binding unit, the target-binding unit comprising a first peptide sequence containing a first binding region and a second peptide sequence containing a second binding region; wherein the first binding region and the second binding region are separated by a spacer and are spaced apart such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule comprising an antigen-binding domain of an antibody; wherein the nucleic acid sequence encodes the first peptide sequence and / or the second peptide sequence.

[0108] In some embodiments, the target-binding moiety comprises a first peptide and a second peptide, wherein the first peptide and the second peptide are spaced apart such that the first peptide and the second peptide simultaneously bind to a single analyte molecule.

[0109] The target-binding moiety can comprise a linear chain or a branched chain.

[0110] In some cases, the target-binding moiety comprises a linear chain, wherein the target-binding units are adjacent in a single peptide chain. As described above, in such cases, the first peptide sequence, the second peptide sequence, and the spacer of each target-binding unit are adjacent and are included in a single peptide chain. In some embodiments, the target-binding moiety is a linear polymer chain. In some embodiments, the target-binding moiety is a linear polypeptide chain. The first peptide, the second peptide, and the spacer of the target-binding unit can be adjacent in the linear polypeptide chain. In some embodiments, the target-binding moiety is a linear polypeptide chain comprising one or more target-binding units, wherein the one or more target-binding units are adjacent in the linear polypeptide chain.

[0111] In some other cases, the target-binding moiety comprises a branched polymer chain having a backbone and two or more side chains. The backbone can be a scaffold. The two or more side chains can comprise peptide sequences. The two or more side chains can be linked to the scaffold via a bond or a linker. In some embodiments, the scaffold of the target-binding moiety is a polypeptide chain. In some embodiments, the scaffold of the target-binding moiety is a polyalkylene glycol chain. In some other embodiments, the scaffold of the target-binding moiety is a polynucleotide chain.

[0112] In some embodiments, the target binding portion comprises a scaffold. In some embodiments, the scaffold comprises a polypeptide chain. The polypeptide chain can be of any length. For example, the length of the polypeptide chain can be from 2 to 50,000 amino acid residues. In some embodiments, the length of the polypeptide chain can be at least about 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid residues. In some embodiments, the length of the polypeptide chain can be at least about 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500 or more amino acid residues. In some embodiments, the length of the polypeptide chain can be at most about 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500 or fewer amino acid residues.

[0113] In some embodiments, the total length of the polypeptide chain is at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 amino acids. In some embodiments, the total length of the polypeptide chain is at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1200, at least 1500, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, or at least 10000 amino acid residues. In some embodiments, the total length of the polypeptide chain is at least 10000, at least 20000, at least 30000, at least 40000, or at least 50000 amino acid residues.

[0114] In some embodiments, the total length of the polypeptide chain is at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, at most 14, at most 15, at most 16, at most 17, at most 18, at most 19, at most 20, at most 21, at most 22, at most 23, at most 24, at most 25, at most 26, at most 27, at most 28, at most 29, at most 30, at most 40, at most 50, at most 60, at most 70, at most 80, at most 90, at most 100, at most 150, at most 200, at most 250, at most 300, at most 350, at most 400, at most 450, or at most 500 amino acids. In some embodiments, the total length of the polypeptide chain is at most 500, at most 600, at most 700, at most 800, at most 900, at most 1000, at most 1200, at most 1500, at most 2000, at most 3000, at most 4000, at most 5000, at most 6000, at most 7000, at most 8000, at most 9000, or at most 10000 amino acid residues. In some embodiments, the total length of the polypeptide chain is at most 10000, at most 20000, at most 30000, at most 40000, or at most 50000 amino acid residues.

[0115] In some embodiments, the scaffold comprises a polynucleotide chain. The polynucleotide chain can be of any length. In some embodiments, the length of the polynucleotide chain is from 1 nucleotide to about 3,000,000 nucleotides, from 1 nucleotide to about 2,500,000 nucleotides, from 1 nucleotide to about 2,000,000 nucleotides, from 1 nucleotide to about 1,500,000 nucleotides, from 1 nucleotide to about 1,000,000 nucleotides, from 1 nucleotide to about 500,000 nucleotides, from 1 nucleotide to about 250,000 nucleotides, from 1 nucleotide to about 200,000 nucleotides, or from 1 nucleotide to about 150,000 nucleotides. Examples of polynucleotides can also be from 1 nucleotide to about 100,000 nucleotides, from 1 nucleotide to about 10,000 nucleotides, from 1 nucleotide to about 5,000 nucleotides, from 4 nucleotides to about 2,000 nucleotides, from 6 nucleotides to about 2,000 nucleotides, from 10 nucleotides to about 1,000 nucleotides, from 10 nucleotides to about 500 nucleotides, from 10 nucleotides to about 300 nucleotides, from 10 nucleotides to about 200 nucleotides, or from 20 nucleotides to about 100 nucleotides, and any range or value therebetween, whether overlapping or not. In some embodiments, the length of the polynucleotide chain can be 5 to 20 nucleotides, 10 to 30 nucleotides, 15 to 35 nucleotides, 20 to 40 nucleotides, 25 to 50 nucleotides, 30 to 60 nucleotides, 50 to 100 nucleotides, 60 to 150 nucleotides, 100 to 200 nucleotides, 200 to 300 nucleotides, 300 to 400 nucleotides, or 400 to 500 nucleotides.

[0116] In some embodiments, the scaffold is a polyalkylene glycol chain. As used herein, the term “polyalkylene glycol” or “poly(alkylene glycol)” refers to polyethylene glycol (poly(ethylene glycol)), polypropylene glycol, polybutylene glycol, and derivatives thereof. For example, other exemplary embodiments are listed in commercial vendor catalogs, e.g., the catalog “Polyethylene Glycol and Derivatives for Biomedical Applications” (2001) of Shearwater Corporation.

[0117] A variety of linkers can be used to attach a peptide sequence to the scaffold. As described above, in the case of a target binding unit, a variety of linkers can be used to attach a peptide sequence to a spacer. In some embodiments, the linker is formed by a first reactive group on the peptide sequence and a second reactive group on the scaffold or spacer.

[0118] In some embodiments, the linker is an organic moiety that connects two portions of the compound. Such linkers can include a direct bond or atoms such as oxygen or sulfur, units such as NH, C(O), C(O)NH, SO, SO2, SO2NH, SS, or chains of atoms such as substituted or unsubstituted C1-C6 alkyl, substituted or unsubstituted C2-C6 alkenyl, substituted or unsubstituted C2-C6 alkynyl, substituted or unsubstituted C6-C12 aryl, substituted or unsubstituted C5-C12 heteroaryl, substituted or unsubstituted C5-C12 heterocyclic, substituted or unsubstituted C3-C12 cycloalkyl, where one or more methylenes can be interrupted or terminated by O, S, S(O), SO2, NH, C(O).

[0119] The linker can be a cleavable linker. Examples of cleavable linking groups include, but are not limited to, redox-cleavable linking groups (e.g., —S—S— and —C(R)2—S—S—, where R is H or C1-C6 alkyl and at least one R is C1-C6 alkyl such as CH3 or CH2CH3); phosphate-based cleavable linking groups (e.g., —O—P(O)(OR)—O—, —O—P(S)(OR)—O—, —O—P(S)(SR)—O—, —S—P(O)(OR)—O—, —O—P(O)(OR)—S—, —S—P(O)(OR)—S—, —O—P(S)(ORk)-S—, —S—P(S)(OR)—O—, —O—P(O)(R)—O—, —O—P(S)(R)—O—, —S—P(O)(R)—O—, —S—P(S)(R)—O—, —S—P(O)(R)—S—, —O—P(S)(R)—S—, —O—P(O)(OH)—O—, —O—P(S)(OH)—O—, —O—P(S)(SH)—O—, —S—P(O)(OH)—O—, —O—P(O)(OH)—S—, —S—P(O)(OH)—S—, —O—P(S)(OH)—S—, —S—P(S)(OH)—O—, —O—P(O)(H)—O—, —O—P(S)(H)—O—, —S—P(O)(H)—O—, —S—P(S)(H)—O—, —S—P(O)(H)—S— and —O—P(S)(H)—S—, where R is optionally substituted linear or branched C1-C 10 alkyl); acid-cleavable linking groups (e.g., hydrazone, ester, and esters of amino acids, —C═NN— and —OC(O)—); ester-based cleavable linking groups (e.g., —C(O)O—); peptide-based cleavable linking groups (e.g., linking groups cleaved by enzymes such as peptidases and proteases, e.g., —NHCHR A C(O)NHCHR B C(O)—, where R Aand R B are the R groups of two adjacent amino acids). Peptide-based cleavable linking groups include two or more amino acids. In some embodiments, the peptide-based cleavage linkage includes an amino acid sequence that is a substrate for a peptidase or protease found in a cell.

[0120] In addition to covalent linkages, the two portions of a compound can be joined together by an affinity binding pair. The term "affinity binding pair" or "binding pair" refers to a first molecule and a second molecule that specifically bind to each other. One member of the binding pair is conjugated to the first portion to be linked, and the second member is conjugated to the second portion to be linked. As used herein, the term "specifically bind" means that the first member of the binding pair binds to the second member of the binding pair with greater affinity and specificity than to other molecules.

[0121] Exemplary binding pairs include any hapten or antigenic compound in combination with its corresponding antibody or binding portion or fragment thereof (e.g., digoxin and anti-digoxin; murine immunoglobulins and goat anti-murine immunoglobulins) and non-immune binding pairs (e.g., biotin-avidin, biotin-streptavidin, biotin-neutravidin, hormones (e.g., thyroxine and cortisol-hormone binding protein), receptor-receptor agonist, receptor-receptor antagonist (e.g., acetylcholine receptor-acetylcholine or analogs thereof), IgG-protein A, IgG-protein G, IgG-synthetic protein AG, lectin-carbohydrate, enzyme-enzyme cofactor, enzyme-enzyme inhibitor, and complementary oligonucleotide pairs capable of forming nucleic acid duplexes), etc. The binding pair can also include a first molecule with a negative charge and a second molecule with a positive charge.

[0122] An example of conjugation using a binding pair is biotin-avidin, biotin-streptavidin, or biotin-neutravidin conjugation. In this method, one molecule (e.g., a peptide) can be biotinylated, and the other (e.g., a scaffold) can be conjugated to avidin or streptavidin. As another example, the scaffold is biotinylated, and the peptide is conjugated to avidin or streptavidin. Many commercial kits are available for biotinylating molecules.

[0123] Another example of conjugation using binding pairs is the biotin-sandwich method. See, e.g., Davis et al., Proc. Natl. Acad. Sci. USA, 103:8155-60 (2006). Two molecules to be conjugated together can be biotinylated and then conjugated together using at least one tetravalent avidin-like molecule (e.g., avidin, streptavidin, or neutravidin) as a linker. Thus, in some embodiments, both the peptide and the scaffold can be biotinylated and then linked together using an avidin-like molecule (e.g., avidin, streptavidin, or neutravidin). In some embodiments, neutravidin and / or streptavidin are used as linkers to bridge the biotinylated molecules together.

[0124] Another example of conjugation using binding pairs is double-stranded nucleic acid conjugation. In this method, a first portion to be joined can be conjugated to a first strand of a double-stranded nucleic acid, and a second portion to be joined can be conjugated to a second strand of the double-stranded nucleic acid. Nucleic acids can include, but are not limited to, defined sequence segments and sequences that include nucleotides, ribonucleotides, deoxyribonucleotides, nucleotide analogs, modified nucleotides, and nucleotides, groups, or bridges that include backbone modifications, branch points, and nucleotide residues.

[0125] Another example of conjugation using binding pairs is coiled-coil conjugation. In this method, a first portion to be joined is conjugated to a first peptide chain that folds into an α-helix, and a second portion to be joined is conjugated to a second peptide chain that folds into an α-helix. The two α-helices interact to form a coiled-coil structure to join the two portions together.

[0126] Other examples of binding pair conjugation include HaloTag, CLIP-Tag, and SNAP-Tag.

[0127] In some embodiments, the linker can be a linker molecule. Examples of linker molecules can include, but are not limited to, polymers, sugars, nucleic acids, peptides, proteins, hydrocarbons, lipids, polyethylene glycols, cross-linking agents, or combinations thereof.

[0128] Non-limiting examples of crosslinking agents that can be used as linker molecules can include, but are not limited to, amine-amine crosslinking agents (e.g., but not limited to those based on NHS-ester and / or imidoester reactive groups), amine-thiol crosslinking agents, carboxyl-amine crosslinking agents (e.g., but not limited to carbodiimide crosslinking agents such as DCC and / or EDC (EDAC); and / or N-hydroxysuccinimide (NHS)), photoreactive crosslinking agents (e.g., but not limited to aryl azides, bisaziridines, and any photoreactive (photoactivatable) chemical crosslinking agents recognized in the art), thiol-carbohydrate crosslinking agents (e.g., but not limited to those based on maleimide and / or hydrazide reactive groups), thiol-hydroxy crosslinking agents (e.g., but not limited to those based on maleimide and / or isocyanate reactive groups), thiol-thiol crosslinking agents (e.g., but not limited to maleimide and / or pyridyl disulfide reactive groups), sulfo-SMCC crosslinking agents, sulfo-SBED biotin labeling transfer reagents, thiol-based biotin labeling transfer reagents, photoreactive amino acids (e.g., but not limited to bisaziridine analogs of leucine and / or methionine), NHS-azide Staudinger ligation reagents (e.g., but not limited to activated azide compounds), NHS-phosphine Staudinger ligation reagents (e.g., but not limited to activated phosphine compounds), and any combination thereof.

[0129] Examples of suitable reactive groups include electrophiles or nucleophiles that can react with corresponding nucleophiles or electrophiles, respectively, on a matrix of interest to form covalent bonds. Non-limiting examples of suitable electrophilic reactive groups can include, for example, esters including activated esters (e.g., succinimide esters), amides, acrylamides, acyl azides, acyl halides, acyl nitriles, aldehydes, ketones, haloalkanes, alkyl sulfonates, acid anhydrides, aryl halides, aziridines, borates, carbodiimides, diazoalkanes, epoxides, haloacetamides, haloplatinates, halotriazines, imidoesters, isocyanates, isothiocyanates, maleimides, phosphoramidites, silyl halides, sulfonates, sulfonyl halides, etc. Non-limiting examples of suitable nucleophilic reactive groups can include, for example, amines, anilines, thiols, alcohols, phenols, hydrazine, hydroxylamine, carboxylic acids, ethylene glycol, heterocycles, etc.

[0130] Encoding portion

[0131] In various embodiments, provided herein is a coding portion linked to a target binding moiety. The coding portion can be a polynucleotide or nucleic acid sequence. The coding portion linked to the target binding moiety can refer to a "polynucleotide barcoded target binding moiety".

[0132] According to one aspect, the coding portion is a single-stranded polynucleotide. According to another aspect, the coding portion is a double-stranded polynucleotide. In some cases, the coding portion can include a single-stranded first polynucleotide portion and a double-stranded second polynucleotide portion. The coding portion can be deoxyribonucleic acid, ribonucleic acid, or a combination thereof.

[0133] The length of the coding portion can vary. According to certain aspects, the length of the coding portion can be from 1 nucleotide to about 3,000,000 nucleotides, from 1 nucleotide to about 2,500,000 nucleotides, from 1 nucleotide to about 2,000,000 nucleotides, from 1 nucleotide to about 1,500,000 nucleotides, from 1 nucleotide to about 1,000,000 nucleotides, from 1 nucleotide to about 500,000 nucleotides, from 1 nucleotide to about 250,000 nucleotides, from 1 nucleotide to about 200,000 nucleotides, or from 1 nucleotide to about 150,000 nucleotides. Examples of the target polynucleotide can also be from 1 nucleotide to about 100,000 nucleotides, from 1 nucleotide to about 10,000 nucleotides, from 1 nucleotide to about 5,000 nucleotides, from 4 nucleotides to about 2,000 nucleotides, from 6 nucleotides to about 2,000 nucleotides, from 10 nucleotides to about 1,000 nucleotides, from 10 nucleotides to about 500 nucleotides, from 10 nucleotides to about 300 nucleotides, from 10 nucleotides to about 200 nucleotides, or from 20 nucleotides to about 100 nucleotides, and any range or value therebetween, whether overlapping or not. In some cases, the length of the coding portion can be equal to or greater than about 100, 200, 300, 400, 500, 600, 700, 800, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 30,000, or more nucleotides. In some cases, the length of the coding portion can be equal to or greater than about 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more nucleotides. For a double-stranded polynucleotide, the "nucleotide length" refers to the "number of base pairs" or the length of one of the two strands in the double-stranded polynucleotide.

[0134] In some embodiments, the coding portion is a polynucleotide that includes a barcode. The barcode can be used to determine the identity of the corresponding peptide or peptidomimetic sequence on the target binding portion. By sequencing or otherwise detecting the barcode, the identity (e.g., sequence) of the peptide or peptidomimetic that makes up the target binding region can be determined.

[0135] The encoding portion can encode a peptide sequence of the target-binding portion. In various embodiments, the encoded peptide sequence is a binding element. In some embodiments, the encoded peptide sequence is a binding element fused to a peptide linker. In some embodiments, the encoding portion is a polynucleotide encoding a binding element. In some embodiments, the encoding portion is a polynucleotide encoding a binding element and a linker such that the encoded binding element can be fused to a peptide linker after expression. In some embodiments, the encoding portion is a polynucleotide encoding a peptide sequence. In such cases, the polynucleotide can be transcribed and / or translated into a polypeptide sequence. And in such cases, the target-binding portion is transcribed and / or translated from the polynucleotide. In some embodiments, polynucleotide barcoded target-binding portions can be made by RNA display. In some embodiments, polynucleotide barcoded target-binding portions can be made by in vitro compartmentalization, e.g., by using emulsions.

[0136] Manufacture of the target binding portion

[0137] Target-binding portions with or without an attached encoding portion can be made by several different methods. In some cases, the target-binding portion is not attached to the encoding portion, and in such cases, various methods can be used to attach a peptide or peptidomimetic sequence to a scaffold as described herein, e.g., via a linker and a binding pair. In some cases, the target-binding portion is attached to a solid surface, and in such cases, various methods can be used to attach the target-binding portion to the solid surface. These methods include, but are not limited to, using a linker generated from two functional groups to attach two different entities. In some other cases, the target-binding portion is attached to an encoding portion, e.g., a polynucleotide, to form a polynucleotide barcoded target-binding portion. Three non-limiting examples of making polynucleotide barcoded target-binding portions are described herein, including RNA display, in vitro compartmentalization, and split-and-pool synthesis.

[0138] As described herein, in some embodiments, the target binding moiety is a polypeptide chain. To prepare a library of target binding moieties, a library of nucleic acid templates can first be prepared. The library of nucleic acid templates can then be used for in vitro transcription and / or translation to prepare a library of target binding moieties. For example, a library of DNA templates can be prepared and then transcribed into RNA templates, which are then translated into polypeptide chains, each polypeptide chain having two or more binding elements. As another example, a library of RNA templates can be prepared and then translated into polypeptide chains, each polypeptide chain having two or more binding elements. The library of nucleic acid templates can be prepared using the methods shown in the Examples section (e.g., Examples 3 and 11). In some cases, each nucleic acid template in the nucleic acid template library can include two or more subsequences that have the same nucleic acid sequence or encode the same peptide sequence. As used herein, "library size" can be used to indicate how many unique species (e.g., target binding moieties or unique binding element sequences) a nucleic acid template library can produce. A unique target binding moiety can refer to a target binding moiety having a unique binding element sequence. In a library of target binding moieties, the sequences of the spacer and the scaffold can be the same in all target binding moieties. For example, the library size of the nucleic acid templates can be at least about 10 3 species, at least about 10 4 species, at least about 10 5 species, at least about 10 6 species, at least about 10 7 species, at least about 10 8 species, at least about 10 9 species, at least about 10 10 species, at least about 10 11 species, at least about 10 12 species, at least about 10 13 species, at least about 10 14 species, at least about 10 15 species, at least about 10 16 species, at least about 10 3 species, at least about 10 4 species, at least about 10 5 species, at least about 10 6 species, at least about 10 7 species, at least about 10 8 species, at least about 10 9 species, at least about 10 10 species, at least about 10 11 species, at least about 10 12 species, at least about 1013 species, at least about 10 14 species, at least about 10 15 species, at least about 10 16 species or more species.

[0139] RNA display

[0140] Compositions containing the target-binding moieties described herein can be prepared by RNA display. Using RNA display, a nucleic acid sequence encoding a peptide sequence can be physically linked to the peptide sequence it encodes. In RNA display, the encoded polypeptide (e.g., a protein or peptide) can be covalently attached to the RNA, e.g., using 3'-puromycin-labeled RNA. Puromycin is a translation inhibitor that can enter the ribosome during translation and form a stable covalent bond with the nascent protein / peptide. This can allow for the formation of a stable covalent linkage between the RNA display template and the protein / peptide it encodes, resulting in the RNA-displayed protein / peptide.

[0141] In RNA display, members of an RNA library can be directly attached to the polypeptide of interest they encode, e.g., by stable covalent linkage to puromycin, an antibiotic that can mimic the aminoacyl terminus of tRNA. Puromycin is an aminonucleoside antibiotic that is active against both prokaryotes and eukaryotes and is derived from Streptomyces alboniger. During translation that occurs in the ribosome, premature chain termination can inhibit peptide synthesis. A portion of the molecule can serve as an analogue of the 3'-end of tyrosyl-tRNA, where a part of its structure can mimic an adenosine molecule and another part can mimic a tyrosine molecule. It can enter the A site and transfer to the growing chain, resulting in the formation of a puromycylated nascent chain and premature chain release. The 3'-position can contain an amide bond rather than the normal ester bond of tRNA, making the molecule more resistant to hydrolysis and preventing progression along the ribosome.

[0142] Other puromycin analogues or derivative inhibitors of protein synthesis include O-demethylpuromycin, O-propargyl-puromycin, 9-{3'-deoxy-3'-[(4-methyl-L-phenylalanyl)amino]-β-D-ribofuranosyl}-6-(N,N'-dimethylamino)purine [L-(4-Me)-Phe-PANS], and 6-dimethylamino-9-[3-(p-azido-L-β-phenylalanyl-amino)-3-deoxy-β-ribofuranosyl]purine.

[0143] Members of the RNA library can be linked to puromycin via a linker (such as, but not limited to, a polynucleotide or a chemical linker, e.g., polyethylene glycol). In some embodiments, the polynucleotide linker is linked to puromycin at the 3'-end. In other embodiments, the PEG linker is linked to puromycin. When the puromycin linked to the 3'-end of the RNA molecule enters the ribosome, due to the peptidyl transferase activity in the ribosome, it can establish a covalent bond with the nascent protein (encoded by the RNA molecule). In turn, a stable amide bond can be formed between the protein and the O-methyltyrosine moiety of puromycin. The RNA library of the RNA-puromycin fusion can be translated in vitro as described herein to produce RNA-puromycin-protein complexes (e.g., polynucleotide barcoded target-binding moieties).

[0144] Affinity selection can be performed on a peptide library displayed on RNA to screen for proteins / peptides with a given property, such as specific binding to an analyte. RNA display can be performed in solution or on a solid support. The selected RNA-displayed proteins / peptides of interest can be purified by standard methods known in the art, such as affinity chromatography. The RNA can be cloned, PCR amplified, and / or sequenced to determine the coding sequence of the selected protein / peptide of interest. In some aspects, members of a nucleic acid member library are linked to puromycin, where each member encodes an expected protein / peptide of interest, and the primary amino acid sequence of the protein / peptide is different from other proteins / peptides encoded by other nucleic acid members. The RNA display system can contain a population or mixture of different complexes such that each complex has a different RNA linked to puromycin, ribosome, and the expected protein / peptide of interest.

[0145] As provided herein, RNA display can be used to generate a library of nucleic acid-peptide fusion molecules. A library of a variety of nucleic acid-peptide fusion molecules can be provided, where each molecule of the library can include a coding nucleic acid linked to an encoded peptide, and where one or more peptides of the library can contain at least one unnatural amino acid residue. As used herein, the term "coding nucleic acid" refers to a nucleotide sequence that includes two or more codons that can be translated into a peptide. Similarly, the term "encoded peptide" refers to an amino acid sequence that can be translated from the coding nucleic acid. In some cases, the coding nucleic acid is an RNA molecule, which can be any RNA molecule that can serve as a translation template for the peptides disclosed herein, e.g., an mRNA molecule transcribed from a DNA or RNA template, or a chemically synthesized RNA molecule.

[0146] According to one aspect, provided herein is a method that includes translating an RNA sequence of an RNA, where the RNA sequence encodes a peptide sequence, where the RNA is linked at its 3′ end to a peptide receptor, and linking the peptide receptor to amino acid residues of a translated peptide that includes the peptide sequence, thereby forming a nucleic acid-peptide fusion molecule. In some embodiments, the nucleic acid-peptide fusion molecule includes a polynucleotide barcoded target-binding moiety that includes: (i) a first peptide sequence that includes a first binding region, and (ii) a second peptide sequence that includes a second binding region; where the first binding region and the second binding region (i) are separated by a spacer region, and (ii) are spaced a distance such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule that includes an antigen-binding domain of an antibody. In some instances, the method further includes providing a DNA, where the DNA encodes the RNA. In some instances, the method further includes transcribing the DNA. In some instances, the method can further include reverse transcribing the RNA. A DNA strand produced by reverse transcribing single-stranded RNA can hybridize to the single-stranded RNA to stabilize the single-stranded RNA. In some embodiments, the DNA molecule and / or the RNA sequence encodes the first peptide sequence, the second peptide sequence, and the spacer region. In some embodiments, the nucleic acid-peptide fusion molecule includes a plurality of nucleic acid-peptide fusion molecules. In some instances, each nucleic acid-peptide fusion molecule of the plurality of nucleic acid-peptide fusion molecules includes a unique nucleic acid sequence. The first binding region and the second binding region can have the same sequence. The first binding region and the second binding region can have the same structure that is recognized by a single molecule. The spacer region of each nucleic acid-peptide fusion molecule can be the same or include the same amino acid sequence. In some embodiments, the spacer region includes a folded polypeptide, a secondary structure, and / or a tertiary structure. For example, the spacer can include a coiled-coil structure or a β-sheet structure. In some instances, the first peptide sequence and the second peptide sequence have the same sequence. In some instances, the sequence of the first peptide sequence is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, or 100% identical to the sequence of the second peptide sequence. The antigen-binding domain can be a fragment of an antibody. The antigen-binding domain can be an scFv, a Fab, or an F(ab)2. In some instances, the first peptide sequence and the second peptide sequence are 4 to 30 amino acids in length, 5 to 20 amino acids in length, 6 to 10 amino acids in length, 8 to 11 amino acids in length, or 10 to 20 amino acids in length. In some instances, the first peptide sequence or the second peptide sequence is at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more amino acids in length. In some embodiments, the translation includes in vitro translation.

[0147] Advantages of the RNA display system can be that the encoding nucleic acid is translationally linked to the peptide it encodes. As used herein, the term "translationally linked" refers to the linkage of the encoding nucleic acid to the peptide it encodes during peptide translation due to the catalytic activity of ribosomal peptidyl transferase. In some cases, the encoding nucleic acid molecule can be directly or indirectly linked to the encoded peptide during peptide translation through its 3' end (e.g., at the C-terminus of the peptide). The peptide and nucleic acid can be translationally linked using a peptide acceptor. As used herein, the term "peptide acceptor" refers to a molecule that can be added to the C-terminus of a growing (nascent) peptide chain due to the catalytic activity of ribosomal peptidyl transferase function. A peptide acceptor such as puromycin can contain a nucleotide or nucleotide-like moiety linked to an amino acid or its analog or derivative (e.g., O-methyltyrosine; see also Ellman et al., Meth. Enzymol. 202:301, 1991, which is incorporated herein by reference), such as adenosine or an adenosine analog (e.g., adenosine dimethylated at the N-6 amino position), where the linkage is, for example, an ester, amide, or ketone linkage. The peptide acceptor can also contain a nucleophile, which can be, for example, an amino, hydroxyl, or sulfhydryl group. The peptide acceptor can also include nucleotide mimetics, amino acid mimetics, or mimetics of a combined nucleotide-amino acid structure.

[0148] The peptide acceptor can be located at the 3' end of the encoding nucleic acid molecule. Thus, the peptide acceptor molecule can be positioned immediately following the final codon of the peptide coding sequence or can be separated from the final codon by a linker (e.g., an intervening nucleotide sequence, which can be DNA or RNA). In some cases, the 3' end of the peptide coding sequence or the linker (if present) includes a translation pause site. The peptide acceptor can be covalently bound to the peptide coding sequence of the nucleic acid or can be non-covalently linked, e.g., by using a second nucleotide sequence that hybridizes selectively at or near the 3' end of the peptide coding sequence and that binds to the peptide acceptor molecule itself or that bridges and hybridizes selectively at or near the 3' end of the peptide coding sequence and at or near the first end of the second nucleotide sequence, where the peptide acceptor is linked at or near the second end of the second nucleotide sequence.

[0149] An exemplary peptide acceptor is puromycin, which is similar to tyrosyl adenosine and functions to attach the growing peptide to its encoding mRNA (see U.S. Patent No. 6,281,344). Puromycin is an antibiotic that can act as a chain terminator. As an analog of aminoacyl-tRNA, puromycin can bind to the A site of the translation complex, accept the growing peptide chain, and dissociate from the ribosome (with K d = 10 -4M; Traut and Monro, J. Mol. Biol. 10:63, 1964; Smith et al., J. Mol Biol. 13:617, 1965) acts as a general inhibitor of protein synthesis. Puromycin can form a stable amide bond with the growing peptide chain, and thus produces a more stable fusion than other peptide acceptors that form, for example, less stable ester bonds. The peptidyl-puromycin molecule can contain a stable amide bond between the C-terminus of the nascent peptide (i.e., the peptide still bound in the translation complex) and the O-methyltyrosine moiety of puromycin. The O-methyltyrosine can be linked via a stable amide bond to the 3'-amino group of the modified adenosine moiety of puromycin. Accordingly, methods for translationally linking an encoding nucleic acid and an encoding peptide are disclosed herein, and which include, for example, using a peptide acceptor (such as puromycin) to effect the linkage, which can be linked at or near the 3'-end of the encoding nucleic acid such that it can enter the ribosomal complex during translation and be incorporated into the C-terminus of the growing (nascent) peptide, thereby terminating translation and linking the encoding nucleic acid and the encoding peptide.

[0150] Additional peptide acceptors useful for translationally linking an encoding nucleic acid and an encoding peptide include, for example, tRNA-like structures at the 3'-end of mRNA, and other compounds that act in a manner similar to puromycin, such as, for example, compounds that include an amino acid residue linked to an adenine or adenine-like compound (e.g., phenylalanyl-adenosine, tyrosyl adenosine, and alanyl adenosine), and amide-linked compounds such as phenylalanyl-3'-deoxy-3'-aminoadenosine, alanyl-3'-deoxy-3'-aminoadenosine, and tyrosyl-3'-deoxy-3'-aminoadenosine. An example of an adenine-like compound with a functional 3'-amino acid attachment is 7-deaza-adenosine (tubercidin) (see Krayevsky and Kukhanova, Prog. Nucl. Acids Res. 23:2-51, 1979, which is incorporated herein by reference). Such peptide acceptors can contain naturally occurring L-amino acids or contain analogs or derivatives thereof, provided that the peptide acceptor can translationally link the encoding nucleic acid and the encoding peptide. A combined tRNA-like 3'-structure-puromycin conjugate can also be used as a peptide acceptor.

[0151] In vitro compartmentalization

[0152] The compositions containing a target-binding moiety described herein can be prepared by in vitro compartmentalization as an alternative strategy to RNA display.

[0153] Many high-throughput display selection methods based on the physical linkage between a gene and its encoded protein / peptide can be used (Griffiths, A.D. and Tawfik, D.S. (2000). Curr Opin Biotechnol 11, 338-53). These methods provide a way to select proteins or peptides that bind to any given analyte. The present disclosure provides an in vitro system for compartmentalizing a larger molecular library (e.g., a library including target-binding moieties of peptides), and provides methods for selecting and isolating target-binding moieties having a given activity (e.g., binding antibodies) from such libraries. In the methods provided herein, a unique peptide sequence can be linked to a unique nucleic acid sequence in a single compartment, such that the unique nucleic acid sequence can be used to identify the peptide sequence after pooling. In some cases, the method of linking the peptide sequence and the nucleic acid sequence can be to compartmentalize the nucleic acid sequence encoding the peptide into a separate compartment and perform in vitro transcription and translation to produce two or more copies of the peptide. In some embodiments, the nucleic acid sequence is a single molecule. In some embodiments, the single molecule includes one or more cloned copies of the sequence. In some embodiments, the nucleic acid is two or more cloned copies of a single molecule. In some embodiments, the nucleic acid includes two or more cloned copies of the sequence. In some embodiments, two or more cloned copies of a single molecule can be linked to a solid support, e.g., beads. As used herein, a "cloned copy" refers to a copy derived from a single molecule template during an amplification process. In some embodiments, the nucleic acid can be RNA, and in such cases, translation is performed only in vitro. And the nucleic acid sequence can be linked to a scaffold having functional groups at each end. After two or more peptide copies are produced in the compartment, two copies of the peptide can be linked to the scaffold through the two functional groups of the scaffold. The peptide can include functional groups that can react with the functional groups on the scaffold to form a covalent bond.

[0154] In some embodiments, the compartment includes a nucleic acid sequence, wherein the nucleic acid sequence is further linked to a scaffold. In some embodiments, the scaffold has a first linkage site and a second linkage site, wherein the first linkage site is linked to a first peptide sequence and the second linkage site is linked to a second peptide sequence. The linker can be any type of suitable linker. For example, the linker can be covalent or non-covalent. For example, the linker can be a bond or a molecule. For another example, the linker can be a crosslinker or a binding pair conjugate. The peptide can be produced by in vitro transcription and translation (or only in vitro translation) of the nucleic acid sequence.

[0155] In some embodiments, multiple compartments are provided herein, each of the multiple compartments including a unique nucleic acid sequence linked to a scaffold.

[0156] The target-binding moieties linked to the coding moiety produced by in vitro compartmentalization can also be pooled together for downstream applications.

[0157] In some embodiments, provided herein is a method that includes (a) expressing a first peptide sequence encoded by a nucleic acid and a second peptide sequence encoded by a nucleic acid in each of a plurality of containers, wherein each of the plurality of containers includes a scaffold that includes a first attachment site and a second attachment site separated by a spacer; and (b) binding the first peptide sequence to the first attachment site and binding the second peptide sequence to the second attachment site, wherein the first peptide sequence bound to the first attachment site and the second peptide sequence bound to the second attachment site are spaced apart such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule that includes an antigen-binding domain of an antibody, thereby forming a plurality of target-binding moieties. In some cases, the nucleic acid is linked to the scaffold via a linker. In some cases, the 5' end or the 3' end of the nucleic acid is linked to the scaffold via a linker. In some cases, the nucleic acid is a single nucleic acid molecule. The nucleic acid can be linked to the scaffold before or after generating the plurality of containers. In some cases, the nucleic acid molecule is double-stranded or single-stranded. In some cases, the nucleic acid molecule is DNA, RNA, or a combination thereof. Expression can include transcription and / or translation. In some cases, where the template is DNA, expression includes transcription and translation. In some cases, where the template is RNA, expression includes translation. In some cases, the first peptide sequence and the second peptide sequence include the same sequence. In some cases, the sequence of the first peptide sequence is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, or 100% identical to the sequence of the second peptide sequence. In some cases, the method further includes pooling the plurality of polynucleotide-barcode-tagged target-binding moieties from the containers. In some cases, the containers are droplets. For example, the droplets are water-in-oil droplets.

[0158] The method can further include barcoding the target-binding moieties among the plurality of target-binding moieties. In some cases, barcoding includes attaching a barcode to the target-binding moiety. In some embodiments, the scaffold is attached to the barcoded polynucleotide before expression. In some embodiments, the scaffold is attached to the nucleic acid encoding the first peptide sequence and / or the nucleic acid encoding the second peptide sequence. In some embodiments, the scaffold is attached to the nucleic acid encoding the first peptide sequence and / or the nucleic acid encoding the second peptide sequence before expression.

[0159] As used herein, a compartment can be a container or a droplet.

[0160] In some cases, the methods provided herein include partitioning a composition into compartments such that in some compartments, there can be only one nucleic acid sequence in a single compartment. The nucleic acid sequence can be transcribed and / or translated into a peptide sequence. In some embodiments, at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% of the compartments contain zero or only one nucleic acid sequence. The number of partitions or compartments employed can vary depending on the application. For example, the number of partitions or compartments can be about 5, 10, 50, 100, 250, 500, 750, 1000, 1500, 2000, 2500, 5000, 7500 or 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, 2000000, 3000000, 4000000, 5000000, 10000000, 20000000 or more. The number of partitions or compartments can be at least about 1, 5, 10, 50, 100, 250, 500, 750, 1000, 1500, 2000, 2500, 5000, 7500 or 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, 2000000, 3000000, 4000000, 5000000, 10000000, 20000000 or more. The number of partitions or compartments can be less than about 5, 10, 50, 100, 250, 500, 750, 1000, 1500, 2000, 2500, 5000, 7500 or 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, 2000000, 3000000, 4000000, 5000000, 10000000 or less than about 20000000.The number of partitions or compartments can be from 5 to about 10,000,000, from 5 to about 5,000,000, from 5 to about 1,000,000, from 10 to about 10,000, from 10 to about 5,000, from 10 to about 1,000, from about 1,000 to about 6,000, from about 1,000 to about 5,000, from about 1,000 to about 4,000, from about 1,000 to about 3,000 or from about 1,000 to about 2,000.

[0161] The number of nucleic acid molecules partitioned into compartments can be about 1, 2, 3, 4, 5, 10, 50, 100, 250, 500, 750, 1,000, 1,500, 2,000, 2,500, 5,000, 7,500 or 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000, 10,000,000, 20,000,000 or more. The number of nucleic acid molecules partitioned into compartments can be at least about 1, 5, 10, 50, 100, 250, 500, 750, 1,000, 1,500, 2,000, 2,500, 5,000, 7,500 or 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000, 10,000,000, 20,000,000 or 10 8 、10 9 、10 10 、10 11 、10 12 、10 13 、10 14 、10 15 、10 16or more. The number of nucleic acid molecules partitioned into compartments can be less than about 2, 5, 10, 50, 100, 250, 500, 750, 1000, 1500, 2000, 2500, 5000, 7500 or 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, 2000000, 3000000, 4000000, 5000000, 10000000, 20000000 or 10 8 、10 9 、10 10 、10 11 、10 12 、10 13 、10 14 、10 15 or less than 10 16 。The number of nucleic acid molecules partitioned into compartments can be from 5 to about 10000000, from 5 to about 5000000, from 5 to about 1000000, from 10 to about 10000, from 10 to about 5000, from 10 to about 1000, from about 1000 to about 6000, from about 1000 to about 5000, from about 1000 to about 4000, from about 1000 to about 3000 or from about 1000 to about 2000. In some embodiments, each nucleic acid molecule has a unique sequence. In some embodiments, two or more nucleic acid molecules can have the same sequence.

[0162] In some embodiments, the partitioning is an emulsion formed passively using a microfluidic device. These methods can involve extrusion, dripping, jetting, tip-streaming, tip-multi-breaking or similar methods. Passive microfluidic droplet generation can be regulated by altering the competitiveness of two different fluids to control the number, size and diameter of the particles. These forces can be capillary forces, viscosity and / or inertial forces when the two solutions are mixed.

[0163] In some embodiments, the compartments are wells in a standard microtiter plate, separated by sorting assistance. In some embodiments, the sorter is a fluorescence-activated cell sorter (FACS). Additionally, as in the case of Fluidigm C1, the partitioning can be coupled with automated library generation in a separated microfluidic chamber.

[0164] In some embodiments, the partitioning is subnanoliter wells and the particles are sealed by a semipermeable membrane.

[0165] In some embodiments, the partitions are microfluidic droplets formed by active control of a microfluidic chip. In active control, the generation of droplets can be manipulated by applying an external force (e.g., an electric, magnetic, or centripetal force). A popular method of controlling the active manipulation of droplets in a microfluidic chip is to change the internal force by adjusting the fluid velocities of two mixed solutions (e.g., oil and water).

[0166] Split - and - pool synthesis

[0167] The target - binding moiety or the polynucleotide - barcoded target - binding moiety can also be manufactured using split - and - pool synthesis.

[0168] In some cases, a polynucleotide - barcoded target - binding moiety comprising two or more binding elements is synthesized by split - and - pool. In this method, for example, a scaffold having two binding - element initiators can be provided, where the scaffold can also be linked to a polynucleotide initiator. Different chemical building blocks can be used to synthesize two identical binding elements from the two binding - element initiators in a combinatorial manner. When synthesizing the two binding elements, short nucleotide fragments can be used as building blocks to synthesize a polynucleotide chain from the polynucleotide initiator. Each short nucleotide fragment having a unique sequence corresponds to a unique chemical building block. Since the binding element and the polynucleotide chain can be synthesized simultaneously, the complete sequence of the polynucleotide chain can be used to determine the additional order of each building block, and thus to identify the identity of the entire binding element. In some cases, the binding element is a peptide. The building blocks can be at least 5, 10, 15, 20, or more different amino acids. Each amino acid can correspond to a short DNA sequence and vice versa. After split - and - pool synthesis, the polynucleotide sequence can be used to determine the peptide sequence of the binding element. In some other cases, the binding element is a peptidomimetic, and the peptidomimetic sequence can be synthesized using split - and - pool as described herein. In some other cases, the binding element can be a chemical substance other than a peptide or a peptidomimetic, and can also be synthesized using split - and - pool synthesis as described herein.

[0169] Barcode

[0170] As described herein, the coding moiety identifies or corresponds to the target - binding moiety. For example, the coding moiety can identify or correspond to the peptide or peptidomimetic sequence of the target - binding moiety.

[0171] The encoding portion described herein can be used as a barcode. In some embodiments, the entire encoding portion serves as a barcode. In some embodiments, a portion of the encoding portion serves as a barcode. The barcode or barcode sequence can be a natural or synthetic nucleic acid sequence composed of polynucleotides, which allows for the unambiguous identification of the polynucleotide and other sequences with the same barcode sequence (e.g., a peptide sequence linked to the polynucleotide). The barcode can be of any suitable length, such as a length of 2 to 100 nucleotides. The barcode can have a random sequence or a predetermined sequence. The barcode sequence can include a sequence of at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95 or at least 100 consecutive nucleotides. The barcode sequence can include a sequence of at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, 300, 400, 500, 600 or more consecutive nucleotides. The barcode sequence can include a randomly assembled sequence of nucleotides. The barcode sequence can be a degenerate sequence. The barcode sequence can be a known sequence. The barcode sequence can be a predefined sequence. The barcode can include one or more barcode segments, where one or more of the barcode segments are contiguous or separated by one or more predefined sequences. The barcode can be single-stranded or double-stranded nucleic acid.

[0172] The length of the barcode can range from 2 to 36 nucleotides, or from 4 to 36 nucleotides, or from 6 to 30 nucleotides, or from 8 to 20 nucleotides, or from 2 to 20 nucleotides or from 4 to 20 nucleotides, or from 6 to 20 nucleotides. In some embodiments, the length of the barcode can range from 3 to 10 nucleotides, or from 10 to 50 nucleotides, or from 50 to 100 nucleotides. In certain aspects, the melting temperatures of the barcodes within a set are within 10 °C of each other, within 5 °C of each other or within 2 °C of each other. In certain aspects, the melting temperatures of the barcodes within a set are not within 10 °C of each other, within 5 °C of each other or within 2 °C of each other. In other aspects, the barcode is a member of a minimal cross-hybridizing set. For example, the nucleotide sequence of each member of such a set can be sufficiently different from the nucleotide sequence of each other member of the set such that under stringent hybridization conditions, no member can form a stable duplex with the complementary sequence of any other member. In some embodiments, the nucleotide sequence of each member of the minimal cross-hybridizing set differs from the nucleotide sequence of each other member by at least two nucleotides. Barcode technology is described in Winzeler et al. (1999) Science 285:901; Brenner (2000) Genome Biol. 1:1 Kumar et al. (2001) Nature Rev. 2:302; Giaever et al. (2004) Proc. Natl. Acad. Sci. USA 101:793; Eason et al. (2004) Proc. Natl. Acad. Sci. USA 101:11046; and Brenner (2004) Genome Biol. 5:240.

[0173] In some cases, the barcode sequence can be flanked on the 5' and / or 3' side of the barcode sequence by a predefined sequence. In some cases, the length of the barcode sequence can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 45, at least 50 or more nucleotides. In some cases, the length of the barcode sequence can be 2 to 4, 3 to 10 nucleotides, 5 to 10, 6 to 12, 10 to 15, 15 to 20, 20 to 30, 30 to 40, 40 to 50 or 10 to 50 nucleotides. The length of the predefined sequence flanking the barcode can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 45, at least 50 or more nucleotides. In some cases, the length of the predefined sequence flanking the barcode sequence can be 2 to 4, 3 to 10 nucleotides, 5 to 10, 6 to 12, 10 to 15, 15 to 20, 20 to 30, 30 to 40, 40 to 50 or 10 to 50 nucleotides.

[0174] In some embodiments, each barcode of the plurality of barcodes has at least 2 nucleotides. For example, the length of each barcode of the plurality of barcodes can be at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 60, 70, 80, 90, 100, 200, 500, or 1000 nucleotides. In some embodiments, each barcode of the plurality of barcodes has at most about 1000 nucleotides. For example, the length of the barcodes in the plurality of barcodes can be at most about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 60, 70, 80, 90, 100, 200, 500, or 1000 nucleotides. In some embodiments, each barcode of the plurality of barcodes has the same nucleotide length. For example, the length of the barcodes in the plurality of barcodes can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 60, 70, 80, 90, 100, 200, 500, or 1000 nucleotides. In some embodiments, one or more barcodes of the plurality of barcodes have nucleotides of different lengths.For example, one or more first barcodes among multiple barcodes can have approximately or at least approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 60, 70, 80, 90, 100, 200, 500, or 1000 nucleotides, and one or more second barcodes among multiple barcodes can have approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 60, 70, 80, 90, 100, 200, 500, or 1000 nucleotides, where the number of nucleotides of one or more first barcodes is different from that of one or more second barcodes.

[0175] Barcodes in a barcode population can have at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more different sequences. For example, barcodes in the population can have at least approximately 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, or more different sequences.

[0176] Method of use

[0177] The present disclosure provides methods of using the compositions described herein. In one aspect, the present disclosure provides methods of identifying a target binding unit or moiety that can bind to an analyte obtained from a healthy or diseased sample. In another aspect, the present disclosure provides methods of profiling an analyte in a sample obtained from a subject using the identified target binding unit or moiety. In various embodiments, the analyte is an antibody.

[0178] In some embodiments, the methods provided herein include contacting a mixture of antibodies with a population of target binding moieties, wherein each target binding moiety in the population comprises a target binding unit having a first binding region and a second binding region, wherein the first binding region and the second binding region are separated by a spacer and are spaced apart such that the first binding region and the second binding region simultaneously bind to a single antibody molecule in the mixture. In some instances, the target binding unit further comprises a first peptide and / or peptidomimetic sequence having a first binding region and a second peptide and / or peptidomimetic sequence having a second binding region. In some instances, the first binding region and the second binding region have the same sequence. In some instances, the first binding region and the second binding region have the same structure recognized by a single molecule.

[0179] In some embodiments, the present disclosure provides a method for selecting a set of peptides, comprising: (a) providing two or more copies of an array comprising a first array and a second array; (b) obtaining a first antibody mixture from a diseased subject and a second antibody mixture from a healthy subject; (c) contacting the first mixture with the first array and contacting the second mixture with the second array, wherein each array has at least 10 4 discrete regions, wherein each discrete region has a unique type of target binding moiety having a first binding region and a second binding region, wherein the first binding region and the second binding region are separated by a spacer and are spaced apart such that the first binding region and the second binding region simultaneously bind to a single antibody molecule in the mixture; (d) removing the unbound antibody fractions on the first array and the second array; (e) quantifying the amount of bound antibody on each discrete region of the first array and the second array; and (f) identifying the peptides in the second array that did not bind to the antibodies on the first array.

[0180] According to some other embodiments, provided herein are methods of selecting a set of peptides, comprising: (a) providing a first solution and a second solution; (b) obtaining a first antibody mixture from a diseased subject and a second antibody mixture from a healthy subject; (c) contacting the first mixture with the first solution and the second mixture with the second solution, wherein each of the first solution and the second solution comprises a plurality of polynucleotide barcoded target binding moieties, and each polynucleotide barcoded target binding moiety of the plurality of polynucleotide barcoded target binding moieties comprises (i) a nucleic acid sequence linked via a linker to (ii) a target binding unit, the target binding unit comprising a first peptide sequence containing a first binding region and a second peptide sequence containing a second binding region; wherein the first binding region and the second binding region are separated by a spacer and are spaced apart such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule comprising an antigen binding domain of an antibody; (d) capturing a plurality of antibody-bound polynucleotide barcoded target binding moieties in the first solution and the second solution; (e) sequencing the nucleic acid sequences captured from the first solution and the second solution; and (f) selecting the set of peptides.

[0181] Identifying on an array

[0182] In some applications, a population of target binding moieties is provided on a solid support. In such cases, methods of using the target binding moieties can be similar to methods of using peptide arrays. For example, a peptide microarray (also referred to as a peptide chip or peptide epitope microarray) is a collection of peptides displayed on a solid surface, e.g., a glass or plastic chip. The assay principle of a peptide microarray is similar to an ELISA protocol. Peptides (e.g., tens of thousands of copies) can be linked to the surface of a glass chip having the size and shape of a microscope slide. Such a peptide chip can be incubated directly with a variety of different biological samples, such as purified enzymes, antibodies, patient or animal sera, or cell lysates, and then detected in a label-dependent manner, e.g., by a primary antibody that targets a binding protein or modifies a substrate. After several washing steps, a secondary antibody with the desired specificity (e.g., anti-IgG human / mouse or anti-phosphotyrosine or anti-myc) can be applied. The secondary antibody can be labeled with a fluorescent label that can be detected by a fluorescence scanner. Other label-dependent detection methods that can be used include, but are not limited to, chemiluminescence, colorimetry, or autoradiography. In some cases, a population of target binding moieties is immobilized on a solid support. The solid support can have at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 10 3 、at least about 10 4 、at least about 10 5 or at least about 106 and at least about 10 7 and at least about 10 8 and at least about 10 9 and at least about 10 10 and at least about 10 11 and at least about 10 12 and at least about 10 13 and at least about 10 14 and at least about 10 15 and at least about 10 16 and at least about 10 17 and at least about 10 18 and at least about 10 19 and at least about 10 20 or more discrete regions. In some cases, each discrete region has a different target binding moiety from a population immobilized thereon. In some cases, the different target binding moieties from the population include the same binding region. In some cases, the different target binding moieties from the population include the same peptide and / or peptidomimetic sequence. In some cases, each discrete region has two or more copies of the different target binding moiety. The method can further include removing unbound antibodies from the mixture. For example, the removing can include washing the surface of the solid support with a buffer. The method can further include quantifying the amount of antibody bound at each discrete region. The quantifying can include detecting a fluorescence signal, an electrochemical signal, a chemiluminescence signal, a chromogenic signal, or a combination thereof. A variety of methods can be used herein to quantify the amount of bound antibody, e.g., by an antibody binder having a detectable label. The antibody binder can be a polynucleotide, a polypeptide (e.g., a protein, an enzyme, and a glycoprotein), an aptamer, a peptidomimetic, or a small molecule (e.g., a sugar and a lipid). The detectable label can be an optical label. In some cases, the optical label can be selected from small molecule dyes, fluorescent molecules or proteins, quantum dots, colorimetric reagents, chromogenic molecules or proteins, Raman labels, chromophores, and any combination thereof. In some cases, the detectable label or the optical label can be a fluorescent molecule or a protein. The detectable tag can generate a signal signature. The type of signal signature can vary. For example, the detectable label can include an optical molecule or label that generates an optical signature. Examples of the optical signature can include, but are not limited to, a fluorescence color (e.g., an emission spectrum under one or more excitation spectra), visible light, colorless or non-light, a signature of a color (e.g., a color defined by visible light wavelengths), a Raman signature, and any combination thereof. In some embodiments, the optical signature can include one or more fluorescence colors, one or more visible lights, one or more colorless or non-lights, one or more signatures of a color, one or more Raman signatures, or any combination thereof. For example, the optical label can include multiple (e.g., at least 2 or more) fluorescence colors (e.g., fluorescent dyes). The optical signature can be detected by optical imaging or spectroscopy.

[0183] In some cases, the method further includes obtaining a blood sample from a subject. In some cases, the method further includes preparing a serum sample from the blood sample, wherein the serum sample comprises an antibody mixture. In some cases, the subject includes a diseased subject and / or a healthy subject. In some cases, the method further includes obtaining a first serum sample from a diseased subject and a second serum sample from a healthy subject, wherein quantifying comprises quantifying the amount of antibody bound at each discrete region in the first serum sample on a first solid support and in the second serum sample on a second solid support. In some cases, the method further includes comparing the amount of antibody bound at each discrete region on the first solid support with the amount of antibody bound at each discrete region on the second solid support. In some cases, the method further includes selecting a set of peptides of discrete regions that differ in the amount of antibody bound on the first solid support and the second solid support.

[0184] Identifying in solution

[0185] In some applications, a population of target-binding moieties is provided in solution. In some cases, each of the population of target-binding moieties is linked to an encoding moiety. The encoding moiety can be a polynucleotide. For example, the polynucleotide is deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a combination thereof. In some cases, the polynucleotide is single-stranded DNA, single-stranded RNA, double-stranded DNA, double-stranded RNA, or a double-stranded DNA-RNA hybrid. In some cases, the encoding moiety includes a nucleic acid sequence that identifies the sequence of the target-binding moiety. In some cases, the nucleic acid sequence encodes the peptide sequence of two or more target-binding moieties. In some cases, the population of target-binding moieties includes at least about 100, at least about 10 3 、at least about 10 4 、at least about 10 5 、at least about 10 6 、at least about 10 7 、at least about 10 8 、at least about 10 9 、at least about 10 10 、at least about 10 11 、at least about 10 12 、at least about 10 13 、at least about 10 14 、at least about 10 15 、at least about 10 16 、at least about 10 17 、at least about 10 18 、at least about 10 19 、at least about 10 20One or more different types of target-binding moieties. As used herein, each type of target-binding moiety from a population comprises the same (or unique) target-binding region. In some cases, each type of target-binding moiety from a population comprises the same peptide and / or peptidomimetic sequence. In some cases, the coding moiety is unique for each type of target-binding moiety in the population. In some cases, the method further comprises capturing, enriching or isolating the antibody-binding fraction of the target-binding moieties in the population. As used herein, "capturing", "enriching", "isolating" or equivalent terms may be used interchangeably and may refer to purifying a fraction from the remainder of the population. A variety of methods can be used to capture, enrich or isolate antibodies or antibody-bound target-binding moieties, for example, using immunoglobulin-binding proteins such as protein A. In addition to protein A, other immunoglobulin-binding proteins such as protein G, protein A / G and protein L can be used for purifying, immobilizing or detecting immunoglobulins. In some cases, the method further comprises amplifying the coding moiety of the antibody-binding fraction of the target-binding moieties. In some cases, the method further comprises quantifying the coding moiety of the antibody-binding fraction or a copy of the coding moiety. In some cases, quantifying comprises sequencing the coding moiety of the antibody-binding fraction of the target-binding moieties or a copy of the coding moiety. In some cases, the method further comprises obtaining a blood sample from a subject. In some cases, the method further comprises preparing a serum sample from the blood sample, wherein the serum sample comprises a mixture of antibodies. In some cases, the subject comprises a diseased subject and / or a healthy subject. In some cases, the method further comprises obtaining a first serum sample from a diseased subject and a second serum sample from a healthy subject, wherein quantifying comprises quantifying the amount of the antibody-binding fraction of the target-binding moieties in the first serum sample and the amount of the antibody-binding fraction of the target-binding moieties in the second serum sample. In some cases, the method further comprises comparing the amount of the antibody-binding fraction of the target-binding moieties in the first serum sample and the amount of the antibody-binding fraction of the target-binding moieties in the second serum sample. A difference between the amounts of the antibody-binding fraction of the target-binding moieties in the first serum sample and the second serum sample can be observed. In some cases, the method further comprises selecting a set of peptides, wherein each peptide in the set has a difference in the amount of the antibody-binding fraction of the target-binding moieties in the first serum sample and the second serum sample. In some cases, each peptide in the set is identified in large amounts in the first serum sample but not in the second serum sample, or each peptide in the set is identified in large amounts in the second serum sample but not in the first serum sample. In some cases, the method further comprises preparing a population of target-binding moieties each linked to a coding moiety by RNA display. And in this case, each of the population of target-binding moieties further comprises a puromycin moiety or a variant thereof.

[0186] In some embodiments, the sequence of the first peptide and / or peptidomimetic sequence is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98% or at least 99% or 100% identical to the sequence of the second peptide and / or peptidomimetic sequence. In some cases, the first peptide sequence, the second peptide sequence, and the spacer are linked by peptide bonds. In some cases, the first peptide and / or peptidomimetic sequence, the second peptide and / or peptidomimetic sequence, and the spacer are linked by non-peptide bonds, e.g., in a branched form. In some cases, the target binding moiety comprises two or more target binding units. In some cases, the length of the first binding region and / or the second binding region is 4 to 30 residues, 5 to 20 residues, 6 to 10 residues, 8 to 11 residues, or 10 to 20 residues. In some cases, the length of the first peptide / peptidomimetic sequence and / or the second peptide / peptidomimetic sequence is at least 5 residues. In some cases, the spacer comprises a polymer. In some cases, the spacer comprises a pre-designed amino acid sequence. In some cases, the spacer of each target binding moiety comprises the same amino acid sequence. The spacer can be a polypeptide, a polynucleotide, or polyethylene glycol. In some cases, the spacer comprises a folded polypeptide, secondary structure, and / or tertiary structure. In some cases, the folded polypeptide comprises a coiled-coil structure or a β-sheet. In some cases, the coiled-coil structure is formed by two separate peptide chains, wherein each of the two separate peptide chains folds into an α-helix. In some cases, the coiled-coil structure is formed by a single peptide chain that comprises at least two regions that fold into α-helices. In some cases, the single peptide chain comprises four regions that fold into α-helices. In some cases, the coiled-coil structure is formed by a first peptide chain, a second peptide chain, and a third peptide chain, wherein the second peptide chain interacts with a first portion of the first peptide chain, and the third peptide chain interacts with a second portion of the first peptide chain. In some cases, the second peptide chain and the third peptide chain are linked to opposite ends of a double-stranded polynucleotide. In some cases, the spacer comprises double-stranded deoxyribonucleic acid. In some cases, the antibody mixture is a mixture of monoclonal antibodies, a mixture of polyclonal antibodies, or a combination thereof. In some cases, the antibody mixture comprises a biomarker.

[0187] Antibody profiling

[0188] The target binding units or moieties provided herein can be used for profiling antibodies in a sample. For example, the target binding units or moieties can be used to distinguish the serum antibody repertoires of diseased or healthy subjects.

[0189] In some embodiments, provided herein is a method for profiling an antibody mixture from a subject, the method comprising: contacting the antibody mixture with an array having at least 10 discrete regions, wherein each discrete region has a unique type of target binding moiety, wherein each unique type of target binding moiety comprises a first binding region and a second binding region, wherein the first binding region and the second binding region are separated by a spacer and spaced apart such that the first binding region and the second binding region simultaneously bind to a single antibody molecule in the mixture. In some instances, the method further comprises removing the unbound fraction of the antibodies in the mixture. In some instances, the method further comprises detecting the bound fraction of the antibodies in the mixture on the array, wherein a signal is observed at each discrete region to which an antibody has bound, thereby generating a signal pattern on the array. In some instances, the method further comprises identifying a disease of the subject. In some instances, the disease is an autoimmune disease, cancer, or an infectious disease.

[0190] In some embodiments, each unique type of target-binding moiety further comprises a first peptide and / or peptidomimetic sequence having a first binding region and a second peptide and / or peptidomimetic sequence having a second binding region. In some cases, the first binding region and the second binding region have the same sequence. In some cases, the first binding region and the second binding region have the same structure recognized by a single molecule. In some cases, the sequence of the first peptide and / or peptidomimetic sequence is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98% or at least 99% or 100% identical to the sequence of the second peptide and / or peptidomimetic sequence. In some cases, the first peptide sequence, the second peptide sequence, and the spacer are linked by peptide bonds. In some other cases, the first peptide and / or peptidomimetic sequence, the second peptide and / or peptidomimetic sequence, and the spacer are linked by non-peptide bonds. In some cases, the target-binding moiety comprises two or more binding regions. In some cases, the length of the first peptide and / or peptidomimetic sequence, or the second peptide and / or peptidomimetic sequence, is at least 5 residues. In some cases, the spacer comprises a polymer. In some cases, the spacer comprises the same amino acid sequence. The spacer can be a polypeptide, a polynucleotide, or polyethylene glycol. In some cases, the spacer may not have a structure. In some other cases, the spacer comprises a folded polypeptide, a secondary structure, and / or a tertiary structure. The folded polypeptide can comprise a coiled-coil structure or a β-sheet. In some cases, the coiled-coil structure is formed by two separate peptide chains, where each of the two separate peptide chains folds into an α-helix. In some cases, the coiled-coil structure is formed by a single peptide chain that comprises at least two regions that fold into α-helices. In some cases, the single peptide chain comprises four regions that fold into α-helices. In some cases, the coiled-coil structure is formed by a first peptide chain, a second peptide chain, and a third peptide chain, where the second peptide chain interacts with a first portion of the first peptide chain, and the third peptide chain interacts with a second portion of the first peptide chain. In some cases, the second peptide chain and the third peptide chain are linked to opposite ends of a double-stranded polynucleotide. In some cases, the spacer comprises double-stranded deoxyribonucleic acid. The antibody mixture can be any type of antibody. For example, the antibody mixture is a mixture of monoclonal antibodies, a mixture of polyclonal antibodies, or a combination thereof. In some cases, the antibody mixture comprises a biomarker.

[0191] Attached to a solid support

[0192] A variety of methods can be used to attach the target-binding moiety to a solid support. For example, in various embodiments, the target-binding moiety can be attached to the solid support via a linker. The linker can be formed by a first reactive group on the target-binding moiety and a second reactive group immobilized on the solid support. In some cases, the reactive group of the target-binding moiety is on a scaffold. In some cases, the scaffold is a polynucleotide, and the reactive group of the target-binding moiety is conjugated to the polynucleotide. In some cases, the scaffold is a polypeptide, and the reactive group of the target-binding moiety is conjugated to the polypeptide. A variety of methods can be used to conjugate the reactive group to the polynucleotide or polypeptide.

[0193] A variety of suitable reactive groups can be used. Such suitable reactive groups can include, but are not limited to, for example, amino, hydroxy, carboxy, carboxylate, aldehyde, ester, ether (e.g., thioether), amide, amine, nitrile, vinyl, sulfide, sulfonyl, phosphoryl, or similar chemical reactive groups. Other suitable reactive groups include, but are not limited to, maleimide, N-hydroxysuccinimide, sulfo-N-hydroxysuccinimide, nitrilotriacetic acid, activated hydroxy, haloacetyl (e.g., bromoacetyl, iodoacetyl), activated carboxy, hydrazide, epoxy resin, aziridine, sulfonyl chloride, trifluoromethyldiaziridine, pyridyl disulfide, N-acylimidazole, imidazole carbamate, vinyl sulfone, succinimidyl carbonate, aromatic azide, acid anhydride, diazoacetate, benzophenone, isothiocyanate, isocyanate, imidoester, fluorobenzene.

[0194] In some embodiments, one of the reactive groups is an electrophilic moiety and the second reactive group is a nucleophilic moiety. The nucleophilic moiety or the electrophilic moiety can be attached to the target-binding moiety. The reactive groups can then be used in a reaction to couple the target-binding moiety to the solid support. Suitable electrophilic moieties can be used to react with the nucleophilic moiety to form a covalent bond. Such electrophilic moieties include, but are not limited to, for example, carbonyl, sulfonyl, aldehyde, ketone, hindered ester, thioester, stable imine, epoxy, aziridine, etc.

[0195] The reaction product between the nucleophile and the electrophile can incorporate atoms originally present in, for example, the nucleophilic moiety. In some embodiments, the electrophile is an aldehyde or ketone having a nucleophilic moiety, including reaction products such as oxime, amide, hydrazone, reduced hydrazone, carbohydrazone, thiosemicarbazone, sufonylhydrazone, semicarbazone, thiosemicarbazone, or similar functional groups, depending on the nucleophilic moiety used and the electrophilic moiety (e.g., aldehyde, ketone, etc.) reacting with the nucleophilic moiety. The bonding to a carboxylic acid can be referred to as a carbohydrazide or hydroxamic acid. The bonding to a sulfonic acid can be referred to as a sulfohydrazide or N-sulfonylhydroxylamine. The resulting bond can then be stabilized by chemical reduction.

[0196] In some embodiments, one of the reactive groups is an electrophile, e.g., an aldehyde or a ketone, and the second reactive group is a nucleophilic moiety. Either the nucleophilic moiety or the electrophilic moiety can be attached to the target binding moiety; then the remaining reactive group is attached to the solid support. Suitable nucleophilic moieties that can react with aldehydes and ketones to form covalent bonds include, for example, aliphatic or aromatic amines such as ethylenediamine. In other embodiments, the reactive group is —NR 1 —NH2 (hydrazide), —NR 1 (C═O)NR 2 NH2 (semicarbazide), —NR 1 (C═S)NR2NH2 (thiosemicarbazide), —(C═O)NR 1 NH2 (carbohydrazide), —(C═S)NR 1 NH2 (thiocarbohydrazide), —(SO2)NR 1 NH2 (sulfohydrazide), —NR 1 NR 2 (C═O)NR 3 NH2 (carbazide), —NR 1 NR 2 (C═S)NR 3 NH2 (thiocarbazide) or —O—NH2 (hydroxylamine), where each R 1 、R 2 and R 3 is independently H, or an alkyl group having 1-6 carbon atoms, preferably H. In some cases, the reactive group is a hydrazide, hydroxylamine, carbohydrazide or sulfohydrazide.

[0197] The reaction group chemistry method used in this article is not limited to the chemical methods listed above. For example, in other embodiments, the reaction between the first and second reaction groups can be carried out by a dipolarophile reaction. For example, the first reaction group can be an azide, and the second reaction group can be an alkyne. Alternatively, the first reaction group can be an alkyne, and the second reaction group can be an azide. The unique reactivity of azide and alkyne functional groups can make them useful reactants for selectively coupling polypeptides to arrays and other solid supports. Organic azides, especially aliphatic azides and alkynes, can be stable under common reaction chemical conditions. Since the Huisgen cycloaddition reaction involves a selective cycloaddition reaction (see, e.g., Huisgen, 1,3-DIPOLAR CYCLOADDITION CHEMISTRY, (ed. Padwa, A., 1984), p. 1-176), rather than nucleophilic substitution, the incorporation of unnatural encoded amino acids with azide- and alkyne-containing side chains can allow the resulting polypeptides to be modified with extremely high selectivity. Both azide and alkyne functional groups can be inert to the 20 common amino acids found in naturally occurring polypeptides. However, when in close proximity, the "spring-loaded" nature of azide and alkyne groups can be revealed, and they can react selectively and efficiently through the Huisgen [3+2] cycloaddition reaction to form the corresponding triazoles. See, e.g., Chin et al., Science 301:964-7 (2003); Wang et al., J. Am. Chem. Soc., 125, 3192-3193 (2003); Chin et al., J. Am. Chem. Soc., 124:9026-9027 (2002). The cycloaddition reaction involving azide- or alkyne-containing polypeptides can be carried out at room temperature, in aqueous conditions, by adding Cu(II) (e.g., in the form of catalytic amounts of CuSO4) in the presence of a catalytic amount of a reducing agent (used to reduce Cu(II) in situ to Cu(I)). See, e.g., Wang et al., J. Am. Chem. Soc. 125, 3192-3193 (2003); Tornoe et al., J. Org. Chem. 67:3057-3064 (2002); Rostovtsev, Angew. Chem. Int. Ed. 41:2596-2599 (2002). Reducing agents include, but are not limited to, ascorbate, metallic copper, quinine, hydroquinone, vitamin K, glutathione, cysteine, Fe 2+ 、Co 2+ and applied electric potential.

[0198] Other reaction chemistries that can be used include, but are not limited to, Staudinger ligation and olefin metathesis chemistry (see, e.g., Mahal et al., (1997) Science 276:1125-1128).

[0199] In some embodiments, the attachment between the target binding moiety and the solid support is a non-covalent attachment. For example, a target binding moiety having a suitable acidic group will form a strong association with a solid support bearing a hydroxyl group or other negatively charged group. In other variants of this system, other types of moieties that have a strong affinity for each other can be incorporated into reactive groups on the target binding moiety and the solid support. For example, the target binding moiety can be conjugated to biotin via a suitable reactive group, and the solid support can be coated with avidin, resulting in a strong non-covalent binding between the target binding moiety and the solid support.

[0200] Alternative non-covalent coupling systems can be used.

[0201] The solid supports (e.g., arrays) suitable for the present disclosure are not limiting. The present disclosure is not intended to be limited to any particular type of solid support material or array configuration.

[0202] The solid support can be flat or planar, or can have a substantially different configuration. For example, the solid support can exist in the form of particles, beads, wires, precipitates, gels, sol-gels, sheets, tubes, spheres, containers, capillaries, pads, slices, films, plates, dipsticks, slides, etc. Magnetic beads or particles, such as magnetic latex beads and iron oxide particles, are examples of solid matrices that can be used in the methods of the present disclosure. Magnetic particles are described, for example, in U.S. Patent No. 4,672,040 and can be purchased from, for example, PerSeptive Biosystems, Inc. (Framingham Mass.), Ciba Corning (Medfield Mass.), Bangs Laboratories (Carmel Ind.), and BioQuest, Inc. (Atkinson N.H.). The solid support is selected to maximize the signal-to-noise ratio, primarily by minimizing background binding, to facilitate washing and reduce costs. In addition, certain solid supports such as beads can be readily used in conventional fluid handling systems such as microtiter plates.

[0203] Examples of solid supports include glass or other ceramics, plastics, polymers, metals, metalloids, alloys, composites, organic materials, etc. For example, solid supports can include materials selected from the following: silicon, silica, quartz, glass, controlled pore glass, carbon, alumina, titanium dioxide, tantalum oxide, germanium, silicon nitride, zeolites, and gallium arsenide. A variety of metals, such as gold, platinum, aluminum, copper, titanium, and their alloys, can also be used as solid supports. In addition, many ceramics and polymers can be used as solid supports. Polymers that can be used as solid supports include, but are not limited to, the following: polystyrene; polytetrafluoroethylene (PTFE); polyvinylidene fluoride; polycarbonate; polymethyl methacrylate; polyvinyl ethylene; polyethyleneimine; polyetheretherketone; polyoxymethylene (POM); polyvinylphenol; polylactide; polymethacrylimide (PMI); polyatkenesulfone (PAS); polypropylene; polyethylene; 2-hydroxyethyl methacrylate (HEMA); polydimethylsiloxane; polyacrylamide; polyimide, and block copolymers. In some cases, the matrix for the array includes silicon, silica, glass, and polymers. The solid support can be composed of a single material (e.g., glass), a mixture of materials (e.g., copolymer), or multiple layers of different materials (e.g., metal coated with a monolayer of small molecules, glass coated with BSA, etc.).

[0204] The configuration of the solid support can be in any suitable form. For example, it can include beads, spheres, particles, spots, gels, sol-gels, self-assembled monolayers (SAMs), or surfaces (which can be flat or can have shape features). The term "solid support" includes semi-solid supports. The surface of the solid support can be flat, substantially flat, or non-planar. The solid support can be porous or non-porous and can have swelling or non-swelling properties. The solid support can be configured in the form of pores, depressions, or other containers, vessels, features, or locations. Multiple solid supports can be configured in an array at multiple positions, addressable for robotic reagent delivery or addressable by detection means, including scanning by laser or other irradiation and CCD, confocal, or deflected light collection.

[0205] For example, in some embodiments, the solid support can be in the form of a glass slide. The glass slide can be made of any material. In some cases, the glass slide can include a plastic or glass substrate. The glass slide can be used to support the solid-phase deposition of compounds (e.g., polypeptides) and can be prepared to include a very large number of addressable locations, e.g., thousands of locations. The process of placing compounds on the glass slide for analysis can be referred to as "printing". The glass slide system can utilize fluorescent dye labeling to detect interactions and can be created using automated machinery capable of depositing very small spots and placing them in positions very close to each other with high precision. For example, the spot diameter can be in the range of 100 microns, and 10,000 - 30,000 spots can be placed on a standard 1″×3″ glass slide. Glass slide arrays may tend to have a large number and very high density of addressable coordinates.

[0206] In some embodiments, the glass slide or other solid support can include a self-assembled monolayer (SAM), which can form due to the affinity interaction and / or covalent bonding of SAM molecules at the surface interface. The SAM can be assembled in a manner similar to the bilayer structure of soap bubbles or cell membranes, but forms a monolayer at the solid interface. The SAM can be assembled from molecules having an interfacial binding group attached to a terminal group. The SAM can be a collection of molecules such as alkanethiols, silanes, fatty acids, or phosphonates. The driving force for SAM assembly can be the affinity interaction between the interfacial binding group and the surface group. The polar orientation of the molecules on the surface can be further enhanced by the interaction of the terminal group with the external environment. The interactions driving the assembly can be, for example, hydrophobic interactions, hydrophilic interactions, ionic attraction, chelation, etc.

[0207] In some embodiments, the solid support is in the form of beads (synonymous with particles), e.g., used in a liquid-phase array system (sometimes referred to as a bead array). These systems can employ a microtiter plate (sometimes referred to as a "microtiter dish") that has any number of wells for accommodating liquid volumes. Examples of micro-well configurations include, but are not limited to, 96-well plates, 384-well plates, and 1536-well plates. Each well can accommodate a specific component being used in a parallel analysis, e.g., beads. The beads can be made of any matrix material, including biological, non-biological, organic, inorganic, polymeric, metallic, or any combination of these. The surface of the beads can be chemically modified and subjected to any type of treatment or coating, e.g., a coating containing reactive groups that allow binding interactions with target-binding moieties.

[0208] In some embodiments, the beads can be produced in a manner that facilitates their rapid separation and / or purification. For example, magnetic beads can be manipulated by applying a magnetic field to rapidly isolate the magnetic beads from the liquid phase within the plate wells.

[0209] In some embodiments, the solid support comprises or consists of a sol-gel. Sol-gel techniques are well known and are described, for example, in Kirk-Othmer Encyclopedia of Chemical Technology (3rd and 4th Editions, particularly Volume 20), Martin Grayso, Executive Editor, Wiley-Interscience, John Wiley and Sons, NY (e.g., Volume 22) and references cited therein. A sol can be a dispersion of colloidal particles (usually nano-scale elements) in a liquid such as water or a solvent. The sol particles can be small enough to remain suspended in the liquid, e.g., by Brownian motion. A gel can be a viscoelastic body with interconnected pores of sub-micron size. Sol-gels can be used to prepare glasses, ceramics, composites, plastics, etc. by preparing a sol, sol-gelation, and removing the liquid suspending the sol. The process can be used in many relatively low-temperature processes for constructing fibers, films, aerogels, etc. (any of which can be a solid support in the present disclosure). In some embodiments, gelation of the colloidal particle dispersion is carried out. In some embodiments, hydrolysis and polycondensation of an alkoxide or metal salt precursor is carried out. In some embodiments, hydrolysis and polycondensation of an alkoxide precursor is carried out, followed by aging and drying at room temperature.

[0210] The surface of the solid phase support can be prepared to generate suitable reactive groups to which a linker can be attached or to which a target binding moiety can be attached directly. Techniques for placing reactive groups on a substrate as listed above by mechanical, physical, electrical, or chemical means (see, e.g., U.S. Patent No. 4,681,870) can be used.

[0211] In addition to directly reacting the target binding moiety with the chemical moiety on the solid support, other tethering mechanisms for linking a polypeptide or polynucleotide to the arrays of the present disclosure can be used. Such tethering methods include: chemical tethering, biotin-mediated binding, crosslinking to the solid support matrix (e.g., UV or fluorescence-activated crosslinking), and using a "soluble" matrix such as PEG, which can be precipitated with EtOH or other solvents to recover the bound material (see also Wentworth, P., 1999, Trends in Biotechnolgy 17:448452).

[0212] Sequencing

[0213] As described herein, in some embodiments, the target binding moiety is provided in solution for peptide identification. In some cases, the target binding moiety is linked to a coding moiety, where the coding moiety identifies or corresponds to the target binding moiety. The coding moiety can identify or correspond to the peptide sequence linked to the target binding moiety.

[0214] The peptide identification method provided herein includes contacting a mixture of antibodies with a plurality of target binding moieties, wherein each target binding moiety includes a first peptide and a second peptide, and wherein the first peptide and the second peptide are separated by a spacer such that the first peptide and the second peptide bind to a single antibody molecule. In some embodiments, each of the plurality of target binding moieties can be linked to an encoding moiety. After contact, the target binding moieties that bind to the antibody can be captured. There are various methods to capture the antibody-bound target binding moieties, for example, using an immunoglobulin-binding protein.

[0215] The encoding moiety linked to the captured target binding moiety can be amplified and sequenced.

[0216] Various sequencing methods can be used in the present disclosure. In some embodiments, the methods described herein employ next-generation sequencing technology (NGS), wherein cloned amplified DNA templates or single DNA molecules are sequenced in a massively parallel manner in a flow cell (e.g., as described in Volkerding et al. Clin Chem 55:641-658

[2009] ; Metzker M Nature Rev 11:31-46

[2010] ). In addition to high-throughput sequence information, NGS can also provide digital quantitative information because each sequence read is a countable "sequence tag" representing a single cloned DNA template or a single DNA molecule. This quantification extends the digital PCR concept of counting free DNA molecules to NGS (Fan et al., Proc Natl Acad Sci USA 105:16266-16271

[2008] ; Chiu et al., Proc Natl Acad Sci USA 2008; 105:20458-20463

[2008] ). The sequencing technologies of NGS include pyrosequencing, reversible dye terminator synthesis sequencing, oligonucleotide probe ligation sequencing, and real-time sequencing.

[0217] Some sequencing technologies are commercially available, such as the sequencing-by-hybridization platform from Affymetrix Inc. (Sunnyvale, Calif.), and the sequencing-by-synthesis platforms from 454 Life Sciences (Bradford, Conn.), Illumina / Solexa (Hayward, Calif.), and Helicos Biosciences (Cambridge, Mass.), and the sequencing-by-ligation platform from Applied Biosystems (Foster City, Calif.), as described below. In addition to single molecule sequencing using the sequencing-by-synthesis of Helicos Biosciences, the methods of the present disclosure also include other single molecule sequencing technologies, and include the SMRT TM technology of Pacific Biosciences, Ion Torrent TM technology, and nanopore sequencing (e.g., of Oxford Nanopore Technologies). Although the automated Sanger method is considered a "first generation" technology, the methods of the present disclosure can also employ Sanger sequencing, including automated Sanger sequencing. The methods of the present disclosure also include other sequencing methods, including the use of developing nucleic acid imaging technologies, such as atomic force microscopy (AFM) or transmission electron microscopy (TEM).

[0218] In some embodiments, the method uses Illumina's sequencing-by-synthesis and reversible terminator-based sequencing chemistries to perform massively parallel sequencing of millions of DNA fragments (e.g., as described in Bentley et al., Nature 6:53-59

[2009] ). Illumina's sequencing technology relies on the attachment of template DNA (e.g., the coding portion or a portion thereof) to a planar, optically transparent surface to which oligonucleotide anchors are bound. The template DNA can be end-repaired to generate 5'-phosphorylated blunt ends, and the polymerase activity of the Klenow fragment can be used to add a single A base to the 3' end of the phosphorylated DNA template at the blunt end. This addition can prepare the DNA template for ligation to an oligonucleotide adaptor that has a protruding end with a single T base at its 3' end to enhance ligation efficiency. The adaptor oligonucleotide can be complementary to the anchors of the flow cell. Under limited dilution conditions, the adaptor-modified single-stranded template DNA can be added to the flow cell and immobilized by hybridization to the anchors. The attached DNA template can be extended and bridge amplified to create an ultra-high density sequencing flow cell with hundreds of millions of clusters, each containing approximately 1,000 copies of the same template. In some embodiments, prior to cluster amplification of the coding portion or a portion thereof, amplification is performed using nucleic acid amplification (e.g., PCR). Alternatively, a no-amplification library preparation can be used, and only cluster amplification can be used to enrich the template DNA (e.g., the coding portion or a portion thereof) (Kozarewa et al., Nature Methods 6:291-295

[2009] ). The template can be sequenced using four-color DNA sequencing-by-synthesis technology that employs reversible terminators with removable fluorescent dyes. High-sensitivity fluorescence identification can be achieved using laser excitation and total internal reflection optics. The sequencing reads can be analyzed using any available data analysis pipeline software. After the first read is completed, the template can be regenerated in situ to perform a second read from the opposite end of the fragment. Thus, single-end or paired-end sequencing of the DNA template can be used according to the method.

[0219] The sequence of the coding portion can be used to identify sequence identity of the peptide sequence of the target-binding portion. In some embodiments, the coding portion is a nucleic acid sequence encoding a peptide sequence, and in such cases, the peptide sequence can be obtained by translating the nucleic acid sequence into an amino acid sequence. In some embodiments, the coding portion functions as a barcode but may not encode the peptide of the target-binding portion, and in such cases, each unique peptide sequence is pre-designed and pre-assigned a unique nucleic acid sequence. For example, a reference library of target-binding portions can be prepared, each target-binding portion including a unique pre-designed peptide sequence, and each target-binding portion can be linked to a unique coding portion of known sequence. After sequencing the coding portion in a sample, the sequence can be used to identify the peptide sequence according to the reference library.

[0220] Subjects and diseases

[0221] The compositions and methods provided herein can be used to identify or quantify biomarkers in a sample from a subject having a disease / condition. The biomarker can be an antibody or a fragment thereof. In some embodiments, a blood sample is obtained from a subject having a disease. In some embodiments, a serum sample is obtained from a subject having a disease. In some embodiments, a mixture of antibodies is obtained from a subject having a disease. In some embodiments, a mixture of antibodies is obtained from a healthy subject as a control. In some embodiments, the disease is cancer. In some other embodiments, the disease is an autoimmune disease. In some embodiments, the disease is an infectious disease.

[0222] The term "subject" as used herein can refer to an individual, host, or patient. In some embodiments, the subject is a human. In some embodiments, the subject is a mammal. In some other embodiments, the subject can be any animal subject, including laboratory animals, livestock, and household pets. The subject can be diagnosed as or suspected of being at high risk for a disease. In some cases, the subject is not necessarily diagnosed as or suspected of being at high risk for a disease.

[0223] The subject can be, for example, a mammal, a human, a pregnant woman, an elderly person, an adult, an adolescent, a pre-adolescent, a child, a toddler, an infant, a neonate, or a neonate less than four weeks old. The subject can be a patient. In some cases, the subject can be a human. In some cases, the subject can be a child (i.e., a young person below the age of puberty). In some cases, the subject can be an infant. In some cases, the subject can be a formula-fed infant. In some cases, the subject can be an individual participating in a clinical study. In some cases, the subject can be a laboratory animal, e.g., a mammal or a rodent. In some cases, the subject can be an obese or overweight subject.

[0224] Cancer can include solid tumors and hematological malignancies. Cancer includes but is not limited to gynecological cancers, ovarian cancer, fallopian tube cancer, peritoneal cancer, breast cancer, cervical cancer, endometrial cancer, prostate cancer, testicular cancer, pancreatic cancer, esophageal cancer, head and neck cancer, gastric cancer, bladder cancer, lung cancer (e.g., adenocarcinoma, NSCLC, and SCLC), bone cancer (e.g., osteosarcoma), colon cancer, rectal cancer, thyroid cancer, brain and central nervous system cancer, glioblastoma, neuroblastoma, neuroendocrine cancer, rhabdomyosarcoma, keratoacanthoma, epidermoid carcinoma, seminoma, melanoma, sarcoma (e.g., liposarcoma), bladder cancer, liver cancer (e.g., hepatocellular carcinoma), kidney cancer (e.g., renal cell carcinoma), myeloid disorders (e.g., AML, CML, myelodysplastic syndromes, and promyelocytic leukemia), and lymphatic disorders (e.g., leukemia, multiple myeloma, mantle cell lymphoma, ALL, CLL, B-cell lymphoma, T-cell lymphoma, Hodgkin lymphoma, non-Hodgkin lymphoma, hairy cell lymphoma). Cancer includes but is not limited to ovarian cancer, breast cancer, cervical cancer, endometrial cancer, prostate cancer, testicular cancer, pancreatic cancer, esophageal cancer, head and neck cancer, gastric cancer, bladder cancer, lung cancer, bone cancer, colon cancer, rectal cancer, thyroid cancer, brain and central nervous system cancer, glioblastoma, neuroblastoma, neuroendocrine cancer, rhabdomyosarcoma, keratoacanthoma, epidermoid carcinoma, seminoma, melanoma, sarcoma, bladder cancer, liver cancer, kidney cancer, myeloma, lymphoma, and combinations thereof.

[0225] In addition, the subjects provided herein may have conditions including benign or malignant tumors (e.g., cancers of the adrenal, liver, kidney, bladder, breast, stomach, ovary, colorectal, prostate, pancreas, lung, thyroid, liver, cervix, endometrium, esophagus, and uterus; sarcomas; glioblastoma; and various head and neck tumors); leukemias and lymphoid malignancies; other conditions such as those of neurons, glia, astrocytes, hypothalamus and other glands, macrophages, epithelium, stroma, and blastocysts; and inflammatory, angiogenic, immune system conditions, and conditions caused by pathogens. Although "subject" or "patient" is preferably a human, as used herein, the term is expressly considered to include any mammalian species.

[0226] In some embodiments, the subject can have a neoplastic condition. The neoplastic condition can be selected from including, but not limited to, adrenal tumors, AIDS-related cancers, alveolar soft part sarcoma, astrocytic tumors, bladder cancer (squamous cell carcinoma and transitional cell carcinoma), bone cancer (adamantinoma, aneurysmal bone cyst, osteochondroma, osteosarcoma), brain and spinal cord cancer, metastatic brain tumors, breast cancer, carotid body tumor, cervical cancer, chondrosarcoma, chordoma, chromophobe renal cell carcinoma, clear cell carcinoma, colon cancer, colorectal cancer, benign fibrous histiocytoma of the skin, desmoplastic small round cell tumor, ependymoma, Ewing's tumor, extraskeletal myxoid chondrosarcoma, fibrous dysplasia of bone, fibrous dysplasia of bone, gallbladder and bile duct cancer, gestational trophoblastic disease, germ cell tumor, head and neck cancer, islet cell tumor, Kaposi's sarcoma, kidney cancer (nephroblastoma, papillary renal cell carcinoma), leukemia, lipoma / benign lipoma, liposarcoma / malignant lipoma, liver cancer (hepatoblastoma, hepatocellular carcinoma), lymphoma, lung cancer (small cell carcinoma, adenocarcinoma, squamous cell carcinoma, large cell carcinoma, etc.), medulloblastoma, melanoma, meningioma, multiple endocrine neoplasia, multiple myeloma, myelodysplastic syndrome, neuroblastoma, neuroendocrine tumor, ovarian cancer, pancreatic cancer, papillary thyroid cancer, parathyroid tumor, pediatric cancer, peripheral nerve sheath tumor, pheochromocytoma, pituitary tumor, prostate cancer, posterior uveal melanoma, rare hematological diseases, renal metastatic cancer, rhabdomyoma, rhabdomyosarcoma, sarcoma, skin cancer, soft tissue sarcoma, squamous cell carcinoma, stomach cancer, synovial sarcoma, testicular cancer, thymic carcinoma, thymoma, thyroid metastatic cancer, and uterine cancer (cervical cancer, endometrial cancer, and leiomyoma).

[0227] In certain embodiments, the proliferative disease includes solid tumors, including but not limited to, cancers and sarcomas of the adrenal, liver, kidney, bladder, breast, stomach, ovary, cervix, uterus, esophagus, colorectum, prostate, pancreas, lung (small cell and non-small cell), and thyroid, glioblastoma, and various head and neck tumors. In other embodiments, the disease is small cell lung cancer (SCLC) and non-small cell lung cancer (NSCLC) (e.g., squamous cell non-small cell lung cancer or squamous cell small cell lung cancer). In one embodiment, the lung cancer can be refractory, recurrent, or resistant to platinum-based drugs (e.g., carboplatin, cisplatin, oxaliplatin, topotecan) and / or taxanes (e.g., docetaxel, paclitaxel, larotaxel, or cabazitaxel).

[0228] In some embodiments, the disease can be a tumor with a neuroendocrine feature or phenotype, including neuroendocrine tumors. True or canonical neuroendocrine tumors (NETs), which arise from the diffuse endocrine system, are relatively rare, occurring in 2-5 cases per 100,000 people, but are highly invasive. Neuroendocrine tumors occur in the kidneys, urogenital tract (bladder, prostate, ovary, cervix, and endometrium), gastrointestinal tract (colon, stomach), thyroid (medullary thyroid carcinoma), and lungs (small cell lung cancer and large cell neuroendocrine carcinoma).

[0229] Regarding hematological malignancies, the disease can be B-cell lymphoma, including low-grade / NHL follicular cell lymphoma (FCC), mantle cell lymphoma (MCL), diffuse large cell lymphoma (DLCL), small lymphocyte (SL) NHL, intermediate / follicular NHL, intermediate diffuse NHL, high-grade immunoblastic NHL, high-grade lymphoblastic NHL, high-grade small non-cleaved cell NHL, large mass NHL, Waldenstrom macroglobulinemia, lymphoplasmacytic lymphoma (LPL), mantle cell lymphoma (MCL), follicular lymphoma (FL), diffuse large cell lymphoma (DLCL), Burkitt lymphoma (BL), AIDS-related lymphoma, monocytoid B-cell lymphoma, angioimmunoblastic lymphadenopathy; small lymphocyte, follicular, diffuse large cell, diffuse small cleaved cell, large cell immunoblastic lymphoblastoma; small, non-cleaved, Burkitt and non-Burkitt, follicular, predominantly large cell lymphoma; follicular, predominantly small cleaved cell lymphoma; and follicular, mixed small cleaved and large cell lymphoma.

[0230] Examples of autoimmune diseases include, but are not limited to, achalasia, Addison's disease, adult Still's disease, agammaglobulinemia, alopecia areata, amyloidosis, ankylosing spondylitis, anti-GBM / anti-TBM nephritis, antiphospholipid syndrome, autoimmune angioedema, autoimmune familial dysautonomia, autoimmune encephalomyelitis, autoimmune hepatitis, autoimmune inner ear disease (AIED), autoimmune myocarditis, autoimmune oophoritis, autoimmune orchitis, autoimmune pancreatitis, autoimmune retinopathy, autoimmune urticaria, axonal and neuronal neuropathy (AMAN), Baló disease, Behçet's disease, benign mucous membrane pemphigoid, bullous pemphigoid, Castleman disease (CD), celiac disease, Chagas disease, chronic inflammatory demyelinating polyneuropathy (CIDP), chronic recurrent multifocal osteomyelitis (CRMO), Churg-Strauss syndrome (CSS) or eosinophilic granulomatosis with polyangiitis (EGPA), cicatricial pemphigoid, Cogan syndrome, cold agglutinin disease, congenital heart block, coxsackievirus myocarditis, CREST syndrome, Crohn's disease, dermatitis herpetiformis, dermatomyositis, Devic's disease (neuromyelitis optica), discoid lupus, Dressler syndrome, endometriosis, eosinophilic esophagitis (EoE), eosinophilic fasciitis, erythema nodosum, essential mixed cryoglobulinemia, Evans syndrome, fibromyalgia, fibrosing alveolitis, giant cell arteritis (temporal arteritis), giant cell myocarditis, glomerulonephritis, Goodpasture syndrome, granulomatosis with polyangiitis, Graves' disease, Guillain-Barré syndrome, Hashimoto's thyroiditis, hemolytic anemia, Henoch-Schönlein purpura (HSP), herpes gestationis or pemphigoid gestationis (PG), hidradenitis suppurativa (HS) (acne inversa), hypogammaglobulinemia, IgA nephropathy, IgG4-related sclerosing disease, immune thrombocytopenic purpura (ITP), inclusion body myositis (IBM), interstitial cystitis (IC), juvenile arthritis, juvenile diabetes (type 1 diabetes), juvenile myositis (JM), Kawasaki disease, Lambert-Eaton syndrome, leukocytoclastic vasculitis, lichen planus, lichen sclerosus, ligneous conjunctivitis, linear IgA disease (LAD), lupus, chronic Lyme disease, Ménière's disease, microscopic polyangiitis (MPA), mixed connective tissue disease (MCTD), Mooren's ulcer, Mucha-Habermann disease, multifocal motor neuropathy (MMN) or MMNCB, multiple sclerosis, myasthenia gravis, myositis, narcolepsy, neonatal lupus, neuromyelitis optica, neutropenia, ocular cicatricial pemphigoid, optic neuritis, relapsing rheumatism (PR), PANDAS, paraneoplastic cerebellar degeneration (PCD),Paroxysmal nocturnal hemoglobinuria (PNH), Parry-Romberg syndrome, pars planitis (peripheral uveitis), Parsonnage-Turner syndrome, pemphigus, peripheral neuropathy, perivenous encephalomyelitis, pernicious anemia (PA), POEMS syndrome, polyarteritis nodosa, polyglandular syndrome type I, polyglandular syndrome type II, polyglandular syndrome type III, polymyalgia rheumatica, polymyositis, post-myocardial infarction syndrome, post-pericardiotomy syndrome, primary biliary cirrhosis, primary sclerosing cholangitis, progesterone dermatitis, psoriasis, psoriatic arthritis, pure red cell aplasia (PRCA), pyoderma gangrenosum, Raynaud phenomenon, reactive arthritis, reflex sympathetic dystrophy, relapsing polychondritis, restless legs syndrome (RLS), retroperitoneal fibrosis, rheumatic fever, rheumatoid arthritis, sarcoidosis, Schmidt syndrome, scleritis, scleroderma, Sjogren syndrome, sperm and testicular autoimmunity, stiff person syndrome (SPS), subacute bacterial endocarditis (SBE), Susac syndrome, sympathetic ophthalmia (SO), Takayasu arteritis, temporal arteritis / giant cell arteritis, thrombotic thrombocytopenic purpura (TTP), Tolosa-Hunt syndrome (THS), transverse myelitis, type 1 diabetes, ulcerative colitis (UC), undifferentiated connective tissue disease (UCTD), uveitis, vasculitis, vitiligo, Vogt-Koyanagi-Harada disease, and Wegener granulomatosis (or granulomatosis with polyangiitis (GPA)).

[0231] Examples of infectious diseases include, but are not limited to, Acute Flaccid Myelitis (AFM), anaplasmosis, anthrax, babesiosis, botulism, brucellosis, Burkholderia mallei infection (glanders), Burkholderia pseudomallei infection (melioidosis), campylobacteriosis (Campylobacter), carbapenem-resistant infections (CRE / CRPA), chancroid, chikungunya virus infection (chikungunya), chlamydia, ciguatera, Clostridioides difficile infection, Clostridium perfringens disease (epsilon toxin), coccidioidal fungal infection (valley fever), Creutzfeldt-Jakob disease (transmissible spongiform encephalopathy, CJD), cryptosporidiosis (Crypto), cyclosporiasis, dengue, diphtheria, Escherichia coli infection (E. coli), Eastern equine encephalitis (EEE), Ebola hemorrhagic fever (Ebola), ehrlichiosis, encephalitis (arbovirus or parainfectious), enterovirus infection (non-polio enterovirus), enterovirus D68 (EV-D68), giardiasis (Giardia), gonococcal infection (gonorrhea), granuloma inguinale, Haemophilus influenzae disease (type B Hib or influenza type H), hantavirus pulmonary syndrome (HPS), hemolytic uremic syndrome (HUS), hepatitis A (Hep A), hepatitis B (Hep B), hepatitis C (Hep C), hepatitis D (Hep D), hepatitis E (Hep E), herpes, herpes zoster, herpes zoster VZV (shingles), histoplasmosis, human immunodeficiency virus / acquired immunodeficiency syndrome (HIV / AIDS), human papillomavirus (HPV), influenza (Flu), lead poisoning, legionellosis (legionnaires' disease), leprosy (Hansen's disease), leptospirosis, listeriosis (Listeria), Lyme disease, lymphogranuloma venereum infection (LVG), malaria, measles, meningitis (viral meningitis), meningococcal disease (bacterial meningitis), Middle East respiratory syndrome coronavirus (MERS-CoV), mumps, norovirus, paralytic shellfish poisoning (paralytic shellfish poisoning, ciguatera), pediculosis (lice, head lice and body lice), pelvic inflammatory disease (PID), pertussis (whooping cough), plague (bubonic plague, septicemic plague, pneumonic plague), pneumococcal disease (pneumonia), poliomyelitis (polio), Powassan encephalitis, psittacosis, Pthiriasis (crabs;Pubic lice infestation), pustular diseases (smallpox, monkeypox, cowpox), Q fever, rabies, ricin poisoning, rickettsiosis (Rocky Mountain spotted fever), rubella (German measles), Salmonella gastroenteritis (Salmonella), scabies infestation (scabies), scombroid poisoning, Severe Acute Respiratory Syndrome (SARS), Shigella gastroenteritis (Shigella), smallpox, methicillin-resistant Staphylococcus aureus infection (MRSA), Staphylococcus food poisoning (Staph food poisoning), vancomycin-intermediate Staphylococcus aureus infection (VISA), vancomycin-resistant Staphylococcus aureus infection (VRSA), Group A streptococcal disease (Strep A), Group B streptococcal disease (Strep-B), streptococcal toxic shock syndrome (STSS, TSS), syphilis (primary, secondary, early latent, late latent or congenital), tetanus infection (lockjaw), trichomycosis infection (trichomycosis), tuberculosis (TB), latent tuberculosis infection (LTBI), tularemia (rabbit fever), typhoid, typhus, vaginitis (yeast infection), varicella (chickenpox), Vibrio cholerae disease (cholera), vibriosis (Vibrio), viral hemorrhagic fevers (Ebola, Lassa virus, Marburg virus), West Nile virus, yellow fever, Yersinia infection (Yersinia) and Zika virus infection (Zika Z).;

[0232] Kit

[0233] The present disclosure also relates to compositions and kits or reagent systems for practicing the methods described herein.

[0234] The compositions of the present disclosure can be provided in a kit for peptide / peptidomimetic identification. In the case of identification in solution, the kit can also include reagents for amplification and / or sequencing, e.g., primers and buffers. In the case of identification on an array, one or more arrays can be provided in the kit. The arrays provided in the kit can include the target-binding moieties described herein. The kit can include one or more of the target-binding moieties described herein. The kit can include one or more polynucleotide-barcodeylated target-binding moieties. The kit can include one or more components for the user to prepare the target-binding moiety or the polynucleotide-barcodeylated target-binding moiety. The kit can include a target-binding unit. The kit can also include additional reagents for the user to perform analyte identification using the target-binding moiety. One or more of the reagents in the kit can be provided in one or more containers. In some cases, one or more of the reagents can be provided in a mixture.

[0235] The presently disclosed compositions can be provided in a kit for antibody profiling. In such a case, the kit can include a set of specific peptide sequences that are known to be able to distinguish a sample from a patient and a sample from a healthy subject.

[0236] The kit may include a reagent combination, which includes the elements required for performing assays according to the methods disclosed herein. The reagent system may be presented in the form of a commercial package, presented as a composition or mixture where reagent compatibility permits, presented in a test device configuration, or presented as a test kit. The kit may include a packaging combination of one or more containers, devices, etc. that house the necessary reagents. The kit may include written instructions regarding the performance of one or more assays. The kits of the present disclosure may be suitable for assays of any configuration and may include compositions for performing any of the various assay formats described herein. The kit may include a composition that includes a primer set for amplifying the nucleic acid sequence of the coding portion, and, when applicable, a reagent for purifying the target-binding portion bound to an antibody, or a reagent for purifying a blood sample. The kit may include multiple primer sets for amplifying multiple sequences. The kit may include other reagents and / or information for peptide identification in solution or on an array and / or antibody profiling in a sample (e.g., buffers, nucleotides, instructions). The kit may also include multiple containers containing appropriate buffers and reagents.

[0237] Detailed description of the drawings

[0238] Figure 1 Two exemplary configurations of the target-binding portion described in the present disclosure are shown. As shown in (a), the target-binding portion may have a branched chain or configuration. As shown in (b), the target-binding portion may have a straight chain or configuration. In both configurations, the target-binding portion includes one or more target-binding units, or two or more binding elements. In configuration (a), the backbone for connecting the binding elements is a scaffold. In configuration (b), the scaffold and the binding elements are included in the same chain. The two binding elements are separated by a spacer such that they can simultaneously bind to a single molecule including the antigen-binding domain of an antibody.

[0239] Figure 2 Two exemplary configurations of the target-binding unit are shown. The target-binding unit may be the smallest unit for target binding. The target-binding unit includes two binding elements (e.g., peptides or peptidomimetics).

[0240] Figure 3Four exemplary structures of the spacer are shown. As shown in (a), the spacer is an unstructured peptide. As shown in (b), the spacer includes two peptide chains that fold into α-helices, and the two peptide chains further interact to form a coiled-coil structure. As shown in (c), the spacer includes two coiled-coil structures, and in this case, the two structured coiled-coils are folded from different regions of a single peptide chain. As shown in (d), the spacer includes a coiled-coil structure formed by three peptide chains, which include one longer peptide chain and two shorter peptide chains. The two shorter peptide chains can be further linked to two opposite ends of a double-stranded polynucleotide.

[0241] Figure 4 An example of an array having a number of discrete regions is shown. Each discrete region has one or more copies of a unique type of target-binding moiety attached thereto. For example, one or more copies of the target-binding moiety have the same sequence of the target-binding region, or have the same sequence of the binding element. In this example, the spacer is a double-stranded polynucleotide.

[0242] Figures 5A - 5C An exemplary scheme for generating polynucleotide-barcodeylated target-binding moieties by in vitro compartmentalization is shown. As Figure 5A shown, the coding portion (IVC01-001) can be compartmentalized in droplets that have a first soluble primer and beads having multiple copies of a second primer for generating cloned copies of the coding portion. The coding portion can encode a binding element and a peptide linker that can fold into an α-helix (LZ-A). As Figure 5B shown, multiple cloned copies of the coding portion are attached to the beads. As Figure 5C shown, each cloned copy of the coding portion is linked to a scaffold that has an α-helix (LZ-B) linked at each end. LZ-A and LZ-B can interact to form a leucine zipper. For more details, see Example 6.

[0243] Figures 6A - 6B An exemplary structure is shown that can be used to generate polynucleotide-barcodeylated binding elements by split-and-pool synthesis. As Figure 6A shown, structures known in the art can be used to generate polynucleotide-barcodeylated chemical libraries. For example, DNA can be synthesized from the 3'-end (as indicated by the arrow), and chemical building blocks can be synthesized from the free NH2 group. As Figure 6B shown, a structure designed in the present disclosure having two free NH2 groups can be used to generate a polynucleotide-barcodeylated chemical library having two identical binding elements. Such a polynucleotide-barcodeylated chemical library includes a library of polynucleotide-barcodeylated target-binding moieties.

[0244] Figures 7A - 7EAn exemplary experimental procedure for generating a library of target-binding moieties is shown (e.g., a library for generating HOP as described in Example 3).

[0245] Figure 8A An exemplary experimental procedure for generating a library of target-binding moieties is shown (e.g., a library for generating HOP and PEHOP as described in Example 11).

[0246] Figure 8B Shows during Figure 8A Experimental gel images of products generated during different steps of the procedure shown, and demonstrates that a library of nucleic acid templates of the desired size (536 bp) can be prepared using the Figure 8A procedure described. The nucleic acid template is DNA that can be further transcribed into RNA, which is used to generate a library of PEHOP.

[0247] Figure 9 A denaturing polyacrylamide gel image is shown that displays the presence of products of the desired size generated by further processing the DNA product (536 bp) of D14 during different steps of RNA display to prepare RNA-peptide fusion molecules, as described in Example 11.

[0248] Figure 10A A schematic diagram showing the use of RNA display to generate polynucleotide-peptide fusion molecules is shown.

[0249] Figure 10B Examples of polynucleotide-barcode target-binding moieties are shown. A DNA or RNA template can be linked to a rigid spacer. Various spacers described in the present invention can be used, e.g., double-stranded DNA or a coiled coil formed by two peptide chains. The two ends of the rigid spacer can be further linked to two binding elements (e.g., peptides). The binding elements can be encoded by the DNA or RNA template.

[0250] Figure 10C Examples of polynucleotide-barcode target-binding moieties with a flexible spacer are shown, which can be further manipulated to generate a rigid spacer. A DNA or RNA template encoding a unique binding element sequence can be fused with other DNA or RNA templates encoding the same unique binding element sequence (step 3*). In step (4), using RNA display, the fused templates can be used to generate polynucleotide-peptide fusion molecules (e.g., polynucleotide-barcode target-binding moieties). The polynucleotide-peptide fusion molecules include a peptide flexible spacer. The peptide flexible spacer can be manipulated to become a rigid spacer, e.g., by adding additional peptide chains to form a coiled coil structure with the peptide flexible spacer.

[0251] Examples

[0252] Example 1: Design of the target binding portion

[0253] Here, we describe a series of methods for preparing and using target-binding moieties and / or polynucleotide-barcodeylated target-binding moieties and other related molecular constructs to profile analytes (e.g., antibodies) in a sample. To simplify the description in the "Examples" section, "PIRM" is used as an example of a "binding element", "HOP" is used as an example of a target-binding moiety having at least two binding elements linked by a spacer, and "PEHOP" is used as an example of a polynucleotide-barcodeylated target-binding moiety. Thus, as described in the "Examples" section, PIRMs can be assembled into soluble oligomers, which we refer to as homo-oligomeric PIRM (HOP). In each HOP, multiple PIRMs (usually all PIRMs) can bind to the same immunoreceptor (e.g., the antigen-binding domain of an antibody). In some embodiments, two PIRMs in a HOP can bind to two identical Fab domains of the same antibody molecule. In some embodiments, all PIRMs in a HOP have the same structure. When a HOP stably associates with a polynucleotide (e.g., DNA or RNA) whose sequence reflects the PIRM sequence, the resulting molecule is called a polynucleotide-barcodeylated HOP (PEHOP).

[0254] Branched (PE)HOP and linear (PE)HOP

[0255] HOP can adopt a branched configuration or a linear configuration. In the branched configuration, there is a soluble scaffold molecule that includes multiple PIRM attachment sites. In some embodiments, the scaffold is a high molecular weight polymer or a dendrimer. In the linear configuration, the spacer and the PIRM can be part of a continuous linear polymer. In the case where both the PIRM and the spacer are polypeptides, the continuous linear polymer can be a continuous linear polypeptide.

[0256] Unfolded PIRM and folded PIRM

[0257] PIRM can be unfolded. For example, when short (e.g., <40-aa) peptides with randomly generated sequences are used as PIRMs, the PIRMs are likely to be unfolded, unstructured, and flexible. Alternatively, PIRM can be folded and structured. For example, antibody domains (e.g., single-chain Fv or V H domain) can be PIRM. Other protein types commonly used to design affinity reagents, such as DARPin, affibody, avidin can also be used as PIRM.

[0258] Peptide-based PIRM can contain unnatural amino acids

[0259] The site - specific incorporation of unnatural amino acids into peptides and proteins using in vitro translation systems, including RNA display, can be used to prepare PIRMs containing at least one unnatural amino acid.

[0260] Domain - level description of polynucleotide sequences

[0261] In the "Examples" section, polynucleotide sequences are sometimes described at the domain level. The name of each domain corresponds to a specific polynucleotide sequence. For example, domain "A" may have the sequence 5’-TATTCCC-3’, domain "B" may have the sequence 5’-AGGGAC-3’, and domain "C" may have the sequence 5’-GGGAAGA-3’. In this case, a polynucleotide having a sequence formed by the tandem connection of domains A, B, and C can be written as [A|B|C}. The symbol "[" represents the 5’ end, the symbol "}" represents the 3’ end, and the symbol "|" separates the domain names. An asterisk indicates sequence complementarity. For example, domain "B*" is the reverse - complementary sequence of domain "B".

[0262] Example 2: Preparation of individual PEHOP species using RNA display (linear, dimer)

[0263] This example shows how to use RNA display to prepare PEHOP with two copies of an antibody - binding peptide sequence.

[0264] Step A: A DNA template named gBlock01 encoding HOP can be prepared by standard DNA synthesis and gene assembly. gBlock01 will contain a T7 promoter sequence, a ribosome - binding site, and the coding sequence of HOP (including an amino - acid spacer and two peptide - based PRIM copies). The spacer will have a CC - B sequence that can form a coiled - coil with peptide CC - A. gBlock01 can be commercially obtained, such as in the form of gBlock from Integrated DNA Technology.

[0265] Step B: gBlock01 can be amplified using conventional PCR. The products of such PCR can be detected by running a 1% agarose gel electrophoresis.

[0266] Step C: The PCR products can be purified by Agencourt AMPure XP beads.

[0267] Step D: High - yield mRNA transcripts can be synthesized by in vitro transcription using T7 RNA polymerase (such as the Hiscribe T7 High Yield RNA Synthesis Kit, NEB#E2040S).

[0268] Step E: The mRNA transcript can be further purified by Agencourt AMPure XP beads. The concentration of the purified RNA can be obtained by UV absorption at 260 nm.

[0269] Step F: To attach the puromycin moiety to the 3'-end of the mRNA, an adaptor can be used. The adaptor can be formed by reacting a modified oligonucleotide named RTL1 with a modified oligonucleotide named DBCO.F.Puro (the sequence is shown in Table 1) to form RTL1-DBCO.F.Puro. For example, 100 μM RTL1 and 100 μM DBCO can be mixed in a 1:1 ratio. The conjugation reaction can be carried out overnight at 37 °C.

[0270] Step G: The RTL1-DBCO.F.Puro adaptor can then be purified by gel extraction. Specifically, the conjugation reaction product can be resolved on a 15% polyacrylamide gel containing 7 M urea, and the band containing conjugated RTL1-DBCO.F.Puro can be cut from the gel and frozen at -80 °C for 20 minutes. After that, the sample can be centrifuged at 15,000 g for 10 minutes. The supernatant can be taken out and further purified using a MicroSpin G25 column (GE27-5325-01 SIGMA).

[0271] Step H: The purified RTL1-DBCO.F.Puro can be mixed with the purified mRNA transcript prepared above at a ratio of 3:1 at 55 °C for 5 minutes for efficient annealing.

[0272] Step I: After the annealing reaction, a ligation reaction can be carried out between RTL1-DBCO.F.Puro and the mRNA transcript using T4 RNA ligase 1 (NEB#M0204S) at 37 °C for 1 hour. The product of this ligation can be named mRNA-RTL1-DBCO.F.Puro.

[0273] Step J: The ligation reaction can be further purified by Agencourt AMPure XP to remove RTL1-DBCO.F.Puro that does not hybridize with the mRNA transcript. The concentration of the purified mRNA-RTL1-DBCO.F.Puro can be obtained from UV absorption at 260 nm.

[0274] Step K: mRNA-RTL1-DBCO.F.Puro can be translated in vitro using the PURExpress in vitro protein synthesis kit (NEB #6800) to form an mRNA-peptide conjugate. Specifically, 10 μL of solution A, 7.5 μL of solution B, 1 μL of RnaseOut recombinant ribonuclease inhibitor (lot number 1872556, Invitrogen) can be mixed with the purified mRNA-RTL1-DBCO.F.Puro at 37 °C for 2 hours. The fusion between the mRNA and the nascent peptide can be promoted by adding salts to the translation reaction such that the final concentration of Mg ++ is 60 mM and the final concentration of K + is 600 mM. After adding the salts, the reaction can be frozen at -80 ° overnight to further promote the fusion. The translated peptide is an example of a linear dimer HOP, and the mRNA-peptide fusion described herein is an example of a linear dimer PEHOP. PEHOP can be further purified using various methods known to skilled biochemists. The purified PEHOP can be treated with reverse transcriptase and dNTPs to convert the mRNA into a cDNA:mRNA duplex. During this process, the RTL1 portion can serve as an RT primer.

[0275] Step L: Optionally, the peptide CC-A (sequence see Table 1) can be produced by standard recombinant protein expression techniques. The CC-A peptide can be mixed with the above mRNA-peptide fusion, and a coiled coil with an amino acid spacer (having the CC-B sequence) will form between two PIRMs and provide enhanced rigidity to the spacer.

[0276] Table 1 - Sequences used in Examples 2 - 5

[0277]

[0278]

[0279] Example 3. Preparation of a dsDNA molecular library encoding a library of linear dimer HOP (programmed DNA replication) Library

[0280] To prepare a linear dimer PEHOP library using RNA display, a library of the coding sequences of such HOPs (e.g., a nucleic acid template library) can be created first. Libraries encoding 10 3 to 10 5A library of oligonucleotides based on peptide - based PIRMs (each typically less than 200 nt). However, it is not easy to convert them into a library of two identical copies of long (e.g., >500 bp) dsDNA each containing the PIRM - encoding sequence. This example shows a method to achieve this. This process can be called programmed DNA replication ( Figures 7A - 7E ).

[0281] Step A: First, a gBlock named gBlock02 containing the sequence [b|E|MID|f} on the sense strand can be obtained from Integrated DNA Technologies (IDT). Specifically, [MID|f} encodes one strand of a coiled - coil, which we call CC - B (see Table 1 for the sequences of all domains).

[0282] Step B: A pair of primers for amplifying gBlock02 and introducing deoxyuridine can be commercially obtained. The forward primer is named SG0020. It has the sequence [b} and contains several deoxyuridine (dU) modifications. The reverse primer is named SG0013 and has the sequence [f*}. This pair of primers can be used to amplify gBlock02 using Taq.

[0283] Step C: The PCR product ( Figure 7A SG01 - 001) can be treated with the USER enzyme mixture (NEB#M5508) at 37 °C for 30 minutes to generate long 3' overhangs (domain b*), as shown in SG01 - 002.

[0284] Step D: A mixture of 10,000 oligonucleotide species (collectively called the SG0012_library) each having the sequence [E*|a|P|b} can be obtained from CustomArray or Twist Biosciences. Note that the sequence of P is variable and each oligonucleotide species has a unique P sequence. The P sequence can be designed randomly or based on biological sequences. An example of P is 5’ - GACTACAAAGACGATGACGATAAA - 3’. The SG0012_library can be annealed with SG01 - 002, where the domain b of the SG0012_library hybridizes with the newly exposed domain b* of SG01 - 002.

[0285] Step E: Phi29 DNA polymerase (NEB#M0269S) can be used to convert the hybridized product into a dsDNA library (SG01 - 003).

[0286] Step F: The dsDNA library can be amplified using primers SG0011 and SG0013 (see Table 1) to form SG01 - 004.

[0287] Step G: Then SG01-004 can be treated with T7 exonuclease (NEB#M0263S) to remove the bottom strand. The phosphorothioate modification can protect the top strand (SG01-005) from degradation.

[0288] Step H: SG01-005 can be purified with AMPure Bead. The purified SG01-005 can be treated with USER enzyme mixture (NEB#M5508) at 37 °C for 30 minutes to remove the 5'-phosphorothioate-containing region and generate a defined product (SG01-006).

[0289] Step I: Then, a 5'-phosphorylated splint named SG0014 (with the sequence [a*|E|f*}) and a photocleavable oligonucleotide named SG0008 (with the sequence [E*|a}, having a photocleavable linker between domain E* and domain a) can be obtained from IDT. SG0014 and SG0008 can be annealed at a ratio of 1:1.5, and the annealing product with approximately 10 nM SG0014 can be added to SG01-006 such that the domain f* of SG0014 binds to the domain f of SG01-006.

[0290] Step J: Unbound SG0014 and SG0014:SG0008 duplexes can be removed using AMPure beads. Then the purified product can be diluted in 50 μL of ligation buffer, and UV irradiation can be applied to cleave SG0008. The residue of SG0008 will spontaneously separate from SG0014. Then the product can be incubated at approximately 50 °C such that SG01-006 can be cyclized to form SG01-007.

[0291] Step K: The mixture can be concentrated at 4 °C using an Amicon Ultra 3K (EMD Millipore) filter. The concentrated DNA can be subjected to a ligation reaction with T4 DNA ligase (NEB#M0202) at room temperature for 30 minutes to form circular ssDNA SG01-008. Then the reaction can be incubated at 65 °C for 10 minutes to inactivate T4 DNA ligase.

[0292] Step L: The ligation product is heated to 95 °C for 5 minutes in the presence of a large excess of the competitor oligomer SG0015 with the sequence [f|E*|a}, which will enter the reaction to bind free oligomer SG0014. Then, SPRI purification is performed using AMPure beads to remove all free short oligomers such as SG0014 and SG0015. The reaction is then treated with Escherichia coli exonuclease I (NEB#M0293S) to degrade all single-stranded DNA using its 3’-5’ single-stranded nuclease activity.

[0293] Step M: Oligomer SG0014 is then added back into the reaction as a primer. Then, a DNA polymerase such as the Klenow fragment (exo-) will be used to extend on SG0014 and form a nicked double-stranded circular DNA, which will be further processed with T4 DNA ligase to form circular dsDNA SG01-009.

[0294] Step N: SG01-009 is then treated with the Nt.BstNBI enzyme (NEB#0607S) to produce nicked gene fragments. Then, Phi29 DNA polymerase (NEB#M0269S) will be used to extend the products to form full-length dsDNA, the top strand of which has the sequence [a|P|b|E|MID|f|E*|a|P|b} (SG01-010).

[0295] Step O: Perform PCR amplification on SG01-010 using two modified primers containing deoxyuridine (dU)-modified SG0007 (basically having the sequence [a}) and SG0004 (basically having the sequence [b*}) in the presence of two blocking oligomers SG0009 (having the sequence [E*|a}) and SG0006 (having the sequence [E*|b*}). Both blocking oligomers should have a 3' reverse dT modification. In the PCR reaction, the concentrations of both SG0007 and SG0004 are about 400 nM; while the concentrations of both SG0009 and SG0006 are about 40 nM. The PCR conditions are as follows: (1) denaturation temperature for 2 minutes; (2) blocking temperature for 5 minutes; (3) priming temperature for 1 minute; (4) extension temperature for 1 minute; go to (1) and cycle 1 more time. The denaturation temperature is about 95 °C; the blocking temperature is about 60 °C; the priming temperature is about 50 °C; the extension temperature is about 72 °C. The exact temperature can be optimized according to the buffer conditions and the enzyme used. A basic principle of this PCR reaction is that at the blocking temperature, the blocking oligomers will bind to the middle of the DNA template rather than the ends of the DNA template because the individual domains a and b cannot stably bind their complementary sequences at the blocking temperature, but at the priming temperature, the individual domains a and b do stably bind their complementary sequences; and at this temperature, the primers will kinetically compete with the blocking oligomers for binding to the ends of the DNA template, thus starting priming. The PCR reaction can be allowed to proceed for 2 cycles.

[0296] Step P: Treat the PCR product with USER enzyme (NEB#M5508) at 37 °C for 30 minutes to produce SG01-011 with long 3' sticky ends. Then, add two ssDNAs SG0010 (having the sequence [C|a}) and \SG0005 (having the sequence [D*|b*}) to SG01-011. Then, raise the temperature to 60 °C for 5 minutes and slowly cool to room temperature to allow the two ssDNAs to anneal to the USER-treated product. Then add T4 DNA ligase (NEB#M0202) to the reaction to ligate the gene fragments. Then incubate the reaction at 65 °C for 10 minutes to inactivate the T4 DNA ligase. Finally, use Phi29 enzyme (NEB#M0269S) to extend the product to form a dsDNA (SG01-012) having the sequence [C|a|P|b|E|MID|f|E*|a|P|b|D}. This final product can be PCR amplified using standard methods.

[0297] Example 4: Generation of linear, dimer PEHOP libraries using RNA display

[0298] Using the RNA display method described in Example 2, the DNA library SG01-012 created in Example 3 can be made into a PEHOP library.

[0299] Example 5: Generation of linear, multimer PEHOP libraries using RNA display

[0300] The above example shows how to create a PEHOP library, where each HOP contains exactly two PIRMs. Here, we show an example of creating a linear PEHOP library, where each HOP includes more than two PIRMs.

[0301] Single-stranded circular DNA SG01-008 can anneal with primer XC0001 having the sequence [D*|f*}, and the annealing product can be extended by Phi29 DNA polymerase to initiate a rolling circle amplification (RCA) reaction. The RCA product will be purified by AMPure beads and will anneal with ssDNA XC0002 having the sequence [C|MID5}, where domain MID5 is the first approximately 20 bases of domain MID. The annealing product will be extended by Phusion DNA polymerase. The product will be further purified by AMPure beads and PCR amplified with primers XC0003 (having a T7 promoter sequence) and XC0004 (having the sequence [D*}).

[0302] The lengths of the PCR products will vary and can be resolved by agarose gel electrophoresis. Products having the desired length can be purified by agarose gel purification. For example, here the length of each [a|P|b|MID|f} repeat is approximately 536 bp. Thus, a product having approximately 10 repeats will be approximately 5.5 kb. Therefore, if approximately 10 repeats are needed, the gel portion corresponding to approximately 5.5 kb can be excised and the contents eluted. The gel-purified DNA template can be further PCR amplified and used for in vitro transcription to produce mRNA transcripts, which can then be used in mRNA display as described in Example 1 to produce a PEHOP library.

[0303] The length and sequence of the spacer (in this case, the CC-B peptide) can be adjusted as needed.

[0304] Example 6: Generation of linear, dimer PEHOP libraries using in vitro compartmentalization

[0305] This example is described with reference to Figure 5. A double-stranded DNA template library (IVC01-001) each having a T7 promoter, a ribosome binding site, and a coding sequence of a peptide-based PIRM (which we call an antibody-binding peptide or ABP) fused to LZ-A will be prepared using standard gene synthesis methods. As described above, the ABP sequence is variable. LZ-A can form a stable leucine zipper with LZ-B. A solution containing IVC01-001, primer-modified magnetic beads (IVC01-003), and primer IVC01-002 will be emulsified to produce water-in-oil droplets such that a large number of droplets contain only 1 molecule of IVC01-001 (seeFigure 5A ) Here, primers IVC01 - 002 and IVC01 - 004 can be used to amplify IVC01 - 001. Primer IVC01 - 002 has a 5' azide modification and will be used in the conjugation reaction described below. The emulsion will be subjected to PCR. As a result, multiple copies of IVC01 - 001 will be covalently linked to magnetic beads ( Figure 5B ). The beads generated by this emulsion PCR method each contain multiple cloned copies of dsDNA and have been used in many sequencing technologies such as Roche 454 and SOLiD.

[0306] Next, the emulsion is demulsified, and the IVC01 - 001 molecules on the beads are covalently conjugated to a scaffold molecule ( Figure 5C ). Here, the scaffold consists of two DNA oligonucleotides IVC01 - 005 and IVC01 - 006, which can hybridize to form a dsDNA of approximately 30 bp. The 3' ends of both IVC01 - 005 and IVC01 - 006 are covalently modified with LZ - B peptide. In addition, IVC01 - 005 also has a 5' DBCO modification that can react with the azide group on the dsDNA (introduced by IVC01 - 002) to form a covalent bond.

[0307] The beads are washed, resuspended in an in vitro transcription and translation mixture, and emulsified again to form water - in - oil droplets (see Tawfik and Griffiths, 1998, Nature Biotechnology, Vol. 16, p. 652). In these droplets, IVC01 - 001 will be transcribed and translated to form a fusion peptide comprising ABP and LZ - A. The LZ - A domain will form a stable leucine zipper with LZ - B on the scaffold, thereby generating HOP. Then the emulsion can be demulsified, and the beads can be washed.

[0308] Finally, the IVC01 - 001 on the beads (now covalently linked to HOP) can be released from the beads by various methods. For example, primer IVC01 - 002 can contain deoxyuridine within the first approximately 5 bases, in which case, the HOP - modified IVC01 - 001 can be released from the beads by treatment with a USER enzyme mixture. The released HOP - modified IVC01 - 001 molecules will be a soluble PEHOP library.

[0309] Example 7: In vitro transcription

[0310] This article provides methods for Example protocol for the kit (ThermoFisher). SP6, T3, and T7 phage RNA polymerases are widely used for in vitro synthesis of RNA transcripts from DNA templates. The template usually has a double-stranded promoter of 19 - 23 bases upstream of the sequence to be transcribed. The template is then mixed with the corresponding RNA polymerase, rNTPs, and transcription buffer, and the reaction mixture is incubated at 37 °C for 10 minutes to 1 hour. The RNA polymerase first binds to its double-stranded DNA promoter, then separates the two DNA strands, and synthesizes complementary 5' to 3' (runaway transcription) at the end of the DNA template using the 3' to 5' strand as a template.

[0311] The initiation of transcription is the rate-limiting step of the in vitro transcription reaction; the elongation of transcription is very rapid. Phage RNA polymerases have high specificity for their respective promoters. Many multi-purpose cloning vectors contain two or more independent phage promoters flanking the multiple cloning site. Since RNA polymerases have high promoter specificity, either strand of the template can be transcribed with little "crosstalk" with the promoter on the opposite strand. The kit can also be used for transcribing DNA templates generated from PCR. In fact, DNA from PCR can be directly used in the MAXIscript kit without any pretreatment or purification.

[0312] Example 8: In vitro translation

[0313] The most commonly used cell-free translation systems consist of extracts from rabbit reticulocytes, wheat germ, and Escherichia coli. All can be prepared as crude extracts containing all the macromolecular components required for translating foreign RNA (70S or 80S ribosomes, tRNAs, aminoacyl-tRNA synthetases, initiation, elongation, and termination factors, etc.). To ensure efficient translation, each extract can be supplemented with amino acids, energy sources (ATP, GTP), energy regeneration systems (phosphocreatine and creatine phosphokinase for eukaryotic systems, and phosphoenolpyruvate and pyruvate kinase for E. coli lysates), and other cofactors (Mg 2+ 、K + etc.).

[0314] There are two methods of in vitro protein synthesis based on the starting genetic material: RNA or DNA. Standard translation systems, such as reticulocyte lysates and wheat germ extracts, use RNA as a template; while "coupled" and "linked" systems start from a DNA template, which is transcribed into RNA and then translated.

[0315] Example 9: Generation of linear, dimer PEHOP libraries using split - pool synthesis

[0316] The creation of DNA-encoded small molecule libraries using a split-and-pool method has been reported. For example, Clark et al. (2009 Nat Chem Biol., Vol. 5, p. 647) described in detail how to prepare a DNA-encoded small molecule library with far more than 5x10 6 species. A PEHOP library can be created using a similar strategy. For example, "AOP-Headpiece" (see the supplement of Clark et al., Figure 2 , which is reproduced here as Figure 6A ) can be replaced by a similar molecule, but has two amino groups on which other parts can be built. The two amino groups can be separated by a spacer part made of approximately 30-bp dsDNA( Figure 6B ). An example of such a spacer is formed by the modified oligonucleotides SnP01-001 and SnP01-002, shown in Figure 6B . Both of these oligonucleotides can have amine modifications at the 3' end. In addition, the 5' end of SnP01-002 can be conjugated to the loop of the original hairpin structure in AOP-Headpiece through various chemical bonds. For example, the amino group of AOP-Headpiece can be replaced by an azide group, and the 5' end of SnP01-002 can be modified with a DBCO group, which can form a covalent bond with the azide group through copper-free click chemistry.

[0317] Example 10: Disease diagnosis using HOP or PEHOP libraries

[0318] The examples provided here are a strategy for developing a diagnostic test for a specific disease.

[0319] Phase 1: Use the PEHOP library to identify a set of peptides that can distinguish serum from diseased subjects and healthy subjects

[0320] (1) Prepare a PEHOP library with at least 10 5 members (each member has many copies)

[0321] (2) Collect serum samples from approximately 100 diseased patients (referred to as "disease serum") and serum samples from approximately 100 non-diseased subjects (referred to as "non-disease serum")

[0322] (3) For each serum sample, mix it with an aliquot of the PEHOP library, remove the PEHOPs that do not bind to the antibodies in the serum, and sequence the coding sequences of the PEHOPs that bind to the antibodies

[0323] (4) Identify a set of approximately 50 peptides that frequently bind to disease serum and rarely bind to non-disease serum.

[0324] Phase 1 (Alternative Strategy): Use a HOP array to identify a set of peptides that can distinguish sera from diseased and healthy subjects

[0325] (1) Prepare a high-density HOP array with at least 10 4 discrete regions, each discrete region having a peptide sequence

[0326] (2) Collect serum samples from approximately 100 diseased patients (referred to as "disease sera") and serum samples from approximately 100 non-diseased subjects (referred to as "non-disease sera")

[0327] (3) For each serum sample, apply it to the HOP array, remove unbound antibodies, and quantify the remaining antibodies on each feature

[0328] (4) Identify a set of approximately 50 peptides that bind frequently to disease sera and rarely to non-disease sera.

[0329] Phase 2: Use the set

[0330] (1) Print a low-density HOP array with approximately 50 features, each feature being a peptide sequence from the set identified in Phase 1.

[0331] (2) Contact the HOP array with a serum sample from a subject with an unknown disease state, wash away unbound antibodies, and quantify the retained antibodies from each feature to form a profile, where the profile indicates the disease state of the subject.

[0332] Example 11: Generation of HOP or PEHOP libraries

[0333] In this example, a modified process for generating a HOP or PEHOP library is described. Some of the steps are similar to those described in the previous examples (Examples 3 - 5).

[0334] Construction of DNA Replication

[0335] The gBlock gene fragment named SG0017_f-middle-E-b contains a sequence encoding an amino acid spacer and was ordered from Integrated DNA Technologies (IDT). PCR was performed using the forward primer SG0020_b primer_dU (containing deoxyuridine (dU) at its 5' end) and the reverse primer SG0013_f* primer (at Figure 8ANamed D1) in []. The PCR conditions were as follows: initial denaturation (95°C for 30 seconds), 20 cycles (95°C for 30 seconds, 51°C for 30 seconds, and 68°C for 30 seconds), and finally extension at 68°C for 5 minutes. The PCR product was treated with USER enzyme (NEB#M5508) at 37°C for 30 minutes. Then, an oligomer containing the sequence encoding the FLAG peptide (SG0012_[E*|a|P|b}_HPLC) was added to anneal to the USER-treated product at the 5' end of the upper strand ( Figure 8A ). The reaction was extended at 30°C for 30 minutes by Phi29 DNA polymerase (NEB#M0269S). Then, PCR was performed using the forward primer SG0011_[A*G*|dU|E*|a} (phosphorylated modified oligomer) and the reverse primer SG0013_f* primer (named D4 in Figure 8A ). The PCR product was treated with T7 exonuclease and then with USER enzyme to generate a single-stranded gene fragment. The single-stranded DNA was further circularized using CircLigase ssDNA ligase (Epicentre#CL4115K) (named D7 in Figure 8A ). Five staple oligomers were also added to facilitate the cyclization reaction, and uncircularized products were removed using exonuclease V (NEB#M0345S). Then, a double-stranded circular product was developed by first annealing the single-stranded circular product (D7 in Figure 8A ) with SG0014_[a*|E|f*} and then extending and ligating using the NEBNext Second Strand Synthesis Enzyme Mix (NEB#E6112). The ligated double-stranded circular product was nicked using Nt.BstNBI enzyme (NEB#R0607S) and further extended using Bst 2.0 DNA polymerase (NEB#M0537S) (named D12 in Figure 8A ). After extension, as in Figure 8AAs shown, the sequence encoding the FLAG peptide was replicated. To further introduce sequences including the T7 promoter, ribosome binding site, heterologous reading frame stop codon, etc., DNA replication and several other steps were continued. First, the extension product was amplified by PCR, and deoxyuridine (dU) was introduced at the 5' end of each strand. Specifically, the SG0007_a primer_dU (0.2 μM) and the SG0004_b* primer_dU (0.2 μM) were used as the forward and reverse primers, respectively, for the PCR reaction. Two blockers containing reverse dT at the 3' end (SG0009_[E*|a} blocker and SG0006_[E*|b*} blocker) were also added to the reaction, but at a much lower concentration (20 nM). The specific PCR conditions were as follows: initial denaturation (95°C for 30 s), 10 cycles (95°C for 30 s, 64°C for 1 min to allow blocker binding, and 50°C for 1.5 min to allow primer binding, then 72°C for 30 s), and finally extension at 72°C for 5 min. The PCR product was further treated with USER enzyme (NEB#M5508), and then annealed with the SG0010[C|a} primer (containing the T7 promoter and ribosome binding site) and the SG0005[D*|b*} primer (containing the heterologous reading frame stop codon). The annealed product was ligated and extended using the NEBNext Second Strand Synthesis Enzyme Mix (NEB#E6112). Finally, the extension product was amplified by PCR using the forward primer SG0001_T7_F and the reverse primer SG0002_ frame stop_R. 8M urea denaturing 5% polyacrylamide gel electrophoresis was used to analyze the final product (named D14 in Figure 8A ), as well as other intermediate products ( Figure 8B ). The full length of the DNA replication construct was 536 bp, which contained the sequence of the T7 promoter, ribosome binding site, and the coding sequence for HOP including the amino acid spacer and two FLAG peptides, labeled as D14 (536 bp) in Figure 8B .

[0336] Ligation of mRNA to puromycin-linker DNA.

[0337] DNA replication mRNA transcripts were synthesized using the HiScribe T7 High Yield RNA Synthesis Kit (NEB #E2040S). The RTL1 adapter (containing an azide group) and the puromycin adapter (containing a dibenzocyclooctyne moiety (DBCO)) were ordered from IDT, and the conjugation reaction between the two adapters was carried out by mixing the two adapters and incubating overnight at 37 °C at a ratio of puromycin (DBCO):RTL1 (3:1). The annealing reaction between the RTL1 adapter and the mRNA 3'-end was carried out by mixing the RTL1-Puro(DBCO) conjugation product with mRNA at a ratio of 3:1. The reaction was incubated at 55 °C for 5 minutes and then slowly cooled to room temperature. The reaction was incubated at 37 °C for 1 hour using T4 RNA ligase (#NEB M0204s) to ligate the mRNA to the RTL1-Puro(DBCO) DNA adapter. In addition, the RTL1 adapter was also directly annealed and ligated to the mRNA as a control for downstream pull-down assays. Finally, all ligated products were analyzed using 8M urea denaturing 5% polyacrylamide gel electrophoresis( Figure 9 ). The ligation products (named D14(RNA)-RTL1 and D14(RNA)-RTL1-Puro, respectively) were confirmed by the size changes in lanes 3 and 4 of Figure 9 .

[0338] Cell-free translation

[0339] The mRNA-puromycin conjugate was mixed with Solution A and Solution B of the PURExpress In Vitro Protein Synthesis Kit (NEB #6800), and the mixture was incubated at 37 °C for 2 hours. The reaction was terminated by incubating at 4 °C for 10 minutes. To enhance the formation of the fusion peptide, the post-translation product was incubated at -20 °C overnight in the presence of high salt (final concentrations of KCl and MgCl2 were 600 mM and 60 mM, respectively).

[0340] Pull-down of the mRNA-protein fusion

[0341] EDTA and urea were added to the post-translation product at final concentrations of 125 mM and 4 M, respectively. The sample was then heated to 95 °C for 5 minutes and subsequently purified using Agencourt AMPure XP beads (BECKMAN COULTER, #A63881). After purification, monoclonal The BioM2 antibody (Sigma Aldrich, #F9291) was added to the solution, and the reaction was incubated at room temperature for 30 minutes. Then, the Dynabead magnetic beads (Invitrogen, MyOne Streptavidin C1) were washed three times with 1X phosphate-buffered saline (PBS) buffer. The washed beads were added to a mixture containing the mRNA-protein fusion and the monoclonal anti-FLAG antibody. The reaction was incubated at room temperature for 30 minutes, and then the beads were washed twice with 1X PBS and eluted with 95% formamide. The final mRNA-protein fusion was obtained by heating the sample at 95 °C for 5 minutes and separating the mRNA-protein fusion and the beads using a magnetic separator (Permagen Labware). The size shift of the fusion protein was as Figure 9 shown (lane 5).

[0342] Table 2 - Sequences used in Example 11

[0343]

[0344]

Claims

1. A composition comprising a plurality of polynucleotide - barcoded target - binding moieties, wherein each polynucleotide - barcoded target - binding moiety of the plurality of polynucleotide - barcoded target - binding moieties comprises: (a) a nucleic acid sequence that is linked via a linker to (b) a target - binding unit, the target - binding unit comprising (i) a first peptide sequence comprising a first binding region, and (ii) a second peptide sequence comprising a second binding region; wherein the first binding region and the second binding region (i) are separated by a polypeptide spacer, and (ii) are spaced such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule comprising an antigen - binding domain of an antibody; wherein the nucleic acid sequence comprises a unique barcode sequence and sequences encoding the first peptide sequence, the second peptide sequence, and the polypeptide spacer; wherein the unique barcode sequence comprises at least 4 consecutive nucleotides; wherein the linker comprises puromycin or a derivative thereof; wherein the polypeptide spacer comprises at least 20 amino acids, and the polypeptide spacer comprises flexible polypeptides and / or folded polypeptides; and wherein the composition is soluble.

2. The composition according to claim 1, wherein the polypeptide spacer comprises the same amino acid sequence between the plurality of polynucleotide - barcoded target - binding moieties.

3. The composition according to claim 1 or 2, wherein the single molecule comprises a first antigen - binding domain and a second antigen - binding domain; wherein the first binding region and the second binding region are spaced such that the first binding region binds to the first antigen - binding domain and the second binding region binds to the second antigen - binding domain.

4. The composition according to claim 1 or 2, wherein the first binding region and the second binding region have the same sequence.

5. The composition according to claim 1 or 2, wherein the first binding region and the second binding region have the same structure recognized by the single molecule.

6. The composition according to claim 1 or 2, wherein the nucleic acid sequence is single - stranded.

7. The composition according to claim 1 or 2, wherein the nucleic acid sequence is double - stranded.

8. The composition according to claim 1 or 2, wherein the nucleic acid sequence is deoxyribonucleic acid (DNA).

9. The composition according to claim 1 or 2, wherein the nucleic acid sequence is ribonucleic acid (RNA).

10. The composition according to claim 1 or 2, wherein the nucleic acid hybridizes with a primer.

11. The composition according to claim 1 or 2, wherein each polynucleotide - barcoded target - binding moiety comprises a single target - binding unit.

12. The composition according to claim 1 or 2, wherein each polynucleotide - barcoded target - binding moiety comprises two or more target - binding units.

13. The composition according to claim 1 or 2, wherein the composition comprises a plurality of at least 1,000 different polynucleotide - barcoded target - binding moieties.

14. The composition according to claim 1 or 2, wherein at least one of the plurality of polynucleotide barcoded target binding moieties comprises a single target binding unit.

15. The composition according to claim 1 or 2, wherein two or more of the plurality of polynucleotide barcoded target binding moieties comprise a single target binding unit.

16. The composition according to claim 1 or 2, wherein at least one of the plurality of polynucleotide barcoded target binding moieties comprises two or more target binding units.

17. The composition according to claim 1 or 2, wherein two or more of the plurality of polynucleotide barcoded target binding moieties comprise two or more target binding units.

18. The composition according to claim 1 or 2, wherein the composition comprises 2 to 1000 target binding units.

19. The composition according to claim 1 or 2, wherein the composition comprises three or more target binding units.

20. The composition according to claim 1 or 2, wherein the polypeptide spacer comprises a coiled coil structure or a β-sheet structure.

21. The composition according to claim 1 or 2, wherein the polypeptide spacer comprises two or more separate peptide chains, wherein at least one of the two or more separate peptide chains comprises an α-helix or a β-strand.

22. The composition according to claim 1 or 2, wherein the polypeptide spacer comprises a single peptide chain folded into at least two α-helices or at least two β-strands.

23. The composition according to claim 22, wherein the single peptide chain comprises four α-helices or four β-strands.

24. The composition according to claim 1 or 2, wherein the polypeptide spacer comprises a first peptide chain, a second peptide chain and a third peptide chain, wherein the second peptide chain interacts with a first portion of the first peptide chain, and wherein the third peptide chain interacts with a second portion of the first peptide chain.

25. The composition according to claim 24, wherein the first peptide chain, the second peptide chain and / or the third peptide chain fold into an α-helix.

26. The composition according to claim 24, wherein the second peptide chain and the third peptide chain are attached to opposite ends of a double-stranded polynucleotide.

27. The composition according to claim 1 or 2, wherein the first binding region and the second binding region comprise the same epitope.

28. The composition according to claim 1 or 2, wherein the sequence of the first peptide sequence is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98% or at least 99% or 100% identical to the sequence of the second peptide sequence.

29. The composition according to claim 1 or 2, wherein the nucleic acid sequence is a double-stranded DNA-RNA hybrid.

30. The composition according to claim 1 or 2, wherein each polynucleotide-barcode-target-binding moiety is a linear polymer chain.

31. The composition according to claim 1 or 2, wherein the first peptide sequence, the second peptide sequence, and the polypeptide spacer are contiguous in a single polypeptide chain.

32. The composition according to claim 1 or 2, wherein the antigen-binding domain is a scFv, Fab, or F(ab)2.

33. The composition according to claim 1 or 2, wherein the first peptide sequence and the second peptide sequence are at least 5 amino acid residues in length.

34. The composition according to claim 1 or 2, wherein each polynucleotide-barcode-target-binding moiety of the plurality of polynucleotide-barcode-target-binding moieties is within a container.

35. The composition according to claim 1 or 2, wherein each polynucleotide-barcode-target-binding moiety of the plurality of polynucleotide-barcode-target-binding moieties is within a different one of a plurality of containers.

36. The composition according to claim 35, wherein the container is a droplet.

37. The composition according to claim 36, wherein the droplet is a water-in-oil droplet.

38. The composition according to claim 1 or 2, wherein the antibody is a biomarker.

39. The composition according to claim 1 or 2, wherein the polypeptide spacer comprises a secondary structure and / or a tertiary structure.

40. A method for preparing a target-binding moiety according to any one of claims 1-39 by RNA display.

41. A method for preparing a target-binding moiety according to any one of claims 1-39 by in vitro compartmentalization.

42. A plurality of nucleic acid molecules, wherein each nucleic acid molecule of the plurality of nucleic acid molecules encodes a polypeptide target-binding moiety, the polypeptide target-binding moiety comprising: (i) a first peptide sequence comprising a first binding region, and (ii) a second peptide sequence comprising a second binding region; wherein the first binding region and the second binding region are separated by a polypeptide spacer, and are spaced such that the first peptide sequence and the second peptide sequence simultaneously bind to a single molecule comprising an antigen-binding domain of an antibody; wherein the polypeptide spacer comprises at least 20 amino acids, and the polypeptide spacer comprises a flexible polypeptide and / or a folded polypeptide; and Wherein the plurality of nucleic acid molecules comprises at least 10 6 unique sequences.

43. The plurality of nucleic acid molecules according to claim 42, wherein the plurality of nucleic acid molecules are a plurality of double-stranded DNA molecules.

44. The plurality of nucleic acid molecules according to claim 42, wherein the plurality of nucleic acid molecules are a plurality of single-stranded RNA molecules.

45. The plurality of nucleic acid molecules according to claim 42, wherein each nucleic acid molecule of the plurality of nucleic acid molecules is a circular molecule.

46. The plurality of nucleic acid molecules according to claim 42, wherein the polypeptide spacer comprises a secondary structure and / or a tertiary structure.

47. A method for preparing a library of polynucleotide-peptide fusion molecules, comprising: (a) providing a plurality of nucleic acid molecules according to any one of claims 42-46; (b) Express the plurality of nucleic acid molecules in an RNA display assay to produce a plurality of peptides, wherein each nucleic acid molecule is linked via a linker to the peptide expressed from the nucleic acid molecule, and wherein the linker comprises puromycin or a derivative thereof.

Citation Information

Patent Citations

  • Magnetic particles for use in separations

    US4672040A

  • Protein A-silica immunoadsorbent and process for its production

    US4681870A

  • Magnetic particle based electrochemiluminescent detection apparatus and method

    US6133043A

  • Nucleic acid-protein fusion molecules and libraries

    US6281344B1

  • Protective casing for electric cables or wires.

    US630634A