Antibody protein product expression constructs for high-throughput sequencing

JP2025506122A5Pending Publication Date: 2026-02-16AMGEN INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024547046
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-02-10
Filing Date
2023-02-08
Publication Date
2026-02-16

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure provides polynucleotide molecules encoding an antibody protein product, the polynucleotide molecules comprising i) a nucleotide sequence specific for addition of a molecular barcode by template-switching reverse transcriptase (RT) and a unique molecular identifier (UMI) barcode, and a nucleotide sequence specific for a universal RT primer to facilitate high throughput sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 308,922, filed February 10, 2022, which is incorporated by reference in its entirety herein.

[0002] The present disclosure provides polynucleotide molecules that encode an antibody protein product, the polynucleotide molecules comprising a nucleotide sequence specific for addition of a molecular barcode by template-switching reverse transcriptase (RT) and a unique molecular identifier (UMI) barcode, as well as a nucleotide sequence specific for a universal RT primer to facilitate high throughput sequencing. [Background technology]

[0003] During the development of therapeutic antibodies, a lot of time and resources are invested in identifying potentially problematic physicochemical properties of the antibody to reduce the risk of expensive late-stage failure in the clinic. High-throughput sequencing (also known as next-generation sequencing) is one method of early evaluation of clones expressing antibody protein products. For example, high-throughput sequencing can be used to evaluate the diversity of antibodies encoded in combinatorial libraries. In addition, high-throughput sequencing allows for the selection of antibody protein products, allowing the identification of sequences that code for proteins with desired properties, such as high-affinity antibodies.

[0004] Fluidic systems can be used to perform high-throughput sequencing of antibody protein products. The use of optical barcodes and fluidic optics provides high-throughput single-cell screening capabilities based on nanofluidic and optoelectronic positioning technology. This technology is based on light-induced electrical motion that produces directed forces in both solid and fluidic structures (Jorgolli et al., Biotechnol Bioeng 2019, 116(9), 2393-2411). For example, commercially available fluidic devices such as the integrated technology of the Berkeley Lights (BLI) Beacon® Optofluidic System (Emeryville, CA) are flexible and have a wide range of application capabilities applicable to commercial large molecule drug development, including antibody discovery, clonal selection, gene editing, phenotype-genotype correlation, and cell line development. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Jorgolli et al.,Biotechnol Bioeng 2019,116(9),2393-2411 Summary of the Invention [Problem to be solved by the invention]

[0006] The use of optofluidic systems allows for rapid CEPA cycling of large panels (10,000 molecules) and high-throughput correlation of phenotype to genotype by machine learning. Because antibody protein products are large, complex molecules, efficient execution of high-throughput sequencing of expression constructs is challenging. For example, conventional optical barcodes typically allow for pooled export of up to 12 beads at a time for off-chip sequencing. Due to the size of constructs expressing antibody protein products, only the ends of the polynucleotide molecules are sequenced using standard construct and sequencing methods, and the entire polynucleotide molecule can only be sequenced if exported one bead at a time, not a pool. Therefore, off-chip sequencing is limited to thousands of sequences if pooled and hundreds if unpooled. Therefore, there is a need for the development of antibody protein product expression constructs that can be efficiently sequenced using high-throughput or next-generation sequencing technologies. [Means for solving the problem]

[0007] The present disclosure provides polynucleotide expression constructs for high throughput screening of clones in the development of antibody protein products. The polynucleotide sequences of the present disclosure include a universal molecular identification (UMI) barcode and a molecular barcode, such as an optical barcode during cloning. By reading both the molecular (e.g., optical) barcode and the UMI barcode together, the identity of the molecular sequence can be determined using an optofluidic device, for example, to associate the barcode with a specific isolation pen or fluidic device chamber.

[0008] The present disclosure provides a polynucleotide molecule encoding an antibody protein product, the polynucleotide molecule comprising: i) a nucleotide sequence specific for addition of a molecular barcode by a template-switching reverse transcriptase (RT); ii) a unique molecular identifier (UMI) barcode that is distinct from the molecular barcode; iii) a nucleotide sequence encoding a light chain polypeptide of the antibody protein product; iv) a nucleotide sequence encoding a heavy chain polypeptide of the antibody protein product; and v) a nucleotide sequence specific for a universal RT primer.

[0009] Template switching reverse transcriptase (RT) refers to an RT that adds several non-templated nucleotides after reaching the 5' end of an RNA template (e.g., 5'CCC). The non-templated nucleotides can anneal to a template switching oligo (TSO) of known sequence, and the reverse transcriptase easily switches the template from RNA to the TSO. For example, the TSO can contain three riboguanosines at its 3' end, which can anneal to the 5'CCC sequence of the RNA template. The resulting cDNA contains a known sequence (complementary to the sequence of the TSO) attached to the 3' end of the cDNA. In the polynucleotide molecules of the present disclosure, "a nucleotide sequence specific for the addition of a molecular barcode by template switching RT" refers to a nucleotide sequence that can anneal to a TSO, which contains a molecular barcode.

[0010] "Molecular barcode" refers to a unique nucleotide sequence that is a unique identifier for a polynucleotide sequence. For example, a molecular barcode refers to a string of random, partially degenerate, or defined nucleotides that can be used to identify a target molecule during DNA processing and / or sequencing. Molecular barcodes are useful for quantitative sequencing applications and also for genomic variant detection. The molecular barcode information, together with the alignment coordinates, allows for grouping of sequencing data into read families that represent individual sample DNA or RNA fragments.

[0011] The term "unique molecular identifier" refers to a molecular barcode used to identify the clone that produced the antibody sample product. In addition, the molecular barcode inserted by template switching RT refers to a nucleotide sequence that provides an additional level of identification, such as a nucleotide sequence that identifies the pen, well, or chamber from which the clone that produced the antibody protein product originated.

[0012] A "universal RT primer" is a primer that has a sequence complementary to a nucleotide sequence that is highly prevalent in a particular set of DNA molecules and cloning vectors. Thus, a universal RT primer can be used to reverse transcribe a collection of different DNA molecules that all share a nucleotide sequence common to the universal RT primer. This primer is capable of binding to a variety of DNA templates. A "nucleotide sequence specific to a universal RT primer" is a nucleotide sequence that anneals to a universal RT primer.

[0013] In any of the present polynucleotide molecules, the molecular barcode (e.g., UMI or optical barcode) comprises a random 6mer, 7mer, 8mer, 9mer, 10mer, 11mer, 12mer, 13mer, 14mer, or 15mer (including ranges between any two of the listed values, e.g., 6mer to 16mer, 10mer to 16mer, or 12mer to 16mer). In some embodiments, the optical barcode comprises a random nucleic acid sequence that is at least 6mer, 7mer, 8mer, 9mer, 10mer, 11mer, 12mer, 13mer, 14mer, or 15mer in length.

[0014] In any of the polynucleotide molecules described herein, the polynucleotide molecule further comprises a promoter sequence.

[0015] In addition, in any of the polynucleotide molecules described herein, the polynucleotide molecule further comprises at least two internal ribosome entry site (IRES) sequences. For example, an IRES can be located between two different open reading frames, such as between an open reading frame for a heavy chain polypeptide and an open reading frame for a light chain polypeptide, or between an open reading frame for a polypeptide of an antibody protein product and an open reading frame for a polypeptide of a selection gene product.

[0016] In addition, in any of the polynucleotide molecules described herein, the polynucleotide molecule further comprises at least an IRES sequence and at least one promoter. The present disclosure also provides any of the polynucleotide molecules further comprising at least two promoters.

[0017] In addition, in any of the polynucleotide molecules described herein, the polynucleotide molecule further comprises nucleotides encoding a selected gene product, such as puromycin-N-acetyltransferase.

[0018] In the exemplary polynucleotide molecules described herein, the nucleotide sequence specific for the universal RT primer is located between the nucleotide sequence encoding the light chain polypeptide and the nucleotide sequence encoding the heavy chain polypeptide.

[0019] In another exemplary polynucleotide molecule described herein, the nucleotide sequence specific for the universal RT primer is downstream of the nucleotide sequence encoding the light chain polypeptide of the antibody protein product.

[0020] In some embodiments, in any of the polynucleotide molecules of the present disclosure, the nucleotide sequence specific for addition of a molecular barcode by template switching reverse transcriptase is adjacent to the UMI barcode. The term "adjacent" refers to an element located directly downstream or directly upstream of a reference element. For example, in any of the polynucleotide molecules of the present disclosure, the nucleotide sequence specific for addition of a molecular barcode by template switching reverse transcriptase can be directly upstream of the UMI barcode.

[0021] In some embodiments, in any of the polynucleotide molecules of the present disclosure, the UMI barcode is directly upstream of the nucleotide sequence encoding the light chain polypeptide of the antibody protein product.

[0022] The present disclosure provides a polynucleotide molecule, wherein a nucleotide sequence specific for a RT universal primer is immediately downstream of a nucleotide sequence encoding a light chain polypeptide of an antibody protein product.

[0023] In any of the polynucleotide molecules of this disclosure, the polynucleotide comprises an IRES immediately downstream of the sequence specific for the RT universal primer.

[0024] In any of the polynucleotide molecules of this disclosure, the polynucleotide comprises an IRES immediately downstream of the nucleotide sequence encoding the heavy chain polypeptide of the antibody protein product.

[0025] In any of the polynucleotide molecules of this disclosure, the heavy chain polypeptide comprises a heavy chain variable region, and the light chain polypeptide comprises a light chain variable region.

[0026] In any of the polynucleotide molecules of the disclosure, (a) the nucleotide sequence encoding the light chain polypeptide of the antibody protein product encodes a light chain variable region and a light chain constant region immediately downstream of the light chain variable region, and (b) the nucleotide sequence encoding the heavy chain polypeptide of the antibody protein product encodes a heavy chain variable region and a heavy chain constant region immediately downstream of the heavy chain variable region.

[0027] In some of the polynucleotide molecules of this disclosure, a nucleotide sequence encoding a light chain polypeptide of an antibody protein product is upstream of a nucleotide sequence encoding a heavy chain polypeptide of the antibody protein product.

[0028] In some of the polynucleotide molecules of this disclosure, the nucleotide sequence encoding the light chain polypeptide of an antibody protein product is downstream of the nucleotide sequence encoding the heavy chain polypeptide of the antibody protein product.

[0029] The present disclosure provides a polynucleotide molecule encoding an antibody protein product comprising, in a 5' to 3' direction: i) a promoter; ii) a nucleotide sequence specific for addition of a molecular barcode by a template switching reverse transcriptase; iii) a unique molecular identifier (UMI) barcode that is distinct from the molecular barcode; iv) a nucleotide sequence encoding a first polypeptide of the antibody protein product; v) a nucleotide sequence specific for a reverse transcriptase (RT) universal primer; vi) a first IRES; vii) a nucleotide sequence encoding a second polypeptide of the antibody protein product; viii) a nucleotide sequence encoding a constant domain of a heavy chain of the antibody protein product; ix) a second IRES; and x) a nucleotide sequence encoding a select gene product, such as puromycin-N-acetyltransferase.

[0030] In addition, the present disclosure provides a polynucleotide molecule encoding an antibody protein product, the polynucleotide sequence comprising: i) an annealing site for a sequencing primer, ii) a unique molecular identifier barcode, iii) a nucleotide sequence encoding a light chain polypeptide of the antibody protein product, iv) a nucleotide sequence encoding a heavy chain polypeptide of the antibody protein product, and v) a nucleotide sequence specific for a universal reverse transcriptase (RT) primer.

[0031] In any of the polynucleotide molecules of the present disclosure, the molecular barcode is an optical barcode. Optical barcode refers to a molecular barcode that is visually detectable. In some embodiments, the optical barcode is identifiable (i) by determining its polynucleotide sequence, and (ii) by annealing to a polynucleotide probe that includes one or more optical moieties. The optical moieties can be fluoroprobes, fluorescent probes, or colorimetric probes. Exemplary optical barcodes are shown in Figures 4 and 5. Thus, a polynucleotide molecule that includes an optical barcode can be annealed to a polynucleotide probe that includes one or more optical moieties, and the identity of the molecular barcode can be identified in situ, for example, by light absorption or emission.

[0032] In any of the polynucleotide molecules of the present disclosure, the nucleotide sequence specific for attachment of an optical barcode is configured to receive an optical barcode conjugated to a solid support.

[0033] Additionally, in any of the polynucleotide molecules of this disclosure, the nucleic acid specific for addition of a molecular barcode by a template-switching reverse transcriptase is configured for addition of an optical barcode by a template-switching oligonucleotide from a template optical barcode conjugated to a solid support.

[0034] In any of the polynucleotide molecules of this disclosure, the 5' end of the nucleotide sequence specific for the attachment of a molecular barcode comprises the polynucleotide sequence CCC. In any of the polynucleotide molecules of this disclosure, the 5' end of the nucleotide sequence specific for the attachment of a molecular barcode comprises a polynucleotide sequence that is the reverse complement to TSO.

[0035] In any of the polynucleotide molecules of this disclosure, the molecular barcode is configured for reverse transcription by a template-switching reverse transcriptase primed by a RT universal primer.

[0036] In some embodiments, any of the polynucleotide molecules of the present disclosure, nucleotide sequences specific for a reverse transcriptase (RT) universal primer are configured to anneal to a free universal primer that is not disposed on a solid support, and the annealed universal primer is configured for reverse transcription of the nucleotide barcode by reverse transcriptase.

[0037] In various embodiments, the polynucleotide molecule of the present disclosure, wherein the solid support is a bead or microsphere, a membrane, a nanofiber, a nanotube, a resin, or agarose. For example, the solid support can be polymer-based, such as polystyrene beads, poly(lactide-co-glycolide) (PLGA) beads, polyethylene oxide (PEO) beads, polyethylene glycol (PEG) beads, polyvinyl alcohol (PVA) beads, or metal-based, such as gold beads. In addition, the solid support can be chitosan, dextran, alginate, gadolinium-based, carbon-based, silica-based, or iron-based. When the solid support is a resin, the resin can be a polymeric resin, such as cellulose, polystyrene, agarose, polyacrylamide, or agarose.

[0038] In any of the polynucleotide molecules of the disclosure, (a) the first polypeptide of the antibody protein product comprises a light chain polypeptide and the second polypeptide of the antibody protein product comprises a heavy chain polypeptide, or (b) the first polypeptide of the antibody protein product comprises a heavy chain polypeptide and the second polypeptide of the antibody protein product comprises a light chain polypeptide.

[0039] For example, in some polynucleotide molecules of the present disclosure, the light chain polypeptide comprises a light chain variable region and the heavy chain polypeptide comprises a heavy chain variable region.

[0040] In addition, in some of the polynucleotide molecules of the disclosure, the nucleotide sequence encoding the light chain polypeptide of the antibody protein product encodes a light chain variable region and a light chain constant region immediately downstream of the light chain variable region, and the nucleotide sequence encoding the heavy chain polypeptide of the antibody protein product encodes a heavy chain variable region and a heavy chain constant region immediately downstream of the heavy chain variable region.

[0041] In any of the polynucleotide molecules of the present disclosure that encode an antibody protein product, the antibody protein product comprises or consists of a larger peptide, an antibody, an antibody fragment, an antibody fusion peptide, or an antigen-binding fragment thereof. For example, the antibody is a polyclonal antibody or a monoclonal antibody.

[0042] The present disclosure also provides a method of screening clones expressing antibody protein products, comprising pooling clones containing any of the polynucleotide molecules disclosed herein, polymerizing the DNA of the pooled clones, and sequencing the DNA in a fluidic device. For example, the polynucleotide molecule may comprise RNA. For example, polymerizing the DNA of the pooled clones comprises reverse transcribing the polynucleotide molecule with template-switching reverse transcriptase.

[0043] The disclosure also provides a method comprising annealing any of the polynucleotide molecules disclosed herein to a template (e.g., the template is conjugated to a solid support) comprising an optical barcode, annealing a universal reverse transcriptase primer to the polynucleotide molecule, and extending the annealed universal reverse transcriptase primer with a template-switching reverse transcriptase, thereby generating a cDNA of the polynucleotide molecule comprising a UMI barcode and an optical barcode. For example, the method is performed in a fluidic device, and optionally, the fluidic device comprises or consists of a microfluidic chip or an isolation pen.

[0044] Also provided is any of the methods of the present disclosure, further comprising detecting the presence of a molecular barcode (e.g., an optical barcode) and / or a UMI barcode.

[0045] In any of the methods of the present disclosure, the method may be performed using any fluidic system, fluidic device, or fluidic apparatus known in the art. For example, the method may be performed in situ in a fluidic system, fluidic device, or fluidic apparatus. A fluidic device (or fluidic apparatus) is an element that includes one or more separate circuits configured to hold a fluid, each circuit being composed of fluidically interconnected circuit elements. The circuit elements, including but not limited to regions, flow paths, channels, chambers, and / or pens, and at least one port, are configured such that fluid can flow into and / or out of the fluidic device. The fluidic circuit may be configured to have a first end that is fluidically coupled to a first port (e.g., an inlet) in the fluidic device, and a second end that is fluidically coupled to a second port (e.g., an outlet) in the fluidic device or coupled to a second fluidic device or a second region, flow path, channel, chamber, or pen in the fluidic device. The fluidic device may be a microfluidic device, although other scales, such as nanoscale, may also be appropriate. For example, the fluidic device can be a microfluidic chip, a microfluidic channel, a microfluidic cell, a nanofluidic chip, a nanofluidic channel, a nanofluidic cell, or an isolation pen. In some embodiments, the fluidic system includes a multiwell plate, such as a 96 or 384 well plate. The multiwell plate can be in fluid communication with the circuit.

[0046] For a microfluidic device, the circuit includes a fluid region, which may include a microfluidic channel and at least one chamber, and holds a volume of fluid of less than about 1 mL (e.g., less than about 750, 500, 250, 200, 150, 100, 75, 50, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, or 2 μL). In certain embodiments, the circuit holds about 1-2, 1-3, 1-4, 1-5, 2-5, 2-8, 2-10, 2-12, 2-15, 2-20, 5-20, 5-30, 5-40, 5-50, 10-50, 10-75, 10-100, 20-100, 20-150, 20-200, 50-200, 50-250, or 50-300 μL. The circuit can be configured to have a first end that is fluidly coupled to a first port (e.g., an inlet) in the microfluidic device and a second end that is fluidly coupled to a second port (e.g., an outlet) in the microfluidic device.

[0047] As used herein, a "nanofluidic device" or "nanofluidic apparatus" is a type of fluidic device having a fluidic circuit including at least one circuit element configured to hold a volume of fluid of less than about 1 μL (e.g., less than about 750, 500, 250, 200, 150, 100, 75, 50, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 nL or less). A nanofluidic device can include a plurality of circuit elements (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10,000 or more). In certain embodiments, one or more (e.g., all) of the at least one circuit element are configured to hold a volume of fluid of about 100 pL to 1 nL, 100 pL to 2 nL, 100 pL to 5 nL, 250 pL to 2 nL, 250 pL to 5 nL, 250 pL to 10 nL, 500 pL to 5 nL, 500 pL to 10 nL, 500 pL to 15 nL, 750 pL to 10 nL, 750 pL to 15 nL, 750 pL to 20 nL, 1 to 10 nL, 1 to 15 nL, 1 to 20 nL, 1 to 25 nL, or 1 to 50 nL. In other embodiments, one or more (e.g., all) of the at least one circuit element are configured to hold a volume of fluid of about 20 nL-200 nL, 100-200 nL, 100-300 nL, 100-400 nL, 100-500 nL, 200-300 nL, 200-400 nL, 200-500 nL, 200-600 nL, 200-700 nL, 250-400 nL, 250-500 nL, 250-600 nL, or 250-750 nL.

[0048] "Fluid channel" or "flow channel" as used herein refers to a flow region of a fluidic device that is significantly longer in length than both the horizontal and vertical dimensions. For example, a flow channel can be at least 5 times longer, e.g., at least 10 times longer, at least 25 times longer, at least 100 times longer, at least 200 times longer, at least 500 times longer, at least 1,000 times longer, at least 5,000 times longer, or longer, in either the horizontal or vertical dimension. In some embodiments, the length of the flow channel ranges from about 50,000 microns to about 500,000 microns, including any range therebetween. In some embodiments, the horizontal dimension ranges from about 100 microns to about 1000 microns (e.g., about 150 to about 500 microns) and the vertical dimension ranges from about 25 microns to about 200 microns (e.g., about 40 to about 150 microns). It is noted that flow channels can have a variety of different spatial arrangements in a fluidic device and are therefore not limited to being perfectly straight elements. For example, a flow channel may include one or more sections having any of the following configurations: curved, bent, spiraling, sloping, descending, forking (e.g., multiple different flow paths), and any combination thereof. In addition, a flow channel may have different cross-sectional areas along its path that expand and contract to provide a desired flow path therein. [Brief description of the drawings]

[0049] [Figure 1] 1 shows a schematic diagram of a polynucleotide molecule for 5' sequencing. [Figure 2A] FIG. 1 shows a schematic diagram of a polynucleotide molecule designed for on-chip sequencing. [Figure 2B] FIG. 1 shows a schematic diagram of a polynucleotide molecule designed for on-chip sequencing. [Figure 3A] 1 shows a schematic diagram of a polynucleotide molecule designed for 3' sequencing. [Figure 3B] 1 shows a schematic diagram of a polynucleotide molecule designed for 3' sequencing. [Figure 4] 1 shows an exemplary 5' optical barcode chemistry. [Diagram 5] 1 shows an exemplary 3' optical barcode chemistry. [Figure 6] FIG. 1 shows a schematic diagram of a conventional polynucleotide molecule that is subject to limitations when attempting to sequence using a 3′ sequencing kit. [Figure 7] FIG. 1 shows a schematic diagram of a conventional polynucleotide molecule that is subject to limitations when attempting to sequence using a 5′ sequencing kit. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0050] The present disclosure provides polynucleotide molecules that encode an antibody protein product, the polynucleotide molecules comprising i) a nucleotide sequence specific for addition of a molecular barcode by a template-switching reverse transcriptase (RT) and a unique molecular identifier (UMI) barcode (distinct from the molecular barcode) and a nucleotide sequence specific for a universal RT primer to facilitate high throughput sequencing.

[0051] The polynucleotide sequences of the present disclosure have the advantage of functioning in commercially available constructs. Conventional template switching reverse transcriptases limit reverse transcription to a short number of nucleotides, for example, about 500 to about 1000 nucleotides. The polynucleotide molecules disclosed herein position a universal RT primer downstream of the polynucleotide sequence encoding the polypeptide of an antibody protein product (e.g., the constant domain of the light chain). This allows the custom primer to be flowed into the fluidic device, thus eliminating the need for the primer to be present on beads or a solid support. Beads containing oligo-dT may be used to capture the mRNA of the polynucleotide sequence. The custom primer then binds to the captured mRNA and initiates extension near the 5' end of the polynucleotide molecule.

[0052] For example, commercially available 3' sequencing kits do not effectively sequence commercially available landing pad constructs that have a polyadenylation signal (pA) in the cell host landing pad too far downstream of the cloning insert junction. The UMI barcode inserted at the insert junction is not close enough to the optical barcode that becomes part of the dT oligo downstream of the pA (see, e.g., FIG. 6). In addition, commercially available 5' sequencing kits do not effectively sequence commercially available landing pad constructs because the primer is too far away from the molecular barcode (see, e.g., FIG. 7). The sequencing depth of the polynucleotide encoding the variable domain of the heavy chain of the antibody protein product is reduced by about 50% compared to the sequencing depth of the polynucleotide sequence encoding the variable domain of the light chain of the antibody protein product. A more significant reduction in sequencing depth would be expected if the reverse transcriptase were to attempt to transcribe a 4 kB insert with an IRES or promoter sequence in the middle of the polynucleotide molecule.

[0053] The fluidic device may allow single cells in a chamber or isolation pen to be grown and expanded, allowing clonal selection of cells that produce the antibody protein product to be sequenced. Clonal selection allows for selection of clones for large-scale protein production and purification during drug discovery and biologic drug manufacturing (e.g., antibody production). The disclosed method also allows for continuous analysis of cells while they are expanded, and the assay may be repeated with the same expanded cells.

[0054] A colony of biological cells is a "clone" if all living cells capable of reproduction in the colony are daughter cells derived from a single progenitor cell. In certain embodiments, all of the daughter cells in a clonal colony are derived from a single progenitor cell by no more than 10 divisions. In other embodiments, all of the daughter cells in a clonal colony are derived from a single progenitor cell by no more than 14 divisions. In other embodiments, all of the daughter cells in a clonal colony are derived from a single progenitor cell by no more than 17 divisions. In other embodiments, all of the daughter cells in a clonal colony are derived from a single progenitor cell by no more than 20 divisions. The term "clonal cells" refers to cells of the same clonal colony.

[0055] As used herein, a "colony" of biological cells refers to two or more cells (e.g., about 2 to about 20, about 4 to about 40, about 6 to about 60, about 8 to about 80, about 10 to about 100, about 20, about 200, about 40, about 400, about 60, about 600, about 80, about 800, about 100, about 1000, or more than 1000 cells).

[0056] As used herein, the term "maintaining cells" refers to providing an environment that includes both liquid and gaseous components, and optionally surfaces that provide the conditions necessary for cells to be maintained to survive and / or expand.

[0057] As used herein, the term "expand," when referring to cells, refers to an increase in cell number.

[0058] Antibody Protein Products The present disclosure provides polynucleotide molecules encoding antibody protein products. Antibody protein products include antibodies, bispecific T cell engager (BiTE®) molecules, antibody fragments, antibody fusion peptides or antigen-binding fragments thereof, or peptides. In related embodiments, the antibody is a polyclonal antibody or a monoclonal antibody. As used herein, the term "antibody protein product" refers to antibodies and any one of several antibody surrogates that are based on the structure of antibodies in various instances but are not found in nature.

[0059] "Antibody" is a subgenus of antibody protein products. It refers to immunoglobulins of any isotype with specific binding to a target antigen, including, for example, monoclonal antibodies. Antibodies can be of any suitable host species (e.g., chimeric, humanized, fully human, fully mouse, fully rabbit, or fully llama). Antibodies generally comprise two full-length heavy chains and two full-length light chains. For example, human antibodies can be of any isotype, including IgG (including IgG1, IgG2, IgG3, and IgG4 subtypes), IgA (including IgA1 and IgA2 subtypes), IgM, and IgE. In some embodiments, the antibody protein product has a molecular weight in the range of at least about 12 kDa to 10 MDa, such as at least about 12 kDa to 5 MDa, 12 kDa to 1 MDa, 12 kDa to 750 kDa, at least about 12 kDa to 250 kDa, or at least about 12 kDa to 150 kDa. In certain embodiments, antibody protein products have a valency (n) ranging from monomers (n=1) to dimers (n=2), trimers (n=3), and tetramers (n=4), unless there are higher valencies. Antibody protein products are in some embodiments based on the complete antibody structure and / or mimic antibody fragments that retain complete antigen-binding capacity (e.g., scFv, Fab, and VHH / VH (discussed below)). The smallest antigen-binding antibody fragment that retains a complete antigen-binding site is the Fv fragment, consisting entirely of the variable (V) region. Soluble flexible amino acid peptide linkers are used to link the V region to scFv (single chain fragment variable) fragments to stabilize the molecule, or constant (C) domains are added to the V region to generate Fab fragments (fragment, antigen-binding). Both scFv and Fab fragments can be easily produced in host cells (e.g., prokaryotic host cells). Other antibody protein products include disulfide bond stabilized scFv (ds-scFv), single chain Fab (scFab) and dimeric and multimeric antibody formats, such as diabodies, triabodies and tetrabodies or minibodies (miniAbs), including different formats comprising scFv linked to oligomerization domains.The smallest fragments are the VHH / VH of camelized heavy chain Abs and single domain Abs (sdAbs), including the UniDab® construct-containing molecules and the UniAb® constructs (TeneoBio). The building blocks most frequently used to create new antibody formats are single-chain variable (V) domain antibody fragments (scFv), which contain V domains (VH and VL domains) from heavy and light chains linked by a peptide linker of about 15 amino acid residues. Peptibodies or peptide-Fc fusions are yet another antibody protein product. The structure of a peptibody comprises a bioactive peptide grafted onto the Fc domain. Peptibodies have been well described in the art. See, for example, Shimamoto et al., mAbs 4(5):586-591 (2012). Other antibody protein products include single chain antibodies (SCAs), diabodies, triabodies, tetrabodies, bispecific antibodies or triabodies. Bispecific antibodies can be classified into five major classes: BsIgG, adduct IgG, BsAb fragments, bispecific fusion proteins and BsAb conjugates. See, for example, Spiess et al., Molecular Immunology 67(2)Part A:97-106(2015). In an exemplary embodiment, the antibody protein product comprises or consists of a Bispecific T Cell Engager (BiTE®) molecule, which is an artificial bispecific monoclonal antibody. A BiTE® molecule is a fusion protein that contains two scFvs of different antibodies. One binds CD3 and the other binds the target antigen. BiTE® molecules are known in the art. See, e.g., Huehls et al., Immuno Cell Biol 93(3):290-296(2015); Rossi et al., MAbs 6(2):381-91(2014); Ross et al., PLoS One 12(8):e0183390.

[0060] Polynucleotide molecules The term "recombinant" indicates that a material (e.g., a nucleic acid or polypeptide) has been artificially or synthetically (i.e., non-naturally) altered by human intervention. This alteration may be performed on the material in its natural environment or condition, or on material removed from this environment or condition. For example, a "recombinant nucleic acid" is one made by recombining nucleic acids, e.g., during cloning, DNA shuffling, or other known molecular biology procedures. Examples of such molecular biology procedures are found in Maniatis et al., Molecular Cloning. A Laboratory Manual. Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1982). A "recombinant DNA molecule" is composed of segments of DNA joined together by such molecular biology techniques. The term "recombinant protein" or "recombinant polypeptide" as used herein refers to a protein molecule expressed using a recombinant DNA molecule. A "recombinant host cell" is a cell that contains and / or expresses a recombinant nucleic acid.

[0061] The term "polynucleotide" or "nucleic acid" includes both single-stranded and double-stranded nucleotide polymers containing two or more nucleotide residues. The nucleotide residues that make up a polynucleotide can be ribonucleotides or deoxyribonucleotides or modified forms of either type of nucleotide. The modifications include base modifications such as bromouridine and inosine derivatives, ribose modifications such as 2',3'-dideoxyribose, and modifications of internucleotide bonds such as phosphorothioates, phosphorodithioates, phosphoroselenoates, phosphorodiselenoates, phosphoroanilothioates, phosphoroaniladates, and phosphoroamidates.

[0062] The term "oligonucleotide" refers to a polynucleotide comprising 200 or fewer nucleotide residues. In some embodiments, the oligonucleotide is 10-60 bases in length. In other embodiments, the oligonucleotide is 12, 13, 14, 15, 16, 17, 18, 19, or 20-40 nucleotides in length. The oligonucleotide can be single-stranded or double-stranded, e.g., for use in constructing mutant genes. The oligonucleotide can be a sense oligonucleotide or an antisense oligonucleotide. The oligonucleotide can include a label, including an isotopic label to facilitate quantification or detection (e.g., 125 I, 14 C. 13 C. 35 S, 3 H, 2 H, 13 N, 15 N, 18 O. 17 O, etc.), fluorescent labels, haptens or antigenic labels for detection assays. Oligonucleotides can be used, for example, as PCR primers, reverse transcription primers, cloning primers or hybridization probes.

[0063] "Polynucleotide sequence", or "nucleotide sequence", or "nucleic acid sequence", as used interchangeably herein, is a primary sequence of nucleotide residues in a polynucleotide (e.g., oligonucleotides, DNA and RNA, nucleic acids) or a character string representing the primary sequence of nucleotide residues, depending on the context. Either a given nucleic acid or a complementary polynucleotide sequence can be determined from any specified polynucleotide sequence. Included are DNAs or RNAs of genomic or synthetic origin that can be single-stranded or double-stranded and represent the sense or antisense strand. Unless otherwise specified, the left-hand end of any single-stranded polynucleotide sequence discussed herein is the 5' end and the left-hand direction of a double-stranded polynucleotide sequence is referred to as the 5' direction. The direction in which the nascent RNA transcript is added 5' to 3' is referred to as the transcription direction, and the regions of sequences present on the DNA strand having the same sequence as the RNA transcript and which are 5' to the 5' end of the RNA transcript are referred to as "upstream" sequences, and the regions of sequences present on the DNA strand having the same sequence as the RNA transcript and which are 3' to the 3' end of the RNA transcript are referred to as "downstream" sequences.

[0064] "Orientation" refers to the order of nucleotides in a given DNA sequence. For example, the orientation of a DNA sequence in reverse relative to another DNA sequence is one in which the 5' to 3' order of the sequence relative to the other sequence is reversed when compared to a reference point of the DNA from which the sequence was derived. Such reference points may include the direction of transcription of another specified DNA sequence in the source DNA and / or the origin of replication of a replicable vector containing the sequence. The 5' to 3' DNA strand is called the "sense", "plus" or "coding" strand with respect to a given gene. The complementary 3' to 5' strand to the "plus" strand is described as "antisense", "minus" or "non-coding".

[0065] As used herein, an "isolated nucleic acid molecule" or "isolated nucleic acid sequence" is a nucleic acid molecule that is (1) identified and separated from at least one contaminant nucleic acid molecule with which it is ordinarily associated in the natural source of the nucleic acid, or (2) that has been cloned, amplified, tagged, or otherwise distinguished from background nucleic acids such that the sequence of the nucleic acid of interest may be determined. An isolated nucleic acid molecule is different from the form or setting in which it is found in nature. However, an isolated nucleic acid molecule includes a nucleic acid molecule that is contained in a cell that normally expresses a polypeptide (e.g., an oligopeptide or antibody), e.g., a nucleic acid molecule that is in a chromosomal location different from that of natural cells.

[0066] As used herein, the terms "nucleic acid molecule encoding," "DNA sequence encoding," and "DNA encoding" refer to the order or sequence of deoxyribonucleotides along a deoxyribonucleic acid strand. The order of these deoxyribonucleotides determines the order of ribonucleotides along an mRNA strand, which in turn determines the order of amino acids along a polypeptide (protein) strand. Thus, the DNA sequence codes for the RNA sequence and the amino acid sequence.

[0067] The term "gene" is used broadly to refer to any nucleic acid associated with a biological function. A gene typically includes a coding sequence and / or regulatory sequences required for expression of such a coding sequence. The term "gene" applies to a particular genomic or recombinant sequence and the cDNA or mRNA encoded by this sequence. A "fusion gene" includes a coding region that encodes a polypeptide having portions from different proteins that are not found together in nature or that are not found together in nature in the same sequence as present in the encoded fusion protein (i.e., a chimeric protein). Genes also include non-expressed nucleic acid segments that, for example, form recognition sequences for other proteins. Non-expressed regulatory sequences, such as transcriptional regulatory elements, to which regulatory proteins, such as transcription factors, bind, resulting in transcription of adjacent or nearby sequences.

[0068] "Expression of a gene" or "expression of a nucleic acid" means, as indicated by the context, transcription of DNA into RNA (optionally including modification of the RNA, such as splicing), translation of RNA into a polypeptide (which may include subsequent post-translational modification of that polypeptide), or both transcription and translation.

[0069] As used herein, the term "coding region" or "coding sequence," when used in reference to a structural gene, refers to a nucleotide sequence that encodes the amino acids found in a nascent polypeptide as a result of translation of an mRNA molecule. In eukaryotes, the coding region is bounded at the 5' end by the nucleotide triplet "ATG," which encodes an initiator methionine, and at the 3' end by one of three triplets that specify a stop codon (i.e., TAA, TAg, TGA).

[0070] The term "control sequence" or "control signal" refers to a polynucleotide sequence that can affect the expression and processing of a coding sequence to which it is ligated in a particular host cell. The nature of such control sequences may depend on the host organism. In particular embodiments, control sequences for prokaryotes may include a promoter, a ribosomal binding site, and a transcription termination sequence. Control sequences for eukaryotes may include a promoter containing one or more recognition sites for transcription factors, a transcription enhancer sequence or element, a polyadenylation site, and a transcription termination sequence. "Control sequences" may include leader sequences and / or fusion partner sequences. Promoters and enhancers consist of short sequences of DNA that interact specifically with cellular proteins involved in transcription (Maniatis, et al., Science 236:1237 (1987)). Promoters and control elements have been isolated from a variety of eukaryotic sources, such as genes in yeast cells, insect cells, mammalian cells, and viruses (analogous control elements (i.e., promoters) are also found in prokaryotes). The selection of a particular promoter and enhancer depends on which cell type is to be used to express the protein of interest. Some eukaryotic promoters and enhancers have a broad host range, while others function in a limited subset of cell types (see Voss, et al., Trends Biochem. Sci., 11:287 (1986) and Maniatis, et al., Science 236:1237 (1987); Magnusson et al., Sustained, high transgene expression in liver with plasmid vectors using optimized promoter-enhancer combinations, Journal of Gene Medicine 13(7-8):382-391 (2011); Xu et al., Optimization of transcriptional regulatory elements for constructing plasmid vectors, Gene. 272(1-2):149-156 (2001)).Enhancers are generally cis-acting and are naturally located on chromosomes up to a million base pairs away from expressed genes. In some cases, the orientation of an enhancer can be inverted without affecting its function.

[0071] The term "vector" refers to any molecule or entity (eg, nucleic acid, plasmid, bacteriophage, or virus) used to introduce protein-coding information into a host cell.

[0072] The term "expression vector" or "expression construct" as used herein refers to a recombinant DNA molecule that contains a desired coding sequence and appropriate nucleic acid control sequences necessary for the expression of the operably linked coding sequence in a particular host cell. Expression vectors may include, but are not limited to, sequences that affect or control transcription, translation, and, if introns are present, RNA splicing of the coding region operably linked thereto. Nucleic acid sequences necessary for expression in prokaryotes include a promoter, optionally an operator sequence, a ribosome binding site, and optionally other sequences. Eukaryotic cells are known to utilize promoters, enhancers, and termination and polyadenylation signals. A secretory signal peptide sequence may also be optionally encoded by the expression vector and operably linked to the coding sequence of interest, so that the expressed polypeptide may be secreted by the recombinant host cell so that the polypeptide of interest can be more easily isolated from the cell, if desired. Such modifications are known in the art. (See, e.g., Goodey, Andrew R.; et al., Peptide and DNA sequences, U.S. Pat. No. 5,302,697; Weiner et al., Compositions and methods for protein secretion, U.S. Pat. Nos. 6,022,952 and 6,335,178; Uemura et al., Protein expression vectors and utilization thereof, U.S. Pat. No. 7,029,909; Ruben et al., 27 human secreted proteins, U.S. Patent Application Publication No. 2003 / 0104400A1).

[0073] An expression vector contains one or more expression cassettes. An "expression cassette" contains at least a promoter, a foreign gene of interest ("GOI") to be expressed, and a polyadenylation site and / or other suitable terminator sequence. A promoter typically contains a suitable TATA box or GC-rich region 5' to the promoter, but need not be immediately adjacent to the transcription start site.

[0074] The terms "in operable combination," "in operable order," and "operably linked," as used interchangeably herein, refer to the linking of two or more nucleic acid sequences such that a nucleic acid molecule capable of inducing transcription of a given gene and / or synthesis of a desired protein molecule is generated. The term also refers to the linking of amino acid sequences in such a way that a functional protein is produced. For example, a control sequence in a vector "operably linked" to a protein coding sequence is ligated to the protein coding sequence such that expression of the protein coding sequence is achieved under conditions compatible with the transcriptional activity of the control sequences. For example, a promoter and / or enhancer sequence (including any combination of cis-acting transcriptional control elements) is operably linked to a coding sequence if it stimulates or modulates the transcription of the coding sequence in an appropriate host cell or other expression system. A promoter control sequence operably linked to a transcribed gene sequence is physically contiguous with the transcribed sequence, whereas a cis-acting control element sequence operably linked to a promoter and / or transcribed gene sequence can be operably linked to it even if the control elements are non-contiguous with the promoter sequence and / or transcribed gene sequence. In some useful embodiments of the invention, the control element may be located 5' to the GAPDH promoter-driven expression cassette, while in other useful embodiments, the enhancer may be positioned 3' to the GAPDH promoter-driven expression cassette.

[0075] The term "host cell" refers to a cell that has been transformed or can be transformed with a nucleic acid, thereby expressing a gene of interest. The term includes the progeny of a parent cell, whether or not the morphology or genetic make-up of the progeny is identical to the original parent cell, so long as the gene of interest is present. Any of a large number of available and known host cells may be used in the practice of the invention, although CHO cell lines are preferred. The selection of a particular host depends on many factors recognized in the art, including, for example, compatibility with the selected expression vector, toxicity of the peptide encoded by the DNA molecule, rate of transformation, ease of peptide recovery, expression characteristics, biological safety, and cost. These factors must be balanced, with the understanding that not all hosts may be equally effective in expressing a particular DNA sequence. Within these general guidelines, useful microbial host cells for culture include bacteria (e.g., Escherichia coli sp.), yeast (e.g., Saccharomyces sp.) and other fungal cells, insect cells, plant cells, mammalian (e.g., human) host cells, such as CHO cells and HEK-293 cells. Modifications can also be made at the DNA level. The DNA sequence encoding the peptide can be altered to codons more compatible with the selected host cell. In the case of E. coli, optimized codons are known in the art. Codons can be substituted to remove restriction sites or to include silent restriction sites, which can facilitate processing of the DNA in the selected host cell. The transformed host is then cultured and purified. The host cells can be cultured under conventional fermentation conditions such that the desired compound is expressed. Such fermentation conditions are known in the art.

[0076] The term "transfection" refers to the uptake of foreign or exogenous DNA by a cell; a cell has been "transfected" when exogenous DNA has been introduced inside the cell membrane. Many transfection techniques are known in the art and are disclosed herein. See, e.g., Graham et al., 1973, Virology 52:456; Sambrook et al., 2001, Molecular Cloning: A Laboratory Manual, supra; Davis et al., 1986, Basic Methods in Molecular Biology, Elsevier; Chu et al., 1981, Gene 13:197. Such techniques can be used to introduce one or more exogenous DNA molecules into a suitable host cell.

[0077] The term "transformation" refers to a change in the genetic characteristics of a cell; a cell has been transformed if it has been modified to contain new DNA or RNA. For example, a cell has been transformed if it has been genetically altered from its native state by introducing new genetic material by transfection, transduction or other techniques. Following transfection or transduction, the transforming DNA may recombine with the DNA of the cell by physically integrating into the cell's chromosomes, or may be maintained transiently without replication as an episomal element, or may be independently replicated as a plasmid. A cell is considered to be "stably transformed" if the transforming DNA is replicated with cell division.

[0078] A "domain" or "region" of a protein (used interchangeably herein) is any portion of the entire protein, up to but typically comprising less than the complete protein. A domain may, but need not, fold independently of the rest of the protein chain and / or be associated with a specific biological, biochemical or structural function or location (e.g., a ligand-binding domain or a cytoplasmic, transmembrane or extracellular domain).

[0079] Selection Marker Element A selectable marker gene encodes a polypeptide essential for the survival and growth of transfected cells grown in selective culture medium. Typical selectable marker genes (a) encode a protein that confers resistance to antibiotics or other toxins (e.g., ampicillin, tetracycline, or kanomycin for prokaryotic host cells, and neomycin, hygromycin, or methotrexate for mammalian cells), (b) encode a protein that complements an auxotrophic deficiency of the cell, or (c) encode a protein that supplies a vital nutrient that is unavailable from a complex medium (e.g., a gene encoding D-alanine racemase for the cultivation of Bacilli).

[0080] All of the above elements, as well as other elements useful in the present invention, are known to those of skill in the art and are described, for example, in Sambrook et al. (Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY

[1989] ) and Berger et al. eds. (Guide to Molecular Cloning Techniques, Academic Press, Inc., San Diego, Calif.

[1987] ).

[0081] Construction of cloning vectors The cloning vectors most useful for amplifying the gene cassettes useful for preparing the recombinant expression constructs of the present invention are those compatible with prokaryotic cell hosts, however, eukaryotic cell hosts and vectors compatible with these cells are within the scope of the present invention.

[0082] In certain cases, some of the various elements contained in the cloning vector may already be present in commercially available cloning or amplification vectors (e.g., pUC18, pUC19, pBR322, pGEM vectors (Promega Corp, Madison, Wis.), pBluescript™ vectors (e.g., pBIISK+ / - (Stratagene Corp., La Jolla, Calif.) and the like), all of which are suitable for prokaryotic cell hosts. In this case, it is only necessary to insert the gene of interest into the vector.

[0083] However, if one or more of the elements used are not previously present in the cloning or amplification vector, they may be obtained separately and ligated into the vector. The methods used to obtain each element and ligate them are known to those of skill in the art and correspond to the methods described above for obtaining a gene of interest (i.e., DNA synthesis, library screening, and the like).

[0084] Vectors used for cloning or amplifying the nucleotide sequence of a gene of interest and / or transfecting a mammalian host cell are constructed using methods known in the art, including, for example, standard techniques of restriction endonuclease digestion, ligation, agarose and acrylamide gel purification of DNA and / or RNA, column chromatographic purification of DNA and / or RNA, phenol / chloroform extraction of DNA, DNA sequencing, polymerase chain reaction amplification, and the like, as described in Sambrook et al., supra.

[0085] The final vector used to practice the invention is typically constructed from a starting cloning or amplification vector, such as a commercially available vector. This vector may or may not contain some of the elements contained in the finished vector. If none of the desired elements are present in the starting vector, each element may be individually ligated to the vector by cutting the vector with an appropriate restriction endonuclease so that the ends of the element to be ligated and the ends of the vector are compatible for ligation. In some cases, it may be necessary to "blunt" the ends to be ligated together to achieve satisfactory ligation. Blunting the ends is accomplished by first filling in the "sticky ends" using Klenow DNA polymerase or T4 DNA polymerase in the presence of all four nucleotides. This procedure is known in the art and described, for example, in Sambrook et al. (supra).

[0086] Alternatively, two or more of the elements to be inserted into the vector may first be ligated to each other (if they are positioned adjacent to each other) and then ligated into the vector.

[0087] Another method of constructing vectors is to carry out all ligations of the various elements simultaneously in one reaction mixture, in which case many non-functional or non-functional vectors will be generated due to improper ligation or insertion of elements, but functional vectors can be identified and selected by restriction endonuclease digestion.

[0088] After the vector is constructed, it can be transfected into a prokaryotic host cell for amplification. Cells typically used for amplification are E. coli DH5-alpha (Gibco / BRL, Grand Island, NY) and other E. coli strains with characteristics similar to DH5-alpha.

[0089] Where mammalian host cells are used, cell lines such as Chinese hamster ovary (CHO cells; Urlab et al., Proc. Natl. Acad. Sci USA, 77:4216

[1980] ) and the human embryonic kidney cell line 293 (Graham et al., J. Gen. Virol., 36:59

[1977] ), as well as other lines, are suitable.

[0090] Transfection of the vector into the selected host cell line for amplification is accomplished using methods such as calcium phosphate, electroporation, microinjection, lipofection, or DEAE-dextran. The method selected will depend in part on the type of host cell to be transfected. These and other suitable methods are known to those of skill in the art and are described in Sambrook et al., supra.

[0091] After culturing the cells long enough to allow sufficient amplification of the vector (usually overnight for E. coli cells), the vector (at this stage often referred to as a plasmid) is isolated and purified from the cells. Typically, the cells are lysed and the plasmid is extracted from the other cellular contents. Suitable methods for plasmid purification include, among others, the alkaline lysis miniprep method (Sambrook et al., supra).

[0092] Recombinant Production of Antibodies and Other Polypeptides The relevant amino acid sequence from the immunoglobulin or polypeptide of interest can be determined by direct protein sequencing, and suitable nucleotide coding sequences can be designed according to a universal codon table. Alternatively, the genome or cDNA encoding the monoclonal antibody can be isolated and sequenced from cells producing such antibodies using conventional procedures (e.g., by using oligonucleotide probes capable of specifically binding to genes encoding the heavy and light chains of the monoclonal antibody). The relevant DNA sequence can be determined by direct or indirect sequencing.

[0093] Cloning of DNA is carried out using standard techniques (e.g., Sambrook et al. (1989) Molecular Cloning: A Laboratory Guide, Vols 1-3, Cold Spring Harbor Press, which is incorporated herein by reference). For example, a cDNA library may be constructed by reverse transcription of polyA+ mRNA (preferably membrane-bound mRNA) and the library screened using a probe specific for human immunoglobulin polypeptide gene sequences. However, in one embodiment, polymerase chain reaction (PCR) is used to amplify the cDNA (or a portion of a full-length cDNA) encoding the immunoglobulin gene segment of interest (e.g., a light or heavy chain variable segment). The amplified sequence may be readily cloned into any suitable vector (e.g., an expression vector, a minigene vector, or a phage display vector). It will be understood that the particular cloning method used is not critical, so long as it is possible to determine the sequence of a portion of the immunoglobulin polypeptide of interest.

[0094] One source of antibody nucleic acid is a hybridoma made by obtaining B cells from an animal immunized with an antigen of interest and fusing them with immortal cells. Alternatively, nucleic acid can be isolated from B cells (or whole spleen) of an immunized animal. Yet another source of nucleic acid encoding an antibody is a library of such nucleic acid, for example, generated by phage display technology. Polynucleotides encoding peptides of interest (e.g., variable region peptides with desired binding properties) can be identified by standard techniques such as panning.

[0095] Although the sequence encoding the entire variable region of an immunoglobulin polypeptide may be determined, sometimes it is sufficient to sequence only a portion of the variable region (e.g., the CDR-encoding portion). Sequencing is performed using standard techniques (e.g., Sambrook et al. (1989) Molecular Cloning: A Laboratory Guide, Vols 1-3, Cold Spring Harbor Press and Sanger, F. et al. (1977) Proc. Natl. Acad. Sci. USA 74:5463-5467, which are incorporated herein by reference). By comparing the sequence of the cloned nucleic acid with the published sequences of human immunoglobulin genes and cDNAs, one skilled in the art can easily determine the sequence of the heavy and light chain variable regions, including (i) the use of germline segments of hybridoma immunoglobulin polypeptides (including heavy chain isotypes), and (ii) the sequences resulting from the process of N-region addition and somatic mutation, depending on the sequenced region. One source of immunoglobulin gene sequence information is the National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Md.

[0096] The isolated DNA can be operably linked to a control sequence or placed into an expression vector, which is then transfected into a host cell that does not otherwise produce immunoglobulin protein to direct the synthesis of monoclonal antibodies in the recombinant host cell. Recombinant production of antibodies is known in the art.

[0097] A nucleic acid is operably linked when it is placed into a functional relationship with another nucleic acid sequence. For example, DNA of a presequence or secretory leader is operably linked to DNA of a polypeptide if it is expressed as a preprotein that participates in the secretion of the polypeptide, or a promoter or enhancer is operably linked to a coding sequence if it affects the transcription of the sequence, or a ribosome binding site is operably linked to a coding sequence if it is positioned such that it promotes translation. Generally, operably linked means that the DNA sequences being linked are contiguous, and in the case of a secretory leader, contiguous and in reading frame. Enhancers, however, need not be contiguous. Linking is accomplished by ligation at convenient restriction sites. If such sites do not exist, synthetic oligonucleotide adaptors or linkers are used in accordance with conventional practice.

[0098] Many vectors are known in the art. Vector components may include one or more of a signal sequence, an origin of replication, one or more selectable marker genes (which may, for example, confer resistance to antibiotics or other drugs, complement auxotrophic deficiencies, or supply essential nutrients not available in the culture medium), control elements, a promoter, and a transcription termination sequence, all of which are known in the art.

[0099] Cells, cell lines, and cell cultures are often used interchangeably, and all such designations herein include progeny. Transformants and transformed cells include the primary subject cell and cultures derived from the primary subject cell regardless of the number of transformations. It is also understood that all progeny may not be precisely identical in DNA content due to deliberate or inadvertent mutations. Mutant progeny that have the same function or biological activity as screened for in the originally transformed cell are included.

[0100] Exemplary host cells include prokaryote, yeast or higher eukaryote cells. Prokaryotic host cells include eubacteria, such as gram-negative or gram-positive microorganisms, for example Enterobacteriaceae, such as Escherichia, e.g., E. coli, Enterobacter, Erwinia, Klebsiella, Proteus, Salmonella, e.g., Salmonella typhimurium, Serratia, e.g., Serratia marcescens and Shigella, and Bacillus, e.g., B. subtilis and B. licheniformis, Pseudomonas and Streptomyces. Eukaryotic microbes, such as filamentous fungi or yeast, are suitable cloning or expression hosts for recombinant polypeptides or antibodies. Saccharomyces cerevisiae, or common baker's yeast, is the most commonly used among lower eukaryotic host microorganisms.However, Pichia, e.g. P. pastoris, Schizosaccharomyces pombe, Kluyveromyces, Yarrowia, Candida, Trichoderma reesia, Neurospora crassa, Schwanniomyces, e.g. Schwanniomyces occidentalis, Many other genera, species and strains are commonly available and useful herein, such as A. occidentalis, and filamentous fungi such as Neurospora, Penicillium, Tolypocladium and Aspergillus hosts, such as A. nidulans and A. niger.

[0101] Host cells for expressing glycosylated antibodies can be derived from multicellular organisms. Examples of invertebrate cells include plant cells and insect cells. Numerous baculovirus strains and variants have been identified, as well as corresponding permissive insect host cells from hosts such as Spodoptera frugiperda (caterpillar), Aedes aegypti (mosquito), Aedes albopictus (mosquito), Drosophila melanogaster (fruit fly) and Bombyx mori. Various virus strains for transfection of such cells are publicly available (e.g., the L-1 variant of Autographa californica NPV and the Bm-5 strain of Bombyx mori NPV).

[0102] Vertebrate host cells are also suitable hosts, and the recombinant production of polypeptides, including antibodies, from such cells is routine. Examples of useful mammalian host cell lines are: Chinese hamster ovary (CHO) cells of any strain, including, but not limited to, CHO-K1 cells (ATCC CCL61), DXB-11, CHO-DG-44, CHO-S, CHO-AM1, CHO-DXB11, and Chinese hamster ovary cells / -DHFR (CHO, Urlaub et al., Proc. Natl. Acad. Sci. USA 77:4216 (1980)); monkey kidney CV1 line transformed by SV40 (COS-7, ATCC CRL 1651); human embryonic kidney line (293 cells or 293 cells subcloned for growth in suspension culture [Graham et al., J. Gen Virol. 36:59 (1977)]; baby hamster kidney cells (BHK, ATCC CCL 10); mouse Sertoli cells (TM4, Mather, Biol. Reprod. 23:243-251 (1980)); monkey kidney cells (CV1 ATCC CCL 70); African green monkey kidney cells (VERO-76, ATCC CRL-1587); human cervical carcinoma cells (HELA, ATCC CCL 2); dog kidney cells (MDCK, ATCC CCL 34); buffalo rat liver cells (BRL 3A, ATCC CRL 1442); human lung cells (W138, ATCC CCL 75); human hepatocellular carcinoma cells (Hep G2, HB 8065); mouse mammary tumor (MMT 060562, ATCC CCL51); TRI cells (Mather et al., Annals NY Acad. Sci. 383:44-68 (1982)); MRC5 cells or FS4 cells; or mammalian myeloma cells.

[0103] Host cells are transformed or transfected with the above-described nucleic acids or vectors for producing polypeptides (including antibodies) and cultured in conventional nutrient media modified as appropriate for inducing promoters, selecting transformants, or amplifying genes encoding the desired sequences. In addition, novel vectors and transfected cell lines having multiple copies of transcription units separated by selectable markers are particularly useful for expressing polypeptides such as antibodies.

[0104] The host cells used to produce the polypeptides useful in the present invention may be cultured in a variety of media. Commercially available media such as Ham's F10 (Sigma), Minimum Essential Medium ((MEM) (Sigma), RPMI-1640 (Sigma) and Dulbecco's Modified Eagle's Medium ((DMEM), Sigma) are suitable for culturing the host cells. In addition, the media described in Ham et al., Meth. Enz. 58:44 (1979), Barnes et al., J. Clin. Microbiol. 1999, 11:1311-1323 (1989) and W. et al., J. Clin. Microbiol. 1999, 11:1311-1323 (1999) are suitable for culturing the host cells. Any of the media described in U.S. Pat. Nos. 4,767,704; 4,657,866; 4,927,762; 4,560,655; or 5,122,469; WO 90103430; 87 / 00195; or U.S. Reissue Patent No. 30,985 may be used as the host cell-like culture medium. Any of these media may optionally contain hormones and / or other growth factors (e.g., insulin, transferrin or epidermal growth factor receptors). Growth factors), salts (e.g., sodium chloride, calcium, magnesium, and phosphate), buffers (e.g., HEPES), nucleotides (e.g., adenosine and thymidine), antibiotics (e.g., the drug Gentamicin™), trace elements (usually defined as inorganic compounds present at final concentrations in the micromolar range), and glucose or an equivalent energy source may be supplemented. Any other necessary nutritional supplements may also be included at appropriate concentrations that would be known to one of skill in the art. Culture conditions such as temperature, pH, etc. will be those previously used with the host cell selected for expression and will be apparent to one of skill in the art.

[0105] When host cells are cultured, the recombinant polypeptide can be produced intracellularly, in the periplasmic space, or directly secreted into the medium. If a polypeptide such as an antibody is produced intracellularly, as a first step, the particulate debris, either host cells or lysed fragments, is removed, for example, by centrifugation or ultrafiltration.

[0106] Antibodies or antibody fragments can be purified, for example, using hydroxylapatite chromatography, cation or anion exchange chromatography, or preferably affinity chromatography using the antigen of interest or protein A or protein G as the affinity ligand. Protein A can be used to purify proteins, including polypeptides based on human gamma 1, gamma 2 or gamma 4 heavy chains (Lindmark et al., J. Immunol. Meth. 62:1-13 (1983)). Protein G is recommended for all mouse isotypes and human gamma 3 (Guss et al., EMBO J. 5:15671575 (1986)). The matrix to which the affinity ligand is attached is most often agarose, although other matrices are available. Physically stable matrices such as controlled pore glass or poly(styrenedivinyl)benzene allow for faster flow rates and shorter processing times than can be achieved by using agarose. If the protein is C H If three domains are involved, Bakerbond ABX™ resin (JT Baker, Phillipsburg, NJ) is useful for purification. Depending on the antibody to be recovered, other techniques for protein purification are possible, such as ethanol precipitation, reversed-phase HPLC, isoelectric focusing, SDS-PAGE, and ammonium sulfate precipitation.

[0107] Antibody generation by phage display technology The development of techniques for constructing repertoires of recombinant human antibody genes and displaying the encoded antibody fragments on the surface of filamentous bacteriophage provides another means of generating human-derived antibodies. Phage display is described, for example, in Dower et al., WO 91 / 17271, McCafferty et al., WO 92 / 01047, and Caton and Koprowski, Proc. Natl. Acad. Sci. USA, 87:6450-6454 (1990), each of which is incorporated herein by reference in its entirety. Antibodies generated by phage technology are usually generated as antigen-binding fragments (e.g., Fv or Fab fragments) in bacteria and therefore lack effector functions. Effector functions can be introduced by one of two strategies: the fragments can be engineered into complete antibodies for expression in mammalian cells, or into bispecific antibody fragments with a second binding site capable of eliciting effector functions.

[0108] Typically, the Fd fragment of an antibody (V H -C H 1) and light chain (V L -C L ) can be cloned separately by PCR and randomly combined in a combinatorial phage display library, which can then be selected for binding to a specific antigen. The antibody fragments are expressed on the phage surface and selection of Fv or Fab (and thus phage containing DNA encoding the antibody fragment) by antigen binding is achieved by several rounds of antigen binding and re-amplification, a procedure called panning. Antigen-specific antibody fragments are enriched and eventually isolated.

[0109] Phage display technology can also be used in an approach for humanization of rodent monoclonal antibodies called "guided selection" (see Jespers, LS, et al., Bio / Technology 12, 899-903 (1994)). In this case, Fd fragments of mouse monoclonal antibodies can be displayed in combination with a human light chain library, and the resulting hybrid Fab library can then be selected with antigen. The mouse Fd fragments thus provide a template to guide the selection. The selected human light chains are then combined with a human Fd fragment library. Selection of the resulting library yields fully human Fabs.

[0110] Various procedures for obtaining human antibodies from phage display libraries have been described (see, e.g., Hoogenboom et al., J. Mol. Biol., 227:381 (1991); Marks et al., J. Mol. Biol, 222:581-597 (1991); U.S. Pat. Nos. 5,565,332 and 5,573,905; Clackson, T., and Wells, JA, TIBTECH 12, 173-184 (1994)). In particular, in vitro selection and evolution of antibodies derived from phage display libraries has become a powerful tool (see Burton, DR, and Barbas III, CF, Adv. Immunol. 57, 191-280 (1994); and Winter, G., et al., Annu. Rev. Immunol. 12, 433-455 (1994); U.S. Patent Application Publication No. 20020004215 and WO 92 / 01047; U.S. Patent Application Publication No. 20030190317 published October 9, 2003; U.S. Patent No. 6,054,287; U.S. Patent No. 5,877,293). Watkins, "Screening of Phage-Expressed Antibody Libraries by Capture Lift," Methods in Molecular Biology, Antibody Phage Display: Methods and Protocols 178:187-193 and U.S. Patent Application Publication No. 20030044772, published March 6, 2003, describe methods for screening phage-expressed antibody libraries or other binding molecules by capture lift, a method that involves immobilization of candidate binding molecules on a solid support.

[0111] Fluid Devices Fluidic devices refer to devices that use small amounts of fluid to perform various types of analysis. Fluidic devices include one or more separate circuits configured to hold fluid, with each circuit being composed of fluidically interconnected circuit elements. The circuit elements, including but not limited to regions, flow paths, channels, chambers, and / or pens, and at least one port, are configured to allow fluid to flow into and / or out of the fluidic device. These devices use chips, cells, channels, or isolated pens that contain the fluid for analysis.

[0112] Fluidic devices, such as microfluidic devices, generally have one or more channels with at least one dimension less than 1 mm. Common fluids used in fluidic devices include whole blood samples, bacterial cell suspensions, protein or antibody solutions, and various buffers. Fluidic devices can be used to obtain a variety of measurements, including molecular diffusion coefficients, fluid viscosity, pH, chemical binding coefficients, and enzyme reaction kinetics. Other applications for fluidic devices include capillary electrophoresis, isoelectric focusing, immunoassays, flow cytometry, sample injection of proteins for analysis via mass spectrometry, PCR amplification, DNA analysis, cell manipulation, cell separation, cell patterning, and chemical gradient formation. Many of these applications have utility for clinical diagnostics.

[0113] The advantages of using fluidic devices include that the volume of fluid in these channels is very small, usually a few nanoliters, and only small amounts of reagents and analytes are used. Furthermore, when analyzing protein-producing cells, a relatively small number of cells (or even a single cell) can produce sufficient amounts and concentrations of protein for analysis, and incubation times for colony expansion growth are reduced or avoided. The fabrication techniques used to build microfluidic devices are relatively inexpensive and highly amenable to both highly sophisticated multiplex devices and also mass production. Fluidic technology allows the manufacture of highly integrated devices to perform several different functions on the same substrate chip.

[0114] Any fluidic device may be used (or adapted for use) in the disclosed methods, including commercially available devices. The fluidic device may be configured for use in a fluidic optical system and may use light to manipulate materials in the fluidic device, such as cells. EXAMPLES

[0115] Example 1 An exemplary polynucleotide molecule expressing a monoclonal antibody of the present disclosure is shown in the schematic diagram of FIG. 1. The polynucleotide molecule is designed to express a monoclonal antibody, and the polynucleotide molecule includes a polynucleotide sequence encoding the variable domain of the light chain of the monoclonal antibody, a polynucleotide sequence encoding the constant domain of the light chain, and a polynucleotide sequence encoding the heavy chain of the monoclonal antibody. The polynucleotide molecule is designed to include an optical barcode and a unique molecular identifier barcode upstream of the polynucleotide sequence encoding the variable domain of the light chain. In the exemplary polynucleotide molecule, the optical barcode (if added) and the unique molecular identifier (UMI) barcode are positioned adjacent to each other and directly upstream of the polynucleotide sequence encoding the variable domain of the light chain of the monoclonal antibody. The exemplary polynucleotide molecule also includes two IRESs, one directly downstream of the specific sequence of the universal primer.

[0116] In the exemplary polynucleotide molecule, the optical barcode is part of a template switching oligonucleotide. This template switching oligonucleotide is conjugated to a dual primer bead that further contains an oligo dT sequence that binds to mRNA. This universal primer starts elongation near the 5' end. In this example, by placing the universal primer downstream of the polynucleotide sequence encoding the constant domain of the light chain of the antibody protein, reverse transcription is restricted to a short number of nucleotides (about 700 nucleotides) that constitute the UMI barcode and the optical barcode.

[0117] Example 2 In addition, some polynucleotide molecules disclosed herein are designed for use in on-chip sequencing by synthesis (see, eg, schematic diagrams in Figures 2A-B).

[0118] In one example, the polynucleotide molecule is designed to express a monoclonal antibody, and the polynucleotide molecule includes a polynucleotide sequence encoding the variable domain of the light chain of the monoclonal antibody, a polynucleotide sequence encoding the constant domain of the light chain, and a polynucleotide sequence encoding the heavy chain of the monoclonal antibody. The polynucleotide molecule is designed to include an oligonucleotide sequence specific for a sequencing primer, directly upstream of a unique molecular identifier barcode, which is directly upstream of a Kozak sequence and a polynucleotide sequence encoding the variable domain of the light chain. The sequencing primer can be annealed to the Kozak sequence. In the exemplary polynucleotide molecule, the sites for the sequencing primer and the unique molecular identifier barcode are located adjacent to each other and directly downstream of the polynucleotide sequence encoding the variable domain of the light chain of the monoclonal antibody. The exemplary polynucleotide molecule also includes two IRES, one directly downstream of the specific sequence of the universal RT primer. The universal RT primer is thus positioned to reverse transcribe the relatively short nucleic acid sequence that constitutes the UMI barcode, and is suitable for on-chip sequencing. See FIG. 2A.

[0119] When sequencing this polynucleotide molecule (see FIG. 2A), the universal oligo-dT primer captures the mRNA and allows reverse transcription including the UMI barcode. The placement of the sequencing primer allows sequencing of 15 or fewer bases, and the reagents diffuse to and from locations on the fluidic device (e.g., isolation pen), allowing on-chip sequencing on the fluidic device. Thus, the sequencing primer sites adjacent to the UMI barcode allow for in situ (e.g., on-chip on the fluidic device).

[0120] This template switching oligonucleotide is part of a dual primer bead that contains an oligo dT sequence that binds to mRNA. The universal primer starts elongation near the 5' end. In this example, reverse transcription is restricted to a short nucleotide number of 700 nucleotides by placing the universal primer downstream of the polynucleotide sequence encoding the constant domain of the light chain of the antibody protein.

[0121] In an exemplary polynucleotide molecule, the sequencing primer anneals downstream of the stop codon (e.g., TAG) and upstream of the UMI. A cloning overhang is positioned downstream of the sequencing primer annealing site between the nucleic acid sequence encoding the heavy chain polypeptide and the IRES (see FIG. 2B).

[0122] Example 3 The polynucleotide molecules described in Example 1 are designed for sequencing reactions in optofluidic devices such as isolation pens. Because cDNA from multiple pens is pooled for export and samples are fragmented during sequencing, current protocols limit sequencing to 500 nt at the 5' or 3' end.

[0123] Alternatively, cDNA is generated and then exported. Long-read sequencing is then performed to verify the whole molecule sequence. This method would not depend on barcoding. Long-read sequencing can be useful for pool cloning strategies.

[0124] Long-read sequencing allows direct sequencing of polynucleotide molecules in real time without the need for amplification. This direct sequencing approach allows for obtaining much longer reads compared to those obtained from short-read sequencing. Instead, the "synthetic" long-read sequencing approach uses modified sample processing and conventional short-read sequencing to computationally reconstruct long reads from shorter sequencing reads.

[0125] Example 4 Additional exemplary polynucleotide molecules described herein are designed with a landing pad construct that has a polyadenylation sequence in the cell host landing pad far downstream of the insert junction. As shown in the schematic diagram of FIG. 3A, a barcode is inserted downstream of the polynucleotide encoding the heavy chain polypeptide. At this position, a custom on-bead RT primer and a custom sequence primer oligonucleotide are inserted. An exemplary insert includes a sequencing primer that includes a stop codon (TAG) with a 100 nucleotide spacer, a unique molecular identifier (UMI) barcode, a 100 nucleotide spacer, and a RT optical barcode. Since this optical barcode is up to 500 bases from the UMI barcode, PCR is then used to perform non-fragmented amplicon PE sequencing (due to the gap between the optical barcode and the UMI barcode, tagging will not be available). In addition, the efficiency of the custom RT primer is expected to be lower compared to that of oligo dT.

[0126] Another exemplary polynucleotide molecule is shown in Figure 3B. In this polynucleotide molecule, a custom oligo is present in the sequence encoding the CH1 domain, upstream of the polynucleotide sequence encoding the light chain polypeptide. The advantage is that the barcode will be adjacent to the polynucleotide sequence encoding the heavy chain polypeptide. However, the disadvantage of this design is that the custom RT primer is present on-bead with the optical barcode. This arrangement is restrictive because the RT primer sequence must be selected from the natural sequence (which places strong constraints on the possible primer sequences) and may be less efficient. Since the optical barcode is up to 500 bases from the UMI barcode, PCR is then used to perform unfragmented amplicon PE sequencing.

Claims

1. 1. A polynucleotide molecule encoding an antibody protein product, comprising: i) a nucleotide sequence specific for addition of a molecular barcode by template-switching reverse transcriptase (RT); ii) a unique molecular identifier (UMI) barcode that is distinct from said molecular barcode; iii) a nucleotide sequence encoding a light chain polypeptide of said antibody protein product; iv) a nucleotide sequence encoding a heavy chain polypeptide of said antibody protein product; and v) a nucleotide sequence specific for a universal RT primer.

2. 2. The polynucleotide molecule of claim 1, wherein the molecular barcode is an optical barcode.

3. 3. The polynucleotide molecule of claim 2, wherein the optical barcode is (i) identifiable by determining its polynucleotide sequence and (ii) identifiable by annealing to a polynucleotide probe comprising one or more optical moieties.

4. 4. The polynucleotide molecule of claim 1, wherein the molecular barcode or the UMI barcode comprises a random 6-mer, 7-mer, 8-mer, 9-mer, 10-mer, 11-mer, 12-mer, 13-mer, 14-mer, or 15-mer.

5. a) a promoter sequence, b) at least two internal ribosome entry site (IRES) sequences; c) at least one internal ribosome entry site IRES sequence and at least one promoter; d) at least two promoters; or e) nucleotides encoding a selected gene product, such as puromycin-N-acetyltransferase; The polynucleotide molecule of any one of claims 1 to 3, further comprising:

6. the nucleotide sequence specific to the universal RT primer is a) located between the nucleotide sequence encoding the light chain polypeptide and the nucleotide sequence encoding the heavy chain polypeptide; or b) downstream of the nucleotide sequence encoding the light chain polypeptide of the antibody protein product; A polynucleotide molecule according to any one of claims 1 to 3.

7. 4. The polynucleotide molecule of any one of claims 1 to 3, wherein the nucleotide sequence specific for addition of a molecular barcode by a template-switching reverse transcriptase is adjacent to or directly upstream of the UMI barcode.

8. 4. The polynucleotide molecule of claim 1, wherein the UMI barcode is immediately upstream of the nucleotide sequence encoding the light chain polypeptide of the antibody protein product.

9. 4. The polynucleotide molecule of claim 1, wherein the nucleotide sequence specific for the RT universal primer is immediately downstream of the nucleotide sequence encoding the light chain polypeptide of the antibody protein product.

10. 4. The polynucleotide molecule of any one of claims 1 to 3, wherein an IRES is immediately downstream of the sequence specific for the RT universal primer or immediately downstream of the nucleotide sequence encoding the heavy chain polypeptide of the antibody protein product.

11. 4. The polynucleotide molecule of claim 1, wherein an IRES is immediately downstream of the nucleotide sequence encoding the heavy chain polypeptide of the antibody protein product.

12. The polynucleotide molecule of any one of claims 1 to 3, wherein the heavy chain polypeptide comprises a heavy chain variable region and the light chain polypeptide comprises a light chain variable region.

13. (a) the nucleotide sequence encoding the light chain polypeptide of the antibody protein product encodes a light chain variable region and a light chain constant region immediately downstream of the light chain variable region; and (b) the nucleotide sequence encoding the heavy chain polypeptide of the antibody protein product encodes a heavy chain variable region and a heavy chain constant region immediately downstream of the heavy chain variable region; A polynucleotide molecule according to any one of claims 1 to 3.

14. 4. The polynucleotide molecule of claim 1, wherein the nucleotide sequence encoding the light chain polypeptide of the antibody protein product is upstream of the nucleotide sequence encoding the heavy chain polypeptide of the antibody protein product.

15. 4. The polynucleotide molecule of any one of claims 1 to 3, wherein the nucleotide sequence encoding the light chain polypeptide of the antibody protein product is downstream of the nucleotide sequence encoding the heavy chain polypeptide of the antibody protein product.

16. 1. A polynucleotide molecule encoding an antibody protein product, comprising, in 5' to 3' direction: i) a promoter; ii) a nucleotide sequence specific for addition of a molecular barcode by a template-switching reverse transcriptase; iii) a unique molecular identifier (UMI) barcode; iv) a nucleotide sequence encoding a first polypeptide of the antibody protein product; v) a nucleotide sequence specific for a reverse transcriptase (RT) universal primer; vi) a first IRES; vii) a nucleotide sequence encoding a second polypeptide of said antibody protein product; viii) a nucleotide sequence encoding a constant domain of a heavy chain of said antibody protein product; ix) a second IRES; and x) a nucleotide sequence encoding a select gene product, such as puromycin-N-acetyltransferase.

17. 1. A polynucleotide molecule encoding an antibody protein product, wherein the polynucleotide sequence comprises: i) an annealing site for a sequencing primer; ii) a unique molecular identifier barcode; iii) a nucleotide sequence encoding a light chain polypeptide of said antibody protein product; iv) a nucleotide sequence encoding a heavy chain polypeptide of said antibody protein product; and v) a nucleotide sequence specific for a universal reverse transcriptase (RT) primer.

18. a) a promoter sequence, b) at least two internal ribosome entry site (IRES) sequences, and / or c) nucleotides encoding a selected gene product, such as puromycin-N-acetyltransferase; 18. The polynucleotide molecule of claim 17, further comprising: at least two internal ribosome entry site (IRES) sequences

19. the nucleotide sequence specific to the universal RT primer is a) located between the nucleotide sequence encoding the light chain polypeptide and the nucleotide sequence encoding the heavy chain polypeptide; b) downstream of the nucleotide sequence encoding the light chain polypeptide of the antibody protein product; or c) immediately downstream of the nucleotide sequence encoding the light chain polypeptide of the antibody protein product; A polynucleotide molecule according to claim 17 or 18. downstream of the nucleotide sequence encoding the light chain polypeptide of the antibody protein product

20. The method further comprises an IRES, wherein the IRES is a) directly downstream of the sequence specific for the RT universal primer, or b) the polynucleotide molecule of claim 17 or 18, immediately downstream of the nucleotide sequence encoding the heavy chain polypeptide of the antibody protein product.

21. A method for manufacturing a nucleotide sequence specific for an optical barcode, the method comprising: b) the nucleic acid specific for addition of a molecular barcode by a template-switching reverse transcriptase is configured for addition of the optical barcode by a template-switching oligonucleotide from a template optical barcode conjugated to a solid support; 20. The polynucleotide molecule of any one of claims 1, 2, 3, 16, 17, or 18.

22. 20. The polynucleotide molecule of any one of Claims 1, 2, 3, 16, 17, or 18, wherein the 5' end of the nucleotide sequence specific for attachment of a molecular barcode comprises the polynucleotide sequence CCC.

23. 20. The polynucleotide molecule of any one of Claims 1, 2, 3, 16, 17, or 18, wherein the nucleotide barcode is configured for reverse transcription by a template-switching reverse transcriptase primed by the RT universal primer.

24. 19. The polynucleotide molecule of any one of Claims 1, 2, 3, 16, 17, or 18, wherein the nucleotide sequence specific for a reverse transcriptase (RT) universal primer is configured to anneal to a free universal primer that is not disposed on the solid support, and the annealed universal primer is configured for reverse transcription of the nucleotide barcode by the reverse transcriptase.

25. 22. The polynucleotide molecule of claim 21, wherein the solid support is a bead, a resin, or agarose.

26. (a) the first polypeptide of the antibody protein product comprises a light chain polypeptide and the second polypeptide of the antibody protein product comprises a heavy chain polypeptide; or (b) the first polypeptide of the antibody protein product comprises a heavy chain polypeptide, and the second polypeptide of the antibody protein product comprises a light chain polypeptide; A polynucleotide molecule according to claim 17 or 18.

27. 27. The polynucleotide molecule of claim 26, wherein the light chain polypeptide comprises a light chain variable region and the heavy chain polypeptide comprises a heavy chain variable region.

28. the nucleotide sequence encoding the light chain polypeptide of the antibody protein product encodes a light chain variable region and a light chain constant region immediately downstream of the light chain variable region; and the nucleotide sequence encoding the heavy chain polypeptide of the antibody protein product encodes a heavy chain variable region and a heavy chain constant region immediately downstream of the heavy chain variable region; 27. The polynucleotide molecule of claim 26.

29. 19. The polynucleotide molecule of claim 17 or 18, encoding an antibody protein product, wherein the antibody protein product comprises or consists of a large peptide, an antibody, an antibody fragment, an antibody fusion peptide, or an antigen-binding fragment thereof.

30. 30. The polynucleotide molecule of claim 29, wherein the antibody is a polyclonal antibody or a monoclonal antibody.

31. 42. A method for screening clones that express an antibody protein product, the method comprising pooling clones that comprise the polynucleotide molecule of any one of claims 1 to 41, polymerizing the DNA of the pooled clones, and sequencing the DNA in a fluidic device.

32. 32. The method of claim 31 , wherein polymerizing the DNA of the pooled clones comprises reverse transcribing the polynucleotide molecules with a template-switching reverse transcriptase.

33. Annealing the polynucleotide molecule of any one of claims 1 to 3 to a template comprising an optical barcode; annealing a universal reverse transcriptase primer to the polynucleotide molecule; extending the annealed universal reverse transcriptase primer with a template-switching reverse transcriptase, thereby generating a cDNA of the polynucleotide molecule comprising the UMI barcode and the optical barcode; A method comprising:

34. 34. The method of claim 33, wherein the method is performed in a fluidic device.

35. 34. The method of claim 33, further comprising detecting the presence of a molecular barcode and / or a UMI barcode.

36. 35. The method of claim 34, wherein the fluidic device comprises or consists of a microfluidic chip or an isolation pen.