Barcoding for tracking cells for biomanufacturing

WO2026072736A3PCT designated stage Publication Date: 2026-05-07JOHNS HOPKINS UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
JOHNS HOPKINS UNIVERSITY
Filing Date
2025-09-24
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Current cell line development (CLD) strategies lack effective methods for high-throughput measurement of drug yields, leading to low throughput and increased costs due to the commitment to cell lines with low yields.

Method used

A high-throughput assay is developed using barcoded constructs to isolate and select cell clones expressing biologics of interest (BOI) by transfecting host cells with vectors containing barcodes and nucleic acid sequences, culturing, and analyzing performance attributes through sequencing data to identify clones with high BOI expression.

Benefits of technology

The assay enables efficient identification and selection of cell clones with high biologic expression, reducing costs and improving drug production efficiency by optimizing cell lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025047816_07052026_PF_FP_ABST
    Figure US2025047816_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Compositions and methods are directed to a diverse library of barcoded constructs expressing biologies that are known to be challenging to express. A high throughput screening assay is also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

BARCODING FOR TRACKING CELLS FOR BIOMANUFACTURINGThe present application claims the benefit of U.S. provisional application number 63 / 698,541 filed September 24, 2024, which application is incorporated by reference herein in its entirety.BACKGROUND

[0001] The most exciting and lucrative drugs of the past decade are biopharmaceutical therapies such as monoclonal antibodies, gene therapies, and cell therapies. These therapies are produced using engineered cells, and to ensure the consistent quality of these drugs, the Food and Drug Administration (FDA) requires that every dose administered throughout a drug’s 20- year lifetime be derived from a single cell’s descendants. Therefore, using an optimal clone from the beginning of drug development is essential to reduce costs associated with drug production. Cell line development (CLD) is the exercise of making an optimal cell line for each biopharmaceutical.

[0002] Current CLD strategies lack effective methods of measuring drug yields over time, and therefore only 5-10 clones are tested at scale. This low throughput leads drug manufacturers to commit to a cell line with low yields, which can double the costs of drug production.SUMMARY

[0003] Compositions and methods are directed to the development of a diverse library of barcoded constructs expressing biologies that are known to be challenging to express, such as infliximab and etanercept.

[0004] Accordingly, in certain aspects, an assay is provided including a high throughput assay for isolating cell clones expressing a biologic of interest (BOI), the assay comprising: a) transfecting host cells with a i) vector comprising a barcode; and ii) a vector comprising a nucleic acid sequence encoding a biologic of interest (BOI); b) selecting clones expressing high levels of the BOI; c) conducting a bioinformatic analysis based on the selected clones.

[0005] In a further aspect, an assay for analyzing or isolating cell clones expressing biologic of interest (BOI) is provided, the assay comprising: a) transfecting host cells with a i) a vector comprising a barcode; and ii) a vector comprising a nucleic acid sequence encoding a biologic of interest (BOI); b) culturing transfected cells in a polyclonal pool and using sequencing data including the i) vector barcode to determine performance attributes; and c) selecting clones expressing high levels of the BOI. Any one or more of a variety of performance attributes may be assessed ibn the assay, including for example specific productivity, growth rate, carrying capacity, viability and / or batch-to-batch variability.

[0006] In an additional aspect, an assay for analyzing or isolating cell clones expressing biologic of interest (BOI)is provided, the assay comprising: a) transfecting host cells with a i) a vector comprising a barcode; and ii) a vector comprising a nucleic acid sequence encoding a biologic of interest (BOI); b) culturing transfected cells in a polyclonal pool and using sequencing data including the i) vector barcode to determine performance attributes; andc) identifying measured differences in BOI expression or clone performance selecting clones expressing high levels of the BOI. Again, Any one or more of a variety of performance attributes may be assessed ibn the assay, including for example specific productivity, growth rate, carrying capacity, viability and / or batch-to-batch variability.

[0007] In that assay in certain aspects, measured differences in BOI expression or clone performance are identified by steps comprising by i) sorting cells into various groups based on markers; ii) performing sequencing that includes the barcode; and iii) calculating the BOI expression or clone performance attributes using the barcode sequencing data. Any one or more of a variety of performance attributes may be assessed ibn the assay, including for example specific productivity, growth rate, carrying capacity, viability and / or batch-to-batch variability.

[0008] In the above assays, in an embodiment, the i) vector and ii) vector may be a single vector comprising a) a barcode and b) a nucleic acid sequence encoding a biologic of interest (BOI).

[0009] In the above assays, in an embodiment, the i) vector and ii) vector may be two or three separate vectors.

[0010] In an additional aspect, an assay is provided comprising:(a) transfecting host cells with i) a vector comprising a barcode; ii) vector comprising a nucleic acid sequence encoding a biologic of interest (BOI); and iii) a vector comprising a gene editing complex wherein the gene editing complex induces expression of a reporter gene vectors from the library;(b) selecting stable clones expressing enhanced levels of the BOI, for example as compared to a baseline control or other assessment;(c) conducting a bioinformatic analysis, for example based on the selected clones.

[0011] In an embodiment, the i) vector, ii) vector and iii) vector may be a single vector comprising a) a barcode, b) a nucleic acid sequence encoding a biologic of interest (BOI), and c) a gene editing complex wherein the gene editing complex induces expression of a reporter gene vectors from the library.

[0012] In an embodiment, the i) vector, ii) vector and iii) vector may be two or three separate vectors.

[0013] In such embodiment comprising use of multiple separate vectors, multiple transfection steps may be utilized. For example, the assay may comprise: a) generating a vector library that comprises i) a vector comprising a barcode and ii) vector comprising a nucleic acid sequence encoding a biologic of interest (BOI), where the i) vector and ii) vector may be a single vector or separate vectors; b ) transfecting host cells with vectors from the library; c) selecting stable clones expressing high levels of the BOI, as compared to a baseline control; and d) transfecting the clones with a vector comprising a gene editing complex wherein the gene editing complex induces expression of a reporter gene.

[0014] Thus, in certain aspects, an assay is provided including a high throughput assay for isolating cell clones expressing biologic of interest (BOI), the assay comprising:(a) generating a vector library comprises i) a vector comprising a barcode and ii) vector comprising a nucleic acid sequence encoding a biologic of interest (BOI), where the i) vector and ii) vector may be a single vector or separate vectors;(b) transfecting host cells with vectors from the library;(c) selecting stable clones expressing high levels of the BOI, for example as compared to a baseline control or by selecting a high level expression group;(d) transfecting the clones with a vector comprising a gene editing complex wherein the gene editing complex induces expression of a reporter gene;(e) conducting a bioinformatic analysis, for example based on the selected clones.

[0015] Suitably, the library further comprises additional components such as regulatory elements and a reporter gene. Suitably, the assay further comprises selecting the clones expressing a reporter gene and the bioinformatic analysis is conducted based on the selected clones.

[0016] A wide variety of biologies or biologies of interest (BOI) may be contained in a vector and expressed in accordance with the present assays. For instance, antibodies including monoclonal antibodies including antibodies used for therapeutic applications can be preferred biologies of interest (BOI) as disclosed herein. Monoclonal antibodies, fc fusion proteins, bispecific antibodies, scFvs, Fabs, recombinant proteins all may be preferred biologies or biologies of interest (BOI) as disclosed herein and may be contained in a vector and expressed in accordance with the present assays.

[0017] Preferred biologies of interest (BOI) for use in the present assays and systems include for example the following and fragments thereof including functional fragments i.e. provide desired biological activity such as therapeutic efficacy): Trastuzumab, Eternacept, Infliximab, and / or Pembrolizumab. Additional Preferred biologies of interest (BOI) for use include Basiliximab, Obinutuzumab, Daratumumab, Nivolumab, Pembrolizumab, Ipilimumb, Dinutuximab, Trastuzumab, Pertuzumab, Ado-trastuzumab emtansine, Adalimumab,Certolizumab pegol, Golimumab, infliximab biosimilars (Remsina, Inflectra, Flexiabi), etanercept biosimilars (Erelzi, Benepali), adalimumab biosimilars (Amjevita), Bevacizumab,Ramucirumab, Aflibercept, Ranibizumab, Omalizumab, Vedolizumab, Natalizumab, Canakinumab, and / or Rilonacept.

[0018] The assay also may comprise further steps, including for example selecting the clones expressing the reporter gene and conducting the bioinformatic analysis of the selected clones.

[0019] In certain embodiments, the barcodes comprise at least 10 nucleotides up to 40 nucleotides. In certain embodiments, the barcodes comprise at least 15 nucleotides up to 35 nucleotides. In certain embodiments, the barcodes comprise at about 21 nucleotides. In certain embodiments, the nucleic acid sequences encoding the BOIs of interest are inserted into the cell genome. In certain embodiments, the genome of the clones expressing the reporter gene is sequenced. In certain embodiments, the sequencing is targeted to a nucleic acid sequence between a 5 ’restriction site and a 3’ restriction site and includes the bar code sequences.

[0020] In certain embodiments, the gene editing complex comprises guide RNAs (gRNAs) specifically targeting each barcode sequence. In certain embodiments, the gene editing complex is comprised in a vector distinct to the vectors in the vector library. In certain embodiments, the gene editing complex is comprised in the same vector in a vector library. In certain embodiments, the gene editing complex comprises one or more single guide RNAs (sgRNAs), a nuclease and a transcriptional activator. In certain embodiments, the gene editing complex comprises a clustered regularly interspaced short palindromic repeat (CRISPR)- mediated transcriptional activation (CRISPRa) system. In certain embodiments, the CRISPRa system comprises a nuclease. In certain embodiments, the nuclease is a catalytically inactive nuclease. In certain embodiments, the catalytically inactive nuclease is Cas9. In certain embodiments, the catalytically inactive Cas9 (dCas9) nuclease is fused to a transcriptionalactivator. In certain embodiments, the gRNAs are specific for each desired barcode. In certain embodiments, the sgRNAs guide the CRISPRa / dCas9 to the target site, thereby activating transcription and inducing expression of the reporter gene.

[0021] In certain embodiments, a vector is utilized that comprises one or more transcriptional regulators. In certain embodiments, the one or more transcriptional regulators regulate expression of the BOIs.

[0022] In another aspect, a vector library is provided that comprises vectors comprising i) a barcode and ii) a nucleic acid sequence encoding a biologic of interest (BOI).

[0023] In further aspects the vector library may further comprise regulatory elements, a gene editing complex, and a reporter gene or combinations thereof.

[0024] Identification of vectors and biologic of interest may be facilitated by using sample discriminating codes or sequences (also known as barcodes, e.g., synthetic nucleic acid barcodes) that may be embedded within or otherwise associated with the samples.

[0025] In certain embodiments, the barcodes comprise at least 10 nucleotides up to 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 200, 240 260, 2080 or 300 nucleotides. In certain embodiments, the barcodes comprise at least 10 nucleotides up to 150 nucleotides. In certain embodiments, the barcodes comprise at least 10 nucleotides up to 40 nucleotides. In certain embodiments, the barcodes comprise at least 15 nucleotides up to 35 nucleotides. In certain embodiments, the barcodes comprise at least or up to about 21 nucleotides. In certain embodiments, the barcodes comprise at least or up to about 67 nucleotides.

[0026] In certain embodiments, the gene editing complex comprises guide RNAs (gRNAs) specifically targeting each barcode sequence. In certain embodiments, a vector comprises one or more transcriptional regulators. In certain embodiments, the one or more transcriptional regulators regulate expression of the BOIs. In certain embodiments, the gene editing complex comprises one or more single guide RNAs (gRNAs), a nuclease and a transcriptional activator. In certain embodiments, the nuclease is a catalytically inactive nuclease. In certain embodiments, the catalytically inactive nuclease is Cas9. In certain embodiments, the catalytically inactive Cas9 (dCas9) nuclease is fused to a transcriptional activator.

[0027] In certain embodiments, the gRNAs are specific for each desired barcode.

[0028] In certain embodiments, the gene editing complex is comprised in a vector distinct to the other vector in the vector library. In certain embodiments, the gene editing complex is comprised in the same vector comprising a barcode, a nucleic acid sequence encoding a biologic of interest (BOI), regulatory elements; a reporter gene or combinations thereof.

[0029] In another aspect, a vector comprises a barcode, a nucleic acid sequence encoding a biologic of interest (BOI), regulatory elements; a clustered regularly interspaced short palindromic repeat (CRISPR)-mediated transcriptional activation (CRISPRa) system, and a reporter gene.

[0030] In another aspect, a vector comprises a barcode, a nucleic acid sequence encoding a biologic of interest (BOI), regulatory elements, a reporter gene or combinations thereof.

[0031] In another aspect, a host cell comprises one or more vectors as disclosed herein.

[0032] In certain aspects, the gene-editing complex comprises: Argonaute family of endonucleases, clustered regularly interspaced short palindromic repeat (CRISPR) nucleases, zinc- finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), meganucleases, endo- or exo-nucleases, or combinations thereof. In certain embodiments, the gene editing complex comprises guide nucleic acid sequence (gNAS), the gNAS being complementary to a target nucleic acid sequence. In certain embodiments, the gNAS comprises a ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). In some embodiments, the gNAS comprises one or more modified nucleic acid bases or chimeric regions. In certain embodiments, the gene editing complex and the at least one gNAS is encoded by the same vector or separate vectors. In certain embodiments, the guide NAS sequences are in single or multiplex configurations

[0033] In one embodiment, the viral vector is an adenovirus vector, an adeno- associated viral vector (AAV), or derivatives thereof. The adeno-associated viral vector comprises AAV serotypes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, DJ or DJ / 8. In one embodiment, the AAV vector is AAV serotype 9 (AAV9).

[0034] In various embodiments, the vector comprises a viral or a bacterial vector. In various embodiments, the viral vector is selected from the group comprising adenovirus, adeno- associated virus (AAV), herpes simplex virus, lentivirus, retrovirus, gammaretrovirus, alphavirus, flavivirus, rhabdovirus, measles virus, Newcastle disease virus, poxvirus, vaccinia virus, modified Ankara virus, vesicular stomatitis virus, picornavirus, tobacco mosaic virus, potato virus x, comovirus or cucumber mosaic virus. In various embodiments, the virus is anoncolytic virus. In various embodiments the virus is a chimeric virus, a synthetic virus, a mosaic virus or a pseudotyped virus.

[0035] Definitions

[0036] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the disclosure pertains. Although any methods and materials similar or equivalent to those described herein can be used in the practice for testing of the present disclosure, the preferred materials and methods are described herein. In describing and claiming the present disclosure, the following terminology will be used.

[0037] It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.

[0038] All genes, gene names, and gene products disclosed herein are intended to correspond to homologs from any species for which the compositions and methods disclosed herein are applicable. It is understood that when a gene or gene product from a particular species is disclosed, this disclosure is intended to be exemplary only, and is not to be interpreted as a limitation unless the context in which it appears clearly indicates. Thus, for example, for the genes or gene products disclosed herein, are intended to encompass homologous and / or orthologous genes and gene products from other species.

[0039] The articles “a” and “an” are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element. Thus, recitation of “a cell”, for example, includes a plurality of the cells of the same type. Furthermore, to the extent that the terms “including”,“includes”, “having”, “has”, “with”, or variants thereof are used in either the detailed descriptionand / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising.”

[0040] ‘About” as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of + / -20%, + / — 10%, + / — 5%, + / — 1%, or + / — 0.1% from the specified value, as such variations are appropriate to perform the disclosed methods. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude within 5-fold, and also within 2-fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value should be assumed.

[0041] The term “barcode,” as used herein, generally refers to a label, or identifier, that conveys or is capable of conveying information about an analyte. A barcode can be part of an analyte. A barcode can be independent of an analyte. A barcode can be a tag attached to an analyte (e.g., nucleic acid molecule) or a combination of the tag in addition to an endogenous characteristic of the analyte (e.g., size of the analyte or end sequence(s)). A barcode may be . Barcodes can have a variety of different formats. For example, barcodes can include: polynucleotide barcodes; random nucleic acid and / or amino acid sequences; and synthetic nucleic acid and / or amino acid sequences. A barcode can be attached to an analyte in a reversible or irreversible manner. A barcode can be added to, for example, a fragment of a deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) sample before, during, and / or after sequencing of the sample. Barcodes can allow for identification and / or quantification of individual sequencing-reads. Nucleic acids comprising a barcode sequence that are optionallyconfigured to interact with a nucleic acid to generate a barcoded nucleic acid may be referred to as a nucleic acid barcode molecule.

[0042] As used herein, the term “barcoded nucleic acid molecule” generally refers to a nucleic acid molecule that results from, for example, the processing of a nucleic acid barcode molecule with a nucleic acid sequence (e.g., nucleic acid sequence complementary to a nucleic acid guide RNA (gRNA) sequence encompassed by the nucleic acid barcode molecule). The nucleic acid sequence may be a targeted sequence (e.g., targeted by a gRNA sequence) or a nontargeted sequence. A barcoded nucleic acid molecule may serve as a template, such as a template polynucleotide, that can be further processed (e.g., amplified) and sequenced to obtain the target nucleic acid sequence.

[0043] A s used herein, the terms “Cas9,” “Cas9 molecule,” and the like, refers to a Cas9 polypeptide or a nucleic acid encoding a Cas9 polypeptide. A “Cas9 polypeptide” is a polypeptide that can form a complex with a guide RNA (gRNA) and bind to a nucleic acid target containing a target domain and, in certain embodiments, a PAM sequence. Cas9 molecules include those having a naturally occurring Cas9 polypeptide sequence and engineered, altered, or modified Cas9 polypeptides that differ, e.g., by at least one amino acid residue, from a reference sequence, e.g., the most similar naturally occurring Cas9 molecule. A Cas9 molecule may be a Cas9 polypeptide or a nucleic acid encoding a Cas9 polypeptide. A Cas9 molecule may be a nuclease (an enzyme that cleaves both strands of a double-stranded nucleic acid), a nickase (an enzyme that cleaves one strand of a double-stranded nucleic acid), or a catalytically inactive (or dead) Cas9 molecule. A Cas9 molecule having nuclease or nickase activity is referred to as a “catalytically active Cas9 molecule” (a “caCas9” molecule). A Cas9 molecule lacking the abilityto cleave or nick target nucleic acid is referred to as a “catalytically inactive Cas9 molecule” (a “ciCas9” molecule) or a “dead Cas9” (“dCas9”).

[0044] As used herein, the term “complementary” refers to the capacity for precise pairing between two nucleotides. For example, if a nucleotide at a given position of a nucleic acid is capable of hydrogen bonding with a nucleotide of another nucleic acid, then the two nucleic acids are considered to be complementary to one another at that position. Complementarity between two single-stranded nucleic acid molecules may be “partial,” in which only some of the nucleotides bind, or it may be complete when total complementarity exists between the single-stranded molecules. A first nucleotide sequence can be said to be the “complement” of a second sequence if the first nucleotide sequence is complementary to the second nucleotide sequence. A first nucleotide sequence can be said to be the “reverse complement” of a second sequence, if the first nucleotide sequence is complementary to a sequence that is the reverse (i.e., the order of the nucleotides is reversed) of the second sequence. As used herein, the terms “complement”, “complementary”, and “reverse complement” can be used interchangeably. It is understood from the disclosure that if a molecule can hybridize to another molecule it may be the complement of the molecule that is hybridizing, gRNA hybridizing to a barcode sequence.

[0045] As used herein, the terms “comprising,” “comprise” or “comprised,” and variations thereof, in reference to defined or described elements of an item, composition, apparatus, method, process, system, etc. are meant to be inclusive or open ended, permitting additional elements, thereby indicating that the defined or described item, composition, apparatus, method, process, system, etc. includes those specified elements — or, as appropriate,equivalents thereof — and that other elements can be included and still fall within the scope / definition of the defined item, composition, apparatus, method, process, system, etc.

[0046] “Encoding” refers to the inherent property of specific sequences of nucleotides in a polynucleotide, such as a gene, a cDNA, or an mRNA, to serve as templates for synthesis of other polymers and macromolecules in biological processes having either a defined sequence of nucleotides (i.e., rRNA, tRNA and mRNA) or a defined sequence of amino acids and the biological properties resulting therefrom. Thus, a gene encodes a protein if transcription and translation of mRNA corresponding to that gene produces the protein in a cell or other biological system. Both the coding strand, the nucleotide sequence of which is identical to the mRNA sequence and is usually provided in sequence listings, and the non-coding strand, used as the template for transcription of a gene or cDNA, can be referred to as encoding the protein or other product of that gene or cDNA.

[0047] The term “expression” as used herein is defined as the transcription and / or translation of a particular nucleotide sequence driven by its promoter.

[0048] As used herein, the term “gene editing complex” refers to any complex such as clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated protein 9 (Cas9), transcription activator-like effector nucleases (TALENs), zinc-finger nucleases (ZFNs), and homing endonucleases or meganucleases. Each of these complexes can be targeted to specific sites by use of guide nucleic acid sequences, e.g. gRNAs, that are complementary to a desired target sequence. For example, a gene editing complex such as CRISPR comprises two RNAs and one protein component: a CRISPR RNA (crRNA), a short RNA that undergoes complementary binding with the foreign DNA; a tracrRNA that hybridizes with the crRNA; anda Cas9 enzyme that interacts with the DNA:RNA complexes and cleaves the DNA at a specific site.

[0049] ‘Isolated” means altered or removed from the natural state. For example, a nucleic acid or a peptide naturally present in a living animal is not “isolated,” but the same nucleic acid or peptide partially or completely separated from the coexisting materials of its natural state is “isolated.” An isolated nucleic acid or protein can exist in substantially purified form, or can exist in a non-native environment such as, for example, a host cell.

[0050] As used herein, the term “gene therapy” refers to the delivery of a transgene into a cell in order to correct a genetic disorder. In certain embodiments, the gene therapy is mediated by a recombinant viral vector, e.g., a retroviral vector, a lentiviral vector, an adenoviral vector, and an adeno-associated virus vector. In general, a recombinant viral vector comprises a transgene, optionally wherein the transgene is operably linked to a transcriptional regulatory element.

[0051] The term “genome,” as used herein, generally refers to genomic information from a subject, which may be, for example, at least a portion or an entirety of a subject's hereditary information. A genome can be encoded either in DNA or in RNA. A genome can comprise coding regions (e.g., that code for proteins) as well as non-coding regions. A genome can include the sequence of all chromosomes together in an organism. For example, the human genome ordinarily has a total of 46 chromosomes. The sequence of all of these together may constitute a human genome.

[0052] As used herein, the term “guide sequence,” “crRNA,” “guide RNA,” or “single guide RNA,” or “gRNA” refers to a polynucleotide comprising any polynucleotide sequence having sufficient complementarity with a target nucleic acid sequence to hybridize withthe target nucleic acid sequence and to direct sequence-specific binding of a RNA-targeting complex comprising the guide sequence and a CRISPR effector protein to the target nucleic acid sequence. In some example embodiments, the degree of complementarity, when optimally aligned using a suitable alignment algorithm, is about or more than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more. Optimal alignment may be determined with the use of any suitable algorithm for aligning sequences, non-limiting example of which include the Smith- Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows- Wheeler Transform (e.g., the Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies; available at www.novocraft.com), ELAND (Illumina, San Diego, Calif), SOAP (available at soap.genomics.org.cn), and Maq (available at maq. sourceforge.net). The ability of a guide sequence (within a nucleic acid-targeting guide RNA) to direct sequencespecific binding of a nucleic acid-targeting complex to a target nucleic acid sequence may be assessed by any suitable assay. For example, the components of a nucleic acid-targeting CRISPR system sufficient to form a nucleic acid-targeting complex, including the guide sequence to be tested, may be provided to a host cell having the corresponding target nucleic acid sequence, such as by transfection with vectors encoding the components of the nucleic acid-targeting complex, followed by an assessment of preferential targeting (e.g., cleavage) within the target nucleic acid sequence, such as by Surveyor assay as described herein. Similarly, cleavage of a target nucleic acid sequence may be evaluated in a test tube by providing the target nucleic acid sequence, components of a nucleic acid-targeting complex, including the guide sequence to be tested and a control guide sequence different from the test guide sequence, and comparing binding or rate of cleavage at the target sequence between the test and control guide sequence reactions. Other assays are possible, and will occur to those skilled in the art. A guide sequence,and hence a nucleic acid-targeting guide may be selected to target any target nucleic acid sequence. The target sequence may be DNA. The target sequence may be any RNA sequence. In some embodiments, the target sequence is a barcode sequence.

[0053] In embodiments, the polynucleotide (e.g., gRNA) is a single-stranded ribonucleic acid. In aspects, the polynucleotide (e.g., gRNA) is 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more nucleic acid residues in length. In aspects, the polynucleotide (e.g., gRNA) is from 10 to 30 nucleic acid residues in length. In aspects, the polynucleotide (e.g., gRNA) is 20 nucleic acid residues in length. In aspects, the length of the polynucleotide (e.g., gRNA) can be at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57,58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83,84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more nucleic acid residues or sugar residues in length. In aspects, the polynucleotide (e.g., gRNA) is from 5 to 50, 10 to 50, 15 to 50, 20 to 50, 25 to 50, 30 to 50, 35 to 50, 40 to 50, 45 to 50, 5 to 75, 10 to 75, 15 to 75, 20 to 75, 25 to 75, 30 to 75, 35 to 75, 40 to 75, 45 to 75, 50 to 75, 55 to 75, 60 to 75, 65 to 75, 70 to 75, 5 to 100, 10 to 100, 15 to 100, 20 to 100, 25 to 100, 30 to 100, 35 to 100, 40 to 100, 45 to 100, 50 to 100, 55 to 100, 60 to 100, 65 to 100, 70 to 100, 75 to 100, 80 to 100, 85 to 100, 90 to 100, 95 to 100, or more residues in length. In aspects, the polynucleotide (e.g., gRNA) is from 10 to 15, 10 to 20, 10 to 30, 10 to 40, or 10 to 50 residues in length.

[0054] As used herein, a “nucleic acid” refers to a polynucleotide sequence, or fragment thereof. A nucleic acid can comprise nucleotides. A nucleic acid can be exogenous or endogenous to a cell. A nucleic acid can exist in a cell-free environment. A nucleic acid can be a gene or fragment thereof. A nucleic acid can be DNA. A nucleic acid can be RNA. A nucleicacid can comprise one or more analogs (e.g. altered backgone, sugar, or nucleobase). Some nonlimiting examples of analogs include: 5 -bromouracil, peptide nucleic acid, xeno nucleic acid, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, florophores (e.g. rhodamine or flurescein linked to the sugar), thiol containing nucleotides, biotin linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudourdine, dihydrouridine, queuosine, and wyosine. “Nucleic acid”, “polynucleotide, “target polynucleotide”, and “target nucleic acid” can be used interchangeably.

[0055] A nucleic acid can comprise one or more modifications (e.g., a base modification, a backbone modification), to provide the nucleic acid with a new or enhanced feature (e.g., improved stability). A nucleic acid can comprise a nucleic acid affinity tag. A nucleoside can be a base-sugar combination. The base portion of the nucleoside can be a heterocyclic base. The two most common classes of such heterocyclic bases are the purines and the pyrimidines. Nucleotides can be nucleosides that further include a phosphate group covalently linked to the sugar portion of the nucleoside. For those nucleosides that include a pentofuranosyl sugar, the phosphate group can be linked to the 2', the 3', or the 5' hydroxyl moiety of the sugar. In forming nucleic acids, the phosphate groups can covalently link adjacent nucleosides to one another to form a linear polymeric compound. In turn, the respective ends of this linear polymeric compound can be further joined to form a circular compound; however, linear compounds are generally suitable. In addition, linear compounds may have internal nucleotide base complementarity and may therefore fold in a manner as to produce a fully or partially double- stranded compound. Within nucleic acids, the phosphate groups can commonlybe referred to as forming the internucleoside backbone of the nucleic acid. The linkage or backbone of the nucleic acid can be a 3' to 5' phosphodiester linkage.

[0056] A nucleic acid can comprise a modified backbone and / or modified internucleoside linkages. Modified backbones can include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone. Suitable modified nucleic acid backbones containing a phosphorus atom therein can include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates such as 3 '-alkylene phosphonates, 5 '-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates including 3 '-amino phosphoramidate and aminoalkylphosphorami dates, phosphorodiami dates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, selenophosphates, and boranophosphates having normal 3'-5' linkages, 2'-5' linked analogs, and those having inverted polarity wherein one or more internucleotide linkages is a 3' to 3', a 5' to 5' or a 2' to 2' linkage.

[0057] A nucleic acid can comprise polynucleotide backbones that are formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatomic or heterocyclic internucleoside linkages. These can include those having morpholino linkages (formed in part from the sugar portion of a nucleoside); siloxane backbones; sulfide, sulfoxide and sulfone backbones; formacetyl and thioformacetyl backbones; methylene formacetyl and thioformacetyl backbones; riboacetyl backbones; alkene containing backbones; sulfamate backbones; methyleneimino and methylenehydrazino backbones; sulfonate and sulfonamide backbones; amide backbones; and others having mixed N, O, S and CH2 component parts.

[0058] A nucleic acid can comprise a nucleic acid mimetic. The term “mimetic” include, for example, polynucleotides wherein only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups, replacement of only the furanose ring can also be referred as being a sugar surrogate. The heterocyclic base moiety or a modified heterocyclic base moiety can be maintained for hybridization with an appropriate target nucleic acid. One such nucleic acid can be a peptide nucleic acid (PNA). In a PNA, the sugar- backbone of a polynucleotide can be replaced with an amide containing backbone, in particular an aminoethylglycine backbone. The nucleotides can be retained and are bound directly or indirectly to aza nitrogen atoms of the amide portion of the backbone. The backbone in PNA compounds can comprise two or more linked aminoethylglycine units which gives PNA an amide containing backbone. The heterocyclic base moieties can be bound directly or indirectly to aza nitrogen atoms of the amide portion of the backbone.

[0059] A nucleic acid can comprise a morpholino backbone structure. For example, a nucleic acid can comprise a 6-membered morpholino ring in place of a ribose ring. In some of these embodiments, a phosphorodiamidate or other non-phosphodiester internucleoside linkage can replace a phosphodiester linkage.

[0060] A nucleic acid can comprise linked morpholino units (i.e. morpholino nucleic acid) having heterocyclic bases attached to the morpholino ring. Linking groups can link the morpholino monomeric units in a morpholino nucleic acid. Non-ionic morpholino-based oligomeric compounds can have less undesired interactions with cellular proteins. Morpholinobased polynucleotides can be nonionic mimics of nucleic acids. A variety of compounds within the morpholino class can be joined using different linking groups. A further class of polynucleotide mimetic can be referred to as cyclohexenyl nucleic acids (CeNA). The furanosering normally present in a nucleic acid molecule can be replaced with a cyclohexenyl ring. CeNA DMT protected phosphoramidite monomers can be prepared and used for oligomeric compound synthesis using phosphoramidite chemistry. The incorporation of CeNA monomers into a nucleic acid chain can increase the stability of a DNA / RNA hybrid. CeNA oligoadenylates can form complexes with nucleic acid complements with similar stability to the native complexes. A further modification can include Locked Nucleic Acids (LNAs) in which the 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring thereby forming a 2'-C,4'-C-oxymethylene linkage thereby forming a bicyclic sugar moiety. The linkage can be a methylene ( — CH2-), group bridging the 2' oxygen atom and the 4' carbon atom wherein n is 1 or 2. LNA and LNA analogs can display very high duplex thermal stabilities with complementary nucleic acid (Tm=+3 to +10° C ), stability towards 3'-exonucleolytic degradation and good solubility properties.

[0061] A nucleic acid can also, in some embodiments, include nucleobase (often referred to simply as “base”) modifications or substitutions. As used herein, “unmodified” or “natural” nucleobases can include the purine bases, (e.g. adenine (A) and guanine (G)), and the pyrimidine bases, (e.g. thymine (T), cytosine (C) and uracil (U)). Modified nucleobases can include other synthetic and natural nucleobases such as 5-methylcytosine (5-me-C), 5- hydroxymethyl cytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl ( — C=C — CH3) uracil and cytosine and other alkynyl derivatives of pyrimidine bases, 6-azo uracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl,8-hydroxyl and other 8-substituted adenines and guanines, 5-halo particularly 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7- methyladenine, 2-F- adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3 -deazaguanine and 3 -deazaadenine. Modified nucleobases can include tricyclic pyrimidines such as phenoxazine cytidine(lH-pyrimido(5,4-b)(l,4)benzoxazin-2(3H)- one), phenothiazine cytidine (lH-pyrimido(5,4-b)(l,4)benzothiazin-2(3H)-one), G-clamps such as a substituted phenoxazine cytidine (e.g. 9-(2-aminoethoxy)-H-pyrimido(5,4-(b) (l,4)benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido(4,5-b)indol-2-one), pyridoindole cytidine (Hpyrido(3',':4, 5)pyrrolo [2,3-d]pyrimidin-2-one).

[0062] The term “promoter” means a DNA sequence recognized by the synthetic machinery of the cell, or introduced synthetic machinery, required to initiate the specific transcription of a polynucleotide sequence. A “minimal” promoter or “truncated” promoter or “functional fragment” of a promoter includes all essential elements of a promoter for transcriptional activation of, for example, a nucleic acid sequence operably linked or under control of the minimal promoter.

[0063] The term “sequencing,” as used herein, generally refers to methods and technologies for determining the sequence of nucleotide bases in one or more polynucleotides. The polynucleotides can be, for example, nucleic acid molecules such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), including variants or derivatives thereof (e.g., single stranded DNA). Sequencing can be performed by various systems currently available, such as, without limitation, a sequencing system by Illumina®, Pacific Biosciences (PacBio®), Oxford Nanopore®, or Life Technologies (Ion Torrent®). Alternatively, or in addition, sequencing may be performed using nucleic acid amplification, polymerase chain reaction (PCR) (e.g., digital PCR, quantitative PCR, or real time PCR), or isothermal amplification. Such systems mayprovide a plurality of raw genetic data corresponding to the genetic information of a subject (e.g., human), as generated by the systems from a sample provided by the subject. In some examples, such systems provide sequencing reads (also “reads” herein). A read may include a string of nucleic acid bases corresponding to a sequence of a nucleic acid molecule that has been sequenced. In some situations, systems and methods provided herein may be used with proteomic information.

[0064] As used herein, the “sequence identity” or percent “identity,” in the context of two or more nucleic acids or polypeptide sequences, refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same (i.e., about 60% identity, preferably 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity over a specified region, when compared and aligned for maximum correspondence over a comparison window or designated region) as measured using a BLAST or BLAST 2.0 sequence comparison algorithms with default parameters described below, or by manual alignment and visual inspection (see, e.g., NCBI web site ncbi.nlm.nih.gov / BLAST / or the like). Such sequences are then said to be “substantially identical.” This definition also refers to, or may be applied to, the compliment of a test sequence. The definition also includes sequences that have deletions and / or additions, as well as those that have substitutions. As described below, the preferred algorithms can account for gaps and the like. Preferably, identity exists over a region that is at least about 25 amino acids or nucleotides in length, or more preferably over a region that is 50-100 amino acids or nucleotides in length.

[0065] The term “target nucleic acid” refers to a nucleic acid, to which the oligonucleotide, e.g. guide RNA, is designed to specifically hybridize. It is either the presence orabsence of the target nucleic acid that is to be detected, or the amount of the target nucleic acid that is to be quantified. The target nucleic acid has a sequence that is complementary to the nucleic acid sequence of the corresponding oligonucleotide directed to the target, e.g. barcode sequence. The term target nucleic acid may refer to the specific subsequence of a larger nucleic acid to which the oligonucleotide is directed or to the overall sequence (e.g., gene or mRNA). The difference in usage will be apparent from context.

[0066] The term “transfected” or “transformed” or “transduced” means to a process by which exogenous nucleic acid is transferred or introduced into the host cell. A “transfected” or “transformed” or “transduced” cell is one which has been transfected, transformed or transduced with exogenous nucleic acid. The transfected / transformed / transduced cell includes the primary subject cell and its progeny.

[0067] A “vector” is a composition of matter which comprises an isolated nucleic acid and which can be used to deliver the isolated nucleic acid to the interior of a cell. Examples of vectors include but are not limited to, linear polynucleotides, polynucleotides associated with ionic or amphiphilic compounds, plasmids, and viruses. Thus, the term “vector” includes an autonomously replicating plasmid or a virus. The term is also construed to include non-plasmid and non-viral compounds which facilitate transfer of nucleic acid into cells, such as, for example, polylysine compounds, liposomes, and the like. Examples of viral vectors include, but are not limited to, adenoviral vectors, adeno-associated virus vectors, retroviral vectors, and the like. The terms “vector” and “plasmid” may be used interchangeably herein.

[0068] Ranges: throughout this disclosure, various aspects of the disclosure can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on thescope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the range.

[0069] Any compositions or methods provided herein can be combined with one or more of any of the other compositions and methods provided herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0070] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the Office upon request and payment of the necessary fee.

[0071] Figure 1. Overview of high-throughput cell line development platform. By screening more cells in relevant biomanufacturing conditions, an optimal cell line can be selected, reducing the cost of biomanufacturing by up to 50%.

[0072] Figure 2. Complete workflow of one work protocol.

[0073] Figure 3. Schematic of the construct design. Each construct includes a 21 nt barcode sequence (red) located adjacent to the PiggyBac ITRs and upstream of the GFP sequence. Because of its proximity in relation to the PiggyBac ITRs, when the construct is inserted into the genome, we can use targeted amplicon sequencing to track and determine every lineage’s (1) abundance in the bulk population, (2) specific productivity, (3) BOI copy number, and (4) BOI insertion sites. To perform this targeted sequencing, we can digest extracted genomic DNA using an enzyme targeting the restriction site (shown in yellow) between the barcode sequences and GFP sequence. Because restriction sites can have recognition sequencesof 4 nt in length, the digestion is expected to cut genomic DNA connected to the barcode roughly256 nt away.

[0074] Figure d. Schematic of FACS-productivity assay.

[0075] Figure 5. Specific productivity calculation for binned populations.

[0076] FIG. 6 (includes FIGS. 6a-6f) shows results of cold capture assay that demonstrates differential stability of antibody-producing CHO cell lines over extended passage.

[0077] FIG. 7 (includes FIGS. 7a-7e) shows development and validation of barcoded cell line libraries for high-throughput clone tracking.

[0078] FIG. 8 (incudes FIGS 8a- 8f) shows antibody productivity assessment in barcoded cell populations using cold capture assay.

[0079] FIG. 9 (includes FIGS. 9a-9c) shows validation of cold capture assay sorting strategy through phenotype retention analysis.

[0080] FIG. 10 (includes FIGS. 10a- lOe) shows a fed-batch culture performance comparison of infliximab-producing cell pools.

[0081] FIG. 11 (includes FIGS. 1 la- 1 Id) shows high-throughput tracking of growth and productivity dynamics for individual barcoded lineages during fed-batch culture.

[0082] FIG. 12 (includes FIGS. 12a- 12c) shows barcode-mediated linkage connects antibody titer to comprehensive clonal phenotypes for multi-parametric clone selection.

[0083] FIG. 13 shows a schematic representation of the barcoded infliximab expression plasmid construct.

[0084] FIG. 14 shows a copy number analysis of selected infliximab-expressing clones.

[0085] FIG. 15 (includes FIS. 15a-15b) shows cold capture assay confirms production of both heavy and light chain components in infliximab-expressing CHO cells

[0086] FIG. 16 (includes FIGS.16a- 16c) shows temporal dynamics of antibodyproducing populations during extended culture.

[0087] FIG. 17 (includes FIGS. 17a- 17b) shows reproducibility of productivity gate distributions between biological replicates.

[0088] FIG. 18 (includes FIGS. 18a, 18b), FIG. 19 (includes FIGS. 19a, 19b), FIG. 20 (includes FIGS. 20a, 20b), FIG. 21 (includes FIGS. 21a, 21b), FIG. 22 (includes FIGS. 22a, 22b), FIG. 23, FIG. 24 and FIG. 25 show results of Examples 2 and 3 which follow.DETAILED DESCRIPTION

[0089] The disclosure is directed to (1) development of a diverse library of barcoded constructs expressing biologies that are known to be challenging to express, such as infliximab and etanercept, to evaluate how different vector designs affect productivity; (2) optimization and validation of a high-throughput screening platform for selecting stable, high-producing clones through both early- and late-stage studies; (3) utilizing of CRISPRa-based activation and FACS- based selection assays to isolate specific high-performance clones; and (4) conducting detailed bioinformatic and Critical Quality Attribute (CQA) analyses to link vector components to productivity, stability, and quality, providing a robust data-driven approach for cell line development in biomanufacturing.

[0090] COMPOSITIONS

[0091] The present disclosure provides compositions which include host cells, plasmids, barcode sequences, gene editing complexes such as CRISPR systems, biologies of interest (BOIs) and the like.

[0092] Gene Editing Complexes

[0093] In one aspect, compositions of the disclosure include at least one gene editing complex, for example comprising CRISPR-associated nucleases such as Cas9 and Cpfl gRNAs, Argonaute family of endonucleases, clustered regularly interspaced short palindromic repeat (CRISPR) nucleases, zine-finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), meganucleases, other endo- or exo-nucleases, or combinations thereof. See Schiffer, 2012, J Virol 88(17):8920-8936, incorporated by reference.

[0094] The composition can also include C2c2 — the first naturally-occurring CRISPR system that targets only RNA. The Class 2 type VI- A CRISPR-Cas effector “C2c2” demonstrates an RN A-guided RNase function. C2c2 from the bacterium Leptotrichia shahii provides interference against RN A phage. In vitro biochemical analysis show that C2c2 is guided by a single crRNA and can be programmed to cleave ssRNA targets carrying complementary protospacers. In bacteria, C2c2 can be programmed to knock down specific mRNAs. Cleavage is mediated by catalytic residues in the two conserved HEPN domains, mutations in which generate catalytically inactive RNA-binding proteins. These results demonstrate the capability of C2c2 as a new RNA-targeting tools.

[0095] C2 c2 can be programmed to cleave particular RNA sequences in bacterial cells. The RNA-focused action of C2c2 complements the CRISPR-Cas9 system, which targets DNA, the genomic blueprint for cellular identity and function. The ability to target only RNA, which helps carry out the genomic instructions, offers the ability to specifically manipulate RNA in a high-throughput manner- and manipulate gene function more broadly.

[0096] CRISPR / Cpfl is a DNA-editing technology analogous to the CRISPR / Cas9 system, characterized in 2015 by Feng Zhang's group from the Broad Institute and MIT. Cpfl is an RNA-guided endonuclease of a class II CRISPR / Cas system. This acquired immunemechanism is found in Prevotella and Francisella bacteria. It prevents genetic damage from viruses. Cpfl genes are associated with the CRISPR locus, coding for an endonuclease that use a guide RNA to find and cleave viral DNA. Cpfl is a smaller and simpler endonuclease than Cas9, overcoming some of the CRISPR / Cas9 system limitations. CRISPR / Cpfl could have multiple applications, including treatment of genetic illnesses and degenerative conditions. As referenced above, Argonaute is another potential gene editing system.

[0097] Argonautes are a family of endonucleases that use 5' phosphorylated short single- stranded nucleic acids as guides to cleave targets (Swarts, D. C. et al. The evolutionary journey of Argonaute proteins. Nat. Struct. Mol. Biol. 21, 743-753 (2014)). Similar to Cas9, Argonautes have key roles in gene expression repression and defense against foreign nucleic acids (Swarts, D. C. et al. Nat. Struct. Mol. Biol. 21 , 743-753 (2014); Makarova, K. S., et al. Biol. Direct 4, 29 (2009). Molloy, A. Nat. Rev. Microbiol. 11, 743 (2013); Vogel, J. Science 344, 972-973 (2014). Swarts, D. C. et al. Nature 507, 258-261 (2014); Olovnikov, I., et al. Mol. Cell 51, 594-605 (2013)). However, Argonautes differ from Cas9 in many ways Swarts, D. C. et al. The evolutionary journey of Argonaute proteins. Nat. Struct. Mol. Biol. 21, 743-753 (2014)). Cas9 only exist in prokaryotes, whereas Argonautes are preserved through evolution and exist in virtually all organisms; although most Argonautes associate with single-stranded (ss)RNAs and have a central role in RNA silencing, some Argonautes bind ssDNAs and cleave target DNAs (Swarts, D. C. et al. Nature 507, 258-261 (2014); Swarts, D. C. et al. Nucleic Acids Res. 43, 5120-5129 (2015)). guide RNAs must have a 3' RNA-RNA hybridization structure for correct Cas9 binding, whereas no specific consensus secondary structure of guides is required for Argonaute binding; whereas Cas9 can only cleave a target upstream of a PAM, there is no specific sequence on targets required for Argonaute. Once Argonaute and guides bind, theyaffect the physicochemical characteristics of each other and work as a whole with kinetic properties more typical of nucleic-acid-binding proteins (Salomon, W. E., et al. Cell 162, 84-95 (2015)).

[0098] Accordingly, in certain embodiments, Argonaute endonucleases comprise those which associate with single stranded RNA (ssRN A) or single stranded DNA (ssDNA). In certain embodiments, the Argonaute is derived from Natronobacterium gregoryi. In other embodiments, the Natronobacterium gregoryi Argonaute (NgAgo) is a wild type NgAgo, a modified NgAgo, or a fragment of a wild type or modified NgAgo. The NgAgo can be modified to increase nucleic acid binding affinity and / or specificity, alter an enzymatic activity, and / or change another property of the protein. For example, nuclease (e.g., DNase) domains of the NgAgo can be modified, deleted, or inactivated.

[0099] The wild type NgAgo sequence can be modified. The NgAgo nucleotide sequence can be modified to encode biologically active variants of NgAgo, and these variants can have or can include, for example, an amino acid sequence that differs from a wild type NgAgo by virtue of containing one or more mutations (e.g., an addition, deletion, or substitution mutation or a combination of such mutations). One or more of the substitution mutations can be a substitution (e.g,, a conservative ammo acid substitution). For example, a biologically active variant of an NgAgo polypeptide can have an amino acid sequence with at least or about 50% sequence identity (e.g., at least or about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity) to a wild type NgAgo polypeptide. Conservative amino acid substitutions typically include substitutions within the following groups: glycine and alanine; valine, isoleucine, and leucine; aspartic acid and glutamic acid; asparagine, glutamine, serine and threonine; lysine, histidine and arginine; and phenylalanine and tyrosine. The ammoacid residues in the NgAgo amino acid sequence can be non-naturally occurring ammo acid residues. Naturally occurring ammo acid residues include those naturally encoded by the genetic code as well as non-standard amino acids (e.g., ammo acids having the D-configuration instead of the L- configuration). The present peptides can also include ammo acid residues that are modified versions of standard residues (e.g. pyrrolysine can be used in place of lysine and selenocysteine can be used in place of cysteine). Non-naturally occurring amino acid residues are those that have not been found in nature, but that conform to the basic formula of an amino acid and can be incorporated into a peptide. These include D-alloisoleucine(2R,3S)-2-amino-3- methylpentanoic acid and L-cyclopentyl glycine (S)-2-amino-2-cyclopentyl acetic acid. For other examples, one can consult textbooks or the worldwide web (a site currently maintained by the California Institute of Technology displays structures of non-natural amino acids that have been successfully incorporated into functional proteins).

[0100] Another gene editing complex is human WRN, a RecQ helicase encoded by the Werner syndrome gene. It is implicated in genome maintenance, including replication, recombination, excision repair and DNA damage response. These genetic processes and expression of WRN are concomitantly upregulated in many types of cancers. Therefore, it has been proposed that targeted destruction of this helicase could be useful for elimination of cancer cells. Reports have applied the external guide sequence (EGS) approach in directing an RNase P RNA to efficiently cleave the WRN mRNA in cultured human cell lines, thus abolishing translation and activity of this distinctive 3'-5' DNA helicase-nuclease. RNase P RNA isanother potential endonuclease for use with the present disclosure.

[0101] CRISPR-Associated Endonucleases

[0102] CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) is found in bacteria and is believed to protect the bacteria from phage infection. It has recently been used as a means to alter gene expression in eukaryotic DNA, but has not been proposed as an anti-viral therapy or more broadly as a way to disrupt genomic material. Rather, it has been used to introduce insertions or deletions as a way of increasing or decreasing transcription in the DNA of a targeted cell or population of cells. See for example, Horvath et al., Science (2010) 327: 167- 170; Terns et al., Current. Opinion in Microbiology.1(2011) 14:321-327; Bhaya et al., Annu Rev Genet (2011) 45:273-297; Wiedenheft et al., Nature (2012) 482:331-338); Jmek M et al., Science (2012) 337:816-821; Cong L et al., Science (2013) 339:819-823; Jmek M et al., (2013) eLife 2:e00471; Mali P et al. (2013) Science 339:823-826; Qi L S et al. (2013) Cell 152: 1173-1183; Gilbert L A et al. (2013) Cell 154:442-451; Yang H et al. (2013) Cell 154: 1370-1379; and Wang H et al. (2013) Cell 153:910-918).

[0103] CRISPR methodologies employ a nuclease, CRISPR-associated (Cas), that complexes with small RNAs as guides (gRNAs) to cleave DNA in a sequence-specific manner upstream of the protospacer adjacent motif (PAM) in any genomic location. CRISPR may use separate guide RNAs known as the crRNA and tracrRNA. These two separate RNAs have been combined into a single RNA to enable site-specific mammalian genome cuting through the design of a short, guide RNA. Cas and guide RNA (gRNA) may be synthesized by known methods. Cas / guide-RNA (gRNA) uses a non-specific DNA cleavage protein Cas, and an RNA oligonucleotide to hybridize to target and recruit the Cas / gRNA complex. See Chang et al., 2013, Cell Res. 23:465-472; Hwang et al., 2013, Nat. Biotechnol. 31 :227-229; Xiao et al., 2013, Nucl. Acids Res. l-l l.

[0104] In general, the CRISPR / Cas proteins comprise at least one RNA recognition and / or RNA binding domain. RNA recognition and / or RNA binding domains interact with guide RNAs. CRISPR / Cas proteins can also comprise nuclease domains (i.e., DNase or RNase domains), DNA binding domains, helicase domains, RNase domains, protein-protein interaction domains, dimerization domains, as well as other domains.

[0105] Barcodes

[0106] The present disclosure also provides barcodes in the generation of a barcoded cell pool expressing BOIs, determining clonal variability using a single construct.

[0107] A DNA barcode, or barcode for short, is a DNA sequence commonly used to identify a target molecule during DNA sequencing, but may also be used for other purposes such as the disruption of gene function or encoding information into a larger DNA region. Libraries of DNA barcodes have found use in chemical compound screens, the study of clonal diversity, and genomic screens. DNA barcoding technology has shown utility in applications such as the discovery of new drug candidates and elucidating protein interactions in yeast.

[0108] DNA barcode libraries can be categorized into two groups, randomly generated DNA barcode libraries that are generated by physically assembling oligos in pools, and rationally designed DNA barcode libraries that are design in silico and then manufactured. Randomly generated DNA libraries may be used for practical reasons such as cost, however, rationally designing DNA barcode libraries can have the advantage of being more robust against misidentification due to DNA sequencing and synthesis errors. Technological constraints have limited the size of the DNA barcode libraries used in screening experiments with the largest sized rationally designed DNA barcode library to date consisting of 240,000 barcodes. But continuing trends in DNA reading and writing technologies are opening up new applications forlarge-scale DNA barcode libraries. The first trend is an increase in the number of short reads made available by Next Generation Sequencing (NGS) technology. The second is the decreasing cost of manufacturing synthetic DNA. Together these trends make it possible to perform high- throughput experiments using synthetic DNA libraries in combination with screening experiments and NGS technology. A designed DNA sequence and DNA barcode are designed and manufactured together in synthetic DNA. In this case it would be beneficial for the DNA barcode sequence and biomolecule sequence it identifies to be known in advance for identification purposes, and for the barcodes to be designed in a robust manner. Applications include the construction and screening of synthetic large-scale barcoded fragment antibody libraries and barcoded synthetic genomes (Lyons, E., Sheridan, P., Tremmel, G. et al. Large- scale DNA Barcode Library Generation for Biomolecule Identification in High-throughput Screens. Sci Rep 7, 13899 (2017). https: / / doi.org / 10.1038 / s41598-017-12825-2).

[0109] A number of methods for generating customizable DNA barcode libraries have been proposed. Barcode Generator and nxCode are free tools that have been employed by their creators to generate publicly available libraries consisting of up to 96 and 587 barcodes, respectively. More sophisticated barcode library design methods based on error-correcting codes have been proposed for correcting substitutions and indel errors inherent to NGS sequencing technologies. The DNA Barcodes R package implements the error-correcting codes based generation method of Buschmann and Bystrykh (Buschmann, T. & Bystrykh, L. V. Levenshtein error-correcting barcodes for multiplexed DNA sequencing. BMC Bioinformatics 14, 1-10, doi.org / 10.1186 / 1471-2105-14-272 (2013)). The libraries described in the package manual contain from tens to hundreds of barcodes. The authors of the TagGD software package report generating libraries consisting of 100,000 barcodes in 1.5 to 7 hours, depending on theconditions of customization. A method to generate barcode libraries using particle swarm optimization has been described (Waang, B. et al. Constructing DNA Barcode Sets based on Particle Swarm Optimization. IEEE / ACM Transactions on Computational Biology and Bioinformatics 1, 5555 (2017)).

[0110] CRISPR / Cas9-Based System

[0111] The present disclosure also provides DNA targeting systems or compositions of at least one CRISPR / Cas9-based system. These compositions may be used in the expression of BOIs. The composition includes a CRISPR / Cas9-based system that includes a catalytically inactive Cas9 (dCas9) protein or dCas9 fusion protein.

[0112] Multiplex CRISPR / Cas9-Based System

[0113] The present disclosure also provides multiplex CRISPR / Cas9-based systems. These compositions may be used in the expression of BOIs. The composition includes a CRISPR / Cas9-based system that includes a catalytically inactive Cas9 (dCas9) protein or dCas9 fusion protein.

[0114] CRISPRa / dCas9 System

[0115] The present disclosure also provides CRISPRa / dCas9-based systems. These compositions include a system for activating a reporter gene. The reporter gene, e.g. GFP, is placed under a promoter that can be selectively activated by the CRISPRa system using the barcode as the CRISPRa target sequence. Specific single-guide RNAs (sgRNAs) are designed to target the desired barcode sequences, guiding the CRISPRa system composed of a catalytically inactive Cas9 (dCas9) fused to a transcriptional activator (e.g., VP64, p65, or Rta), to the target site, thereby activating transcription and inducing reporter gene expression. In certain embodiments the vectors comprise one or more transcriptional activators.

[0116] The fusion of transcriptional activator domains to a nuclease-dead form ofCas9 (dCas9) enables CRISPR-mediated transcriptional activation (CRISPRa). A host of activator domains and transgene expression systems have been engineered to enable the production of CRISPRa-competent cells.

[0117] The method may include administering to a cell a CRISPR / Cas9-based system, a polynucleotide or vector encoding said CRISPR / Cas9-based system, or DNA targeting systems or compositions of at least one CRISPR / Cas9-based system. The method may include administering a CRISPR / Cas9-based system, such as administering a Cas9 fusion protein containing transcription activation domain or a nucleotide sequence encoding a Cas9 fusion protein. The Cas9 fusion protein may include a transcription activation domain such as aVP16, p65, or Rta protein or a transcription co-activator such as a p300 protein.

[0118] In embodiments, the CRISPR / Cas system can be a type I, a type II, or a type III system. Non-limiting examples of suitable CRISPR / Cas proteins include Cas9, CasX, CasY.l, CasY.2, CasY.3, CasY.4, CasY.5, CasY.6, spCas, eSpCas, SpCas9-HFl, SpCas9-HF2, SpCas9- HF3, SpCas9-HF4, ARMAN 1, ARMAN 4, Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, CasSal, Cas8a2, Cas8b, Cas8c, Cas9, CaslO, CaslOd, CasF, CasG, CasH, Csyl, Csy2, Csy3, Csel (or CasA), Cse2 (or CasB), Cse3 (or CasE), Cse4 (or CasC), Cscl , Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Cszl, Csxl5, Csfl, Cs£2, Csf3, Csf4, and Cui 966.

[0119] The Cas9 can be an orthologous. Six smaller Cas9 orthologues have been used and reports have shown that Cas9 from Staphylococcus aureus (SaCas9) can edit the genome with efficiencies similar to those of SpCas9, while being more than 1 kilobase shorter.

[0120] In addition to the wild type and variant Cas9 endonucleases described, embodiments of the disclsoure also encompass CRISPR systems including newly developed “enhanced-specificity” 5. pyogenes Cas9 variants (eSpCas9), which dramatically reduce off target cleavage. These variants are engineered with alanine substitutions to neutralize positively charged sites in a groove that interacts with the non-target strand of DNA. This aim of this modification is to reduce interaction of Cas9 with the non-target strand, thereby encouraging rehybridization between target and non-target strands. The effect of this modification is a requirement for more stringent Watson-Crick pairing between the gRNA and the target DNA strand, which limits off-target cleavage (Slaymaker, I. M. et al. (2015)DOI: 10.1126 / science.aad5227).

[0121] In certain embodiments, three variants found have the fewest off-target effects: SpCas9 (K855A), SpCas9 (K810A / K1003A / R1060A) (a.k.a. eSpCas9 1.0), and SpCas9 (K848A / K1003A / R1060A) (a.k.a. eSPCas9 1.1) are employed in the compositions. The disclosure is by no means limited to these variants, and also encompasses all Cas9 variants (Slaymaker, I. M. et al. Science. 2016 Jan. 1; 351(6268): 84-8. doi: 10.1126 / science.aad5227. Epub 2015 Dec. 1). The present disclosure also includes another type of enhanced specificity Cas9 variant, “high fidelity” spCas9 variants (HF-Cas9). Examples of high fidelity variants include SpCas9-HFl (N497A / R661 A / Q695A / Q926A), SpCas9-HF2(N497A / R661 A / Q695A / Q926A / D1135E), SpCas9-HF3(N497A / R661 A / Q695A / Q926A / L169A), SpCas9-HF4 (N497A / R661 A / Q695 A / Q926A / Y450A). Also included are all SpCas9 variants bearing all possible single, double, triple and quadruple combinations of N497A, R661A, Q695A, Q926A or any other substitutions (Kleinstiver, B. P. et al., 2016, Nature. DOI: 10.1038 / naturel6526).

[0122] As used herein, the term “Cas” is meant to include all Cas molecules comprising variants, mutants, orthologues, high-fidelity variants and the like.

[0123] In one embodiment, the endonuclease is derived from a type II CRISPR / Cas system. In other embodiments, the endonuclease is derived from a Cas9 protein and includes Cas9, CasX, CasY.l, Cas ¥.2, CasY.3, CasY.4, CasY.5, CasY.6, spCas, eSpCas, SpCas9-HFl, SpCas9-HF2, SpCas9-HF3, SpCas9-HF4, ARMAN 1, ARMAN 4, mutants, variants, high- fidelity variants, orthologs, analogs, fragments, or combinations thereof. The Cas9 protein can be from Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Alicyclobadllus addocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobaderium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium, difficile, Finegoldia tnagna, Natran aero bins thermophilus, Pelolomaculum thermopropionicum, Addilhiobacillus caldus, Addilhiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophil.us, Nitrosococcus walsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arlhrospira sp,, Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africamts, or Acaryochloris marina. Included are Cas9 proteins encoded in genomes of the nanoarchaea ARMAN-1 (Candidatus Micrarchaeum acidiphilum ARMAN- 1) and ARMAN-4(Candidatus Parvarchaeum acidiphilum ARMAN-4), CasY (Kerfeldbacteria, Vogelbacteria, Komeilibacteria, Katanobacteria), CasX (Planctomycetes, Deltaproteobacteria).

[0124] In general, CRISPR / Cas proteins comprise at least one RNA recognition and / or RNA binding domain. RNA recognition and / or RNA binding domains interact with guide RNAs. CRISPR / Cas proteins can also comprise nuclease domains (i.e., DNase or RNase domains), DNA binding domains, helicase domains, RNAse domains, protein-protein interaction domains, dimerization domains, as wed as other domains. Active DNA-targeting CRISPR-Cas systems use 2 to 4 nucleotide protospacer-adjacent motifs (PAMs) located next to target sequences for self versus non-self discrimination. ARMAN-1 has a strong ‘NGG’ PAM preference. Cas9 also employs two separate transcripts, CRISPR RN A (crRNA) and transactivating CRISPR RNA (tracrRNA), for RNA-guided DNA cleavage. Putative tracrRNA was identified in the vicinity of both ARMAN- 1 and ARMAN-4 CRISPR-Cas9 systems (Burstein, D. et al. New' CRISPR-Cas systems from uncultivated microbes. Nature. 2017 Feb. 9;542(7640):237-241. doi: 10.1038 / nature21059. Epub 2016 Dec. 22).

[0125] Embodiments of the disclosure also include a new' type of class 2 CRISPR- Cas system found in the genomes of two bacteria recovered from groundwater and sediment samples. This system includes Cast, Cas2, Cas4 and an approximately ~980 amino acid protein that is referred to as CasX. The high conservation (68% protein sequence identity) of this protein in two organisms belonging to different phyla, Deltaproteobacteria and Planctomycetes, suggests a recent cross-phyla transfer. The CRISPR arrays associated with each CasX has highly similar repeats (86% identity) of 37 nucleotides (nt), spacers of 33-34 nt, and a putative tracrRNA between the Cas operon and the CRISPR array. Distant homology detection and protein modeling identified a RuvC domain near the CasX C-terminal end, with organizationreminiscent of that found in type V CRISPR-Cas systems. The rest of the CasX protein (630 N- terminal amino acids) showed no detectable similarity to any known protein, suggesting this is a novel class 2 effector. The combination of tracrRNA and separate Cast, Cas2 and Cas4 proteins is among type V systems, and phylogenetic analyses indicate that the Cast from the CRISPR- CasX system is distant from those of any other known type V. Further, CasX is considerably smaller than any known type V proteins: 980 aa compared to a typical size of about 1,200 amino acids for Cpfl, C2cl and C2c3 (Burstein, D. et al., 2017 supra).

[0126] Another new class 2 Cas protein is encoded in the genomes of certain candidate phyla radiation (CPR) bacteria. This approximately 1,200 amino acid Cas protein, termed CasY, appears to be part of a minimal CRISPR-Cas system that includes Casl and a CRISPR array. Most of the CRISPR arrays have unusually short spacers of 17-19 nt, but one system, which lacks Casl (CasY.5), has longer spacers (27-29 nt). Accordingly, in some embodiments of the disclosure, the CasY molecules comprise CasY 1 . CasY.2, CasY.3, CasY.4, CasY.5, CasY.6, mutants, variants, analogs or fragments thereof.

[0127] The CRISPR / Cas-like protein can be a wild type CRISPR / Cas protein, a modified CRISPR / Cas protein, or a fragment of a wild type or modified CRISPR / Cas protein. The CRISPR / Cas-like protein can be modified to increase nucleic acid binding affinity and / or specificity, alter an enzymatic activity, and / or change another property of the protein. For example, nuclease (i.e., DNase, RNase) domains of the CRISPR / Cas-like protein can be modified, deleted, or inactivated. Alternatively, the CRISPR / Cas-like protein can be truncated to remove domains that are not essential for the function of the fusion protein. The CRISPR / Cas- like protein can also be truncated or modified to optimize the activity of the effector domain of the fusion protein.

[0128] In some embodiments, the CRISPR / Cas-like protein can be derived from a wild type Cas protein or fragment thereof. In other embodiments, the CRISPR / Cas-like protein can be derived from modified Cas proteins. For example, the amino acid sequence of the Cas9 protein can be modified to alter one or more properties (e.g., nuclease activity, affinity, stability, etc.) of the protein. Alternatively, domains of the Cas9 protein not involved in RNA-guided cleavage can be eliminated from the protein such that the modified Cas9 protein is smaller than the wild type Cas9 protein.

[0129] Guide RNAs

[0130] Guide RNA sequences according to the present disclosure can be sense or anti- sense sequences. The guide RNA sequence generally includes a proto-spacer adjacent motif (PAM). The sequence of the PAM can vary depending upon the specificity' requirements of the CRISPR endonuclease used. In the CRISPR-Cas system derived from S. pyogenes, the target DNA typically immediately precedes a 5'-NGG proto-spacer adjacent motif (PAM). Thus, for the S. pyogenes Cas9, the PAM sequence can be AGG, TGG, CGG or GGG. Other Cas9 orthologs may have different PAM specificities. For example, Cas 9 from S. thermophilus requires 5'-NNAGA A for CRISPR 1 and 5'-NGGNG for CRISPR 3) and Neiseria meningitidis requires 5'-NNNNGATT).

[0131] The guide RNA sequence can be configured as a single sequence or as a combination of one or more different sequences, e.g., a multiplex configuration. Multiplex configurations can include combinations of two, three, four, five, six, seven, eight, nine, ten, or more different guide RNAs.

[0132] The gRNA sequences can include additional 5' and / or 3' sequences that may or may not be complementary to a target sequence. They can have less than 100%complementarity to a target sequence, for example 75% complementarity. The gRNA sequences can be employed as a combination of one or more different sequences, e.g., a multiplex configuration. Multiplex configurations can include combinations of two, three, four, five, six, seven, eight, nine, ten, or more different guide RNAs.

[0133] In some embodiments, the RNA molecules e.g. crRNA, tracrRNA, gRNA are engineered to comprise one or more modified nucleobases. For example, known modifications of RNA molecules can be found, for example, in Genes VI, Chapter 9 (“Interpreting the Genetic Code”), Lewis, ed. (1997, Oxford University Press, New York), and Modification and Editing of RNA, Grosjean and Benne, eds. (1998, ASM Press, Washington D.C.). Modified RNA components include the following: 2'-O-methylcytidine; N4-methylcytidine; N4-2'-O- dimethylcytidine; N4-acetylcytidine; 5-methylcytidine; 5,2'-O-dimethylcytidine; 5- hydroxymethylcytidine; 5-formylcytidine; 2'-O-methyl-5-formaylcytidine; 3 -methylcytidine; 2- thiocytidine; lysidine; 2'-O-methyluridine; 2-thiouridine; 2-thio-2'-O-methyluri dine; 3,2'-O- dimethyluridine; 3-(3-amino-3-carboxypropyl)uridine; 4-thiouridine; ribosylthymine; 5,2'-O- dimethyluridine; 5-methyl-2-thiouridine; 5-hydroxyuridine; 5-methoxyuridine; uridine 5- oxyacetic acid; uridine 5-oxyacetic acid methyl ester; 5 -carboxymethyl uridine; 5- methoxy carbonylmethyluridine; 5-methoxycarbonylmethyl-2'-O-methyluridine; 5- methoxycarbonylmethyl-2'-thiouridine; 5-carbamoylmethyluridine; 5-carbamoylmethyl-2'-O- methyluridine; 5-(carboxyhydroxymethyl)uridine; 5-(carboxyhydroxymethyl) uridinemethyl ester; 5-aminomethyl-2-thiouridine; 5-methylaminomethyluridine; 5-methylaminomethyl-2- thiouridine; 5-methylaminomethyl-2-selenouridine; 5-carboxymethylaminomethyluridine; 5- carboxymethylaminomethyl-2'-O-methyl-uridine; 5-carboxymethylaminomethyl-2-thiouridine; dihydrouridine; dihydroribosylthymine; 2'-methyladenosine; 2-methyladenosine;N6Nmethyladenosme; N6,N6-dimethyladenosine; N6,2'-O-trimethyladenosme; 2 methylthio- N6Nisopentenyladenosine; N6-(cis-hydroxyisopentenyl)-adenosine; 2-methylthio-N6-(cis- hydroxyisopentenyl)-adenosine; N6-glycinylcarbamoyl)adenosine; N6threonylcarbamoyl adenosine; N6-methyl-N6-threonylcarbamoyl adenosine; 2-methylthio-N6-methyl-N6- threonylcarbamoyl adenosine; N6-hydroxynorvalylcarbamoyl adenosine; 2-methylthio-N6- hydroxnorvalylcarbamoyl adenosine; 2'-O-ribosyladenosine (phosphate); inosine; 2'0-methyl inosine; 1 -methyl inosine; l,2'-O-dimethyl inosine; 2'-O-methyl guanosine; 1 -methyl guanosine; N2-methyl guanosine; N2,N2-dimethyl guanosine; N2,2'-O-dimethyl guanosine; N2,N2,2'-O- trimethyl guanosine; 2'-O-ribosyl guanosine (phosphate); 7-methyl guanosine; N2,7-dimethyl guanosine; N2,N2;7-trimethyl guanosine; wyosine; methylwyosine; under-modified hydroxywybutosine; wybutosine; hydroxywybutosine; peroxywybutosine; queuosine; epoxyqueuosine; galactosyl-queuosine; mannosyl-queuosine; 7-cyano-7-deazaguanosine; arachaeosine [also called 7-formamido-7-deazaguanosine]; and 7-aminomethyl-7- deazaguanosine.

[0134] Modified or Mutated Nucleic Acid Sequences

[0135] In some embodiments, any of the nucleic acid sequences may be modified or derived from a native nucleic acid sequence, for example, by introduction of mutations, deletions, substitutions, modification of nucleobases, backbones and the like. The nucleic acid sequences include the vectors, gene-editing agents, gRNAs, etc. Examples of some modified nucleic acid sequences envisioned for this disclosure include those comprising modified backbones, for example, phosphorothioates, phosphotriesters, methyl phosphonates, short chain alkyl or cycloalkyl intersugar linkages or short chain heteroatomic or heterocyclic intersugar linkages. In some embodiments, modified oligonucleotides comprise those withphosphorothioate backbones and those with heteroatom backbones, CH2 — NH — O — CH2, CH, — N(CH3) — O — CH2 [known as a methylene(methylimino) or MMI backbone], CH2 — O — N(CH3)- CH2, CH2- N(CH3 ) N(CH3 ) CH2 and O-N(CH3)-CH2-CH2 backbones. wherein the native phosphodi ester backbone is represented as 0 — -P — O — CH,). The amide backbones disclosed by De Mesmaeker et ai. Acc. Chem. Res. 1995, 28:366-374) are also embodied herein. In some embodiments, the nucleic acid sequences having morpholino backbone structures (Summerton and Weller, U.S. Pat. No. 5,034,506), peptide nucleic acid (PNA) backbone wherein the phosphodiester backbone of the oligonucleotide is replaced with a polyamide backbone, the nucleobases being bound directly or indirectly to the aza nitrogen atoms of the polyamide backbone (Nielsen et al. Science 1991, 254, 1497). The nucleic acid sequences may also comprise one or more substituted sugar moieties. The nucleic acid sequences may also have sugar mimetics such as cyclobutyls in place of the pentofuranosyl group.

[0136] The nucleic acid sequences may also include, additionally or alternatively, nucleobase (often referred to in the art simply as “base”) modifications or substitutions. As used herein, “unmodified” or “natural” nucleobases include adenine (A), guanine (G), thymine (T), cytosine (C) and uracil (U). Modified nucleobases include nucleobases found only infrequently or transiently in natural nucleic acids, e.g,, hypoxanthine, 6-methyladenine, 5-Me pyrimidines, particularly 5-rnethylcytosine (also referred to as 5-methyl-2' deoxycytosine and often referred to in the art as 5-Me-C), 5-hydroxymethylcytosine (HMC), glycosyl HMC and gentobiosyl HMC, as well as synthetic nucleobases, e.g., 2-aminoadenine, 2-(methylamino)adenine, 2-(imidazolylalkyl)adenine, 2-(aminoalklyamino)adenine or other heterosubstituted alkyladenines,2-thiouracil, 2-thiothymine, 5-bromouracil, 5-hydroxymethyluracil, 8-azaguanine, 7- deazaguanine, N6(6-aminohexyl)adenine and 2,6-diaminopurine. Kornberg, A., DNAReplication, W. H. Freeman & Co., San Francisco, 1980, pp 75-77; Gebeyehu, G , et al. Nucl. Acids Res. 1987, 15:4513). A “universal” base known in the art, e.g., inosine may be included. 5- Me-C substitutions have been shown to increase nucleic acid duplex stability by 0.6-1.2° C. (Sanghvi, Y. S., in Crooke, S. T. and Lebleu, B., eds., .Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276-278).

[0137] Another modification of the nucleic acid sequences of the disclosure involves chemically linking to the nucleic acid sequences one or more moieties or conjugates which enhance the activity or cellular uptake of the oligonucleotide. Such moieties include but are not limited to lipid moieties such as a cholesterol moiety, a cholesteryl moiety (Letsinger et al., Proc. Natl. Acad. Set. USA 1989, 86, 6553), cholic acid (Manoharan et al. Bioorg. Med. Chem. Let. 1994, 4, 1053), a thioether, e.g., hexyl-S-tritylthiol (Manoharan et al. Ann. N.Y. Acad. Sei. 1992, 660, 306; Manoharan et al. Bioorg. Med. Chem. Let. 1993, 3, 2765), a thiocholesterol (Oberhauser et al., Nucl. Adds Res. 1992, 20, 533), an aliphatic chain, e.g., dodecandiol or undecyl residues (Saison-Behmoaras et al. EMBO J. 1991, 10, 111; Kabanov et al. FEBSLett. 1990, 259, 327; Svinarchuk et al. Biochimie 1993, 75, 49), a phospholipid, e.g., di- hexadecyl- rac-glycerol or triethylammonium l,2-di-O-hexadecyl-rac-glycero-3-H-phosphonate (Manoharan et al. Tetrahedron Lett. 1995, 36, 3651; Shea et al Nucl. Acids Res. 1990, 18, 3777), a polyamine or a polyethylene glycol chain (Manoharan et al. Nucleosides & Nucleotides 1995, 14, 969), or adamantane acetic acid (Manoharan et al. Tetrahedron Lett. 1995, 36, 3651). It is not necessary for all positions in a given nucleic acid sequence to be uniformly modified, and in fact more than one of the aforementioned modifications may be incorporated in a single nucleic acid sequence or even at within a single nucleoside within a nucleic acid sequence.

[0138] Kits

[0139] The disclosure also provides for kits. The kits can include any one or more components of the compositions disclosed herein.

[0140] Examples

[0141] Example 1: The goals of this disclosure are to (1) develop a diverse library of barcoded constructs expressing biologies that are known to be challenging to express, such as infliximab and etanercept, to evaluate how different vector designs affect productivity; (2) optimize and validate our high-throughput screening platform for selecting stable, high- producing clones through both early- and late-stage studies; (3) utilize CRISPRa-based activation and FACS-based selection assays to isolate specific high-performance clones; and (4) conduct detailed bioinformatic and Critical Quality Attribute (CQA) analyses to link vector components to productivity, stability, and quality, providing a robust data-driven approach for cell line development in biomanufacturing.

[0142] Generation of Barcoded Library Cell Line Expressing BOI

[0143] This task focuses on developing a barcoded cell line (e.g. CHOZN GS cell line) for high-throughput screening of biologic production involving two key steps: constructing a vector library and generating a stable cell line. First, a diverse vector library is created containing a barcode along with randomly assembled gene components, including sequences for the biologic of interest (BOIs) and various regulatory elements (e.g., promoters, UTRs, poly(A) sequences). In a preferred aspect, this library uses the PiggyBac transposon system for integration into host cells. In a preferred aspect, the library is designed for parallel screening of different vector designs to identify optimal combinations for biomanufacturing. A CRISPR- mediated transcriptional activation (CRISPRa) system on the plasmid will enable later selection of high-performing clones by targeting specific barcodes to activate GFP expression, allowingfor fluorescence-based sorting. Next, in a preferred aspects, the library is transfected into selected cells such as CHOZN GS cells, followed by selection e.g. using glutamine-free media and FACS-based productivity assays to isolate high-productivity clones, ultimately generating a stable, barcoded cell line suitable for further analysis and biomanufacturing applications.

[0144] Generation of the Vector Library Containing Barcodes and BOI

[0145] This focuses on creating plasmid constructs, each with a barcode, expressing a BOI such as infliximab (Remicade) or etanercept (Enbrel )( Figure 3, shown in purple). These BOIs were selected because they are known to be hard-to-express (7-4), allowing us to demonstrate that Biolinco’s platform can effectively optimize the expression of a range of recombinant proteins beyond trastuzumab. In one aspect, we use the PiggyBac transposon system to integrate the BOI into the cell line. The plasmid contains the PiggyBac-specific inverted terminal repeat (ITR) sequences, for the transposase to recognize and integrate the transgene into the host genome. Additionally, the plasmid preferably includes a glutamine synthetase (GS) expression cassette under the control of a constitutive promoter (shown in blue), serving as a selectable marker.

[0146] Subsequent library preparation can proceed with standard kits, such as with NEBNext® Ultra™ II DNA Library Prep with Sample Purification Beads. In addition, long-read sequencing such as Oxford Nanopore sequencing for the variant library will be performed by Plasmidsaurus to obtain continuous reads that span entire plasmid constructs. Nanopore sequencing enables the accurate identification of combinations of different vector components within individual plasmids, ensuring precise characterization of the diverse constructs present in the library. Sequencing the library will provide detailed information on the barcode abundance, vector components, copy numbers, and insertion sites. Moreover, after we have identified a clonethat will demonstrate high performance in large-scale manufacturing, we can use the barcode sequence to isolate that specific lineage via targeted CRISPR activation.

[0147] The plasmid design is structured to include varying vector elements that control the transcriptional regulation of the biologies of interest (BOIs) to create a diverse library of expression constructs. This enables us to rapidly screen many different vector designs in parallel to determine the best combination for biomanufacturing. To achieve this, the vector will incorporate variant versions of promoters, 5' untranslated regions (5' UTRs), 3' untranslated regions (3' UTRs), and polyadenylation (poly(A)) sequences. By using Golden Gate Assembly, a highly efficient and modular cloning strategy, these different regulatory elements can be rapidly and randomly assembled into the final construct, enabling the creation of a diverse plasmid library where each construct represents a combination of these regulatory sequences.

[0148] To facilitate the selection of a desired lineage, a CRISPRa system will be employed to activate the green fluorescent protein (GFP) reporter gene (shown in green), enabling us to sort GFP-positive cells via fluorescence-activated cell sorting (FACS). The GFP gene is placed under a promoter that can be selectively activated by the CRISPRa system using the barcode as the CRISPRa target sequence. Specific single-guide RNAs (sgRNAs) are designed to target the desired barcode sequences, guiding the CRISPRa system — composed of a catalytically inactive Cas9 (dCas9) fused to a transcriptional activator (e.g., VP64, p65, or Rta) — to the target site, thereby activating transcription and inducing GFP expression.

[0149] Generation of the Cell Line Expressing the Barcoded Library and BOI

[0150] This involves transfecting cells with the barcoded plasmid library and selecting cells expressing the BOI. We will use CHOZN® GS- / - ZFN-modified CHO cell line, which is a GS- (Glutamine Synthetase-negative) cell line. Chemical transfection methods such asPEI or Lipofectamine will be used. After transfecting the plasmid library, cells will be cultured in a glutamine-free medium, allowing only those that have successfully integrated the GS gene connected to the BOI to survive. Subsequently, FACS will be used to isolate polyclonal cell populations (1000 cells) based on APC fluorescence intensity as described in the cold capture assay (5). Briefly, the cold capture assay uses fluorescence activated cell sorting (FACS) to measure a cell’s specific productivity by staining surface-bound antibody products with an APC- conjugated F(ab’)2 antibody fragment targeting human IgGs, such as Jackson ImmunoResearch Laboratories Allophycocyanin (APC) AffiniPure™ F(ab')2 Fragment Goat Anti- Human IgG, Fey fragment specific (109-136-170). The resulting fluorescence intensity provides a direct readout of protein productivity for each cell. We will call this assay FACS- productivity for clarity (Figure 4).

[0151] High-Throughput Screening for Early- and Late-Stage Clone Selection

[0152] This focuses on refining Biolinco’s high-throughput screening platform. The first subtask involves an early-stage study to evaluate the performance of a diverse barcoded pool of CHOZN cells expressing biologies under batch and fed-batch conditions. Daily samples will be collected for barcode sequencing and FACS-productivity assays to identify clones with high growth rates and productivity. The second subtask is a long-term stability study of the barcoded pool over 60-90 days to identify stable, high-producing clones. By comparing early and latestage results, this task aims to validate the platform's ability to reliably select stable, optimal clones for large-scale biomanufacturing.

[0153] Early-stage high-throughput study to track lineage performance in biomanufacturing environments'. The objective is to cultivate a diverse barcoded pool of theCHOZN cells expressing biologies of interest (BOIs) under both batch and fed-batch cultureconditions. During the culture period, daily samples of the bulk population will be collected for barcode sequencing using an Illumina MiSeq sequencer, and the FACS-Productivity assay will be run daily to determine cell specific productivity levels. Additionally, we will determine the bulk cell density and bulk titer for use in determining every lineage’s abundance and specific productivity.

[0154] More specifically, titer (T), productivity (qP), and cell density (X) are related by the equationBecause of this relationship, either titer or productivity can be used to select a top performing clone (6). Using our workflow, first we can determine the cell density for a particular lineage by targeted amplicon sequencing of the barcodes in the and a simple bulk cell density count. If, for example, the cell density is 10xl06cells / mL and we know that the barcode / lineage ATCG makes up 1% of the bulk population, we know that there are 0.1x106 cells / mL of lineage ATCG in the culture. We can similarly determine the specific productivity from the FACS assay (Figure 5). Measuring the bulk viable cell density and titer will allow us to calculate the bulk specific productivity qP,buik. This is the average qPacross all cells in the population and is equal to the area under the FACS curve. To determine every lineage’s qP, we can calculate the area of each bin by multiplying the bin’s average fluorescence by the fraction of cells in each bin and solving. This can be used to determine the productivity of each lineage by taking the productivity of each lineage weighted by each lineage’s abundance in each bin (see FACS-productivity assay details on how to determine lineage abundance in each bin). With the qPand X for each lineage now in hand, we can use the equation at the start of this paragraph to calculate the final titer for each lineage and select the optimal clone. Therefore, we will be able to identify clonal lineages that exhibit both high growth rates and high productivity under specific culture conditions. Clones with high calculated titers are identified as optimal forbiomanufacturing scale-up, while those with lower calculated titers are categorized as suboptimal.

[0155] Late-stage high-throughput study to identify stable and high producers

[0156] While the early-stage study provides expedited results and validation of the assays, the CLD process requires more extended timelines to assess the stability of clones. This will focus on conducting a long-term stability study by growing the barcoded pool of cells expressing the BOI for 60 days, or 60 population doublings, reflecting the typical duration from thaw to scale-up in a standard CLD process. During this period, the stability of each clone will be monitored to identify barcodes associated with robust and stable expression profiles. After this aging process, another round of batch or fed-batch culture will be conducted to evaluate the productivity and growth performance of the long-term cultured clones using the same protocol described above. The results will successfully identify clones that not only show high growth rates and productivity in the short term but also maintain these characteristics over extended periods, indicating their suitability for large-scale production.

[0157] CRISPRa Design to Selectively Activate Clones

[0158] These experiments will develop a CRISPRa-based system to use GFP to isolate a specific lineage via FACS. To activate the GFP, we will co-express a dCas9 fused with a transcriptional activator and sgRNAs specifically designed to target the top, bottom, and middle 3 barcodes identified above. FACS will be utilized to isolate lineages with those barcodes as indicated by GFP expression. We will call this assay FACS-select for clarity. The isolated cells will be expanded and banked.

[0159] Cloning the dCas9-activator construct'. To generate the selection vector, we will clone the desired gRNA sequence into a CRISPRa vector purchased from OriGene. The all-in- one CRISPRa vectors contain both gRNA cloning sites under a U6 promoter, and a CMV- driven dCas9-VP64 expression cassette.

[0160] After cloning the gRNA sequence targeting a specific barcode, the corresponding lineage can express the dCas9-VP64 fusion protein and gRNA.

[0161] Transfection of the CRISPRa Vector and FACS-Select to isolate the select cells

[0162] The objective of this is to transfect a bank containing young, unaged barcoded CHOZN HOST cells with the validated CRISPRa vector to selectively activate GFP expression in cells containing the specific barcodes.

[0163] The CRISPRa vector will be transfected into the cell population using lipofectamine and upon successful transfection will selectively activate GFP expression in cells with those barcodes. The GFP-positive cells will be sorted out. These cells will then be expanded in culture to generate sufficient cell numbers for further subsequent scale-up and biomanufacturing.

[0164] Batch / Fed-batch study on the isolated clone

[0165] Following the isolation of targeted clones via the FACS-select process, a fed- batch culture study will be conducted on the isolated monoclonal lineages to evaluate their growth dynamics and productivity. This study will mirror the conditions used in the earlier fed- batch studies on the polyclonal pool, allowing us to compare the behavior of isolated clones against their performance within the polyclonal population. Additionally, this will allow us to confirm that identified high and low performing cells are similarly high and low performing in monoclonal culture. By leveraging sequencing data obtained from the polyclonal pool, we will reconstruct the growth and productivity profile for each given clone and compare it with itsperformance as an isolated monoclonal population. This analysis will provide insight into how individual clones behave in polyclonal culture versus in monoclonal culture. The goal is to demonstrate that the best-performing clones in the polyclonal population maintain their high productivity and growth characteristics as monoclonal populations.

[0166] Bioinformatics and CQAs analysis

[0167] The goal is to evaluate the top-performing monoclonal clones for their suitability in large-scale biomanufacturing by focusing on their Critical Quality Attributes (CQAs) and understanding the vector components influencing productivity. The first step involves detailed CQA analysis of the recombinant protein product to ensure it meets industry standards for safety, efficacy, and stability, conducted in collaboration with a CRO. The second focuses on bioinformatic analysis to link productivity data with vector design elements, such as promoters and UTRs, identifying the components that drive high productivity. This integrated approach aims to optimize clone selection and vector design for better biomanufacturing outcomes.

[0168] Characteristics / measurement of CQA for selected clones

[0169] This focuses on the evaluation of CQAs for the top-performing monoclonal lineages identified. CQAs are essential parameters that define the safety, efficacy, and stability of biologies, and include attributes such as protein aggregation, glycosylation patterns, purity, and charge variants. This will involve selecting a subset of clones for detailed CQA analysis to ensure that they meet industry standards for therapeutic development and production. The analysis will be conducted in collaboration with a specialized Contract Research Organization (CRO) such as Charles River Laboratories that has the expertise and facilities to perform advanced characterization studies. More specifically for infliximab, CQAs such as TNA-abinding activities, glycosylation profiles, FcyR-IIIa binding affinities, will be measured (8).Similarly, for etanercept, CQAs including TNF-a binding and neutralization, TNF-0 neutralization, and Fc-mediated functions will be measured (9). In addition, we will be able to compare our antibody against existing biosimilars.

[0170] Bioinformatic analysis to identify vector components contributing to highly productive clones

[0171] The FACS-APC productivity assay will generate a large volume of information about a lineage’s growth rate and specific productivity that can be linked to its barcode. Additionally, we will know the vector components (which promoter, CDS, poly(A) tail, etc.) associated with each barcode. Thus, we can link each construct with the resulting growth rate, specific productivity, copy number, and genomic insertion site. By connecting construct design to these attributes, we can quantify which combination of vector components produces a highly productive cell line. We will do this by using custom scripts, some of which are already used in the Kalhor lab and some of which will be developed throughout this work. This data can be used to inform future designs that further improve cellular productivity.

[0172] Bibliography1. V. Le Fourn, P.-A. Girod, M. Buceta, A. Regamey, N. Mermod, CHO cell engineering to prevent polypeptide aggregation and improve therapeutic protein secretion. Metab Eng 21, 91-102 (2014).2. K. Kaneyoshi, K. Kuroda, K. Uchiyama, M. Onitsuka, N. Yamano-Adachi, Y. Koga, T. Omasa, Secretion analysis of intracellular “difficult-to-express” immunoglobulin G (IgG) in Chinese hamster ovary (CHO) cells. Cytotechnology 71, 305-316 (2019).3. N. Pristovsek, H. G. Hansen, D. Sergeeva, N. Borth, G. M. Lee, M. R. Andersen, H. F. Kildegaard, Using Titer and Titer Normalized to Confluence Are Complementary Strategies for Obtaining Chinese Hamster Ovary Cell Lines with High Volumetric Productivity of Etanercept. Biotechnol J 13 (2018).4. P. M. O’Callaghan, M. E. Berthelot, R. J. Young, J. W. A. Graham, A. J. Racher,D. Aldana, Diversity in host clone performance within a Chinese hamster ovary cell line. Biotechnol Prog 31, 1187-1200 (2015).5. K. V. Meyer, I. G. Siller, J. Schellenberg, A. Gonzalez Salcedo, D. Solle, J. Matuszczyk,6. T. Scheper, J. Bahnemann, Monitoring cell productivity for the production of recombinant proteins by flow cytometry: An effective application using the cold capture assay. EngLife Sci 21, 288-293 (2021).7. N. Pristovsek, H. G. Hansen, D. Sergeeva, N. Borth, G. M. Lee, M. R. Andersen, H. F. Kildegaard, Using Titer and Titer Normalized to Confluence Are Complementary Strategies for Obtaining Chinese Hamster Ovary Cell Lines with High Volumetric Productivity of Etanercept. Biotechnol J 13 (2018).8. C. A. Gersbach, P. Perez-Pinera, Activating human genes with zinc finger proteins, transcription activator-like effectors and CRISPR / Cas9 for gene therapy and regenerative medicine. Expert Opin Ther Targets 18, 835-839 (2014).9. J. Kang, K. Pisupati, A. Benet, B. T. Ruotolo, S. P. Schwendeman, A. Schwendeman, Infliximab Biosimilars in the Age of Personalized Medicine. Trends Biotechnol 36, 987-992 (2018).10. U.S. Food and Drug Administration, “Clinical Pharmacology and Biopharmaceutics Review(s): Application Number 7610420rigls000” (2016); https: / / www.accessdata.fda.gov / drugsatfda_docs / nda / 2016 / 761042Origls000Cros sR.pdf.

[0173] Example 2: Validation of Cold Capture Assay for Clone Stability Assessment

[0174] To assess the stability of antibody production in CHO cell lines over extended culture periods, a cold capture assay was employed using APC-conjugated secondary antibodies to detect surface-bound antibodies. Flow cytometry analysis confirmed the assay's specificity and sensitivity, with untransfected CHOZN host cells showing minimal background (99.2% APC-negative), while the NIST-CHO reference standard demonstrated robust antibody production (98.9% APC-positive) (in FIG. 6, ).

[0175] Two established monoclonal cell lines were evaluated with known stability characteristics over 29 passages to demonstrate the assay's ability to track productivity changes. Clone A, previously characterized as stable, maintained consistently high antibody productionthroughout extended culture, with only a 1.5% decrease in producing cells from P3 to P32 (In FIG. 6, c-6d). In contrast, Clone B exhibited substantial instability, with the percentage of APC- positive cells declining from 91.1% at P3 to 48.9% at P32, representing a 46% reduction in the producing cell population (In FIG. 6, ). These results demonstrate that the cold capture assay provides a reliable method for tracking antibody production stability and highlight the significant variation in clone performance over extended culture periods typical of biomanufacturing scale- up processes.

[0176] In FIG. 6, Flow cytometry analysis of antibody production using APC- conjugated secondary antibody cold capture assay is shown, (a) CHOZN host cells (untransfected negative control) showing 99.2% APC-negative population, (b) NIST-CHO reference standard (positive control) showing 98.9% APC-positive population. (c,d) Clone A at passage 3 (P3) and passage 32 (P32) demonstrating stable antibody production with 99.6% and 98.1% APC-positive cells, respectively. (e,f) Clone B at P3 and P32 showing progressive loss of antibody production from 91.1% to 48.9% APC-positive cells over 29 passages. APC-positive cells (right gate, FL4- A : : APC-A) indicate antibody-producing cells, while APC-negative cells (left gate) represent nonproducing cells. FL2-A :: 525_50-A represents GFP detection. Numbers shown represent percentage of total cell population based on APC gating.

[0177] Development and Characterization of Barcoded Library Platform

[0178] To enable high throughput tracking of individual clonal lineages, a barcoded plasmid system incorporating unique DNA sequences alongside the gene of interest was produced. The 12,383 bp expression vector utilized PiggyBac inverted terminal repeats for targeted genomic integration, with the cargo region containing triple barcode cassettes (60 bp total), an sfGFP reporter under CMV promoter, infliximab heavy and light chain genes, IRES2 sequence, andglutamine synthetase selection marker (FIG. 7a, In FIG. 13). Two base promoter configurations containing BsmBI restriction sites flanking the barcode insertion location were generated: IF (rEFla driving heavy chain, SV40 driving light chain) (the sequence of this construct is set forth under the heading Sequence of infliximab-F -base-construct at the end of this Example 2) and IS (SV40 driving heavy chain, rEFla driving light chain) (the sequence of this construct is set forth Sequence infliximab-S-base-construct at the end of this Example 2). Library generation employed Golden Gate assembly followed by Gibson assembly and validation through high-throughput sequencing (FIG. 7b). Golden Gate assembly using the BsmBI enzyme is then employed to insert a unique double-stranded DNA barcode into the vector to generate a vector library. The sequences of these constructs containing the barcode are set forth under the headings Sequence of infliximab- F-construct with barcode and Sequence of infliximab-S-construct with barcode at the end of this Example 2. Following those sequences with barcode regions (shown in bold), sequences for two exemplary constructs with restrictions sites shown (shown with restriction sites in bold and italics).

[0179] Cell viability following transfection revealed distinct recovery patterns between experimental conditions and demonstrated successful stable transfection (FIG. 7c). PiggyBac transposase-mediated integration (PB) generally showed superior recovery compared to random integration (RI). The resulting library exhibited substantial barcode diversity, spanning approximately three orders of magnitude in frequency distribution when analyzed on a log 10 scale (FIG. 7d), with randomized nucleotide distribution across all barcode positions confirming the absence of systematic sequence bias (FIG. 7e). Quantitative PCR revealed that selected cell pools contained 2-6 transgene copies per cell, consistent with optimal productivity ranges reported for commercial CHO cell lines (FIG. 14).

[0180] In FIG. 7: (a) Schematic of barcoded plasmid construct design. PiggyBac inverted terminal repeats (ITRs) flank the cargo containing static barcodes, sfGFP reporter, IRES- linked glutamine synthetase selection marker, and infliximab heavy and light chain genes under rEFla or SV40 promoters. Unique barcode sequences enable lineage tracking following genomic integration, (b) Library generation workflow using Golden Gate assembly to create diverse barcoded plasmids from template fragments, followed by Gibson assembly and sequencing validation, (c) Cell viability recovery following transfection and selection over time. Experimental conditions include random integration (RI) vs. PiggyBac transposase-mediated integration (PB), with promoter configurations IF (rEFla heavy chain / SV40 light chain) and IS (SV40 heavy chain / rEFla light chain), compared to non-transfected control (NT), (d) LoglO-transformed barcode frequency distribution in stable infliximab-expressing cell pool, demonstrating library diversity spanning approximately 3 orders of magnitude, (e) Sequence logo representation of barcode nucleotide composition across all positions, showing randomized barcode design.

[0181] Comprehensive sequencing analysis revealed substantial differences in barcode representation between promoter configurations. The IF configuration demonstrated superior barcode retention with 665,294 unique barcodes detected compared to 217,934 in the IS configuration, representing a 3 -fold difference in library complexity. Shannon entropy analysis confirmed this diversity advantage, with IF populations showing 2A501,048 effective diversity compared to 2A162,453 for IS populations. Remarkably, the two populations showed minimal overlap, with only 26 barcodes shared between configurations (0.003% of 883,202 total unique barcodes) as show in the following Table A:

[0182] Table A: Barcode diversity analysis for infliximab-expressing cell populations by promoter configuration. Comparative analysis of barcode representation and diversity metricsfor IF and IS promoter configurations in infliximab-expressing CHO cell pools. Number of barcodes represents total unique barcode sequences detected in each population. Shannon Entropy (2A) indicates barcode diversity within each population, with higher values representing more even barcode distribution. Union (unique barcodes) shows the total number of distinct barcodes across both configurations combined. Intersection (shared barcodes) indicates the number of barcode sequences present in both IF and IS populations.Cold Capture Assay Enables Quantitative and Dynamic Analysis of Antibody ProductionThe cold capture assay was established as a robust method for analyzing antibody production and tracking the dynamics of heterogeneous cell populations. The assay's sensitivity was demonstrated by its ability to clearly resolve dramatic differences in productivity between cells transfected with two different promoter configurations. The IS configuration (SV40 driving heavy chain, rEFla driving light chain) resulted in 14.95% APC-positive cells with 2.84% achieving high producer status (APC++), demonstrating moderate antibody production with clear population heterogeneity (FIG. 8e). In stark contrast, the IF configuration showed a lower antibody production, with only 0.23% of cells registering as APC-positive and minimal high- producer population (FIG. 8f).The assay's analytical capabilities were extended through chain-specific detection, which confirmed that selected infliximab-expressing CHO cells produced both heavy and light antibodycomponents (FIG. 15). Temporal analysis demonstrated the platform's capability to monitor productivity dynamics over extended culture periods. The assay revealed shifts in the population's distribution across productivity gates between day 3 and day 9, showcasing its utility for longitudinal performance tracking (FIG. 16).In FIG 8: Flow cytometry analysis of antibody production using APC-conjugated secondary antibody detection, (a) CHOZN-HOST negative control showing minimal background with 0.12% APC-positive cells, (b) NIST-CHO reference standard showing 98.47% APC- positive cells, including 78.54% high producers (APC++). (c,d) Histogram representations of CHOZN-HOST and NIST-CHO populations, respectively, (e) IS promoter configuration (SV40 heavy chain / rEFla light chain) showing 14.95% APC-positive cells with 2.84% high producers, (f) IF promoter configuration (rEFla heavy chain / SV40 light chain) showing 0.23% APC- positive cells. EGFP-A-Compensated represents enhanced green fluorescent protein detection, while APC-A-Compensated represents antibody productivity. Gate boundaries distinguish nonproducing (APC-), producing (APC+), and high-producing (APC++) populations.To validate the reliability of cold capture assay-based sorting for clone isolation, we performed FACS separation of APC-negative and APC-positive populations from IS-transfected cells. Following 10 days of expansion post-sorting, both populations largely retained their original productivity phenotypes. Cells sorted from the APC-negative gate remained predominantly non-productive, though a small portion transitioned to APC-positive status. Conversely, cells sorted from the APC-positive gate maintained high antibody production levels, with the majority remaining in the producing population (FIG. 9). The observed phenotypic drift reflects the inherent instability of polyclonal populations, underscoring the importance ofscreening approaches that can monitor clonal behavior over time rather than relying on singletimepoint measurements.In FIG. 9, flow cytometry validation of FACS-based separation of antibody-producing and non-producing cell populations using IS promoter configuration, (a) Initial mixed population showing distinct APC-negative (red gate) and APC-positive (yellow gate) populations separated by a buffer zone to ensure clean population separation, (b) Post-sort analysis of cells isolated from the APC-negative gate (red population) after 10 days expansion, demonstrating retention of non-producing phenotype, (c) Post-sort analysis of cells isolated from the APC-positive gate (yellow population) after 10 days expansion, showing maintenance of antibody-producing phenotype. APC-A-Compensated represents antibody productivity detected by cold capture assay. Color coding indicates sorted population identity: red represents APC-negative sorted cells, yellow represents APC-positive sorted cells.Manufacturing-Scale Performance Tracking of Barcoded Populations

[0183] Fed-batch culture experiments conducted over 15 days evaluated antibody production under manufacturing-relevant conditions. All cell populations demonstrated robust growth, with viable cell densities reaching 40-50 x 106cells / mL by days 14-18 and viability remaining consistently high (>95%) throughout most of the culture period (FIG. lOa-lOb). Glucose concentration was maintained between 4-8 g / L to support optimal cell growth (FIG. 10c). Lactate level was monitored daily to assess the metabolic state of the cultures (FIG. lOd). NIST- CHO reference standard achieved 5,763 pg / mL by day 15, while CHOZN host cells showed minimal background production. Antibody production in the three infliximab-expressing pools was first detectable on day 5, with final titers varying among pools: Pool 1 achieved 17.9 pg / mL, Pool3 reached 15.5 pg / mL, and Pool 2 produced 12.1 pg / mL by culture termination (FIG. lOe).

[0184] In FIG. 10, time course analysis of fed-batch culture parameters for NIST-CHO reference standard, CHOZN host cells, and three independent infliximab-expressing cell pools over 20 days, (a) Viable cell density (*106cells / mL). (b) Cell viability percentage, (c) Glucose concentration (g / L). (d) Lactate concentration (g / L). (e) Antibody titer (pg / mL) measured by HPLC, excluding NIST-CHO to highlight productivity differences among experimental pools. Data points represent mean ± standard deviation from biological triplicates. HOST serves as negative control, and Infliximab Pools 1-3 represent independent transfected populations from the barcoded library system.

[0185] High-throughput barcode sequencing enabled the longitudinal tracking of thousands of individual clonal lineages throughout the 15-day fed-batch culture, revealing significant heterogeneity in clonal fitness (FIGS, lla-llb). The comprehensive tracking of thousands of individual barcodes demonstrates the power of population-scale analysis, with some clones expanding dramatically while others are more moderate during extended culture.

[0186] The platform's capability was further demonstrated by assigning a unique productivity trajectory to each individual barcode, allowing for a detailed, clone-level analysis of antibody production (FIGS, llc-lld). This method uncovered a wide, four-order-of-magnitude dynamic range of productivity across the population. It was noted that promoter configuration influenced these profiles, with IS configuration lineages (cyan lines) generally showing higher productivity than IF lineages (red lines). This integrated analysis of growth and productivity phenotypes enables identification of rare high-performing clones that combine robust growth with sustained antibody production, which is a selection precision impossible with traditional methods that assess these parameters independently. The reproducibility of these productivity distributionswas confirmed across biological replicates, with both independent experiments showing consistent population distributions across productivity bins (FIG. 17).

[0187] In FIG. 11, longitudinal analysis of thousands of clonal lineages tracked simultaneously over 15-day fed-batch culture. (a,b) Growth trajectories showing normalized viable cell density (VCD) for individual barcoded lineages shown on linear (a) and log (b) scales. (c,d) Antibody productivity trajectories for individual barcoded lineages stratified by promoter configuration. Productivity measured as fluorescence units from cold capture assay is shown on linear (c) and log (d) scales. Each line represents an individual barcoded lineage, color-coded by vector design: IF (rEFla-heavy chain / SV40-light chain, red), IS (SV40-heavy chain / rEF la-light chain, cyan), and NA (unassigned, gray).

[0188] The platform's core capability is to characterize the overall performance of thousands of individual clones by generating a multidimensional profile for each unique barcode. This comprehensive profile is built from linking a clone's identity to multiple phenotypic measurements. For instance, the platform can resolve single-parameter distributions for the entire population, such as the characteristic right-skewed distribution of antibody titers seen in the histogram and rank-ordered plots (FIGS. 12a-12B).

[0189] The true power of the platform is its ability to integrate these metrics into a multi-parametric performance fingerprint for each clone. This is visualized in the radar plot, which synthesizes five critical manufacturing parameters: Titer, Initial VCD, Growth Rate (p), Max VCD, and Average Productivity (FIG. 12c). This multidimensional view provides a holistic 'fingerprint' of each clone's unique behavior. This approach provides the resolution necessary for rational clone selection based on a complete performance profile, rather than on a single metric like titer alone.

[0190] In FIG. 12, (a) Distribution of relative antibody titers across 600 barcoded clones measured by cold capture assay, (b) Rank-ordered titer plot demonstrating distribution of productivity across the population, with individual clones colored by promoter configuration (IF / red, IS / cyan). (c) Multi-dimensional performance fingerprints for individual clones identified through barcode tracking. Five parameters were extracted by linking barcode identity to phenotypic measurements: Titer, Initial VCD, Growth Rate, Max VCD, and Average Productivity. Each line represents an individual clone, color-coded by promoter configuration, revealing distinct performance profiles for the IF (red) and IS (cyan) populations.

[0191] In FIG. 13, circular plasmid map (12,383 bp) showing the modular design of the barcoded expression vector used for infliximab production and lineage tracking. The construct contains PiggyBac inverted terminal repeats (ITRs) flanking the expression cassette to facilitate targeted genomic integration. Key functional elements include: static barcode sequences for unique clone identification; sfGFP reporter gene under CMV promoter for transfection monitoring; infliximab heavy and light chain genes under rEFla and SV40 promoters (configurations can be swapped between IF and IS variants); IRES2 sequence for polycistronic expression; glutamine synthetase selection marker for stable clone isolation; and various regulatory elements including enhancers, signals, and fragments for optimal expression. The modular design enables simultaneous tracking of clonal identity while maintaining antibody expression and selection capabilities for high-throughput cell line development applications.

[0192] In FIG. 14, quantitative PCR analysis of integrated plasmid copy number per cell for individual clones isolated from IF and IS populations. Clone identifications indicate promoter configuration (IF or IS) followed by clone number. Copy numbers range from approximately 1.8 to 6.3 copies per cell, with IS+11 showing the highest integration level. Datarepresent mean ± standard deviation from triplicate qPCR reactions performed on independent genomic DNA extractions.

[0193] In FIG. 15, flow cytometry analysis using chain-specific detection antibodies demonstrates that selected CHO cell populations produce both (a) heavy chain [APC] and (b) light chain [AlexaFluor 647] components of infliximab. The bimodal distributions indicate heterogeneous expression levels within the population, with clear separation between nonproducing (red) and producing (yellow) cell subsets for both antibody chains, confirming assembly-competent antibody production.

[0194] In FIG. 16, flow cytometry analysis showing antibody production levels over a 9-day time course in IS cells. Histograms display cell count distributions across antibody productivity levels at (a) Day 3, (b) Day 6, and (c) Day 9. Vertical gates define populations by increasing antibody productivity levels: B (lowest), R, MR, M, MW, to WD (highest). The x-axis (mAb X-APC-X*-A) represents antibody surface capture signal intensity detected by cold capture assay.

[0195] In FIG. 17, (a,b) Distribution of barcoded clones across productivity bins at day 12 of fed-batch culture for two independent biological replicates (I1D12 and I2D12). Each bar represents a distinct productivity bin as determined by cold capture assay and FACS analysis, with bins ordered from lowest (left) to highest (right) antibody production. Bar heights indicate the fraction of the total population within each bin.

[0196] Methods

[0197] Plasmid Construction and Library Generation

[0198] Barcoded expression vectors were constructed using a two-step cloning strategy. First, a library of random 20 bp barcodes was generated using Golden Gate assembly tocreate diversity at three tandem barcode positions (60 bp total). The barcode region consisted of three copies of 20 bp random sequences. The initial backbone vector was derived from Addgene plasmid #61883 (trastuzumab expression vector).

[0199] The final 12,383 bp expression vectors were assembled using Gibson assembly to combine 7 fragments: PiggyBac 5' and 3' inverted terminal repeats (ITRs), the triple barcode cassette, sfGFP reporter under CMV promoter, infliximab heavy and light chain genes, IRES2 sequence, and glutamine synthetase (GS) selection marker. Two promoter configurations were generated: IF (rEFla promoter driving heavy chain, SV40 promoter driving light chain) and IS (SV40 promoter driving heavy chain, rEFla promoter driving light chain). All constructs contained an ampicillin resistance gene for bacterial selection and origin or replication.

[0200] Library diversity and correct assembly were verified by Nanopore sequencing (Plasmidsaurus) for structural confirmation. Following successful assembly, the library was transformed and amplified under ampicillin selection before plasmid purification for transfection.

[0201] Cell Culture

[0202] CHO cell lines were maintained in CD CHO Fusion medium (Thermo Fisher Scientific) supplemented with 6 mM GlutaMAX (Thermo Fisher Scientific). CHOZN GS- / - cells (MilliporeSigma) served as the host cell line for transfection experiments. The NIST-CHO reference standard cell line was cultured under identical conditions. All cultures were maintained in 125 mL shake flasks at 37°C with 5% CO2 and orbital shaking at 125 rpm. Routine passaging was performed twice weekly, seeding cells at 0.3 * 106cells / mL.

[0203] For fed-batch culture experiments, cells were seeded at 0.3 x 106cells / mL and cultured for 15 days. Feeding was initiated when viable cell density reached >2 x 106cells / mL (typically day 3-4) using Cellvento ModiFeed Prime COMP (MilliporeSigma) following themanufacturer's low feed strategy: 3% (v / v) on days 3 and 7, 5.5% (v / v) on day 5, and 3% (v / v) on days 10 and 13, for a total feed volume of 17.5%. Culture supernatants were collected daily and analyzed immediately using a YSI 2950 Biochemistry Analyzer according to manufacturer's instructions. Glucose concentrations were maintained between 4-8 g / L through supplementation as needed. Viable cell density and viability were determined daily using a CytoSMART cell counter. Cultures were terminated when cell viability dropped below 70% or at day 15.

[0204] Stable Transfection

[0205] CHOZN cells were seeded at 2 x 106cells / mL in 5 mL of CD CHO Fusion medium supplemented with 6 mM GlutaMAX in T25 flasks 24 hours prior to transfection to ensure log-phase growth. DNA-lipid complexes were prepared in 500 pL OptiMEM serum-free medium using TransIT-PRO Transfection Reagent (Mirus Bio). For random integration (RI) conditions, 5 pg of barcoded expression plasmid was combined with 5 pL TransIT-PRO reagent. For PiggyBac-mediated integration (PB) conditions, 5 pg of barcoded plasmid was co-transfected with 2 pg Super PiggyBac transposase plasmid (System Biosciences PB210PA-1) at a 2.5:1 mass ratio using 5 pL TransIT-PRO reagent. Transfection complexes were incubated for 15-20 minutes at room temperature before addition to cultures.

[0206] Selection timing was optimized based on integration method: medium was changed to glutamine-free CD CHO Fusion medium 24 hours post-transfection for RI conditions and 72 hours post-transfection for PB conditions. Transfected cells were centrifuged at 200 x g for 5 minutes and resuspended in glutamine-free medium to initiate metabolic selection through the glutamine synthetase (GS) system.

[0207] Selection cultures were maintained under standard culture conditions (37°C, 5% CO2, 125 rpm orbital shaking) for 3 weeks. Cell viability and density were monitored daily, withfresh selection medium added every 2-3 days or when viability dropped below 50%, maintaining cells between 0.5-2.0 x 106cells / mL. Stable pools were considered recovered when viable cell density exceeded 2 x 106cells / mL with viability consistently above 90% for at least 3 consecutive passages. Upon recovery, barcoded pools were expanded and cryopreserved in CD CHO Fusion medium containing 7% DMSO at 1 x 107cells / vial. No subcloning was performed to maintain maximum barcode diversity for high-throughput screening.

[0208] Copy Number Analysis by Quantitative PCR

[0209] Genomic DNA was extracted using the DNeasy Blood & Tissue Kit (Qiagen) or Zymo and analyzed by quantitative PCR using plasmid-specific primers and single-copy genomic reference primers. Reactions were performed in triplicate using PowerUp SYBR Green Master Mix (Applied Biosystems) on a QuantStudio Real-Time PCR System. Copy number was calculated using the comparative CT method (AACT) with normalization to endogenous reference genes. Results represent the mean of at least three independent DNA extractions with standard deviation.

[0210] Cold Capture Assay and Flow Cytometry Analysis

[0211] Antibody-producing cells were analyzed using a cold capture assay based on Brezinsky et al (6). The assay exploits temperature-dependent reduction in vesicular transport at 4°C, allowing secreted antibodies to accumulate on the cell surface for fluorescent detection.

[0212] For analysis, 5-10 million cells were harvested and processed on ice. Cells were centrifuged at 300 x g for 5 minutes at 4°C, washed once with 1 mb PBS, and resuspended in PBS containing fluorescent secondary antibodies at 1:25 dilution. Primary detection used APC- conjugated goat anti-human IgG Fey specific antibody (Jackson ImmunoResearch, 109-136-170).For chain-specific analysis, AlexaFluor 647-conjugated anti-light chain antibodies were used at 1:25 dilution.

[0213] Following 30-minute incubation at 4°C in darkness, cells were washed twice with PBS (300xg, 5 minutes, 4°C), filtered through a cell strainer, and analyzed immediately by flow cytometry using either a SONY SH800 or BD FACSMelody™ Cell Sorter. Propidium iodide (10 pL of 2 pL PI in 198 pL PBS working solution) was added immediately before analysis for viability discrimination.

[0214] Antibody production levels were quantified based on fluorescence intensity, with populations categorized as APC- (non-producing), APC+ (producing), and APC++ (high- producing) using untransfected CHOZN host cells (negative control) and NIST-CHO reference standard (positive control) to establish gating parameters.

[0215] Antibody Quantification by HPLC

[0216] Antibody titers were quantified using an Agilent 1260 Infinity II LC System equipped with a Protein A column (Bio-Monolith Protein A, 4.95 x 5.2 mm, Agilent). The mobile phase consisted of 10 mM phosphate buffer with 100 mM sodium chloride (pH 7.2) for binding and 100 mM citric acid (pH 2.5) for elution. Detection was performed at 280 nm. Harvested cell culture supernatants were filtered through 0.22 pm filters prior to analysis. Antibody concentrations were determined using external calibration with purified human IgG standards.

[0217] Amplicon PCR and library preparation

[0218] Barcoded regions from genomic DNA were amplified for next-generation sequencing using a two-stage quantitative PCR (qPCR) protocol, adapted from Leeper et al (2021)(1)

[0219] The initial amplification (PCR1) reactions contained lx Forget-me-not master mix (low ROX) (Biotium), 0.18 pM each of primers trSBS3 and trSBS9, 0.06 pM each of a forward and reverse index primer, and genomic DNA template at a final concentration not exceeding 2.5 ng / pL. Thermocy cling was initiated with a denaturation step at 95°C for 2 minutes, followed by amplification cycles of 95°C for 5 seconds, 62°C for 12 seconds (with fluorescence acquisition), and 72°C for 20 seconds. The reaction was terminated once the change in normalized reporter signal (ARn) reached a value between 0.1 and 1.0. Following amplification, a minimum of 2 pL from each reaction was combined with other uniquely dual-indexed PCR1 products, and the resulting pool was diluted 10-fold.

[0220] The second amplification (PCR2) was performed in 20 pL reactions containing lx Forget-me-not master mix (low ROX), 0.5 pM each of primers P5 and P7, and 0.008 pM each of a forward and reverse index primer. The diluted PCR1 pool constituted 20% (v / v) of the reaction volume, serving as the template. Thermocycling conditions were 95°C for 2 minutes, followed by cycles of 95°C for 5 seconds, 64°C for 15 seconds (with fluorescence acquisition), and 72°C for 12 seconds, again proceeding until the ARn reached the 0.1-1.0 range. Upon completion of PCR2, all reaction products were pooled and purified using a Zymo DNA Clean & Concentrator kit, with the final library eluted in 50 pL of elution buffer.

[0221] Library Quality Control and Sequencing

[0222] The concentration of the purified library was quantified via fluorometry (Qubit), and the library's fragment size was validated by electrophoresis on a 4% E-Gel EX agarose gel (Thermo Fisher). The gel was imaged using an Azure Biosystems cSeries instrument to confirm the correct product size. The quality-controlled amplicon library was then sequenced on anIllumina NovaSeq X Plus platform using a 2x150 bp paired-end configuration. To improve base-calling accuracy for the low-diversity amplicon library, a 20% PhiX control library was spiked in prior to sequencing. Following the run, raw sequencing data was demultiplexed based on the unique dual indices to generate FASTQ files for subsequent computational analysis.

[0223] Bioinformatics Analysis of Barcode Sequences

[0224] Raw paired-end sequencing data were processed using a custom pipeline on a high-performance computing cluster. Read pairs were merged using BBMerge (BBTools v38.90) (2) to generate consensus sequences spanning the complete barcode region. Two-stage demultiplexing was performed with Cutadapt v3.4 (3) : first using sample-specific indices, then extracting the 60 bp barcode cassettes using conserved flanking sequences:5' : CCCTAGAAAGATAGTCTGCGTAAAATTGACGCATG3': GGGGTAGGCGTGTACGGTGGGAGGCCTATA

[0225] Only sequences containing complete barcodes (64-68 bp including minimal flanking regions) were retained.

[0226] Barcode clustering and error correction were performed using Star code vl.4 (102) with a Levenshtein distance threshold of 8 to account for sequencing errors while maintaining discrimination between distinct barcodes. This threshold was selected based on the expected error rate of <5% across the 60 bp barcode region.

[0227] Count normalization was performed in R v4.2.2 using counts per million (CPM) to account for sequencing depth variation. For longitudinal tracking, barcodes present at >10 reads in at least one timepoint were retained. Shannon entropy was calculated to assess population diversity. Barcodes were filtered to remove sequences with >2 ambiguous nucleotides, extreme GC content (<20% or >80%), or homopolymer runs >8 bp. Data visualization employed ggplot2 and ComplexHeatmap (103)packages with loglO-transformation to capture the 3-4 order of magnitude range in barcode frequencies.

[0228] Bayesian Integration of Barcode Frequency and Productivity Data

[0229] To link individual barcode identities to productivity phenotypes, barcode sequencing data from FACS-sorted productivity bins were integrated using Bayesian inference.The probability that a specific barcode belonged to a given productivity bin was calculated usingBayes' theorem:P(barcode\bin) P(biri)P (bin\barcode)P (barcode) where P(bin\barcode) represents the posterior probability that a barcode belongs to a specific productivity bin, P(barcode\bin) is the frequency of the barcode within that bin as determined by amplicon sequencing, P(bin) is the prior probability of the bin based on the proportion of cells in each productivity gate from flow cytometry analysis, and P(barcode) is the overall frequency of the barcode across all bins.

[0230] This approach enabled assignment of productivity phenotypes to individual barcoded lineages by incorporating both the relative abundance of each barcode within sorted populations and the overall population distribution across productivity gates. Barcodes with probability scores >0.5 for a specific bin were assigned to that productivity category for downstream analysis.

[0231] Statistical Analysis

[0232] Data analysis and visualization were performed using GraphPad Prism software. All experiments were conducted with biological triplicates (n=3), representing independent cultures maintained simultaneously under identical conditions. Data are presented as mean ± standard deviation unless otherwise indicated.

[0233] For flow cytometry analysis, samples were collected from each biological replicate and analyzed individually. For time-course experiments, measurements were taken from each replicate culture at specified timepoints throughout the culture duration.

[0234] Barcode abundance quantification was normalized using beads to enable absolute quantification across samples and timepoints. Barcode frequency data underwent log transformation for visualization and comparative analysis. No formal statistical significance testing was performed; all comparisons are descriptive.

[0235] Sequence of Infliximab-S-Construct — bar code sequence portion is shown below in bold type and larger font.

[0237] Sequence of Infliximab-F-construct - bar code sequence portion is shown in bold type and referred to above.

[0238] cgcccgccccacgacccgcagcgcccgaccgaaaggagcgcacgaccccatgcatcgaacaaaagaaaag gggactggaagggctaattcactcccaacgaagacaagatatcataacttcgtatagcatacattatacgaagttatcggctagccttttccccgtAdditional Infliximab-S-base-construct - restriction site shown in bold italicsAdditional Infliximab-F-base-construct - restriction site shown in bold italicsReferences for Example 21. K. Leeper, K. Kalhor, A. Vernet, A. Graveline, G. M. Church, P. Mali, R. Kalhor, Lineage barcoding in mice with homing CRISPR. NatProtoc 16, 2088-2108 (2021).2. B. Bushnell, BBMap: A Fast, Accurate, Splice-Aware Aligner. Lawrence Berkeley National Laboratory v38.90 [Preprint] (2014). https: / / sourceforge.net / projects / bbmap / .3. M. Martin, Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet J 17, 10 (2011).4. E. Zorita, P. Cusco, G. J. Filion, Starcode: sequence clustering based on all-pairs search.Bioinformatics 31, 1913-1919 (2015).5. Z. Gu, R. Eils, M. Schlesner, Complex heatmaps reveal patterns and correlations in multidimensional genomic data. Bioinformatics 32, 2847-2849 (2016).

[0239] Example 3: Additional biologies

[0240] Additional biologies were evaluated. For each biologic, polyclonal pools of CHO cells expressing the molecule were cultured as described in Example 2 above with infliximab. In this case, three flasks were cultured for each molecule, containing a mixture of lineages expressing a single BOI (“E” for etanercept, “I” for infliximab, or “T” for trastuzumab) from a mixture of several vector designs (“R”, “S”, “F”). For example, etanercept polyclonal pools ER and ES were cultured together in the same three flasks.

[0241] Analysis was conducted with infliximab as described above in Example 2.Briefly, cells were cultured in shake flasks, stained every three days using a fluorescently labeled F(ab’)2 antibody fragment, and sorted into several bins based on their signal intensity. DNA was extracted from the sorted cells, and amplicon sequencing was used to measure the abundance of each barcode in the sorted samples. Computational analysis of the sequencing data enabled us to compare productivity across several different pools, molecules, and vector designs. Analysis is shown in FIG. 18. Each line corresponds to a single barcode, and the color corresponds to a vector design. Left figure: linear scale of productivity over time. Right figure: log scale of productivity over time.

[0242] Example 4: Additional biologies

[0243] Analysis conducted with infliximab as described in Example 2 above. Briefly, bulk cell culture samples were taken daily over the course of a fed-batch culture and their DNAextracted for amplicon sequencing of the DNA barcodes. Coupled with growth curve data, barcode abundance was used to generate growth curves for each lineage. Results are shown in FIG. 19 where each line corresponds to a single barcode, and the color corresponds to a vector design. Left figure: linear scale of partial VCD (million cells / mL) over time. Right figure: log scale of partial VCD over time.

[0244] In FIG. 20, sequencing data from productivity and growth analyses were combined to determine the expected titer and other performance attributes for each lineage. Each dot or line corresponds to a barcode, and each color corresponds to a vector design.

[0245] In FIG. 21, comparison of estimated growth rate parameter for each barcode from two separate infliximab fed-batch culture runs. As expected, estimates are tightly centered around 1, the typical doubling time for CHO cells, (right) Comparison of estimated carrying capacity parameter for each barcode from two separate infliximab fed-batch culture runs. Each lineage generally aligns with the carrying capacity of the overall polyclonal pool.

[0246] In FIG. 22 the initial abundance of each barcode from two separate infliximab fed-batch culture runs, (left) linear scale comparing initial abundances, (right) log scale comparing initial abundances.

[0247] In FIG. 23, correlation of average productivity across two separate infliximab runs is shown. More abundant species (yellower and larger circles) tend to be more highly correlated than less abundant species.

[0248] In FIG. 24, correlation of titer calculations across two separate infliximab runs is shown. More abundant species (yellower and larger circles) tend to be more highly correlated than less abundant species. The diagonal line represents the line y=x.

[0249] In FIG. 25, rank correlation of clones across two separate infliximab runs overlaid on kernel density estimates of data is shown. More abundant species (yellower and larger circles) tend to be more highly correlated than less abundant species. Additionally, higher and lower ranks tend to be more dense and more correlated than middle ranks.OTHER EMBODIMENTS

[0250] From the foregoing description, it will be apparent that variations and modifications may be made to the disclosure described herein to adopt it to various usages and conditions. Such embodiments are also within the scope of the following claims.

[0251] All citations to sequences, patents and publications in this specification are herein incorporated by reference to the same extent as if each independent patent and publication was specifically and individually indicated to be incorporated by reference. By their citation of various references in this document, Applicants do not admit any particular reference is “prior art” to their disclosure.

Claims

1. What is claimed:

1. An assay for isolating cell clones expressing biologic of interest (BOI), the assay comprising:(a) transfecting host cells with a i) vector comprising a barcode; and ii) a vector comprising a nucleic acid sequence encoding a biologic of interest (BOI);(b) selecting clones expressing high levels of the BOI;(c) conducting a bioinformatic analysis based on the selected clones.

2. An assay for analyzing or isolating cell clones expressing biologic of interest (BOI), the assay comprising:(a) transfecting host cells with a i) a vector comprising a barcode; and ii) a vector comprising a nucleic acid sequence encoding a biologic of interest (BOI);(b) culturing transfected cells in a polyclonal pool and using sequencing data including the i) vector barcode to determine performance attributes; and(c) selecting clones expressing high levels of the BOI.

3. An assay for analyzing or isolating cell clones expressing biologic of interest (BOI), the assay comprising:(a) transfecting host cells with a i) a vector comprising a barcode; and ii) a vector comprising a nucleic acid sequence encoding a biologic of interest (BOI);(b) culturing transfected cells in a polyclonal pool and using sequencing data including the i) vector barcode to determine performance attributes; and(c) identifying measured differences in BOI expression or clone performance selecting clones expressing high levels of the BOI.

4. The assay of claim 3 wherein measured differences in BOI expression or clone performance are identified by steps comprising by i) sorting cells into various groups based on markers; ii) performing sequencing that includes the barcode; and iii) calculating the BOI expression or clone performance attributes using the barcode sequencing data.

5. An assay of any one of claims 1 to 4 wherein the i) vector and ii) vector is a single vector comprising a) a barcode and b) a nucleic acid sequence encoding a biologic of interest (BOI).

6. An assay of any one of claims 1 to 4 wherein host cells are transfected with a i) vector comprising a barcode; ii) a vector comprising a nucleic acid sequence encoding a biologic of interest (BOI); and iii) a vector comprising a gene editing complex wherein the gene editing complex induces or represses expression of a reporter gene vectors.

7. An assay of claim 6 wherein the i) vector, ii) vector and iii) vector is a single vector comprising a) a barcode, b) a nucleic acid sequence encoding a biologic of interest (BOI), and c) a gene editing or targeting complex wherein the gene editing or targeting complex induces or represses expression of a reporter gene.

8. An assay of clam 6 or 7 wherein the i) vector, ii) vector and iii) vector are two, three or more separate vectors.

9. An assay of any one of claims 1 to 8 comprising multiple transfecting steps.

10. An assay of any one of claims 1 to 8 wherein clones are selected that express high levels of the BOI as compared to a control.

11. An assay for isolating cell clones expressing biologic of interest (BOI), comprising:(a) generating a vector library comprising a vector comprising i) a barcode and ii) a nucleic acid sequence encoding a biologic of interest (BOI);(b) transfecting host cells with the vector comprising a barcode obtained from the vector library, facilitating tracking of each host cell comprising the barcodes;(c) selecting stable clones expressing high levels of the BOI;(d) transfecting the clones with a vector comprising a gene targeting complex wherein the gene targeting complex induces or represses expression of a reporter gene;(e) conducting a bioinformatic analysis based on the clones.

12. An assay for isolating cell clones expressing a biologic of interest (BOI), comprising:(a) generating a vector library comprising one or more vectors comprising a barcode, a nucleic acid sequence encoding a biologic of interest (BOI), regulatory elements, and a reporter gene;(b) transfecting host cells with vectors comprising a barcode obtained from the vector library, facilitating tracking of each host cell comprising the barcodes;(c) selecting stable clones expressing high levels of the BOI, as compared to a baseline control;(d) transfecting the clones with a vector comprising a gene targeting complex wherein the gene targeting complex induces or represses expression of the reporter gene;(e) selecting the clones expressing the reporter gene; and,(f) conducting a bioinformatic analysis based on the selected clone.

13. The assay of any one of claims 1 to 12 wherein each of the barcodes comprises at least 10 nucleotides up to 150 nucleotides.

14. The assay of any one of claims 1 to 12 wherein each of the barcodes comprises at least 10 nucleotides up to 40 nucleotides.

15. The assay of any one of claims 1 to 12 wherein each of the barcodes comprises at least 20 nucleotides up to 100 nucleotides.

16. The assay of any one of claims 1 to 12 wherein each of the barcodes comprises at least 15 nucleotides up to 35 nucleotides.

17. The assay of any one of claims 1 to 12 wherein the barcodes comprise at about 21 or 67 nucleotides.

18. The assay of any one of claims 1 to 17 wherein the nucleic acid sequences encoding the BOIs of interest are inserted into the cell genome.

19. The assay of any one of claims 1 to 18 wherein the genome of the clones expressing the reporter gene is sequenced.

20. The assay of claim 19 wherein the sequencing is targeted to a nucleic acid sequence between a 5’restriction site and a 3’ restriction site and includes the barcode sequences.

21. The assay of claim 20 wherein the gene editing complex comprises guide RNAs (gRNAs) specifically targeting each barcode sequence.

22. The assay of any one of claims 1 to 21 wherein the gene editing complex is comprised in a vector separate from vector(s) comprising a barcode component and a nucleic acid sequence encoding a biologic of interest (BOI).

23. The assay of any one of claims 1 to 21 wherein the gene editing complex is comprised in a vector together with a barcode component and a nucleic acid sequence encoding a biologic of interest (BOI) VECTOR plasmid distinct to the plasmids in a vector library.

24. The of any one of claims 1 to 23 wherein the gene editing complex comprises one or more single guide RNAs (sgRNAs), a nuclease and a transcriptional activator.

25. The assay of any one of claims 1 to 21 wherein the gene editing complex comprises a clustered regularly interspaced short palindromic repeat (CRISPR)-mediated transcriptional activation (CRISPRa) system.

26. The assay of claim 25 wherein the CRISPRa system comprises a nuclease.

27. The assay of claim 26 wherein the nuclease is a catalytically inactive nuclease.

28. The assay of claim 27 wherein the catalytically inactive nuclease is Cas9.

29. The assay of claim 28 wherein the catalytically inactive Cas9 (dCas9) nuclease is fused to a transcriptional activator.

30. The assay of any one of claims 21 to 29 wherein the gRNAs are specific for each desired barcode.

31. The assay of claim 30 wherein sgRNAs guide the CRISPRa / dCas9 to the target site, thereby activating transcription and inducing expression of the reporter gene.

32. The assay of any one of claims 1 to 31 wherein each vector comprises one or more transcriptional regulators.

33. The assay of claim 32 wherein the one or more transcriptional regulators regulate expression of the BOIs.

34. A vector library comprising vectors comprising a barcode, a nucleic acid sequence encoding a biologic of interest (BOI), regulatory elements; a gene editing complex, a reporter gene or combinations thereof.

35. The vector library of claim 34 wherein the barcodes comprise at least 10 nucleotides up to 150 nucleotides.

36. The vector library of claim 34 wherein the barcodes comprise at least 10 nucleotides up to 40 nucleotides.

37. The vector library of claim 34 wherein the barcodes comprise at least 20 nucleotides up to 10 nucleotides.

38. The vector library of claim 34 wherein the barcodes comprise at least 15 nucleotides up to 35 nucleotides.

39. The vector library of claim 34 wherein the barcodes comprise at about 21 or 67 nucleotides.

40. The vector library of any one of claims 34 to 39 wherein the gene editing complex comprises guide RNAs (gRNAs) specifically targeting each barcode sequence.

41. The vector library of any one of claims 34 to 40 wherein each vector comprises one or more transcriptional regulators.

42. The vector library of claim 41 wherein the one or more transcriptional regulators regulate expression of the BOIs.

43. The vector library of any one of claims 34 to 42 wherein the gene editing complex comprises one or more single guide RNAs (gRNAs), a nuclease and a transcriptional activator.

44. The vector library of any one of claims 34 to 43 wherein the gene editing complex is comprised in a vector separate from vector(s) comprising a barcode component and a nucleic acid sequence encoding a biologic of interest (BOI).

45. The vector library of any one of claims 34 to 44 wherein the gene editing complex is comprised in a vector together with a barcode component and a nucleic acid sequence encoding a biologic of interest (BOI) VECTOR plasmid distinct to the plasmids in a vector library.

46. The vector library of any one of claim 40 to 45 wherein the nuclease is a catalytically inactive nuclease.

47. The vector library of claim 46 wherein the catalytically inactive nuclease is Cas9.

48. The vector library of claim 47 wherein the catalytically inactive Cas9 (dCas9) nuclease is fused to a transcriptional activator.

49. The vector library of any one of claims 43 to 48 wherein the gRNAs are specific for each desired barcode.

50. The vector library of any one of claims 34 to 49 wherein the gene editing complex is comprised in a vector distinct to the other vectors in the vector library.

51. The vector library of any one of claims 34 to 50 wherein the gene editing complex is comprised in the same plasmid comprising a barcode, a nucleic acid sequence encoding a biologic of interest (BOI), regulatory elements, a reporter gene or combinations thereof.

52. A vector comprising a barcode, a nucleic acid sequence encoding a biologic of interest (BOI), regulatory elements; gene editing complex, a reporter gene or combinations thereof.

53. A vector comprising a barcode, a nucleic acid sequence encoding a biologic of interest (BOI), regulatory elements, a reporter gene or combinations thereof.

54. A host cell comprising a vector of any one of claims 34 to 52.

Citation Information

Patent Citations

  • Cell sorting

    US20180127745A1

  • Compositions and methods for producing and characterizing viral vector producer cells for cell and gene therapy

    US20240167055A1