Randomized configuration targeted integration of nucleic acids
The RCTI strategy efficiently identifies host cells for expressing recombinant proteins by simultaneous integration and screening, addressing resource and time challenges in conventional cell line development.
Patent Information
- Application Number
- JP2025073331
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-03-19
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-20
AI Technical Summary
Conventional methods for developing cell lines to express recombinant proteins, such as monoclonal antibodies and bispecific antibodies, are resource-intensive and time-consuming, often failing to identify optimal clones due to random integration methods or requiring extensive screening in targeted integration approaches.
A 'randomized constitutive targeted integration' (RCTI) strategy that allows for high-throughput screening of host cells by introducing multiple sequences of interest simultaneously, using recombinase-mediated or gene editing-mediated integration, and selecting cells based on desirable expression and product quality attributes.
The RCTI method reduces resource and time requirements while maintaining or improving expression levels, product quality, and production culture performance compared to traditional strategies.
Smart Images

Figure 2025121945000001_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 62 / 866,893, filed June 26, 2019, and U.S. Provisional Application No. 62 / 991,708, filed March 19, 2020, the contents of each of which are incorporated herein by reference in their entirety and to which priority is claimed. [Technical Field]
[0002] The subject matter of this disclosure relates to a "randomized constitutive targeted integration" (also referred to herein as "randomized strand targeted integration") (RCTI) strategy for generating and identifying host cells capable of expressing recombinant proteins, such as monoclonal antibodies, and compositions derived therefrom, such as bispecific antibodies, and other complex format proteins, such as membrane protein complexes, and other difficult-to-express molecules.
[0003] Sequence Listing This specification references a Sequence Listing (submitted electronically as a .txt file named "00B206_0926_SL.txt" on June 26, 2020). The 00B206_0926_SL.txt file was created on June 24, 2020 and is 1,296,777 bytes in size. The entire contents of the Sequence Listing are incorporated herein by reference. [Background technology]
[0004] Rapid advances in cell biology and immunology have fueled the demand for developing novel therapeutic recombinant proteins, such as monoclonal antibodies, bispecific antibodies, and proteins in complex formats, for various diseases, including cancer, cardiovascular disease, and metabolic disease. These biopharmaceutical candidates are generally produced by commercially available cell lines capable of expressing the protein of interest. For example, Chinese hamster ovary (CHO) cells have recently been widely adapted to produce therapeutic monoclonal or bispecific antibodies as well as proteins in more complex formats.
[0005] Conventional strategies for developing commercial cell lines generally involve iterative efforts to integrate nucleotide sequences encoding a polypeptide of interest, either randomly or at specific (“targeted”) locations, followed by selection and isolation of cell lines that produce that polypeptide. However, these approaches have their own inherent drawbacks. While random integration methods offer the possibility of obtaining different clones with different compositions and ratios of transgenes targeted for expression, screening hundreds or thousands of clones after transfection to isolate cell lines exhibiting desirable expression levels, product quality attributes, and production culture performance is time-consuming and resource-intensive. Furthermore, clone(s) with optimal ratios of various transgenes may not exist or may not be isolated due to instability or the inability to screen a sufficient number of clones. Furthermore, when producing multi-chain polypeptides, such as monoclonal antibodies, additional screening may be required to address the number and arrangement of nucleic acids encoding such polypeptides. When using random integration methods, it is difficult to determine the number and location of all transgenes introduced into the genome as well as their arrangement, which is an important step in creating a reasonable correlation between transgene arrangement / copy number and desired product attributes.
[0006] In contrast, targeted integration approaches, where transgene(s) are inserted at specific, predetermined locations within the genome, can provide more control over the placement and copy number of each specific transgene. However, such targeted integration approaches are designed to minimize the possibility of generating random combinations and ratios of different transgenes. Therefore, if the designed placement and copy number of transgenes are accidentally suboptimal, all of the resulting clones will similarly have suboptimal product quality and / or titer. To increase the chances of expressing complex or difficult-to-express molecules using targeted integration approaches, multiple different transgene placements and copy number configurations are individually tested, resulting in increased cell line development (CLD) workload and resource requirements. Therefore, there is a need in the art for new cell line development strategies that conserve resources while generating cell lines that exhibit expression levels, product quality attributes, and production culture performance comparable to traditional methodologies. Summary of the Invention
[0007] The subject matter of this disclosure relates to a "randomized constitutive targeted integration" (also referred to herein as "randomized strand targeted integration") (RCTI) strategy for generating and identifying host cells capable of expressing recombinant proteins, such as monoclonal antibodies, and compositions derived therefrom, such as bispecific antibodies, and other complex format proteins, such as membrane protein complexes, and other difficult-to-express molecules. The RCTI method described in this disclosure can be used to screen the same number of vector constructs, if not more, than standard cell line development methods, while using fewer resources. By using the RCTI method described in this disclosure, fewer individual clones can be screened as well.
[0008] In certain embodiments, the present disclosure provides methods for generating and high-throughput screening a library of targeted integration (TI) host cells expressing at least one sequence of interest (SOI), the method comprising: a) generating a library of TI host cells comprising a plurality of TI host cells expressing one or more SOIs by: i) providing a plurality of TI host cells; ii) contacting the plurality of TI host cells with a plurality of vectors comprising one or more SOIs; and iii) introducing the one or more SOIs into one or more of the plurality of TI host cells; b) isolating the library into single clones; and c) screening the clones for specific cellular or product attributes.
[0009] In certain embodiments, the present disclosure provides a method for generating a library of TI host cells comprising a plurality of exogenous nucleotide SOIs, the method comprising: a) providing a plurality of TI host cells; b) contacting the plurality of TI host cells with a plurality of vectors comprising one or more SOIs; and c) introducing the one or more SOIs into one or more of the plurality of TI host cells. In certain embodiments, one or more of the plurality of TI host cells comprises one or more exogenous nucleotide sequences integrated into one or more loci in the genome of the TI host cell, the exogenous nucleotide sequences comprising at least two recombinase recognition sequences (RRSs) flanking at least one first selectable marker. In certain embodiments, the vector comprises: a) at least two RRSs that match at least two RRSs on the integrated exogenous nucleotide sequences; and b) one or more exogenous SOIs and at least one second selectable marker flanking the RRSs.
[0010] The presently disclosed subject matter also provides methods for the targeted integration of one or more exogenous nucleic acids into a host cell to promote expression of a sequence of interest. In certain embodiments, such methods involve targeted integration of one or more exogenous nucleic acids into a host cell by recombinase-mediated integration or gene editing-mediated integration. In certain embodiments, such methods involve a cell comprising one or more exogenous nucleotide sequences integrated into a locus in the genome of the host cell, the locus comprising a nucleotide sequence that is at least about 90% homologous to a sequence selected from SEQ ID NOs: 1-12.
[0011] In certain embodiments, the present disclosure provides a method for preparing TI host cells expressing one or more SOIs, the method comprising: a) providing a plurality of TI host cells; b) contacting the plurality of TI host cells with a plurality of vectors comprising one or more SOIs; c) introducing the one or more SOIs into one or more of the plurality of TI host cells; and d) selecting TI host cells expressing the one or more SOIs. In certain embodiments, one or more of the plurality of TI host cells comprises one or more exogenous nucleotide sequences integrated into one or more loci in the genome of the TI host cell, the exogenous nucleotide sequences comprising at least two RRSs flanking at least one first selectable marker. In certain embodiments, the vector comprises: a) at least two RRSs that match at least two RRSs on the integrated exogenous nucleotide sequence; and b) one or more exogenous SOIs and at least one second selectable marker flanking the RRSs. In certain embodiments, the method includes introducing one or more recombinases or nucleic acids encoding recombinases that recognize an RRS into one or more of the plurality of TI host cells, and d) selecting TI cells that express a second selection marker, thereby isolating TI host cells that express a sequence of interest. In certain embodiments, the exogenous nucleotide sequence includes a first and a second RRS flanking at least one first selection marker and a third RRS located between the first and second RRS, all of the RRSs being heterospecific; the plurality of vectors includes: i. a first vector that includes at least one first exogenous SOI that matches the first and third RRSs on the integrated exogenous nucleotide sequence and two RRSs that are flanked by at least one second selection marker; ii. a second vector that includes two RRSs that match the second and third RRSs on the integrated exogenous nucleotide sequence and are flanked by at least one second exogenous SOI; and selecting TI cells that express the second selection marker, thereby isolating TI host cells that express the first and second sequences of interest.
[0012] In certain embodiments, the present disclosure provides a method for expressing an SOI, the method comprising: a) providing a plurality of TI host cells; b) contacting the plurality of TI host cells with a plurality of vectors comprising one or more SOIs; c) introducing the one or more SOIs into one or more of the plurality of TI host cells; d) selecting TI host cells that express a sequence of interest; and e) culturing the cells of d) under conditions suitable for expressing the sequence of interest.
[0013] In certain embodiments, the present disclosure provides a method for expressing an SOI, the method comprising: a) providing a plurality of TI host cells, each TI host cell comprising one or more exogenous nucleotide sequences integrated into one or more loci in the genome of the TI host cell, the exogenous nucleotide sequences comprising at least two RRSs flanking at least one first selectable marker; b) contacting the plurality of TI host cells with a plurality of vectors comprising: a. at least two RRSs that match the at least two RRSs on the integrated exogenous nucleotide sequences; and b. one or more exogenous SOIs and at least one second selectable marker flanking the RRSs; c) introducing one or more recombinases or nucleic acids encoding recombinases that recognize the RRSs; d) selecting TI cells that express the second selectable marker, thereby isolating TI host cells that express the sequence of interest; and e) culturing the cells of d) under conditions suitable for expression of the sequence of interest, and recovering the product of the sequence of interest therefrom. In certain embodiments, the exogenous nucleotide sequence comprises a first and a second RRS flanking at least one first selectable marker and a third RRS located between the first and second RRS, all of which are heterospecific; the plurality of vectors comprises: i. a first vector comprising at least one first exogenous SOI and two RRSs flanking at least one second selectable marker that match the first and third RRSs on the integrated exogenous nucleotide sequence; and ii. a second vector comprising two RRSs flanking at least one second exogenous SOI that match the second and third RRSs on the integrated exogenous nucleotide sequence, and TI cells expressing the second selectable marker are selected, thereby isolating TI host cells expressing the first and second sequences of interest.
[0014] In certain embodiments above, the one or more SOIs are introduced into one or more of the plurality of TI host cells by recombinase-mediated integration. In certain embodiments above, the one or more SOIs are introduced into one or more of the plurality of TI host cells by gene editing-mediated integration.
[0015] In certain of the above embodiments, the one or more SOIs are operably linked to one or more regulatable promoters, wherein the one or more regulatable promoters are selected from the group consisting of SV40 and CMV promoters.
[0016] In certain embodiments of the above, the TI host cell is a mammalian host cell. In certain embodiments of the above, the TI host cell is a hamster host cell, a human host cell, a rat host cell, or a mouse host cell. In certain embodiments of the above, the TI host cell is a CHO host cell, a CHO K1 host cell, a CHO K1SV host cell, a DG44 host cell, a DUKXB-11 host cell, a CHOK1S host cell, or a CHO K1M host cell.
[0017] In certain of the above embodiments, the SOI encodes a polypeptide subunit of a multi-subunit protein or a fragment thereof, hi certain of the above embodiments, the SOI encodes a single chain antibody, an antibody light chain, an antibody heavy chain, a single chain Fv fragment (scFv), or an Fc fusion protein.
[0018] In certain embodiments of the above, the particular cell attribute is selected from cell growth, cell titer, specific productivity, volumetric productivity, and clonal stability. In certain embodiments of the above, the particular product attribute is selected from level of glycosylation, level of charge dispersion, reduced mismatch, reduced protein / peptide aggregation, and protein sequence heterogeneity.
[0019] In certain embodiments of the above, one or more SOIs are introduced into one or more cells of the plurality of TI host cells at one or more loci that are at least about 90% homologous to a sequence selected from the following: SEQ ID NOs: 1-12; NW_006874047.1; NW_006884592.1; NW_006881296.1; NW_003616412.1; NW_003615063.1; NW_006882936.1; and NW_003615411.1.
[0020] In certain embodiments of the above, the at least one sequence of interest comprises 2, 3, 4, 5, 6, 7, 8, 9, 10 sequences of interest. [Brief explanation of the drawings]
[0021] [Figure 1A-1B] 1A-1B are schematic diagrams illustrating an exemplary targeted integration (TI) cell line development (CLD) workflow. [Figure 2A-2B] 2A-2B are schematic diagrams illustrating a comparison of standard TI and "randomized construct targeted integration" (also referred to herein as "randomized strand targeted integration") (RCTI) workflows. [Figure 3A-3B] 3A-3B show the distribution of titers and HMWS (%) of clones evaluated in production cultures, as well as the identified constructs. [Figure 4] Figure 4 shows other product quality attributes that are comparable between the standard TI CLD approach and the RCTI approach. [Figure 5] Figure 5 shows the comparable product quality attributes between the standard CLD approach and the RCTI approach. [Figures 6A-6C] Figures 6A-6C show that the RCTI method improves timeline flexibility while maintaining clone performance. [Figures 7A-7B] 7A-7B show that although titers are comparable between the standard CLD and RCTI approaches for molecule Y, the specific productivity of clones from the RCTI approach is higher than that of standard CLD. [Figures 8A-8G] 8A-8G show comparable product quality attributes between the standard CLD approach and the RCTI approach for molecule Y. [Figure 9A-9B] 9A-9B show that while titers are comparable between the standard CLD and RCTI approaches for molecule Z, the specific productivity of clones from the RCTI approach is higher than that of standard CLD. [Figures 10A-10F]10A-10F show comparable product quality attributes between the standard CLD approach and the RCTI approach for molecule Z. [Figures 11A-11C] 11A-11C are schematic diagrams showing exemplary molecules expressed by the methods of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0022] In certain embodiments, the host cells, genetic constructs (e.g., vectors), compositions, and methods described herein can be used in the development and / or use of "randomized constitutive targeted integration" (also referred to herein as "randomized strand targeted integration") (RCTI) strategies for more efficient (e.g., less time- and / or resource-intensive) identification of host cells that exhibit expression levels, product quality attributes, and production culture performance comparable to conventional methodologies.
[0023] For clarity of disclosure, and not by way of limitation, the detailed description is divided into the subsections that follow. 1.Definition 2. Host Cell Preparation and Screening Strategy 3. Exogenous Nucleotide Sequence 4.Host cells 5. Targeted integration 6. Products
[0024] 1.Definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. In case of conflict, the present application, including any definitions herein, will control. Preferred methods and materials are described below, although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the presently disclosed subject matter. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. Additionally, the materials, methods, and examples are illustrative only and are not intended to be limiting.
[0025] As used herein, "comprise(s)," "include(s)," "having," "has," "can," "contain(s)," and variations thereof are intended to be open-ended transitional phrases, terms, or words that do not exclude the possibility of additional acts or constructs. The singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. The present disclosure also contemplates other embodiments that "comprising," "consisting of," and "consisting essentially of" the embodiments or elements presented herein, whether explicitly stated or not.
[0026] For the recitation of numerical ranges herein, each intervening number is expressly contemplated to the same degree of precision. For example, in the range 6 to 9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and in the range 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are expressly contemplated.
[0027] As used herein, the term "about" or "approximately" refers to an acceptable error range for a particular value, as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within 3 or more standard deviations, as is customary in the art. Alternatively, "about" can mean within 20%, preferably 10%, more preferably 5%, and even more preferably 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5-fold, and more preferably within 2-fold of a value.
[0028] As used herein, the term "selection marker" refers to a gene that allows cells carrying a gene to be specifically selected or specifically eliminated in the presence of a corresponding selection agent. For example, but not limited to, a selection marker can allow host cells transformed with the selection marker gene to be positively selected in the presence of the gene, while untransformed host cells cannot grow or survive under the selective culture conditions. A selection marker can be positive, negative, or bifunctional. A positive selection marker can allow for the selection of cells carrying the marker, while a negative selection marker can allow for the selective elimination of cells carrying the marker. A selection marker can confer resistance to a drug in a host cell or complement a metabolic or catabolic defect. In prokaryotic cells, genes that confer resistance to ampicillin, tetracycline, kanamycin, or chloramphenicol, among others, can be used. Resistance genes useful as selectable markers in eukaryotic cells include, but are not limited to, genes for aminoglycoside phosphotransferases (APH) (e.g., hygromycin phosphotransferase (HYG), neomycin, and G418 APH), dihydrofolate reductase (DHFR), thymidine kinase (TK), glutamine synthetase (GS), asparagine synthetase, tryptophan synthase (indole), histidinol dehydrogenase (histidinol D)), and genes encoding resistance to puromycin, blasticidin, bleomycin, phleomycin, chloramphenicol, zeocin, and mycophenolic acid. Additional marker genes are described in WO 92 / 08796 and WO 94 / 28143.
[0029] Selectable markers not only facilitate selection in the presence of the corresponding selection agent, but can also alternatively provide genes encoding molecules not normally present in cells, such as green fluorescent protein (GFP), enhanced GFP (eGFP), synthetic GFP, yellow fluorescent protein (YFP), enhanced YFP (eYFP), cyan fluorescent protein (CFP), mPlum, mCherry, tdTomato, mStrawberry, J-red, DsRed-monomer, mOrange, mKO, mCitrine, Venus, YPet, Emerald, CyPet, mCFPm, Cerulean, and T-Sapphire. Cells carrying such genes can be distinguished from cells not carrying the gene by, for example, detection of fluorescence emitted by the encoded polypeptide.
[0030] As used herein, the term "operably linked" refers to the juxtaposition of two or more components, in a relationship permitting them to function in a desired manner. For example, a promoter and / or enhancer is operably linked to a coding sequence if it functions to regulate the transcription of the coding sequence. In certain embodiments, DNA sequences that are "operably linked" are contiguous and adjacent on a single chromosome. In certain embodiments, where it is necessary to join coding regions for two proteins, such as a secretory leader and a polypeptide, these sequences are contiguous, adjacent, and in reading frame. In certain embodiments, an operably linked promoter may be located upstream of and adjacent to the coding sequence. In certain embodiments, for example, with respect to an enhancer sequence that regulates expression of a coding sequence, two components may be operably linked, yet not adjacent. An enhancer is operably linked to a coding sequence if the enhancer increases the transcription of the coding sequence. An operably linked enhancer can be located upstream, within, or downstream of the coding sequence, and can be located a considerable distance from the promoter of the coding sequence. Operable linkage can be achieved by recombinant methods known in the art, for example, using PCR methods and / or by ligation at convenient restriction sites. If convenient restriction sites are not present, synthetic oligonucleotide adapters or linkers can be used in accordance with conventional techniques. An internal ribosome entry site (IRES) is operably linked to an open reading frame (ORF) if it enables translation to be initiated at a location internal to the ORF independently of the 5' end.
[0031] As used herein, the term "expression" refers to transcription and / or translation. In certain embodiments, the level of transcription of a desired product can be determined based on the amount of corresponding mRNA present. For example, mRNA transcribed from a sequence of interest can be quantified by PCR or Northern hybridization. In certain embodiments, the protein encoded by a sequence of interest can be quantified by various methods, such as by ELISA, by assaying for the biological activity of the protein, or by using assays that are independent of such activity, such as Western blotting or radioimmunoassays that use antibodies that recognize and bind to the protein.
[0032] The term "sequence of interest" is used herein to refer to a polypeptide sequence (or, in certain cases, a nucleic acid encoding a polypeptide sequence), the expression of which is of interest. Such a polypeptide sequence may, in certain embodiments, comprise a subunit of a multi-subunit protein complex. In certain embodiments, such a polypeptide sequence may comprise a fragment of such a subunit. Such a polypeptide sequence may, in certain embodiments, comprise an antibody sequence, e.g., an antibody heavy or light chain sequence. In certain embodiments, such a polypeptide sequence may comprise a fragment of such an antibody sequence.
[0033] As used herein, the term "antibody" is used in the broadest sense and encompasses a variety of antibody structures, including, but not limited to, monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), half antibodies, and antibody fragments, so long as the antibody exhibits the desired antigen-binding activity.
[0034] As used herein, a "reference antibody" and a "reference monoclonal antibody" are antibodies or antibody fragments that have a single binding specificity. In certain embodiments, the single binding specificity of a reference antibody is the result of pairing a heavy chain sequence, or fragment thereof, with a light chain sequence, or fragment thereof.
[0035] As used herein, a "bispecific antibody" or "BsAb" is an antibody that can simultaneously bind to two different epitopes, e.g., two different epitopes on two different antigens or two different epitopes on a single antigen. A BsAb is an antibody that binds to two different epitopes simultaneously, with one "arm" (i.e., a pair of Vs) of the BsAb having the binding specificity of a first parent antibody. H and V L ) and paired variable-weight (V) fragments of two different parent monoclonal antibodies resulting in a second "arm" of the BsAb with the binding specificity of the second parent antibody. H ) and light (V L BsAbs encompass a number of different structures, including those containing a β- or β-domain. BsAbs are a subset of multispecific antibodies, which contain at least two binding specificities (i.e., BsAbs), but also include trispecific antibodies and antibodies with greater numbers of specificities.
[0036] As used herein, the term "antibody fragment" refers to a molecule other than an intact antibody that contains a portion of an intact antibody that binds to the antigen to which the intact antibody binds. Examples of antibody fragments include, but are not limited to, Fv, Fab, Fab', Fab'-SH, F(ab')2; diabodies; linear antibodies; single-chain antibody molecules (e.g., scFv); and multispecific antibodies formed from antibody fragments.
[0037] As used herein, the term "variable region" or "variable domain" refers to the domain of an antibody heavy or light chain that is involved in binding the antibody to an antigen. The variable domains of the heavy and light chains of a naturally occurring antibody (V H and V L ) generally have a similar structure, with each domain containing four conserved framework regions (FR) and three hypervariable regions (HVR). See, for example, Kindt et al., Kuby Immunology, 6th ed. W.H. Freeman and Co., p. 91 (2007). A single V H or V LThe V domain may be sufficient to confer antigen-binding specificity. Furthermore, an antibody that binds to a particular antigen may have a V domain that is sufficient to confer antigen-binding specificity. H or V L The complementary V domains were isolated using L or V H Libraries of domains may be screened. See, e.g., Portolano et al., J. Immunol. 150:880-887 (1993); Clarkson et al., Nature 352:624-628 (1991).
[0038] As used herein, the term "vector" refers to a nucleic acid molecule capable of propagating another nucleic acid to which it is linked. The term includes vectors as self-replicating nucleic acid structures as well as vectors that are integrated into the genome of a host cell into which they are introduced. In certain embodiments, vectors direct the expression of nucleic acids to which they are operably linked. Such vectors are referred to herein as "expression vectors."
[0039] As used herein, the term "homologous sequences" refers to sequences that share significant sequence similarity as determined by sequence alignment. For example, two sequences may be approximately 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 99.9% homologous. Alignment is performed by algorithms and computer programs, including but not limited to BLAST, FASTA, HMME, etc., which compare sequences and calculate the statistical significance of matches based on factors such as sequence length, sequence identity and similarity, and the presence and length of sequence mismatches and gaps. Homologous sequences can refer to both DNA sequences and protein sequences.
[0040] As used herein, the term "adjacent" refers to a first nucleotide sequence being located at either the 5' or 3' end, or both ends, of a second nucleotide sequence. The adjacent nucleotide sequence may be located next to the second nucleotide sequence or at a predetermined distance from it. There is no particular limitation on the length of the adjacent nucleotide sequence. For example, the adjacent sequence may be a few base pairs or several thousand base pairs.
[0041] As used herein, the term "exogenous" refers to a nucleotide sequence that is not native to the host cell and is introduced into the host cell by conventional DNA delivery methods, such as transfection, electroporation, or transformation. The term "endogenous" refers to a nucleotide sequence that is derived from the host cell. An "exogenous" nucleotide sequence may have an "endogenous" counterpart that is identical in terms of base composition, but an "exogenous" sequence may be introduced into the host cell by, for example, recombinant DNA techniques.
[0042] As used herein, an "integration site" includes a nucleic acid sequence in a host cell genome into which an exogenous nucleotide sequence is inserted. In certain embodiments, the integration site is located between two adjacent nucleotides on the host cell genome. In certain embodiments, the integration site comprises a stretch of nucleotide sequence. In certain embodiments, the integration site is located within a specific locus in the genome of the TI host cell. In certain embodiments, the integration site is located within an endogenous gene of the TI host cell.
[0043] As used herein, the term "TI host cell" refers to a cell that contains a genomic locus or loci, i.e., integration site(s), for use in expressing a sequence of interest. In certain embodiments, integration of an SOI into a TI host cell is facilitated by the presence of an exogenous nucleotide sequence at another integration site that contains two or more RRSs. In certain embodiments, integration of an SOI into a TI host cell is facilitated by a genome editing system that can edit the TI host cell genome at one or more integration sites.
[0044] 2. Host Cell Preparation and Screening Strategy Standard CLD strategies generally involve multiple rounds of host cell transfection with one or more specific exogenous nucleic acids encoding the sequence(s) of interest. As illustrated in Figure 1A, preparation of such exogenous nucleic acids, e.g., plasmids containing antibody heavy and light chain coding sequences, is followed by transfection, e.g., by recombinase-mediated cassette exchange ("RMCE"), in the context of targeted integration. Regardless of the integration strategy (e.g., random or targeted integration), the pool of transfected host cells is then recovered prior to selection and single-cell cloning (SCC). Multiple rounds of clonal analysis are then used to narrow down the number of clones for more detailed productivity and product quality assays, e.g., high molecular weight species content ("HMWS(%)"), size variation, acidity variation, and glycosylation variation, e.g., via homogeneous time-resolved fluorescence ("HTRF") titer assays.
[0045] Figure 2A illustrates a standard targeted integration-based CLD strategy that uses multiple CLD cycles to identify cells exhibiting desirable expression levels, product quality attributes, and production culture performance. In this example, each pool of transfected cells is prepared by contacting multiple host cells with a specific combination of two exogenous nucleic acids encoding sequences of interest. While Figure 2A illustrates a strategy involving targeted integration of sequences of interest from a first "front" plasmid and a second "back" plasmid into a specific locus within the host genome (see Section 5.1 below for a general description of two-vector RCME strategies), it should be understood that the targeted integration approach can involve integration of a single sequence of interest or more than two sequences of interest. Furthermore, targeted integration strategies, as outlined herein, can employ not only recombinase-mediated SOI integration as exemplified in Figures 2A and 2B, but also other locus-specific strategies for SOI integration, such as gene editing-mediated SOI integration. However, as noted above, a defining feature of targeted integration approaches is that they are designed to result in the integration of specific sequences within the genome, i.e., specific sequences in a fixed configuration of the transgene. In light of this design feature, it is necessary to perform multiple cycles of CLD, as outlined in Figure 2A, to ensure that a sufficient number of variations in copy number and positioning of the sequence of interest are assayed.
[0046] In contrast to the standard, resource-intensive, multi-cycle targeted integration CLD strategy shown in Figure 2A, the subject matter of this application relates to a "randomized constitutive targeted integration" (also referred to herein as "randomized strand-targeted integration") (RCTI) strategy, which allows for single-cycle assays of variations in copy number and positioning of a sequence of interest, while retaining the locus-specific integration benefits of targeted integration. As outlined in Figure 2B, all of the relevant exogenous nucleic acids encoding the sequences of interest can be evaluated by creating a single transfection pool. By combining the full complement of relevant exogenous nucleic acids encoding the sequences of interest with multiple host cells, a single transfection pool is created from which clones exhibiting desirable expression levels, product quality attributes, and production culture performance can be identified.
[0047] While exemplary Figure 2B depicts an embodiment using three "front" and three "back" plasmids (each containing one or more sequences of interest) that integrate into a front or back cassette of an exogenous nucleotide sequence present at a specific locus within the host cell genome (see Section 5.1 below for a full description of such a "two-vector RMCE" strategy), it should be understood that the subject matter of the present disclosure encompasses a wide variety of variations on such targeted integration strategies. For example, rather than using a first plasmid containing a front cassette and a second plasmid containing a back cassette, the present disclosure also encompasses methods involving the integration of a single cassette. Furthermore, the present disclosure also encompasses methods involving the integration of 3, 4, 5, 6, 7, 8, 9, 10, or more cassettes into a single exogenous nucleotide sequence present at a specific locus in the host cell genome. Furthermore, not only can each cassette contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more individual sequences of interest, but each host genome can contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more exogenous nucleotide sequences present at specific loci in the host cell genome where the cassettes containing the sequences of interest can be integrated. Thus, while exemplary Figure 2B shows a strategy that can generate nine possible transgene constructs of sequences of interest using three front cassettes and three back cassettes, the present disclosure encompasses strategies for generating both fewer and more transgene constructs that include more than one locus within the host genome.
[0048] Use of the RCTI strategy outlined herein can result in significant resource savings, including, but not limited to, cost and time savings, labor savings, and mitigation of automation bottlenecks. As shown in Figures 3-5, these resource savings do not adversely affect the observed distribution of desirable expression levels, product quality attributes, or production culture performance compared to standard CLD strategies.
[0049] The RCTI approach also allows for single-cell cloning at the early, mid-stage, or full recovery stage, as shown in Figures 6A-6C. Because standard CLD strategies generally involve waiting for full transfection pool recovery (e.g., >90% cell viability) before engaging in single-cell cloning, the use of the RCTI strategy can result in significant time savings without affecting clonal performance (% of clones with high titer and desired product quality) or heterogeneity (range of vector configuration, titer, and product quality).
[0050] In certain embodiments, the present disclosure relates to a high-throughput RCTI-based method for screening a library of TI host cells expressing at least one SOI. In certain embodiments, the method includes generating a library of TI host cells, the library comprising a plurality of TI host cells expressing one or more SOIs, by segregating the library into single clones and screening the clones for particular cell and / or particular product attributes. In certain embodiments, generating the library of TI host cells includes providing a plurality of TI host cells, contacting the plurality of TI host cells with a plurality of vectors comprising one or more SOIs, and introducing the one or more SOIs into one or more of the plurality of TI host cells.
[0051] In certain embodiments, a sequence of interest may be introduced into a host cell of the present disclosure by recombinase-mediated integration. In certain embodiments, a sequence of interest may be introduced into a host cell of the present disclosure by gene editing-mediated integration.
[0052] In certain embodiments, the screening methods of the present disclosure may involve screening cells for a cellular attribute, which may be, but is not limited to, any of the following: cell proliferation, cell titer, specific productivity, volumetric productivity, clonal stability. Screening for these cellular attributes may be performed using any technique known in the art.
[0053] In certain embodiments, the screening methods of the present disclosure may include screening cells for product attributes, which may be, but are not limited to, any of the following: level of glycosylation, level of charge dispersion, reduced mismatch, reduced protein / peptide aggregation, protein sequence heterogeneity. Screening for these product attributes may be performed using any technique known in the art, including, but not limited to, size exclusion chromatography (SEC), cDNA sequencing, peptide mapping, CE-SDS or SDS-PAGE, HPLC, CE-based glycan assays, IEF, and ion exchange chromatography.
[0054] In certain embodiments, the present disclosure provides an RCTI-based method for generating a library of TI host cells containing multiple exogenous nucleotide SOIs, the method comprising contacting a plurality of TI host cells with a vector having one or more SOIs, introducing the one or more SOIs into the TI host cells, and selecting TI cells that express the one or more SOIs. In certain embodiments, the TI host cells generated by the methods of the present disclosure comprise one or more exogenous nucleotide sequences integrated into one or more loci in the genome of the TI host cells. In certain embodiments, the exogenous nucleotide sequences may comprise at least two RRSs, which may be flanked by one or more selectable markers.
[0055] In certain embodiments, the vector may contain one or more selectable marker genes to provide a phenotypic trait for selecting transformed host cells, such as dihydrofolate reductase or neomycin resistance for eukaryotic cell culture. In certain embodiments, the host cell may be a higher eukaryotic cell, such as a mammalian cell, or a lower eukaryotic cell, such as a yeast cell, or the host cell may be a prokaryotic cell, such as a bacterial cell. In certain embodiments, the library may be screened for a particular sequence of interest, such as, but not limited to, a polypeptide sequence. In certain embodiments, the resulting library may be screened for clones that exhibit activity against the polypeptide of interest in a phenotypic assay. In certain embodiments, the library may be screened for specific protein, e.g., enzymatic, activity by procedures known in the art.
[0056] In certain embodiments, the present disclosure relates to a RCTI-based method for selecting host cells expressing a sequence of interest, the method comprising: providing a plurality of TI host cells; contacting the plurality of TI host cells with a plurality of vectors comprising one or more SOIs and at least one marker; introducing the one or more SOIs into one or more of the plurality of TI host cells; and selecting TI host cells that express the selectable marker, thereby isolating TI host cells that express the sequence of interest. In certain embodiments, the SOI can encode a polypeptide subunit of a multi-subunit protein or a fragment thereof. In certain embodiments, the SOI can encode a single-chain antibody, an antibody light chain, an antibody heavy chain, a single-chain Fv fragment (scFv), or an Fc fusion protein. In certain embodiments, the locus into which the exogenous nucleotide sequence is integrated is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or at least about 99.9% homologous to a sequence selected from SEQ ID NOs: 1-7.
[0057] In certain embodiments, one or more exogenous nucleotide sequences may be integrated into one or more loci having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or at least about 99.9% homology to a sequence selected from the following: SEQ ID NOs: 1-12; NW_006874047.1; NW_006884592.1; NW_006881296.1; NW_003616412.1; NW_003615063.1; NW_006882936.1; and NW_003615411.1.
[0058] In certain embodiments, the locus into which the exogenous nucleotide sequence is integrated is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or at least about 99.9% homologous to a sequence selected from SEQ ID NOS: 1-4 of U.S. Patent No. 9,816,110, which correspond to SEQ ID NOS: 8-11 of the present disclosure, or SEQ ID NO: 1 of International Application No. PCT / US2017 / 028555, which corresponds to SEQ ID NO: 12 of the present disclosure. In certain embodiments, the one or more sequences of interest may be introduced into a host cell of the present disclosure by recombinase-mediated integration. In certain embodiments, the one or more sequences of interest may be introduced into a host cell of the present disclosure by gene editing-mediated integration.
[0059] In certain embodiments, TI host cells used in the RCTI-based methods for selecting host cells expressing a sequence of interest described herein can each have one or more exogenous nucleotide sequences integrated into one or more loci in the genome of the TI host cell. In certain embodiments, the exogenous nucleotide sequence can include at least two RRSs flanking at least one first selectable marker. In certain embodiments, a vector used in the RCTI-based methods for selecting host cells includes at least two RRSs matching at least two RRSs on the integrated exogenous nucleotide sequence, one or more exogenous SOIs, and at least one second selectable marker flanking these RRSs. In certain embodiments, the exogenous nucleotide sequence includes a first RRS and a second RRS flanking at least one first selectable marker, and a third RRS located between the first RRS and the second RRS, all of which may be heterospecific. In certain embodiments, the plurality of vectors can include a first vector and a second vector. In certain embodiments, the first vector can include two RRSs that match the first and third RRSs on the integrated exogenous nucleotide sequence and flank at least one first exogenous SOI and at least one second selectable marker, and the second vector can include two RRSs that match the second and third RRSs on the integrated exogenous nucleotide sequence and flank at least one second exogenous SOI, In certain embodiments, the plurality of vectors can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more vectors.
[0060] In certain embodiments, the sequence of interest comprises the sequence of one or more subunits of a multi-subunit protein complex. In certain embodiments, such polypeptide sequences can comprise fragments of such subunit sequences. In certain embodiments, the sequence of interest can comprise a combination of such subunit sequences. For example, and without limitation, such combinations can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more first subunit sequences and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, or more subunit sequences. Moreover, in certain embodiments, such combinations may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more different variations of a first subunit sequence, and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more different variations of a second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth or more subunit sequences.
[0061] In certain embodiments, the plurality of sequences of interest comprises one or more antibody heavy chain sequences ("H") and / or one or more antibody light chain sequences ("L"). As used herein, such "H" and "L" sequences can be full-length heavy or light chain sequences, as well as heavy or light chain fragments, including, but not limited to, variable region fragments and complementarity determining region fragments. In certain embodiments, the sequences of interest can comprise combinations of such H and L sequences. For example, and without limitation, such combinations can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more H sequences and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more L sequences. Furthermore, in certain embodiments, such combinations can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more identical or different distinct H' sequences, and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more identical or different distinct L' sequences. The inclusion of 1-10 (or more) additional H and L sequences, e.g., H", H'", L", and L'", is also within the subject matter of the present disclosure. Exemplary embodiments include, but are not limited to, sequences comprising any combination of a variety of different H and a variety of different L sequences, such as: HL; LH; H'L; LH'; H'L'; L'H'; HLL; HHL; LLH; LHH; H'LL; H'HL; L'HH; L'LH; H'H'L; LH'H'; HLLL; LLLH; HLHL; LHLH; H'LLL; L'LLH; H'LHL; L'HLH, etc.
[0062] In certain embodiments, the present disclosure relates to a method of expressing an SOI, the method comprising providing a plurality of TI host cells; contacting the plurality of TI host cells with a plurality of vectors comprising one or more SOIs and at least one marker; introducing the one or more SOIs into one or more of the plurality of TI host cells; selecting TI cells that express a sequence of interest; and culturing the cells under conditions suitable for expression of the sequence of interest and recovering the sequence of interest therefrom.
[0063] In certain embodiments, transfected host cells contain one or more sequences of interest that comprise the sequences of one or more subunits of a multi-subunit protein complex. In certain embodiments, such polypeptide sequences can comprise fragments of such subunit sequences. In certain embodiments, the sequences of interest can comprise combinations of such subunit sequences. For example, and without limitation, such combinations can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more first subunit sequences and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, or more subunit sequences of a sequence. Furthermore, in certain embodiments, such combinations may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more distinct variations of the same or different first subunit sequences, and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more distinct variations of the same or different second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth or more subunit sequences.
[0064] In certain embodiments, transfected host cells comprise one or more sequences of interest, which may include combinations of H and L sequences. For example, and without limitation, such combinations may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more H sequences and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more L sequences. Furthermore, in certain embodiments, such combinations may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more distinct H' sequences and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more distinct L' sequences. The inclusion of 1-10 (or more) additional H and L sequences, e.g., H", H'", L", and L'", is also within the subject matter of the present disclosure. Exemplary embodiments include, but are not limited to, sequences comprising any combination of a variety of different H and a variety of different L sequences, such as: HL; LH; H'L; LH'; H'L'; L'H'; HLL; HHL; LLH; LHH; H'LL; H'HL; L'HH; L'LH; H'H'L; LH'H'; HLLL; LLLH; HLHL; LHLH; H'LLL; L'LLH; H'LHL; L'HLH, etc.
[0065] 3. Exogenous Nucleotide Sequence The presently disclosed subject matter provides host cells suitable for the integration of exogenous nucleotide sequences. In certain embodiments, the exogenous nucleotide sequence serves as an integration site, for example, by including one or more recombinase recognition sequences. In certain embodiments, the exogenous nucleotide sequence encodes a sequence of interest. Thus, in certain embodiments, a host cell includes one or more exogenous nucleotide sequences that facilitate targeted integration of one or more exogenous nucleotide sequences encoding one or more sequences of interest. In certain embodiments, a host cell including an exogenous nucleotide sequence integrated into an integration site on the host cell's genome is referred to as a TI host cell. The exogenous nucleotide sequence encoding one or more sequences of interest can then be introduced into the TI host cell, and integration can be targeted to the integration site. As outlined below, the TI host cell can include multiple integration sites defined by the presence of elements that facilitate integration of the exogenous nucleotide sequence encoding one or more sequences of interest, for example, exogenous nucleotide sequences including recombinase recognition sequences.
[0066] In certain embodiments, the integration site and / or nucleotide sequences adjacent to the integration site can be identified experimentally. In certain embodiments, the integration site and / or nucleotide sequences adjacent to the integration site can be identified by genome-wide screening approaches to isolate host cells that express desired levels of a polypeptide of interest encoded by one or more SOIs integrated into one or more exogenous nucleotide sequences, where the exogenous sequences are themselves integrated into one or more loci within the genome of the host cell. In certain embodiments, the integration site and / or nucleotide sequences adjacent to the integration site can be identified by genome-wide screening approaches after a transposase-based cassette integration event. In certain embodiments, the integration site and / or nucleotide sequences adjacent to the integration site can be identified by brute-force random integration screening. In certain embodiments, the integration site and / or nucleotide sequences adjacent to the integration site can be determined by conventional sequencing approaches such as targeted locus amplification (TLA) followed by next-generation sequencing (NGS) and whole-genome NGS. In certain embodiments, the location of the integration site on the chromosome may be determined by conventional cell biological approaches, such as fluorescence in situ hybridization (FISH) analysis.
[0067] In certain embodiments, the host cell comprises a first exogenous nucleotide sequence integrated at a first integration site within a specific first locus within the genome of the host cell and a second exogenous nucleotide sequence integrated at a second integration site within a specific second locus within the genome, hi certain embodiments, the host cell comprises multiple exogenous nucleotide sequences integrated at multiple integration sites in the genome of the host cell.
[0068] 3.1 Exogenous sequences containing recombinase recognition sequences In certain embodiments, the integrated exogenous nucleotide sequence comprises one or more recombinase recognition sequences (RRSs), which can be recognized by a recombinase. In certain embodiments, the integrated exogenous nucleotide sequence comprises at least two RRSs. In certain embodiments, the integrated exogenous nucleotide sequence comprises two RRSs, and the two RRSs are identical. In certain embodiments, the integrated exogenous nucleotide sequence comprises two RRSs, and the two RRSs are different. In certain embodiments, the integrated exogenous nucleotide sequence comprises three RRSs, and the third RRS is located between the first and second RRSs. In certain embodiments, the first and second RRSs are the same, and the third RRS is different from both the first and second RRSs. In certain embodiments, all three RRSs are different. In certain embodiments, the integrated exogenous nucleotide sequence comprises four, five, six, seven, or eight RRSs. In certain embodiments, the integrated exogenous nucleotide sequence comprises multiple RRSs. In certain embodiments, two or more of the RRSs are identical. In certain embodiments, two or more RRSs are different. In certain embodiments, a subset of the total number of RRSs are the same, and a subset of the total number of RRSs are different. In certain embodiments, the RRS or multiple RRSs may be selected from the group consisting of a LoxP sequence, a LoxP L3 sequence, a LoxP 2L sequence, a LoxFas sequence, a Lox511 sequence, a Lox2272 sequence, a Lox2372 sequence, a Lox5171 sequence, a Loxm2 sequence, a Lox71 sequence, a Lox66 sequence, an FRT sequence, a Bxb1 attP sequence, a Bxb1 attB sequence, a φC31 attP sequence, and a φC31 attB sequence.
[0069] In certain embodiments, the integrated exogenous nucleotide sequence comprises at least one selectable marker. In certain embodiments, the integrated exogenous nucleotide sequence comprises one RRS and at least one selectable marker. In certain embodiments, the integrated exogenous nucleotide sequence comprises a first RRS, a second RRS, and at least one selectable marker. In certain embodiments, the selectable marker is located between the first and second RRS. In certain embodiments, the two RRSs are adjacent to at least one selectable marker, i.e., the first RRS is located 5' upstream of the selectable marker and the second RRS is located 3' downstream of the selectable marker. In certain embodiments, the first RRS is adjacent to the 5' end of the selectable marker and the second RRS is adjacent to the 3' end of the selectable marker.
[0070] In certain embodiments, the selectable marker is located between the first and second RRSs, and the two flanking RRSs are identical. In certain embodiments, both of the two RRSs flanking the selectable marker are LoxP sequences. In certain embodiments, both of the two RRSs flanking the selectable marker are FRT sequences. In certain embodiments, the selectable marker is located between the first and second RRSs, and the two flanking RRSs are different from each other. In certain embodiments, the first flanking RRS is a LoxP L3 sequence, and the second flanking RRS is a LoxP 2L sequence. In certain embodiments, the LoxP L3 sequence is located 5' of the selectable marker, and the LoxP 2L sequence is located 3' of the selectable marker. In certain embodiments, the first flanking RRS is a wild-type FRT sequence, and the second flanking RRS is a mutant FRT sequence. In certain embodiments, the first flanking RRS is a Bxb1 attP sequence, and the second flanking RRS is a Bxb1 attB sequence. In certain embodiments, the first adjacent RRS is a φC31 attP sequence and the second adjacent RRS is a φC31 attB sequence. In certain embodiments, the two RRSs are positioned in the same orientation. In certain embodiments, the two RRSs are both forward or reverse oriented. In certain embodiments, the two RRSs are positioned in opposite orientations.
[0071] In certain embodiments, the selectable marker may be an aminoglycoside phosphotransferase (APH) (e.g., hygromycin phosphotransferase (HYG), neomycin, and G418 APH), dihydrofolate reductase (DHFR), thymidine kinase (TK), glutamine synthetase (GS), asparagine synthetase, tryptophan synthetase (indole), histidinol dehydrogenase (histidinol D), and a gene encoding resistance to puromycin, blasticidin, bleomycin, phleomycin, chloramphenicol, zeocin, or mycophenolic acid. In certain embodiments, the selectable marker can be a GFP, eGFP, synthetic GFP, YFP, eYFP, CFP, mPlum, mCherry, tdTomato, mStrawberry, J-red, DsRed-monomer, mOrange, mKO, mCitrine, Venus, YPet, Emerald, CyPet, mCFPm, Cerulean, or T-Sapphire marker.
[0072] In certain embodiments, the integrated exogenous nucleotide sequence comprises two selectable markers flanked by two RRSs, the first selectable marker being different from the second selectable marker. In certain embodiments, both selectable markers are selected from the group consisting of a glutamine synthetase selectable marker, a thymidine kinase selectable marker, a HYG selectable marker, and a puromycin resistance selectable marker. In certain embodiments, the integrated exogenous nucleotide sequence comprises a thymidine kinase selectable marker and a HYG selectable marker. In certain embodiments, the first selectable marker is an aminoglycoside phosphotransferase (APH) (e.g., hygromycin phosphotransferase (HYG), neomycin, and G418). APH), dihydrofolate reductase (DHFR), thymidine kinase (TK), glutamine synthetase (GS), asparagine synthetase, tryptophan synthetase (indole), histidinol dehydrogenase (histidinol D), and genes encoding resistance to puromycin, blasticidin, bleomycin, phleomycin, chloramphenicol, zeocin, and mycophenolic acid; and the second selection marker is selected from the group consisting of GFP, eGFP, synthetic GFP, YFP, eYFP, CFP, mPlum, mCherry, tdTomato, mStrawberry, J-red, DsRed-monomer, mOrange, mKO, mCitrine, Venus, YPet, Emerald, CyPet, mCFPm, Cerulean, and T-Sapphire. In certain embodiments, the first selection marker is a glutamine synthetase selection marker and the second selection marker is a GFP marker. In certain embodiments, the two RRSs flanking both selection markers are identical. In certain embodiments, the two RRSs flanking both selection markers are different from each other.
[0073] In certain embodiments, the selectable marker is operably linked to a promoter sequence. In certain embodiments, the selectable marker is operably linked to an SV40 promoter. In certain embodiments, the selectable marker is operably linked to a cytomegalovirus (CMV) promoter.
[0074] In certain embodiments, the integrated exogenous nucleotide sequence comprises at least one selectable marker and an IRES, wherein the IRES is operably linked to the selectable marker. In certain embodiments, the selectable marker operably linked to the IRES is selected from the group consisting of GFP, eGFP, synthetic GFP, YFP, eYFP, CFP, mPlum, mCherry, tdTomato, mStrawberry, J-red, DsRed-monomer, mOrange, mKO, mCitrine, Venus, YPet, Emerald, CyPet, mCFPm, Cerulean, and T-Sapphire markers. In certain embodiments, the selectable marker operably linked to the IRES is a GFP marker. In certain embodiments of interest, the integrated exogenous nucleotide sequence comprises an IRES flanked by two RRSs and two selectable markers, wherein the IRES is operably linked to a second selectable marker. In certain embodiments, the integrated exogenous nucleotide sequence comprises an IRES flanked by two RRSs and three selectable markers, where the IRES is operably linked to a third selectable marker. In certain embodiments, the integrated exogenous nucleotide sequence comprises an IRES flanked by two RRSs and three selectable markers, where the IRES is operably linked to a third selectable marker. In certain embodiments, the third selectable marker is different from the first or second selectable marker. In certain embodiments, the integrated exogenous nucleotide sequence comprises a first selectable marker operably linked to a promoter and a second selectable marker operably linked to an IRES. In certain embodiments, the integrated exogenous nucleotide sequence comprises a glutamine synthetase selectable marker operably linked to an SV40 promoter and a GFP selectable marker operably linked to an IRES. In certain embodiments, the integrated exogenous nucleotide sequence comprises a thymidine kinase selectable marker and an HYG selectable marker operably linked to a CMV promoter, and a GFP selectable marker operably linked to an IRES.
[0075] In certain embodiments, the integrated exogenous nucleotide sequence comprises three RRSs. In certain embodiments, the third RRS is located between the first and second RRSs. In certain embodiments, all three RRSs are identical. In certain embodiments, the first and second RRSs are the same, and the third RRS is different from both the first and second RRSs. In certain embodiments, all three RRSs are different.
[0076] In certain embodiments, the exogenous nucleotide sequence that serves as the integration site is present at a site within a specific locus in the genome of the TI host cell. In certain embodiments, the locus into which the exogenous nucleotide sequence is integrated is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or at least about 99.9% homologous to a sequence selected from SEQ ID NOS: 1-7. In certain embodiments, the locus into which the exogenous nucleotide sequence is integrated is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or at least about 99.9% homologous to a sequence selected from SEQ ID NOS: 1-4 of U.S. Patent No. 9,816,110, which correspond to SEQ ID NOS: 8-11 of the present disclosure, or SEQ ID NO: 1 of International Application No. PCT / US 2017 / 028555, which corresponds to SEQ ID NO: 12 of the present disclosure.
[0077] In certain embodiments, the exogenous nucleotide sequence is integrated into a site within a specific locus in the genome of the TI host cell. In certain embodiments, the locus into which the exogenous nucleotide sequence is integrated is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or at least about 99.9% homologous to a sequence selected from Contigs NW_006874047.1, NW_006884592.1, NW_006881296.1, NW_003616412.1, NW_003615063.1, NW_006882936.1, and NW_003615411.1.
[0078] In certain embodiments, the exogenous nucleotide sequence is integrated into an integration site located within a position selected from nucleotides numbered 1 to 1,000 bp, 1,000 to 2,000 bp, 2,000 to 3,000 bp, 3,000 to 4,000 bp, and 4,000 to 4,301 bp of SEQ ID NO:1. In certain embodiments, the exogenous nucleotide sequence is integrated into an integration site located within a position selected from nucleotides numbered 1 to 100,000 bp, 100,000 to 200,000 bp, 200,000 to 300,000 bp, 300,000 to 400,000 bp, 400,000 to 500,000 bp, 500,000 to 600,000 bp, 600,000 to 700,000 bp, and 700,000 to 728785 bp of SEQ ID NO:2. In certain embodiments, the exogenous nucleotide sequence is integrated at an integration site located within a position selected from nucleotides numbered 1 to 100,000 bp, 100,000 to 200,000 bp, 200,000 to 300,000 bp, 300,000 to 400,000 bp, and 400,000 to 413,983 of SEQ ID NO: 3. In certain embodiments, the exogenous nucleotide sequence is integrated at an integration site located within a position selected from nucleotides numbered 1 to 10,000 bp, 10,000 to 20,000 bp, 20,000 to 30,000 bp, and 30,000 to 30,757 bp of SEQ ID NO: 4. In certain embodiments, the exogenous nucleotide sequence is integrated into an integration site located within a position selected from nucleotides numbered 1 to 10,000 bp, 10,000 to 20,000 bp, 20,000 to 30,000 bp, 30,000 to 40,000 bp, 40,000 to 50,000 bp, 50,000 to 60,000 bp, and 60,000 to 68,962 bp of SEQ ID NO:5. In certain embodiments, the exogenous nucleotide sequence is integrated into an integration site located within a position selected from nucleotides numbered 1 to 10,000 bp, 10,000 to 20,000 bp, 20,000 to 30,000 bp, 30,000 to 40,000 bp, 40,000 to 50,000 bp, and 50,000 to 51,326 bp of SEQ ID NO:6.In certain embodiments, the exogenous nucleotide sequence is integrated into an integration site located within a position selected from nucleotides numbered 1 to 10,000 bp, 10,000 to 20,000 bp, and 20,000 to 22,904 bp of SEQ ID NO:7.
[0079] In certain embodiments, the nucleotide sequence immediately 5' to the integrated exogenous sequence is selected from the group consisting of nucleotides 41190-45269 of NW_006874047.1, nucleotides 63590-207911 of NW_006884592.1, nucleotides 253831-491909 of NW_006881296.1, nucleotides 69303-79768 of NW_003616412.1, nucleotides 293481-315265 of NW_003615063.1, nucleotides 2650443-2662054 of NW_006882936.1, or nucleotides 82214-97705 of NW_003615411.1, and sequences at least 50% homologous thereto. In certain embodiments, the nucleotide sequence immediately 5' to the integrated exogenous sequence is nucleotides 41190-45269 of NW_006874047.1, nucleotides 63590-207911 of NW_006884592.1, nucleotides 253831-491909 of NW_006881296.1, nucleotides 69303-79768 of NW_003616412.1, and nucleotides 69303-79768 of NW_003615063. NW_006882936.1, nucleotides 2650443-2662054 of NW_006882936.1, or nucleotides 82214-97705 of NW_003615411.1.
[0080] In certain embodiments, the nucleotide sequence immediately 3' to the integrated exogenous sequence is selected from the group consisting of nucleotides 45270-45490 of NW_006874047.1, nucleotides 207912-792374 of NW_006884592.1, nucleotides 491910-667813 of NW_006881296.1, nucleotides 79769-100059 of NW_003616412.1, nucleotides 315266-362442 of NW_003615063.1, nucleotides 2662055-2701768 of NW_006882936.1, or nucleotides 97706-105117 of NW_003615411.1, and sequences at least 50% homologous thereto. In certain embodiments, the nucleotide sequence immediately 3′ to the integrated exogenous sequence is nucleotides 45270-45490 of NW_006874047.1, nucleotides 207912-792374 of NW_006884592.1, nucleotides 491910-667813 of NW_006881296.1, nucleotides 79769-100059 of NW_003616412.1, nucleotides 80069-90069 of NW_003615063 NW_006882936.1, nucleotides 2662055-2701768 of NW_006882936.1, or nucleotides 97706-105117 of NW_003615411.1, are at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or at least about 99.9% homologous to nucleotides 315266-362442 of NW_006882936.1, nucleotides 2662055-2701768 of NW_006882936.1, or nucleotides 97706-105117 of NW_003615411.1.
[0081] In certain embodiments, the integrated exogenous nucleotide sequence is operably linked to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 1-7 and sequences at least 50% homologous thereto. In certain embodiments, the nucleotide sequence operably linked to the exogenous nucleotide sequence is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or at least about 99.9% homologous to a sequence selected from SEQ ID NOs: 1-7. In certain embodiments, the integrated exogenous nucleotide sequence comprises at least one SOI. In certain embodiments, the operably linked nucleotide sequence increases the expression level of the SOI compared to a randomly integrated SOI. In certain embodiments, the integrated exogenous SOI is expressed at about 20%, 30%, 40%, 50%, 100%, 2-fold, 3-fold, 5-fold, or 10-fold higher than a randomly integrated SOI.
[0082] In certain embodiments, the integrated exogenous sequence is a nucleotide sequence selected from the group consisting of nucleotides 41190-45269 of NW_006874047.1, nucleotides 63590-207911 of NW_006884592.1, nucleotides 253831-491909 of NW_006881296.1, nucleotides 69303-79768 of NW_003616412.1, nucleotides 293481-315265 of NW_003615063.1, nucleotides 2650443-2662054 of NW_006882936.1, and nucleotides 82214-97705 of NW_003615411.1, and sequences at least 50% homologous thereto. Adjacent to the 5' side of the sequence and adjacent to the 3' side of a nucleotide sequence selected from the group consisting of nucleotides 45270-45490 of NW_006874047.1, nucleotides 207912-792374 of NW_006884592.1, nucleotides 491910-667813 of NW_006881296.1, nucleotides 79769-100059 of NW_003616412.1, nucleotides 315266-362442 of NW_003615063.1, nucleotides 2662055-2701768 of NW_006882936.1, and nucleotides 97706-105117 of NW_003615411.1, and sequences at least 50% homologous thereto.In certain embodiments, the nucleotide sequence 5′ flanking the integrated exogenous nucleotide sequence is nucleotides 41190-45269 of NW_006874047.1, nucleotides 63590-207911 of NW_006884592.1, nucleotides 253831-491909 of NW_006881296.1, nucleotides 69303-79768 of NW_003616412.1, nucleotides 79303-89768 of NW_003616412.1, nucleotides 80303-80309 of NW_003616412.1, nucleotides 90303-90309 of NW_003616412.1, nucleotides 100303-100309 of NW_003616412.1, nucleotides 110303-110309 of NW_003616412.1, nucleotides 120303-120309 of NW_003616412.1, nucleotides 130303-130309 of NW_003616412.1, nucleotides 140303-140309 of NW_003616412.1, nucleotides 150303-150309 of NW_003616412.1, nucleotides 160303-160309 of NW_003616412.1, nucleotides 170303-170309 of NW_003616412.1, nucleotides NW_006882936.1, nucleotides 2650443-2662054 of NW_003615411.1, and nucleotides 82214-97705 of NW_003615411.1. The nucleotide sequences flanking the 3' side of the integrated exogenous nucleotide sequence are nucleotides 45270-45490 of SEQ ID NO: NW_006874047.1, nucleotides 207912-792374 of NW_006884592.1, nucleotides 491910-667813 of NW_006881296.1, nucleotides 79769-100059 of NW_003616412.1, and nucleotides 80069-90069 of NW_003615 NW_006882936.1, nucleotides 2662055-2701768 of NW_006882936.1, and nucleotides 97706-105117 of NW_003615411.1.
[0083] In certain embodiments, the integrated exogenous nucleotide is integrated into a locus immediately adjacent to all or a portion of a sequence selected from the group consisting of sequences at least about 90% homologous to a sequence selected from SEQ ID NOs: 1-7.
[0084] In certain embodiments, the integrated exogenous nucleotide sequence is flanked by a nucleotide sequence selected from the group consisting of SEQ ID NOs: 1-7 and sequences at least 50% homologous thereto. In certain embodiments, the integrated exogenous nucleotide sequence is within about 100 bp, about 200 bp, about 500 bp, or about 1 kb of a sequence selected from the group consisting of SEQ ID NOs: 1-7 and sequences at least 50% homologous thereto. In certain embodiments, the nucleotide sequence flanking the exogenous nucleotide sequence is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or at least about 99.9% homologous to a sequence selected from SEQ ID NOs: 1-7.
[0085] In certain embodiments, the exogenous nucleotide sequence is integrated into an integration site adjacent to a position selected from nucleotides numbered 1 to 1,000 bp, 1,000 to 2,000 bp, 2,000 to 3,000 bp, 3,000 to 4,000 bp, and 4,000 to 4,301 bp of SEQ ID NO: 1. In certain embodiments, the exogenous nucleotide sequence is integrated into an integration site adjacent to a position selected from nucleotides numbered 1 to 100,000 bp, 100,000 to 200,000 bp, 200,000 to 300,000 bp, 300,000 to 400,000 bp, 400,000 to 500,000 bp, 500,000 to 600,000 bp, 600,000 to 700,000 bp, and 700,000 to 728785 bp of SEQ ID NO:2. In certain embodiments, the exogenous nucleotide sequence is integrated into an integration site adjacent to a position selected from nucleotides numbered 1 to 100,000 bp, 100,000 to 200,000 bp, 200,000 to 300,000 bp, 300,000 to 400,000 bp, and 400,000 to 413,983 of SEQ ID NO: 3. In certain embodiments, the exogenous nucleotide sequence is integrated into an integration site adjacent to a position selected from nucleotides numbered 1 to 10,000 bp, 10,000 to 20,000 bp, 20,000 to 30,000 bp, and 30,000 to 30,757 bp of SEQ ID NO: 4. In certain embodiments, the exogenous nucleotide sequence is integrated into an integration site adjacent to a position selected from nucleotides numbered 1 to 10,000 bp, 10,000 to 20,000 bp, 20,000 to 30,000 bp, 30,000 to 40,000 bp, 40,000 to 50,000 bp, 50,000 to 60,000 bp, and 60,000 to 68,962 bp of SEQ ID NO:5. In certain embodiments, the exogenous nucleotide sequence is integrated into an integration site adjacent to a position selected from nucleotides numbered 1 to 10,000 bp, 10,000 to 20,000 bp, 20,000 to 30,000 bp, 30,000 to 40,000 bp, 40,000 to 50,000 bp, and 50,000 to 51,326 bp of SEQ ID NO:6.In certain embodiments, the exogenous nucleotide sequence is integrated into an integration site adjacent to a position selected from nucleotides numbered 1 to 10,000 bp, 10,000 to 20,000 bp, and 20,000 to 22,904 bp of SEQ ID NO:7.
[0086] In certain embodiments, the locus comprising the integration site of the exogenous nucleotide sequence does not encode an open reading frame (ORF). In certain embodiments, the locus comprising the integration site of the exogenous nucleotide sequence comprises cis-acting elements, such as promoters and enhancers. In certain embodiments, the locus comprising the integration site of the exogenous nucleotide sequence does not comprise any cis-acting elements, such as promoters and enhancers, that enhance gene expression.
[0087] In certain embodiments, the exogenous nucleotide sequence is integrated into an integration site within an endogenous gene selected from the group consisting of LOC107977062, LOC100768845, ITPR2, ERE67000.1, UBAP2, MTMR2, and XP_003512331.2, including wild-type and all homologous sequences of the LOC107977062, LOC100768845, ITPR2, ERE67000.1, UBAP2, MTMR2, and XP_003512331.2 genes. In certain embodiments, the homologous sequences of the LOC107977062, LOC100768845, ITPR2, ERE67000.1, UBAP2, MTMR2, and XP_003512331.2 genes may be at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or at least about 99.9% homologous to the wild-type LOC107977062, LOC100768845, ITPR2, ERE67000.1, UBAP2, MTMR2, and XP_003512331.2 genes. In certain embodiments, the LOC107977062, LOC100768845, ITPR2, ERE67000.1, UBAP2, MTMR2, and XP_003512331.2 genes are wild-type mammalian LOC107977062, LOC100768845, ITPR2, ERE67000.1, UBAP2, MTMR2, and XP_003512331.2 genes. In certain embodiments, the LOC107977062, LOC100768845, ITPR2, ERE67000.1, UBAP2, MTMR2, and XP_003512331.2 genes are wild-type human LOC107977062, LOC100768845, ITPR2, ERE67000.1, UBAP2, MTMR2, and XP_003512331.2 genes.In certain embodiments, the LOC107977062, LOC100768845, ITPR2, ERE67000.1, UBAP2, MTMR2, and XP_003512331.2 genes are wild-type hamster LOC107977062, LOC100768845, ITPR2, ERE67000.1, UBAP2, MTMR2, and XP_003512331.2 genes.
[0088] In certain embodiments, the integration site is operably linked to an endogenous gene selected from the group consisting of LOC107977062, LOC100768845, ITPR2, ERE67000.1, UBAP2, MTMR2, XP_003512331.2, and sequences at least about 90% homologous thereto. In certain embodiments, the integration site is adjacent to an endogenous gene selected from the group consisting of LOC107977062, LOC100768845, ITPR2, ERE67000.1, UBAP2, MTMR2, XP_003512331.2, and sequences at least about 90% homologous thereto.
[0089] Table 1 provides exemplary TI host cell integration sites. JPEG2025121945000002.jpg68170
[0090] 3.2 Exogenous nucleotide sequence containing the sequence of interest In certain embodiments, the integrated exogenous nucleotide sequence comprises at least one exogenous SOI. In certain embodiments, the integrated exogenous nucleotide sequence comprises at least one selectable marker and at least one exogenous SOI. In certain embodiments, the integrated exogenous nucleotide sequence comprises at least one selectable marker, at least one exogenous SOI, and at least one RRS. In certain embodiments, the integrated exogenous nucleotide sequence comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more SOIs. In certain embodiments, the SOIs are identical. In certain embodiments, the SOIs are different.
[0091] As noted above, in certain embodiments, the SOI encodes one or more subunits of a multi-subunit protein complex. In certain embodiments, such polypeptide sequences can comprise fragments of such subunit sequences. In certain embodiments, the sequence of interest can comprise a combination of such subunit sequences. For example, and without limitation, such combinations can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more first subunit sequences and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, or more subunit sequences of the sequence. Moreover, in certain embodiments, such combinations may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more different variations of a first subunit sequence, and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more different variations of a second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth or more subunit sequences.
[0092] In certain embodiments, the SOI encodes a single-chain antibody or a fragment thereof. In certain embodiments, the SOI encodes an antibody heavy chain sequence or a fragment thereof. In certain embodiments, the SOI encodes an antibody light chain sequence or a fragment thereof. In certain embodiments, the incorporated exogenous nucleotide sequence comprises an SOI encoding an antibody heavy chain sequence or a fragment thereof, and an SOI encoding an antibody light chain sequence or a fragment thereof. In certain embodiments, the incorporated exogenous nucleotide sequence comprises an SOI encoding a first antibody heavy chain sequence or a fragment thereof, an SOI encoding a second antibody heavy chain sequence or a fragment thereof, and an SOI encoding an antibody light chain sequence or a fragment thereof. In certain embodiments, the incorporated exogenous nucleotide sequence comprises an SOI encoding a first antibody heavy chain sequence or a fragment thereof, an SOI encoding a second antibody heavy chain sequence or a fragment thereof, an SOI encoding a first antibody light chain sequence or a fragment thereof, and a second SOI encoding an antibody light chain sequence or a fragment thereof. In certain embodiments, the number of SOIs encoding heavy and light chain sequences can be selected to achieve a desired expression level of heavy and light chain polypeptides, e.g., to achieve a desired amount of bispecific antibody production. In certain embodiments, the individual SOIs encoding heavy and light chain sequences can be integrated into, for example, a single exogenous nucleic acid sequence present at a single integration site, multiple exogenous nucleic acid sequences present at a single integration site, or multiple exogenous nucleic acid sequences integrated into different integration sites within the TI host cell.
[0093] In certain embodiments, the integrated exogenous nucleotide sequence comprises at least one selectable marker, at least one exogenous SOI, and one RRS. In certain embodiments, the RRS is located adjacent to the at least one selectable marker or at least one exogenous SOI. In certain embodiments, the integrated exogenous nucleotide sequence comprises at least one selectable marker, at least one exogenous SOI, and two RRSs. In certain embodiments, the integrated exogenous nucleotide sequence comprises at least one selectable marker and at least one exogenous SOI located between the first and second RRSs. In certain embodiments, the two RRSs flanking the selectable marker and the exogenous SOI are identical. In certain embodiments, the two RRSs flanking the selectable marker and the exogenous SOI are different. In certain embodiments, the first flanking RRS is a LoxP L3 sequence, and the second flanking RRS is a LoxP 2L sequence. In certain embodiments, a LoxP L3 sequence is located 5' of the selectable marker and a LoxP 2L sequence is located 3' of the selectable marker and the exogenous SOI.
[0094] In certain embodiments, the integrated exogenous nucleotide sequence comprises three RRSs and two exogenous SOIs, with the third RRS located between the first and second RRSs. In certain embodiments, the first SOI is located between the first and third RRSs, and the second SOI is located between the third and second RRSs. In certain embodiments, the first and second SOIs are different. In certain embodiments, the first and second RRSs are the same, and the third RRS is different from both the first and second RRSs. In certain embodiments, all three RRSs are different. In certain embodiments, the first RRS is a LoxP L3 site, the second RRS is a LoxP 2L site, and the third RRS is a LoxFas site. In certain embodiments, the integrated exogenous nucleotide sequence comprises three RRSs, one exogenous SOI, and one selectable marker. In certain embodiments, the SOI is located between the first and third RRSs, and the selectable marker is located between the third and second RRSs. In certain embodiments, the integrated exogenous nucleotide sequence comprises three RRSs, two exogenous SOIs, and one selectable marker. In certain embodiments, the first SOI and selectable marker are located between the first and third RRSs, and the second SOI is located between the third and second RRSs.
[0095] In certain embodiments, the exogenous SOI encodes a polypeptide of interest, including, but not limited to, an antibody, an enzyme, a cytokine, a growth factor, a hormone, a viral protein, a bacterial protein, a vaccine protein, or a protein with therapeutic function. In certain embodiments, the exogenous SOI encodes an antibody or an antigen-binding fragment thereof. In certain embodiments, the exogenous SOI encodes a single-chain antibody, an antibody light chain, an antibody heavy chain, a single-chain Fv fragment (scFv), or an Fc fusion protein. In certain embodiments, the exogenous SOI is operably linked to at least one cis-acting element, such as a promoter or enhancer. In certain embodiments, the exogenous SOI is operably linked to a CMV promoter.
[0096] In certain embodiments, the incorporated exogenous nucleotide sequence comprises two RRSs and at least two exogenous SOIs located between the two RRSs. In certain embodiments, an SOI encoding one heavy chain and one light chain of an antibody is located between the two RRSs. In certain embodiments, an SOI encoding one heavy chain and two light chains of an antibody is located between the two RRSs. In certain embodiments, an SOI encoding a different combination of copies of the heavy and light chains of an antibody is located between the two RRSs.
[0097] In certain embodiments, the incorporated exogenous nucleotide sequence comprises three RRSs and at least two exogenous SOIs, with the third RRS located between the first and second RRSs. In certain embodiments, at least one SOI is located between the first and third RRSs and at least one SOI is located between the third and second RRSs. In certain embodiments, the first and second RRSs are the same and the third RRS is different from both the first and second RRSs. In certain embodiments, all three RRSs are different. In certain embodiments, an SOI encoding one heavy chain and one light chain of the first antibody is located between the first and third RRSs, and an SOI encoding one heavy chain and one light chain of the second antibody is located between the third and second RRSs. In certain embodiments, the SOI encoding one heavy chain and two light chains of the first antibody is located between the first and third RRSs, and the SOI encoding one heavy chain and one light chain of the second antibody is located between the third and second RRSs. In certain embodiments, the SOI encoding one heavy chain and three light chains of the first antibody is located between the first and third RRSs, and the SOIs encoding one light chain of the first antibody and one heavy chain and one light chain of the second antibody are located between the third and second RRSs. In certain embodiments, the SOI encoding one heavy chain and three light chains of the first antibody is located between the first and third RRSs, and the SOIs encoding two light chains of the first antibody and one heavy chain and one light chain of the second antibody are located between the third and second RRSs. In certain embodiments, SOIs encoding different combinations of copies of the heavy and light chains of multiple antibodies are located between the first and third RRSs and between the third and second RRSs.
[0098] 4.Host cells In certain embodiments, the host cell is a eukaryotic host cell. In certain embodiments, the host cell is a mammalian host cell. In certain embodiments, the host cell is a hamster host cell, a human host cell, a rat host cell, or a mouse host cell. In certain embodiments, the host cell is a Chinese hamster ovary (CHO) host cell, a CHO K1 host cell, a CHO K1SV host cell, a DG44 host cell, a DUKXB-11 host cell, a CHOK1S host cell, or a CHO K1M host cell.
[0099] In certain embodiments, the host cell is an SV40-transformed monkey kidney CV1 line (COS-7), a human embryonic kidney line (e.g., 293 or 293 cells as described in Graham et al., J. Gen Virol. 36:59 (1977)), baby hamster kidney cells (BHK), mouse Sertoli cells (e.g., TM4 cells as described in Mather, Biol. Reprod. 23:243-251 (1980)), monkey kidney cells (CV1), African green monkey kidney cells (VERO-76), human cervical carcinoma cells (HELA), canine kidney cells (MDCK), buffalo rat liver cells (BRL 3A), human lung cells (W138), human liver cells (Hep G2), mouse mammary tumor (MMT060562), e.g., Mather et al., Annals The cells are selected from the group consisting of TRI cells, MRC5 cells, FS4 cells, Y0 cells, NS0 cells, Sp2 / 0 cells, and PER.C6® cells, as described in NYAcad. Sci. 383:44-68 (1982).
[0100] In certain embodiments, the host cell is a cell line. In certain embodiments, the host cell is a cell line that has been cultured for a certain number of generations. In certain embodiments, the host cell is a primary cell.
[0101] In certain embodiments, expression of a polypeptide of interest is stable if the expression level is maintained at a particular level and increases or decreases by less than 20% over 10, 20, 30, 50, 100, 200, or 300 generations. In certain embodiments, expression of a polypeptide of interest is stable if the culture can be maintained without selection. In certain embodiments, expression of a polypeptide of interest is high if the polypeptide product of the gene of interest reaches about 1 g / L, about 2 g / L, about 3 g / L, about 4 g / L, about 5 g / L, about 10 g / L, about 12 g / L, about 14 g / L, or about 16 g / L.
[0102] In certain embodiments, the polypeptide of interest is produced and secreted into the cell culture medium. In certain embodiments, the polypeptide of interest is expressed and retained within the host cell. In certain embodiments, the polypeptide of interest is expressed, inserted into, and retained in the host cell membrane.
[0103] An exogenous nucleotide or vector of interest can be introduced into a host cell by conventional cell biology methods, including, but not limited to, transfection, transduction, electroporation, or injection. In certain embodiments, an exogenous nucleotide or vector of interest is introduced into a host cell by chemical-based transfection methods, including lipid-based transfection, calcium phosphate-based transfection, cationic polymer-based transfection, or nanoparticle-based transfection. In certain embodiments, an exogenous nucleotide of interest is introduced into a host cell by virus-mediated transduction, including, but not limited to, lentivirus-, retrovirus-, adenovirus-, or adeno-associated virus-mediated transduction. In certain embodiments, an exogenous nucleotide or vector of interest is introduced into a host cell by gene gun-mediated injection. In certain embodiments, both DNA molecules and RNA molecules are introduced into a host cell using the methods described herein.
[0104] 5. Targeted integration Targeted integration method allows exogenous nucleotide sequence to be integrated into one or more predetermined sites of host cell genome.In certain embodiments, targeted integration is mediated by recombinase that recognizes one or more RRS.In certain embodiments, targeted integration is mediated by homologous recombination.In certain embodiments, targeted integration is mediated by exogenous site-specific nuclease, followed by HDR and / or NHEJ.
[0105] 5.1. Targeted Integration by Recombinase-Mediated Recombination A "recombinase recognition sequence" (RRS) is a nucleotide sequence that is recognized by a recombinase and is both necessary and sufficient for a recombinase-mediated recombination event. An RRS can be used to define the location in a nucleotide sequence where a recombination event is expected to occur.
[0106] In certain embodiments, the RRS is selected from the group consisting of a LoxP sequence, a LoxP L3 sequence, a LoxP 2L sequence, a LoxFas sequence, a Lox511 sequence, a Lox2272 sequence, a Lox2372 sequence, a Lox5171 sequence, a Loxm2 sequence, a Lox71 sequence, a Lox66 sequence, an FRT sequence, a Bxb1 attP sequence, a Bxb1 attB sequence, a φC31 attP sequence, and a φC31 attB sequence.
[0107] In certain embodiments, the RRS can be recognized by Cre recombinase. In certain embodiments, the RRS can be recognized by FLP recombinase. In certain embodiments, the RRS can be recognized by Bxb1 integrase. In certain embodiments, the RRS can be recognized by φC31 integrase.
[0108] In certain embodiments, when the RRS is a LoxP site, the host cell requires Cre recombinase to perform recombination. In certain embodiments, when the RRS is an FRT site, the host cell requires FLP recombinase to perform recombination. In certain embodiments, when the RRS is a Bxb1 attP or Bxb1 attB site, the host cell requires Bxb1 integrase to perform recombination. In certain embodiments, when the RRS is a φC31 attP or φC31 attB site, the host cell requires φC31 integrase to perform recombination. Recombinases can be introduced into host cells using expression vectors containing the coding sequences for the enzymes.
[0109] The Cre-LoxP site-specific recombination system is widely used in many biological experimental systems. Cre is a 38-kDa site-specific DNA recombinase that recognizes 34-bp LoxP sequences. Cre is derived from bacteriophage P1 and belongs to the tyrosine family of site-specific recombinases. Cre recombinase can mediate both intramolecular and intermolecular recombination between LoxP sequences. The LoxP sequence consists of an 8-bp nonpalindromic core region flanked by two 13-bp inverted repeats. Cre recombinase binds to the 13-bp repeats, thereby mediating recombination within the 8-bp core region. Cre-LoxP-mediated recombination occurs with high efficiency and does not require any other host factors. When two LoxP sequences are located in the same nucleotide sequence and in the same orientation, Cre-mediated recombination excises the DNA sequence located between the two LoxP sequences into a covalently closed circle. If two LoxP sequences are located in opposite positions on the same nucleotide sequence, Cre-mediated recombination will reverse the orientation of the DNA sequence located between the two sequences. LoxP sequences can also be located on different chromosomes to facilitate recombination between different chromosomes. If two LoxP sequences are located on two different DNA molecules and one of the DNA molecules is circular, Cre-mediated recombination will result in the integration of the circular DNA sequence.
[0110] In certain embodiments, the LoxP sequence is a wild-type LoxP sequence. In certain embodiments, the LoxP sequence is a mutant LoxP sequence. The mutant LoxP sequence was developed to increase the efficiency of Cre-mediated integration or replacement. In certain embodiments, the mutant LoxP sequence is selected from the group consisting of LoxP L3 sequence, LoxP 2L sequence, LoxFas sequence, Lox511 sequence, Lox2272 sequence, Lox2372 sequence, Lox5171 sequence, Loxm2 sequence, Lox71 sequence, and Lox66 sequence. For example, the Lox71 sequence has a 5-bp mutation in the left 13-bp repeat. The Lox66 sequence has a 5-bp mutation in the right 13-bp repeat. Both wild-type and mutant LoxP sequences can mediate Cre-dependent recombination.
[0111] The FLP-FRT site-specific recombination system is similar to the Cre-Lox system. It contains flippase (FLP) recombinase derived from the 2 μm plasmid of the yeast Saccharomyces cerevisiae. FLP also belongs to the tyrosine family of site-specific recombinases. The FRT sequence is a 34-bp sequence consisting of two 13-bp palindromic sequences, each flanked by an 8-bp spacer. FLP binds to the 13-bp palindromic sequences and mediates DNA cleavage, exchange, and ligation within the 8-bp spacer. As with Cre recombinase, the position and orientation of the two FRT sequences determine the outcome of FLP-mediated recombination. In certain embodiments, the FRT sequence is a wild-type FRT sequence. In certain embodiments, the FRT sequence is a mutant FRT sequence. Both wild-type and mutant FRT sequences can mediate FLP-dependent recombination. In certain embodiments, the FRT sequence is fused to a responsive receptor domain sequence, such as, but not limited to, a tamoxifen-responsive receptor domain sequence.
[0112] Bxb1 and φC31 belong to the serine recombinase family. Both are derived from bacteriophages and are used by these bacteriophages to establish lysogeny and promote site-specific integration of the phage genome into the bacterial genome. These integrases catalyze site-specific recombination events between short (40-60 bp) DNA substrates, called attP and attB sequences, which are the original attachment sites located on the phage DNA and bacterial DNA, respectively. After recombination, two new sequences are formed, called attL and attR sequences, each containing half sequences derived from attP and attB. Recombination can also occur between the attL and attR sequences, excising the integrated phage from the bacterial DNA. Both integrases can catalyze recombination without the aid of additional host factors. In the absence of any auxiliary factors, these integrases mediate unidirectional recombination between attP and attB with 80% efficiency. Due to the short DNA sequences that can be recognized by these integrases and the unidirectional recombination, these recombination systems were developed as a complement to the Cre-LoxP and FRT-FLP systems that are widely used for genetic purposes.
[0113] The term "matched RRS" indicates that recombination occurs between two RRSs. In certain embodiments, the two matched RRSs are the same. In certain embodiments, both RRSs are wild-type LoxP sequences. In certain embodiments, both RRSs are mutant LoxP sequences. In certain embodiments, both RRSs are wild-type FRT sequences. In certain embodiments, both RRSs are mutant FRT sequences. In certain embodiments, the two matched RRSs are different sequences from each other but can be recognized by the same recombinase. In certain embodiments, the first matched RRS is a Bxb1 attP sequence and the second matched RRS is a Bxb1 attB sequence. In certain embodiments, the first matched RRS is a φC31 attB sequence and the second matched RRS is a φC31 attB sequence.
[0114] In certain embodiments, a "single-vector RMCE" strategy is used. For example, in certain embodiments, the integrated exogenous nucleotide sequence comprises two RRSs, and the vector comprises two RRSs that match the two RRSs on the integrated exogenous nucleotide sequence; i.e., the first RRS on the integrated exogenous nucleotide sequence matches the first RRS on the vector, and the second RRS on the integrated exogenous nucleotide sequence matches the second RRS on the vector. In certain embodiments, the first RRS on the integrated exogenous nucleotide sequence and the first RRS on the vector are identical to the second RRS on the integrated exogenous nucleotide sequence and the second RRS on the vector. In certain embodiments, the first RRS on the integrated exogenous nucleotide sequence and the first RRS on the vector are different from the second RRS on the integrated exogenous nucleotide sequence and the second RRS on the vector. In certain embodiments, the first RRS on the integrated exogenous nucleotide sequence and the first RRS on the vector are both LoxP L3 sequences, and the second RRS on the integrated exogenous nucleotide sequence and the second RRS on the vector are both LoxP 2L sequences.
[0115] In certain embodiments, a "two-vector RMCE" strategy is used. For example, but not limited to, the integrated exogenous nucleotide sequence can include three RRSs, e.g., an arrangement in which the third RRS ("RRS3") is located between the first RRS ("RRS1") and the second RRS ("RRS2"), where the first vector includes two RRSs that match the first and third RRSs on the integrated exogenous nucleotide sequence, and the second vector includes two RRSs that match the third and second RRSs on the integrated exogenous nucleotide sequence. Such a two-vector RMCE strategy allows for the introduction of 10 SOIs by incorporating an appropriate number of SOIs between each pair of RRSs.
[0116] Both single-vector and two-vector RMCE involve the unidirectional integration of one or more donor DNA molecules into a predetermined site in the host cell genome, allowing for precise exchange of a DNA cassette present on the donor DNA with a DNA cassette on the host genome where the integration site resides. The DNA cassette features at least one selectable marker (although in certain two-vector RMCE examples, a "split selectable marker" can be used as outlined herein) and / or two heterospecific RRSs flanking at least one exogenous SOI. RMCE involves a double recombination crossover event, catalyzed by a recombinase, between two heterospecific RRSs within the target genomic locus and the donor DNA molecule. RMCE is designed to introduce a copy of the SOI or selectable marker into a predetermined locus in the host cell genome. Unlike recombination, which involves only a single crossover event, RMCE can be performed in a way that does not introduce prokaryotic vector sequences into the host cell genome, thereby reducing and / or preventing undesired triggering of host immune or defense mechanisms. The RMCE procedure can be repeated with multiple DNA cassettes; for example, in Figure 2B, the RMCE procedure is used to facilitate the introduction of three separate "front" cassettes and three separate "back" cassettes. However, as noted above, RMCE (and other integration strategies) can be used to introduce as few as one cassette and as many as ten or more cassettes.
[0117] In certain embodiments, targeted integration is achieved by a single crossover recombination event in which one exogenous nucleotide sequence comprising one RRS flanked by at least one exogenous SOI or at least one selectable marker is integrated into a predetermined site in the host cell genome. In certain embodiments, targeted integration is achieved by a single RMCE in which a DNA cassette comprising at least an exogenous SOI or at least one selectable marker flanked by two heterospecific RRSs is integrated into a predetermined site in the host cell genome. In certain embodiments, targeted integration is achieved by two RMCEs in which two different DNA cassettes, each comprising at least an exogenous SOI or at least one selectable marker flanked by two heterospecific RRSs, are integrated into a predetermined site in the host cell genome. In certain embodiments, targeted integration is achieved by multiple RMCEs in which DNA cassettes from multiple vectors, each comprising at least one exogenous SOI or at least one selectable marker flanked by two heterospecific RRSs, are integrated into a predetermined site in the host cell genome. In certain embodiments, the selectable marker may be partially encoded on a first vector and partially encoded on a second vector, such that integration of both RMCEs allows for expression of the selectable marker.
[0118] In certain embodiments, targeted integration via recombinase-mediated recombination results in a selectable marker or one or more exogenous SOIs that originate from the prokaryotic vector and are integrated into one or more predetermined integration sites in the host cell genome. In certain embodiments, targeted integration via recombinase-mediated recombination results in a selectable marker or one or more exogenous SOIs that are integrated into one or more predetermined integration sites in the host cell genome that do not contain sequences from the prokaryotic vector.
[0119] 5.2 Targeted integration via homologous recombination, HDR or NHEJ The subject matter of the present disclosure also relates to targeted integration mediated by homologous recombination or by exogenous site-specific nucleases followed by HDR or NHEJ. In certain embodiments, such integration is referred to herein as "gene editing-mediated integration."
[0120] Homologous recombination is the recombination between DNA molecules that share extensive sequence homology. It can be used to induce error-free repair of double-stranded DNA breaks, generating sequence diversity in gametes during meiosis. Homologous recombination involves the exchange of genetic information between two homologous DNA molecules and does not alter the overall arrangement of genes on chromosomes. During homologous recombination, a nick or break is formed in the double-stranded DNA (dsDNA), followed by invasion of the homologous dsDNA molecule by the single-stranded DNA end, pairing of the homologous sequences, branch migration to form Holliday junctions, and eventual Holliday junction resolution.
[0121] Double-strand breaks (DSBs) are the most severe form of DNA damage, and repair of such DNA damage is essential for maintaining genome integrity in all organisms. There are two main repair pathways for DSB repair. The first is the homologous recombination repair (HDR) pathway, and homologous recombination is the most common form of HDR. Because HDR requires the presence of homologous DNA within the cell, this repair pathway is usually active during the S and G2 phases of the cell cycle, when newly replicated sister chromatids are available as homologous templates. HDR is also the primary repair pathway for repairing collapsed replication forks during DNA replication. HDR is considered a relatively error-free repair pathway. The second repair pathway for DSBs is non-homologous end joining (NHEJ). NHEJ is a repair pathway in which the ends of broken DNA are joined together without the need for a homologous DNA template.
[0122] Targeted integration can be promoted by HDR followed by exogenous site-specific nucleases, such as gene editing nucleases. This is because the introduction of DSBs at specific target genomic sites can increase the frequency of homologous recombination. In certain embodiments, the exogenous nuclease can be selected from the group consisting of zinc finger nucleases (ZFNs), ZFN dimers, transcription activator-like effector nucleases (TALENs), TAL effector domain fusion proteins, RNA-guided DNA endonucleases, engineered meganucleases, and clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) endonucleases.
[0123] The CRISPR / Cas and TALEN systems are two genome editing tools that offer the easiest construction and highest efficiency. CRISPR / Cas was identified as a bacterial immune defense mechanism against invading bacteriophages. Cas is a nuclease that, when guided by a synthetic guide RNA (gRNA), associates with specific nucleotide sequences within a cell and can edit the DNA within or surrounding that nucleotide sequence by, for example, creating one or more single-strand breaks, DSBs, and / or point mutations. TALENs are engineered site-specific nucleases composed of the DNA-binding domain of a TALE (transcription activator-like effector) and the catalytic domain of the restriction endonuclease FokI. By varying the amino acids present in the highly variable residue region of the DNA-binding domain monomer, different artificial TALENs can be created that target various nucleotide sequences. The DNA-binding domain then directs the nuclease to the target sequence, creating DSBs.
[0124] Targeted integration by homologous recombination or HDR involves the presence of a homologous sequence to the integration site. In certain embodiments, the homologous sequence is present on a vector. In certain embodiments, the homologous sequence is present on a polynucleotide.
[0125] In certain embodiments, homologous recombination is performed without the use of any auxiliary factors. In certain embodiments, homologous recombination is facilitated by the presence of an integrating vector. In certain embodiments, the integrating vector is selected from the group consisting of an adeno-associated viral vector, a lentiviral vector, a retroviral vector, and an integrating phage vector.
[0126] 5.3 Adjusted Systems The subject matter of the present disclosure also relates to a controlled system for use in RCTI, known as "controlled targeted integration of randomized constructs" (also referred to herein as "controlled targeted integration of randomized strands") ("RCRTI"). For example, protein expression levels are often suboptimal, primarily due to the encoded protein being difficult to express. Low expression levels of difficult-to-express proteins can have a variety of causes, making them difficult to identify. One possibility is toxicity of the protein expressed in the host cell. In such cases, regulated expression systems can be used to express toxic proteins in which the protein-encoding sequence of interest is under the control of an inducible promoter. In these systems, expression of the difficult-to-express protein is promoted only when a regulator, such as a small molecule such as, but not limited to, tetracycline or its analog, doxycycline (DOX), is added to the culture. Regulating the expression of the toxic protein can mitigate toxic effects, allowing the culture to achieve the desired cell growth prior to production. In certain embodiments, a "controlled targeted integration of randomized constructs" (also referred to herein as "controlled targeted integration of randomized strands") (RCRTI) system comprises an SOI that is integrated into a specific locus, e.g., an exogenous nucleic acid sequence containing one or more RRSs, and transcribed under a regulated promoter operably linked thereto. In certain embodiments, the RCRTI system can be used to determine the underlying cause of low protein expression of difficult-to-express molecules, such as, but not limited to, antibodies. In certain embodiments, the ability to selectively silence expression of the SOI in the RCRTI system can be used to link expression of the SOI to observed adverse effects.
[0127] In certain embodiments, a "controlled targeted integration of randomized constructs" (also referred to herein as "controlled targeted integration of randomized strands") (RCRTI) system can be used to minimize the effects of transcriptional and cell line variation during root cause analysis of difficult-to-express molecules. For example, but not by way of limitation, expression of an SOI in a TI host can be induced by the addition of a regulatory factor, such as doxycycline, to the culture. In certain embodiments, the RCRTI vector utilizes a tetracycline-regulated promoter to express the SOI, which can be integrated into an exogenous nucleic acid sequence containing, for example, an RRS, which itself integrates into an integration site within the genome of the host cell, allowing for regulated expression of the SOI.
[0128] In certain embodiments, the RCRTI system described in this disclosure can be used to successfully determine the underlying cause of low protein expression of an SOI, e.g., a therapeutic antibody, compared to a control cell line. In certain embodiments, once relatively low expression of an SOI, e.g., a therapeutic antibody, in an RCRTI cell line is confirmed, the intracellular accumulation and secretion levels of the SOI can be assessed by utilizing protein translation inhibitor treatments, e.g., Dox and cycloheximide.
[0129] For example, but not by way of limitation, such regulation can be based on a gene switch to block or activate mRNA synthesis by regulated coupling of a transcriptional repressor or activator to a constitutive or minimal promoter. In certain non-limiting embodiments, repression can be achieved, for example, by binding to a repressor protein that sterically blocks transcription initiation, or by actively suppressing transcription via a transcriptional silencer. In certain non-limiting embodiments, activation of a mammalian or viral enhancer-less minimal promoter can be achieved by regulated coupling to an activation domain.
[0130] In certain embodiments, conditional coupling of a transcriptional repressor or transcriptional activator may be achieved by using an allosteric protein that binds to the promoter in response to an external stimulus. In certain embodiments, conditional coupling of a transcriptional repressor or transcriptional activator may be achieved by using an intracellular receptor that can bind to the target promoter upon release from a sequestering protein. In certain embodiments, conditional coupling of a transcriptional repressor or transcriptional activator may be achieved by using a chemically derived dimerizing agent.
[0131] In certain embodiments, the allosteric protein used in the TI system of the present disclosure can be a protein that regulates transcriptional activity in response to a culture parameter, such as an antibiotic, a bacterial quorum-sensing messenger, a catabolite, or temperature, e.g., cold or heat. In certain embodiments, such an RCRTI system can be catabolite-based, for example, in which a bacterial repressor that represses catabolic genes of alternative carbon sources has been introduced into mammalian cells. In certain embodiments, repression of the target promoter can be achieved by cumate-responsive binding of the repressor CymR. In certain embodiments, a catabolite-based system can rely on activation of a chimeric promoter by 6-hydroxynicotine-responsive binding of the prokaryotic repressor HdnoR fused to the herpes simplex VP16 transcription activation domain.
[0132] In certain embodiments, the TI system can be a prokaryotic quorum-sensing-based expression system that manages intra- and inter-population communication through quorum-sensing molecules. These quorum-sensing molecules bind to receptors in target cells, modulating the receptor's affinity for its cognate promoter and initiating specific regulon switches. In certain embodiments, the quorum-sensing molecule can be N-(3-oxo-octanoyl)-homoserine lactone, the presence of which activates expression from a minimal promoter fused to a TraR-specific operator sequence. In certain embodiments, the quorum-sensing molecule can be butyrolactone SCB1 (racemic 2-(1'-hydroxy-6-methylheptyl)-3-(hydroxymethyl)butanolide) in a system based on the ScbR repressor, Streptomyces coelicolor A3(2) SCB1, which binds to its cognate operator, OScbR, in the absence of SCB1. In certain embodiments, the quorum-sensing molecule can be a homoserine-derived inducer used in an RTI system, in which the Pseudomonas aeruginosa quorum-sensing repressors RhlR and LasR are fused to the SV40 T-antigen nuclear localization sequence and herpes simplex VP16 domain to activate promoters containing specific operator sequences (las boxes).
[0133] In certain embodiments, the inducer molecule that regulates the allosteric protein used in the RCRTI system of the present disclosure can be, but is not limited to, coumarate, isopropyl-β-D-galactopyranoside (IPTG), macrolides, 6-hydroxynicotine, doxycycline, streptogramins, NADH, and tetracycline.
[0134] In certain embodiments, the intracellular receptor used in the disclosed RCRTI system can be a cytoplasmic or nuclear receptor. In certain embodiments, the disclosed RCRTI system can utilize the release of transcription factors from protein capture and inhibition using small molecules. In certain embodiments, the disclosed RCRTI system can rely on steroid regulation, where a hormone receptor is fused to a natural or artificial transcription factor that can be released from HSP90 in the cytosol, translocate to the nucleus, and activate a selected promoter. In certain embodiments, mutant receptors regulated by synthetic steroid analogs can be used to avoid crosstalk with endogenous steroid hormones. In certain embodiments, the receptor can be a 4-hydroxytamoxifen-responsive estrogen receptor variant or a RU486-inducible progesterone receptor variant. In certain embodiments, a nuclear receptor-derived rosiglitazone-responsive transcriptional switch based on the human nuclear peroxisome proliferator-activated receptor γ (PPARγ) can be used in the disclosed RCRTI system. In certain embodiments, the steroid-responsive receptor variant can be RheoSwitch, which is based on a modified Choristoneura fumiferana ecdysone receptor and the mouse retinoid X receptor (RXR) fused to a Gal4 DNA-binding domain and a VP16 transactivator. In the presence of synthetic ecdysone, the RheoSwitch variant can bind to and activate a minimal promoter fused to several repeats of the Gal4 response element.
[0135] In certain embodiments, the RCRTI system disclosed herein may utilize chemically induced dimerization of DNA-binding proteins and transcriptional activators for activation of a minimal core promoter fused to a cognate operator. In certain embodiments, the RCRTI system disclosed herein may utilize rapamycin-regulated dimerization of FRB and FKBP. In this system, FRB is fused to the p65 transactivator, and FKBP is fused to a zinc finger domain specific for the cognate operator site placed upstream of an engineered minimal interleukin-12 promoter. In certain embodiments, the FKBP may be mutated. In certain embodiments, the RCRTI system disclosed herein may utilize the bacterial gyrase B subunit (GyrB), which dimerizes in the presence of the antibiotic coumermycin and dissociates with novobiocin.
[0136] In certain embodiments, the RCRTI system of the present disclosure can be used for regulated siRNA expression. In certain embodiments, the regulated siRNA expression system can be a tetracycline, a macrolide, or an OFF-type and ON-type QuoRex system. In certain embodiments, the RTI system can utilize Xenopus terminal oligopyrimidine elements (TOPs), which block translation initiation by forming a hairpin structure in the 5' untranslated region.
[0137] In certain embodiments, the RCRTI systems described in this disclosure may utilize gas-phase regulated expression, such as the acetaldehyde-inducible regulation (AIR) system. The AIR system may use the Aspergillus nidulans AlcR transcription factor to specifically activate a PAIR promoter assembled from an AlcR-specific operator fused to a minimal human cytomegalovirus promoter in the presence of non-toxic concentrations of gaseous or liquid acetaldehyde.
[0138] In certain embodiments, the RCRTI systems of the present disclosure may utilize a Tet-On or Tet-Off system, in which expression of one or more SOIs can be regulated by tetracycline or its analog, doxycycline.
[0139] In certain embodiments, the RCRTI system of the present disclosure may utilize a PIP on or PIP off system, in which expression of the SOI may be regulated by, for example, pristinamycin, tetracycline, and / or erythromycin.
[0140] 6. Products The host cells of the present disclosure can be used for the expression of any molecule of interest. In certain embodiments, the host cells of the present disclosure can be used for the expression of polypeptides, such as mammalian polypeptides. Non-limiting examples of such polypeptides include hormones, receptors, fusion proteins, regulatory factors, growth factors, complement system factors, enzymes, coagulation factors, anticoagulants, kinases, cytokines, CD proteins, interleukins, therapeutic proteins, diagnostic proteins, and antibodies. In certain embodiments, the host cells of the present disclosure can be used for the expression of chaperones, protein-modifying enzymes, shRNAs, gRNAs, or other proteins or peptides, constitutively or regulatedly expressing a therapeutic protein or molecule of interest.
[0141] In certain embodiments, the polypeptide of interest is a bispecific, trispecific, or multispecific polypeptide, such as a bispecific antibody.
[0142] The host cells of the present disclosure can be used to produce large quantities of a molecule of interest in a shorter time frame than cells used in conventional cell culture methods, e.g., non-TI cells. In certain embodiments, the host cells of the present disclosure can be used to improve the quality of a molecule of interest compared to cells used in conventional cell culture methods, e.g., non-TI cells. In certain embodiments, the host cells of the present disclosure can be used to enhance seed train stability by preventing chronic toxicity that can be caused by products that can cause cellular stress and clonal instability over time. In certain embodiments, the host cells of the present disclosure can be used for optimal expression of acutely toxic products.
[0143] In certain embodiments, the host cells and systems of the present disclosure may be used in cell culture process optimization and / or process development.
[0144] In certain embodiments, the host cells of the present disclosure may be used for the constitutive expression of selected subunits of a therapeutic molecule and the regulated expression of other, different subunits of the same therapeutic molecule. In certain embodiments, the therapeutic molecule may be a fusion protein. In certain embodiments, the host cells of the present disclosure may be used to understand the role and effect of each antibody subunit in the expression and secretion of a fully assembled antibody molecule.
[0145] In certain embodiments, the host cells of the present disclosure can be used as an investigative tool. In certain embodiments, the host cells of the present disclosure can be used as a diagnostic tool to determine the root cause of low protein expression of problematic molecules in various cells. In certain embodiments, the host cells of the present disclosure can be used to directly link observed phenomena or cellular behavior to transgene expression in the cell. The host cells of the present disclosure can also be used to demonstrate whether observed behavior is reversible in the cell. In certain embodiments, the host cells of the present disclosure can be used to identify and alleviate problems with transgene transcription and expression in cells.
[0146] In certain embodiments, host cells of the present disclosure may be used to exchange transgene subunits of difficult-to-express molecules, such as, but not limited to, the HC and LC subunits of an antibody, with those of the average molecule in the system to identify problematic subunits. In certain embodiments, amino acid sequence analysis may then be used to narrow down and focus on amino acid residues or regions that may be responsible for low protein expression. [Example]
[0147] Materials and Methods cell culture Stable cell line development was performed by targeted integration of an antibody-encoding cassette into a host cell line derived from the CHO-K1 strain (Crawford Y. et al., Biotechnol Prog 2013, 29, 1307-1315). Cells were cultured in proprietary medium at 37°C and 5% CO2 in 125 mL shake flasks at 150 rpm or in 50 mL tube spin bioreactors at 225 rpm. Cultured cells were cultured at 4 x 10 5 The cells were passaged every 3–4 days at a seeding density of 1000 cells / mL.
[0148] PCR reactions to determine the vector organization of RCTI clones expressing molecule X 2 × 10 using the DNeasy Blood and Tissue Kit (cat#69506, Qiagen) 6 Genomic DNA was extracted from the cells and PCR was performed using LongAmp Taq Master Mix (Cat. No. M 0287, New England Biolabs). Thermal cycling parameters were 94°C for 3 minutes, 40x (94°C for 30 seconds, 60°C for 1 minute, and 65°C for 9.75 minutes) and 65°C for 20 minutes. Diagnostic digests of PCR products were performed using enzymes from New England Biolabs. The following primers were used for PCR: forward primer: GGTTCTCCTTGACCAATACCTCGTAAG; reverse primer: GCGGGACTATGGTTGCTGACTAAT.
[0149] Shake flask fed-batch and Ambr™ bioreactor production assays Feed-batch evaluation of clones expressing monoclonal antibody A ("mAb A") and molecule Z was performed in shake flasks using a proprietary chemically defined medium during a 14-day manufacturing process. Cells were cultured at 1 x 10 cells per day. 6 Cultures were seeded at 1000 cells / mL, and the cultures were temperature shifted from 37°C to 35°C on day 3, and a bolus feed consisting of a proprietary blend of chemically defined components was added on days 3, 7, and 10. Viability and viable cell count were measured using a Vi-Cell XR (Beckman Coulter), and lactate and glucose levels were assayed using a Nova BioProfile 400 (Nova Biomedical).
[0150] For evaluation of clones expressing molecule Y, a 14-day production culture was performed in the ambr™ system (Sartorius) as described (Hsu, WT et al., Cytotechnology 2012, 64, 667-678) with the following modifications: an inoculum was added to a shake flask containing 1 × 10 6 The production culture was inoculated at 2 × 10 cells / mL for 4 days at 37 °C and 150 rpm. 6 Cells were inoculated at 0.05 cells / mL and initially maintained at 36°C, followed by a temperature shift to 35°C on day 6. A bolus feed consisting of a proprietary blend of chemically defined components was added on days 3, 6, 8, and 10. Viability and cell count were measured with a BioProfile Flex 2 (Nova Biomedical). pH, pO2, pCO2, glucose, lactate, and other metabolites were assayed with either a BioProfile Flex 2 (Nova Biomedical) or an ABL90 Flex (Radiometer).
[0151] Analytical assays for titer and product quality Titers were measured using protein A affinity chromatography with UV detection. Size and charge variants were measured by protein A PhyTip purification (PhyNexus) followed by size exclusion chromatography and imaging capillary isoelectric focusing (icIEF), respectively. Phytium-purified samples were treated with carboxypeptidase B before analysis by icIEF. For molecule Y, liquid chromatography-mass spectrometry (LC-MS) was used to quantify the relative amounts of correctly assembled heterodimers, homodimers, half-antibodies, and light chain mispair species, as described (Williams, AJ et al., Ind Eng Chem Res 2017, 56, 1713-1722).
[0152] Comparison of current transfection (TFX) strategies with TFX strategies for RCTI Stable cell lines producing monoclonal antibody A ("mAb A") (Figure 11A) were generated using a "randomized constitutive targeted integration" (also referred to herein as "randomized chain-targeted integration") approach. In this approach, host cells were transfected with a library of three "front" and three "back" expression vectors, each containing one of three possible combinations of mAb A heavy or light chain sequence repeats. As used in this example, references to the "front" and "back" expression vectors refer to the upstream (front) and downstream (back) cassette sites in the context of two-vector RCME-based introduction of exogenous sequences of interest. For further details regarding the two-vector RCME-based strategy, see, for example, Section 5.1 above. Top clones produced titers of 3.4-4.7 g / L and were selected based on titer, productivity, cell culture performance, product quality, and flow cytometry population heterogeneity.
[0153] Compared to previous cell line development (CLD) efforts for mAb A, the RCTI CLD method described herein produced clones with titer, product quality, and production culture performance comparable to clones produced by standard CLD methods, while screening fewer clones in fewer CLD cycles.
[0154] In previous CLD efforts for mAb A, individual combinations of front and back vectors were tested in a separate standard CLD workflow (see Figure 2A). A total of seven combinations of front and back vector configurations were transfected and evaluated by pool production. Of these seven vector combinations, four were carried forward for clonal evaluation based on SCC and pool production titers. Therefore, four independent CLD cycles were performed, along with three additional partial CLD cycles, in which three additional vector combinations were transfected and evaluated by pool production.
[0155] In comparison, a single CLD cycle was performed for the RCTI CLD method. Table 1 summarizes the number of clones evaluated for each method, and Figures 2A and 2B show a comparison of the RCTI and standard CLD workflows for mAb A clone evaluation. Table 1. Summary of screened clones for standard versus RCTI CLD. JPEG2025121945000003.jpg38170
[0156] For the RCTI CLD method, three experimental arms were evaluated, corresponding to transfected pools at three stages of recovery: (i) early (day 5 post-transfection, sorted into selective medium), (ii) mid-stage (clones recovered to approximately 35% viability after selection), or (iii) complete (clones recovered to >90% viability after selection) (see Figure 6A). For the purpose of comparing the RCTI CLD method with the standard CLD method, only clones from the third arm (clones recovered to >90% viability after selection) are considered in this section, since standard CLD designates SCC for pools with >90% viability. For both the standard CLD and RCTI methods, one CLD cycle is defined as the entire process from transfection to clonal production culture evaluation.
[0157] In previous mAb A CLD efforts, the majority of the highest titer clones also exhibited high aggregates, i.e., HMWS (%) exceeding 10-15%. Therefore, when comparing clones produced by the standard CLD method with the RCTI method, we considered both titer and HMWS (%) in selecting the top RCTI CLD clones. For clones produced by the standard CLD method, the distribution of HMWS (%) roughly clustered between the front vector and back vector configurations for some constructs (Figure 3B). For example, clones with the C-2 construct (Figure 3A) generally produced lower HMWS (%), while clones with the B-2 construct (Figure 3A) produced higher HMWS (%).
[0158] For RCTI clones, the range of titers and HMWS (%) was comparable to those produced by standard CLD clones (Figure 3B), although far fewer clones were screened. Because the RCTI method involves transfection of a random library of front and back DNA vectors, long-range PCR was used to determine the front vector structure of clones evaluated in production cultures. Although three of the 24 clones had inconclusive long-range PCR results, the remaining clones clustered similarly in terms of front vector structure and HMWS (%) profile when compared with standard CLD clones. The majority of clones with C front vector structure exhibited lower HMWS (%) than clones with B structure, while clones with A front structure spanned a range of HMWS (%). Because HMWS (%) levels directly correlate with front vector structure alone, determining the back vector structure was not necessary here. Furthermore, the highest day 14 titers of RCTI clones were comparable to those of standard CLD clones (≥4.5 g / L). Note that when only the B-2 construct (Fig. 3A) or the A-1 construct (Fig. 3A) were tested with the standard CLD approach, few, if any, clones with the combination of high titer and low aggregates were obtained.
[0159] Growth, productivity, and other product quality attributes were also comparable between the standard CLD and RCTI CLD methods. Figure 5 summarizes the distribution of various attributes across all clones evaluated in production cultures for both CLD methods. Similarly, the top clones produced by the RCTI CLD method performed comparably in production culture to the top clones produced by the standard CLD method. Of the configurations tested with the standard CLD method, only configurations A-2 and C-2 yielded high titers (≥3.6 g / L) while producing low aggregates (HMWS (%) ≤15.4%). Clones obtained from the RCTI approach had similar aggregation and titer ranges to those obtained with the standard CLD approach. While a representative high-producing clone from the RCTI approach appears to be configuration C in Table 2, high-titer and low-aggregation clones from configuration A were also observed, as shown in Figure 3B, but they were not ranked among the highest-producing clones. Table 2 compares the titer, growth, and product quality attributes of the top clones produced by the standard and RCTI CLD approaches. Table 2. Comparison of top mAb A clones produced using standard and RCTI CLDs JPEG2025121945000004.jpg96170
[0160] Compared to previous CLD efforts on mAb A using standard CLD methods, the RCTI method offers several advantages. First, a single RCTI CLD cycle can be used to screen the same, if not more, number of vector constructs as standard CLD cycles while using fewer resources. For mAb A, the standard CLD method required 4+ CLD cycles to screen seven of the nine possible vector construct combinations, whereas the RCTI method screened all nine combinations in a single CLD cycle. Fewer individual clones can be screened as well. Over all standard CLD cycles for mAb A, a total of 4,224 clones were picked from SCC and narrowed down to 120 clones for screening during production culture. In comparison, a total of 704 clones were picked from SCC for the RCTI approach, and 24 clones were screened in production culture. Due to the reduced number of CLD cycles and clones to screen, the RCTI method allows for a reduction in CLD labor requirements, including both reduced manual labor and the burden on automation resources to perform SCC, imaging, and hit picking. Finally, despite the reduced number of CLD cycles required and the reduced number of clones screened, the RCTI method produced clones with comparable titer, cell culture performance, and product quality distributions to the standard CLD method. The top clones produced from both methods were also comparable.
[0161] Evaluation of RCTI CLD for bispecific antibody expression The RCTI approach was used to evaluate the expression of a bispecific antibody ("molecule Y") for a disease indication with high clinical demand, thus necessitating the isolation of high-titer cell lines during the clone screening process. Expression of bispecific antibodies in a single host can result in the collection of undesired by-products, such as homodimeric and heterodimeric heavy chains with common or mismatched light chain species (Dillon M. et al., mAbs 2017, 9, 213-230; Carter, PJ, Experimental cell research 2011, 317, 1261-1269). Therefore, a critical step during bispecific antibody clone evaluation is the isolation of clones with not only high total titers but also high effective titers (the desired bispecific antibody species) with minimal levels of by-product(s) that are difficult to purify from the target bispecific antibody. For molecule Y (Figure 11B), the most difficult by-product to purify was the heterodimeric heavy chain with a common light chain species (referred to herein as "common LC BsAb"). Therefore, during the standard TI CLD process, four different vector constructs were used in parallel CLD runs, and mass spectrometry was used to measure the common LC BsAb and effective titer levels during pool production evaluation before SCC. The pool with the lowest level of common LC BsAb was selected for SCC, and the top 62 clones with the highest titer and lowest level of common LC BsAb were then evaluated in a production assay (Figure 7A, squares) for the standard TI CLD approach.
[0162] Because we previously established that the recovered pools still maintained heterogeneity in vector configuration in the RCTI CLD approach (Figure 6B and Figure 6C), for the RCTI CLD of molecule Y, all four vector constructs were used to transfect TI hosts, and SCC was performed only after complete pool recovery. Furthermore, during the initial screening step, clones were ranked based solely on their effective titers using rapid combustion mass spectrometry. During this process, the levels of common LC BsAb or other individual byproducts were not considered for clone selection. Based on these criteria, the top 36 RCTI clones were then evaluated in a production assay (Figure 7A, circles). Despite evaluating only half the number of clones in the production assay, we were able to isolate clones from the RCTI CLD approach with relatively high total (approximately 6 g / L) and effective (4-5.5 g / L) titers compared to standard TI CLD clones (Figure 7A). Furthermore, RCTI clones had overall better specific productivity (Figure 7B), better viability (Figure 8A), and slightly lower growth (Figure 8B) profiles compared with standard TI clones. Because it is easier to improve culture growth through process optimization, RCTI clones expressing molecule Y may have even higher total and effective titers than their standard TI clone counterparts. RCTI clones also had very low (approximately 5% or less) common LC BsAb by-products, nearly as low as those observed with standard TI clones (Figure 8C). However, standard TI clones were expected to have lower overall common LC BsAb by-products because they were screened for this trait prior to evaluation in production cultures. The product quality of RCTI clones was comparable to or better than their standard TI counterparts. RCTI clones expressing molecule Y also had relatively lower aggregates (Figure 8D), lower acidic species (%) (Figure 8E), and higher main species (%) (Figure 8F) compared with standard TI clones, all of which are more desirable product quality parameters. The % levels of basic species were relatively comparable between clones isolated from either approach (Figure 8G).
[0163] Evaluation of RCTI CLD for the expression of composite chimeric antibody / ligand molecules Many upcoming therapeutic molecules for various disease indications are no longer simply conventional antibodies but rather have more complex structures and are therefore more difficult to express. Expression of one such complex molecule ("molecule Z") in a standard TI approach required parallel CLD efforts using six different vector constructs followed by screening 96 different clones in production cultures to obtain clones with high effective titers (Figure 9A, squares). Thus, molecule Z, which contains a half-antibody complexed with an Fc fusion ligand (Figure 11C), was considered a challenging molecule to express, and we decided to test its expression using the RCTI CLD approach. Using a mixture of all six different vector constructs to transfect TI host cells and select them in a single pool followed by SCC, evaluation of only 23 RCTI clones in production assays resulted in the isolation of clones with titers and product quality comparable to the top clones isolated by the standard TI CLD approach (Figure 9A). Compared to standard TI clones, RCTI clones had higher overall specific productivity (Figure 9B), comparable viability (Figure 10A), slightly lower growth (Figure 10B), and lower overall aggregate % (Figure 10C). Because of their significantly higher specific productivity (approximately 2-fold higher, Figure 9B), RCTI clones could potentially achieve even higher effective titers if cell growth is further increased through process optimization. Finally, RCTI clones had comparable acidic % (Figure 10D), main % (Figure 10E), and basic % (Figure 10F) charge variants to standard TI clones. Overall, these data support the idea that the RCTI approach can be used to isolate clones with comparable titer and product quality attributes to the standard TI approach, but with significantly fewer resources and effort.
[0164] The foregoing examples are merely illustrative of the subject matter disclosed herein and should not be construed as limiting in any way.
[0165] In addition to the various embodiments depicted and claimed, the presently disclosed subject matter is directed to other embodiments having other combinations of the features disclosed and claimed herein. Thus, specific features presented herein may be combined with each other in other ways within the scope of the presently disclosed subject matter, such that the presently disclosed subject matter includes any suitable combination of features disclosed herein. The foregoing descriptions of specific embodiments of the presently disclosed subject matter have been presented for purposes of illustration and description and are not intended to be exhaustive or to limit the presently disclosed subject matter to the disclosed embodiments.
[0166] It will be apparent to those skilled in the art that various modifications and variations can be made in the compositions and methods of the presently disclosed subject matter without departing from the spirit or scope of the presently disclosed subject matter. Thus, it is intended that the presently disclosed subject matter cover modifications and variations that come within the scope of the appended claims and their equivalents.
[0167] Various publications, patents, and patent applications are cited herein, the contents of which are incorporated herein by reference in their entireties.
Claims
1. 1. A method for generating and high-throughput screening a library of targeted integration (TI) host cells expressing at least one sequence of interest (SOI), comprising: a) providing said library of TI host cells comprising a plurality of TI host cells expressing one or more SOIs; i) providing a plurality of TI host cells; ii) contacting said plurality of TI host cells with a plurality of vectors comprising one or more SOIs; iii) introducing said one or more SOIs into one or more of said plurality of TI host cells; b) separating the library into single clones; c) screening said clones for particular cell or product attributes.
2. 10. The method of claim 1, wherein the one or more SOIs are introduced into one or more of the plurality of TI host cells by recombinase-mediated integration.
3. 10. The method of Claim 1, wherein the one or more SOIs are introduced into one or more of the plurality of TI host cells by gene editing-mediated integration.
4. 10. The method of claim 1, wherein the one or more SOIs are operably linked to one or more regulatable promoters.
5. 5. The method of claim 4, wherein the one or more regulatable promoters are selected from the group consisting of SV40 and CMV promoters.
6. The method of claim 1 , wherein the TI host cell is a mammalian host cell.
7. 7. The method of claim 6, wherein the TI host cell is a hamster host cell, a human host cell, a rat host cell, or a mouse host cell.
8. 8. The method of claim 7, wherein the TI host cell is a CHO host cell, a CHO K1 host cell, a CHO K1SV host cell, a DG44 host cell, a DUKXB-11 host cell, a CHOK1S host cell, or a CHO K1M host cell.
9. The method of claim 1 , wherein the SOI encodes a polypeptide subunit of a multi-subunit protein or a fragment thereof.
10. 2. The method of claim 1, wherein the SOI encodes a single chain antibody, an antibody light chain, an antibody heavy chain, a single chain Fv fragment (scFv), or an Fc fusion protein.
11. 2. The method of claim 1, wherein the specific cell attribute is selected from cell proliferation, cell titer, specific productivity, volumetric productivity, and clonal stability.
12. 2. The method of claim 1, wherein the specific product attribute is selected from the group consisting of level of glycosylation, level of charge dispersion, reduced mismatches, reduced protein / peptide aggregation, and protein sequence heterogeneity.
13. 2. The method of claim 1, wherein the one or more SOIs are introduced into one or more of the plurality of TI host cells at one or more loci that are at least about 90% homologous to a sequence selected from the following: SEQ ID NOs: 1-12; NW_006874047.1; NW_006884592.1; NW_006881296.1; NW_003616412.1; NW_003615063.1; NW_006882936.1; and NW_003615411.
1.
14. 1. A method for generating a library of TI host cells comprising a plurality of exogenous nucleotide SOIs, comprising: a) providing a plurality of TI host cells; b) contacting the plurality of TI host cells with a plurality of vectors comprising one or more SOIs; c) introducing said one or more SOIs into one or more of said plurality of TI host cells.
15. 15. The method of claim 14, wherein one or more of the plurality of TI host cells comprises one or more exogenous nucleotide sequences integrated into one or more loci in the genome of the TI host cell, the exogenous nucleotide sequences comprising at least two RRSs flanking at least one first selectable marker.
16. The vector is a) at least two RRSs that match the at least two RRSs on the integrated exogenous nucleotide sequence; 15. The method of claim 14, wherein the method further comprises: b) one or more exogenous SOIs and at least one second selection marker flanking the RRS.
17. 17. The method of claim 16, wherein the one or more SOIs are introduced into one or more of the plurality of TI host cells by recombinase-mediated integration.
18. 15. The method of Claim 14, wherein the one or more SOIs are introduced into one or more of the plurality of TI host cells by gene editing-mediated integration.
19. 15. The method of claim 14, wherein the one or more SOIs are operably linked to one or more regulatable promoters.
20. 20. The method of claim 19, wherein the one or more regulatable promoters are selected from the group consisting of SV40 and CMV promoters.
21. 15. The method of claim 14, wherein the SOI encodes a polypeptide subunit of a multi-subunit protein or a fragment thereof.
22. 15. The method of claim 14, wherein the SOI encodes a single chain antibody, an antibody light chain, an antibody heavy chain, a single chain Fv fragment (scFv), or an Fc fusion protein.
23. 15. The method of claim 14, wherein the one or more exogenous nucleotide sequences of (a) are integrated at one or more loci that are at least about 90% homologous to a sequence selected from the following: SEQ ID NOs: 1-12; NW_006874047.1; NW_006884592.1; NW_006881296.1; NW_003616412.1; NW_003615063.1; NW_006882936.1; and NW_003615411.
1.
24. 15. The method of claim 14, wherein the at least one sequence of interest comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequences of interest.
25. 15. The method of claim 14, wherein the TI host cell is a mammalian host cell.
26. 26. The method of claim 25, wherein the TI host cell is a hamster host cell, a human host cell, a rat host cell, or a mouse host cell.
27. 27. The method of claim 26, wherein the TI host cell is a CHO host cell, a CHO K1 host cell, a CHO K1SV host cell, a DG44 host cell, a DUKXB-11 host cell, a CHOK1S host cell, or a CHO K1M host cell.
28. 1. A method for preparing a TI host cell expressing one or more SOIs, comprising: a) providing a plurality of TI host cells; b) contacting the plurality of TI host cells with a plurality of vectors comprising one or more SOIs; c) introducing said one or more SOIs into one or more of said plurality of TI host cells; d) selecting a TI host cell that expresses said one or more SOIs.
29. 29. The method of claim 28, wherein the one or more SOIs are introduced into one or more of the plurality of TI host cells by recombinase-mediated integration.
30. 29. The method of Claim 28, wherein the one or more SOIs are introduced into one or more of the plurality of TI host cells by gene editing-mediated integration.
31. 29. The method of claim 28, wherein the one or more SOIs are operably linked to one or more regulatable promoters.
32. 32. The method of claim 31 , wherein the one or more regulatable promoters are selected from the group consisting of SV40 and CMV promoters.
33. 29. The method of claim 28, wherein the TI host cell is a mammalian host cell.
34. 34. The method of claim 33, wherein the TI host cell is a hamster host cell, a human host cell, a rat host cell, or a mouse host cell.
35. 35. The method of claim 34, wherein the TI host cell is a CHO host cell, a CHO K1 host cell, a CHO K1SV host cell, a DG44 host cell, a DUKXB-11 host cell, a CHOK1S host cell, or a CHO K1M host cell.
36. 29. The method of claim 28, wherein the SOI encodes a polypeptide subunit of a multi-subunit protein or a fragment thereof.
37. 29. The method of claim 28, wherein the SOI encodes a single chain antibody, an antibody light chain, an antibody heavy chain, a single chain Fv fragment (scFv), or an Fc fusion protein.
38. 29. The method of claim 28, wherein the one or more SOIs are introduced into one or more of the plurality of TI host cells at one or more loci that are at least about 90% homologous to a sequence selected from the following: SEQ ID NOs: 1-12; NW_006874047.1; NW_006884592.1; NW_006881296.1; NW_003616412.1; NW_003615063.1; NW_006882936.1; and NW_003615411.
1.
39. 29. The method of claim 28, wherein the SOI encodes a polypeptide subunit of a multi-subunit protein or a fragment thereof.
40. 29. The method of claim 28, wherein the SOI encodes a single chain antibody, an antibody light chain, an antibody heavy chain, a single chain Fv fragment (scFv), or an Fc fusion protein, and fragments thereof.
41. 29. The method of claim 28, a) one or more of the plurality of TI host cells comprises one or more exogenous nucleotide sequences integrated into one or more loci in the genome of the TI host cell, the exogenous nucleotide sequences comprising at least two RRSs flanking at least one first selectable marker; b) the vector is a. at least two RRSs that match the at least two RRSs on the integrated exogenous nucleotide sequence; b. one or more exogenous SOIs and at least one second selectable marker flanking said RRS; c) introducing into one or more of said plurality of TI host cells one or more recombinases or nucleic acids encoding recombinases that recognize said RRS; d) selecting for TI cells that express said second selectable marker, thereby isolating TI host cells that express the sequence of interest.
42. 42. The method of claim 41, wherein the SOI encodes a polypeptide subunit of a multi-subunit protein or a fragment thereof.
43. 42. The method of claim 41, wherein the SOI encodes a single chain antibody, an antibody light chain, an antibody heavy chain, a single chain Fv fragment (scFv), or an Fc fusion protein, and fragments thereof.
44. 42. The method of claim 41, wherein the one or more exogenous nucleotide sequences of (a) are integrated at one or more loci that are at least about 90% homologous to a sequence selected from the following: SEQ ID NOs: 1-12; NW_006874047.1; NW_006884592.1; NW_006881296.1; NW_003616412.1; NW_003615063.1; NW_006882936.1; and NW_003615411.
1.
45. 42. The method of claim 41 , a. the exogenous nucleotide sequence comprises a first RRS and a second RRS flanking at least one first selectable marker, and a third RRS located between the first RRS and the second RRS, wherein all of the RRSs are heterospecific; b. the plurality of vectors i. a first vector comprising two RRSs that match the first and third RRSs on the integrated exogenous nucleotide sequence and flank at least one first exogenous SOI and at least one second selectable marker; ii. a second vector comprising two RRSs that match the second and third RRSs on the integrated exogenous nucleotide sequence and flank at least one second exogenous SOI; c. selecting for TI cells that express said second selectable marker, thereby isolating TI host cells that express the first and second sequences of interest.
46. 46. The method of any one of claims 28, 41, and 45, wherein the vector comprising at least one sequence of interest comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequences of interest.
47. 46. The method of any one of claims 28, 41, and 45, wherein the TI host cell is a mammalian host cell.
48. 48. The method of claim 47, wherein the TI host cell is a hamster host cell, a human host cell, a rat host cell, or a mouse host cell.
49. 49. The method of claim 48, wherein the TI host cell is a CHO host cell, a CHO K1 host cell, a CHO K1SV host cell, a DG44 host cell, a DUKXB-11 host cell, a CHOK1S host cell, or a CHO K1M host cell.
50. 1. A method for expressing an SOI, comprising: a) providing a plurality of TI host cells; b) contacting the plurality of TI host cells with a plurality of vectors comprising one or more SOIs; c) introducing said one or more SOIs into one or more of said plurality of TI host cells; d) selecting TI host cells that express the sequence of interest; e) culturing the cells of d) under conditions suitable for expression of said sequence of interest.
51. 51. The method of claim 50, wherein the one or more SOIs are introduced into one or more of the plurality of TI host cells by recombinase-mediated integration.
52. 51. The method of Claim 50, wherein said one or more SOIs are introduced into one or more of said plurality of TI host cells by gene editing-mediated integration.
53. 51. The method of claim 50, wherein the one or more SOIs are operably linked to one or more regulatable promoters.
54. 54. The method of claim 53, wherein the one or more regulatable promoters are selected from the group consisting of SV40 and CMV promoters.
55. 51. The method of claim 50, wherein the TI host cell is a mammalian host cell.
56. 56. The method of claim 55, wherein the TI host cell is a hamster host cell, a human host cell, a rat host cell, or a mouse host cell.
57. 57. The method of claim 56, wherein the TI host cell is a CHO host cell, a CHO K1 host cell, a CHO K1SV host cell, a DG44 host cell, a DUKXB-11 host cell, a CHOK1S host cell, or a CHO K1M host cell.
58. 51. The method of claim 50, wherein the SOI encodes a polypeptide subunit of a multi-subunit protein or a fragment thereof.
59. 51. The method of claim 50, wherein the SOI encodes a single chain antibody, an antibody light chain, an antibody heavy chain, a single chain Fv fragment (scFv), or an Fc fusion protein.
60. 51. The method of claim 50, wherein the one or more SOIs are introduced into one or more of the plurality of TI host cells at one or more loci that are at least about 90% homologous to a sequence selected from the following: SEQ ID NOs: 1-12; NW_006874047.1; NW_006884592.1; NW_006881296.1; NW_003616412.1; NW_003615063.1; NW_006882936.1; and NW_003615411.
1.
61. 51. The method of claim 50, wherein the at least one sequence of interest comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequences of interest.
62. 1. A method for expressing an SOI, comprising: a) providing a plurality of TI host cells, each TI host cell comprising one or more exogenous nucleotide sequences integrated into one or more loci in the genome of said TI host cell, said exogenous nucleotide sequences comprising at least two RRSs flanking at least one first selectable marker; b) subjecting said plurality of TI host cells to one of: a. at least two RRSs that match the at least two RRSs on the integrated exogenous nucleotide sequence; b. contacting the vectors with one or more exogenous SOIs and at least one second selectable marker flanking the RRS; c) introducing one or more recombinases or nucleic acids encoding recombinases that recognize said RRS; d) selecting for TI cells that express the second selectable marker, thereby isolating TI host cells that express the sequence of interest; e) culturing the cells of d) under conditions suitable for expression of said sequence of interest and recovering the product of said sequence of interest therefrom.
63. 63. The method of claim 62, wherein the SOI encodes a polypeptide subunit of a multi-subunit protein or a fragment thereof.
64. 63. The method of claim 62, wherein the one or more SOIs are operably linked to one or more regulatable promoters.
65. 65. The method of claim 64, wherein the one or more regulatable promoters are selected from the group consisting of SV40 and CMV promoters.
66. 63. The method of claim 62, wherein the SOI encodes a single chain antibody, an antibody light chain, an antibody heavy chain, a single chain Fv fragment (scFv), or an Fc fusion protein, and fragments thereof.
67. 63. The method of claim 62, wherein the one or more exogenous nucleotide sequences of (a) are integrated at one or more loci that are at least about 90% homologous to a sequence selected from the following: SEQ ID NOs: 1-12; NW_006874047.1; NW_006884592.1; NW_006881296.1; NW_003616412.1; NW_003615063.1; NW_006882936.1; and NW_003615411.
1.
68. 63. The method of claim 62, a. the exogenous nucleotide sequence comprises a first RRS and a second RRS flanking at least one first selectable marker, and a third RRS located between the first RRS and the second RRS, wherein all of the RRSs are heterospecific; b. the plurality of vectors i. a first vector comprising two RRSs that match the first and third RRSs on the integrated exogenous nucleotide sequence and flank at least one first exogenous SOI and at least one second selectable marker; ii. a second vector comprising two RRSs that match the second and third RRSs on the integrated exogenous nucleotide sequence and flank at least one second exogenous SOI; c. selecting for TI cells that express said second selectable marker, thereby isolating TI host cells that express the first and second sequences of interest.
69. 69. The method of claim 62 or 68, wherein the at least one sequence of interest comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequences of interest.
70. 63. The method of claim 62, wherein the at least one first SOI and the at least one second SOI are operably linked to one or more regulatable promoters.
71. 71. The method of claim 70, wherein the one or more regulatable promoters are selected from the group consisting of SV40 and CMV promoters.
72. 69. The method of claim 62 or 68, wherein the TI host cell is a mammalian host cell.
73. 73. The method of claim 72, wherein the TI host cell is a hamster host cell, a human host cell, a rat host cell, or a mouse host cell.
74. 74. The method of claim 73, wherein the TI host cell is a CHO host cell, a CHO K1 host cell, a CHO K1SV host cell, a DG44 host cell, a DUKXB-11 host cell, a CHOK1S host cell, or a CHO K1M host cell.