Generation of landing pad cell lines
Patent Information
- Application Number
- JP2024539465
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-29
- Filing Date
- 2022-12-28
- Publication Date
- 2026-01-09
AI Technical Summary
Existing methods for generating cell lines with targeted gene integration are inefficient and unpredictable, often resulting in random integration and require extensive time to achieve high titer and stable protein expression.
A method for generating landing pad cells by selecting parent cells with high expression titer and low copy number of a gene of interest, screening for loss of parental plasmid, and integrating a landing pad plasmid at a targeted genomic site using site-specific recombination, such as CRISPR or Cre/Lox technology, to ensure reliable and reproducible protein expression.
This approach allows for the rapid generation of cell lines with high protein expression stability and reduced variability, achieving three to four times higher titers than traditional random integration methods.
Smart Images

Figure 00000135_0000 
Figure 00000136_0000 
Figure 00000136_0001
Abstract
Description
[Technical Field]
[0001] Reference to sequence listings submitted electronically via EFS-WEB The contents of the electronically submitted sequence listing (Name: 3338_196PC02_Seqlisting_ST26.txt, Size: 45,039,473 bytes, and Creation Date: December 27, 2022) are incorporated herein by reference in their entirety.
[0002] The present disclosure provides methods for the generation of landing pad cells suitable for targeted gene integration. [Background technology]
[0003] Historically, cell lines have been generated by transfecting cells with expression plasmid DNA, usually in a linearized state, allowing it to integrate into the cellular genome in an essentially random manner. Only cells containing the expression plasmid survive because the plasmid confers a selective advantage, such as drug resistance or auxotrophic complementation. It is desirable to minimize the time and increase predictability required to generate cell lines expressing a biologic of interest while maintaining acceptable levels of performance, such as high titer, post-translational modifications, expression stability, cell density, and bioreactor viability, among other parameters.
[0004] One way to achieve this goal is to use a technique called targeted integration (TI), in which an expression cassette or expression plasmid is inserted into the same locus of a cell line during cell line construction. The targeted locus is also called a landing pad, and a cell line containing a landing pad is also called a landing pad cell line. The expression cassette or expression plasmid can be integrated into the landing pad by site-directed recombination, such as using Cre / Lox technology, sometimes called recombination-mediated cassette exchange (RMCE) or site-specific recombination (SSR), or by using homologous recombination stimulated by CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats), TALEN, or other such site-specific nucleases.
[0005] A major challenge for TI is identifying the cell line and locus to be used. The cell line itself must be able to function well, and the landing pad must be located at a locus (a "hot spot") where high transcription occurs and transcription of the biologic is not silenced by epigenetic modifications or other mechanisms. TI landing pad host cell lines must have a "hot spot" in the chromosome for high expression, but it is understood that this "hot spot" requires a "hot cell" environment that supports all of the intermediate steps necessary for high protein expression of the biologic. Generating and identifying landing pad cell lines is extremely difficult, in part due to the variability caused by the inherent flexibility of the host cell genome. Summary of the Invention [Problem to be solved by the invention]
[0006] Therefore, there is a need for an efficient method for generating landing pad cells capable of reliable and reproducible protein expression. [Means for solving the problem]
[0007] The present disclosure provides methods for selecting parent cells suitable for generating landing pad cell lines, the methods comprising: (i) screening and selecting cell lines with high expression titers of a gene of interest (GOI); and (ii) further screening the cells of (i) to select cells with a low copy number of a parent plasmid comprising a nucleic acid encoding the GOI, wherein the copy number is 1 or 2. In some embodiments, the parent plasmid comprises two site-specific recombination sites (SSRS), one SSRS, or no SSRS.
[0008] The present disclosure also provides a method for selecting landing pad cells, comprising: (i) screening for loss of the parental plasmid or a portion thereof in the parent cell line, and selecting cells with such loss (deletion); and (ii) further screening the cells of (i) for the presence of a landing pad, and selecting cells in which a landing pad is present.
[0009] Also provided are methods for selecting landing pad cells, comprising: (i) screening for loss of at least one parental plasmid or portion thereof in a parent cell line, and selecting cells with such loss (deletion), and (ii) further screening the cells of (i) for the presence of at least one landing pad, and selecting cells in which a landing pad is present. In some embodiments, the method further comprises screening the landing pad sequences in the landing pad cells for a characteristic selected from the group consisting of: (i) the presence or absence of low-complexity or high-complexity regions, (ii) the presence or absence of retrotransposon sequences, (iii) the presence or absence of Alu repeats, (iv) the presence or absence of long interspersed nuclear elements (LINEs), (v) the presence or absence of CpG islands, (vi) the level of cytosine methylation, (vii) the level of histone acetylation, (viii) the presence or absence of active transcription, and (ix) any combination thereof.
[0010] Also provided are methods for generating landing pad cells, comprising: (i) deleting at least one parental plasmid, or portion thereof, containing a first GOI in a parent cell line; and (ii) after the at least one deletion, introducing a landing pad plasmid, or portion thereof, containing a landing pad into the cell. In some embodiments, the landing pad plasmid, or portion thereof, containing a landing pad is inserted at the site of the deletion in (i). In some embodiments, the landing pad plasmid, or portion thereof, containing a landing pad is inserted at a site other than the site of the deletion in (i).
[0011] The present disclosure also provides methods for generating landing pad cells, comprising integrating landing pad plasmids into the genome of a parent cell at targeted integration sites using homologous recombination, wherein the sequences targeted for homologous recombination are located in parent plasmids, and each landing pad plasmid comprises: (1) a polynucleotide sequence comprising a nucleic acid encoding at least one selectable marker and / or at least one nucleic acid sequence encoding a detectable marker; (2) two site-specific recombination sites (SSRS) flanking the polynucleotide sequence of (1); and (3) two homologous recombination sites located 5′ and 3′ to the SSRS of (2) that are homologous to corresponding homologous recombination sites in the parent plasmid, wherein the homologous recombination sites of the landing pad plasmid recombine with the corresponding homologous recombination sites of the parent plasmid, thereby integrating the landing pad plasmid at a location within the parent plasmid where it was inserted into the parent cell genomic DNA. In some embodiments, the parent plasmids are located at two or more genomic loci.
[0012] The present disclosure also provides a method for identifying a landing pad cell line, the method comprising: (1) removing at least a portion of a first GOI from a parental plasmid integrated into the genomic sequence of a parent cell; (2) integrating the landing pad plasmid at an alternative genomic locus; and (3) screening a library of candidate cells containing at least one copy of the landing pad plasmid integrated at at least one alternative genomic locus, wherein the candidate cell lines are evaluated for one or more of the following characteristics: (a) a cell titer above a predetermined threshold level; (b) a predetermined copy number of the landing pad plasmid or landing pad; (c) an RNA expression level above a predetermined threshold level; (d) a specific plasmid configuration if multiple plasmid copies are present; (e) deletion of at least a portion of the first GOI from the parental plasmid; and (f) the presence of at least one landing pad containing a functional SSRS. In some embodiments, the parent cell is a historical cell line. In some embodiments, the library of candidate cells is a library generated by random integration of landing pad sequences at multiple locations within the genome of the parent cell. In some embodiments, the method selects hot cells containing landing pad sequences integrated at hot spots. In some embodiments, the parent cell line is a CHO cell line.
[0013] The present disclosure also provides a method for generating an expression cell comprising integrating a second GOI plasmid into the genome of a landing pad cell according to any of the methods disclosed above by using site-specific recombinase recombination, wherein the resulting expression plasmid comprises (1) a polynucleotide sequence comprising a nucleic acid encoding the second GOI and (2) two SSRSs flanking the polynucleotide of (1), and wherein the site-specific recombination sites of the landing pad plasmid recombine with corresponding site-specific recombination sites of the second GOI plasmid, thereby integrating the GOI plasmid at a location internal to the landing pad plasmid in the landing pad cell.
[0014] Also provided is a method for producing an expressing cell, comprising: (a) integrating a landing pad plasmid or portion thereof into the genome of a parent cell at a targeted integration site using homologous recombination, wherein sequences targeted for homologous recombination are located in parent plasmids, and each landing pad plasmid or portion thereof comprises: (1a) a polynucleotide sequence comprising a nucleic acid encoding at least one selectable marker and / or at least one nucleic acid sequence encoding a detectable marker; (2a) two SSRSs flanking the polynucleotide sequence of (1a); and (3a) two homologous recombination sites located 5' and 3' to the SSRS of (2a), which are homologous to corresponding homologous recombination sites in the parent plasmid, wherein the homologous recombination sites of the landing pad plasmid or portion thereof recombine with corresponding homologous recombination sites of the parent plasmid within the landing pad at a different genomic locus, thereby integrating the landing pad plasmid or portion thereof at a location internal to the landing pad at a different genomic locus in the parent cell genomic DNA; and (b) integrating a second GOI plasmid into the genome of the landing pad cell using site-specific recombinase recombination, wherein the expression plasmid comprises (1b) a polynucleotide sequence comprising a nucleic acid encoding the GOI and (2b) two SSRSs flanking the polynucleotide of (1b), and wherein the site-specific recombination sites of the landing pad plasmid recombine with corresponding site-specific recombination sites of the GOI plasmid, thereby integrating the GOI plasmid at a location internal to the landing pad plasmid in the landing pad cell. Also disclosed is a method comprising:
[0015] Also provided is a method for generating landing pad cells, comprising: (a) removing at least a portion of the parental plasmid from a first hotspot location in the parental cell line; and (b) integrating landing pad plasmids into a second hotspot location within the genome of the parental cells at targeted integration sites using homologous recombination or random integration, wherein the sequences targeted for homologous recombination or random integration were present in the landing pad plasmids, and each landing pad plasmid comprises: (1b) a polynucleotide sequence comprising a nucleic acid encoding at least one selectable marker and / or at least one nucleic acid sequence encoding a detectable marker; (2b) two site-specific recombination sites (SSRS) flanking the (1b) polynucleotide sequence; and (3b) two homologous recombination sites located 5' and 3' to the SSRS (2b), which are homologous to corresponding homologous recombination sites in the parental cell line genome. Also provided is a method comprising:
[0016] Also provided is a method for producing an expressing cell, comprising: (a) removing the parental plasmid or a portion thereof from a first hotspot location in the parental cell line; (b) integrating the landing pad plasmids into a second hotspot location within the genome of the parental cells at targeted integration sites using homologous recombination, wherein the sequences targeted for homologous recombination were present in the parental cell line, and each landing pad plasmid comprises: (1b) a polynucleotide sequence comprising a nucleic acid encoding at least one selectable marker and / or at least one nucleic acid sequence encoding a detectable marker; (2b) two site-specific recombination sites (SSRS) flanking the polynucleotide sequence of (1b); and (3b) two homologous recombination sites located 5' and 3' to the SSRS of (2b) that are homologous to corresponding homologous recombination sites in the parental cell line, wherein the homologous recombination sites of the landing pad plasmid recombine with the corresponding homologous recombination sites of the parental cell line, thereby integrating the landing pad plasmid at a location within the parental cell genomic DNA; and (c) integrating the GOI plasmid into the genome of the landing pad cell using site-specific recombinase recombination, wherein the expression plasmid comprises (1c) a polynucleotide sequence comprising a nucleic acid encoding a first GOI and (2c) two SSRSs flanking the polynucleotide of (1c), and wherein the site-specific recombination sites of the landing pad plasmid recombine with corresponding site-specific recombination sites of the GOI plasmid, thereby integrating the GOI plasmid at a location internal to the landing pad plasmid in the landing pad cell. Also provided is a method comprising:
[0017] In some embodiments of the above-disclosed methods, the landing pad cells are CG1 / -[P1]-[P2]-[SSRS]-[M]-[SSRS]-[P2]-[P1]- / CG2, CG1 / -[P1]-([P2]-[SSRS]-[M]-[SSRS]-[P2])n-[P1]- / CG2, CG1 / -([P2]-[SSRS]-[M]-[SSRS]-[P2])n- / CG2, CG1 / -[P1]-([P2]-[SSRS]-[M]-[P2])n-[P1]- / CG2, CG1 / -([P2]-[SSRS]-[M]-[P2])n- / CG2, CG1 / -([P1]-([P2]-[SSRS]-[P3]-[SSRS]-[P2])n-[P1])- / CG2, CG1 / -([P2]-[SSRS]-[P3]-[SSRS]-[P2])n- / CG2, CG1 / -([P1]-([P2]-[SSRS]-[P3]-[P2])n-[P1])- / CG2, CG1 / -([P2]-[SSRS]-[P3]-[P2])n- / CG2, CG1 / -[P1]-([P2]-[M]-[SSRS]-[P2])n-[P1]- / CG2, CG1 / -([P2]-[M]-[SSRS]-[P2])n- / CG2, CG1 / -[P1]-[P2]-[P1]- / CG2, CG1 / -[P1]- / CG2, CG1 / -([P1]-([P2])n-[P1])- / CG 2、 -[P1]-[P2]-[P1]-, or -([P1]-([P2])n-[P1])- [where: CG1 and CG2 are parental cell genomic sequences flanking the inserted plasmid; [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [M] is a polynucleotide sequence comprising at least one nucleic acid sequence encoding at least one selectable and / or detectable marker; [SSRS] is a site-specific recombination site (SSRS), and n is an integer between 1 and 10, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.
[0018] In some embodiments of the above-disclosed methods, the topology of the plasmid incorporated into the expression cell is as described: CG1 / -[P1]-([P2]-[SSRS]-[P3]-[SSRS]-[P2])n-[P1]- / CG2, CG1 / -([P2]-[SSRS]-[P3]-[SSRS]-[P2])n- / CG2, CG1 / -[P1]-([P2]-[SSRS]-[P3]-[P2])n-[P1]- / CG2, CG1 / -([P2]-[SSRS]-[P3]-[P2])n- / CG2, CG1 / -[P1]-([P2]-[P3]-[SSRS]-[P2])n-[P1]- / CG 2、 CG1 / -([P2]-[P3]-[SSRS]-[P2])n- / CG2 , [Here CG1 and CG2 are parental cell genomic sequences flanking the inserted plasmid; [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [P3] is a polynucleotide sequence derived from a plasmid containing a gene of interest (GOI), [SSRS] is a site-specific recombination site (SSRS), n is an integer between 1 and 10, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.
[0019] In some embodiments of the methods disclosed above, the homologous recombination is mediated by a CRISPR / Cas system, a TALEN system, or a ZFN system. In some embodiments, the CRISPR / Cas system further comprises a single guide RNA (sgRNA). In some embodiments of the methods disclosed above, the site-specific recombinase recombination site (SSRS) is a Tyr-recombinase site, a Tyr-integrase site, a serine-resolvase / invertase site, or a serine-integrase site. In some embodiments, the Tyr-recombinase site comprises a Cre, Dre, Flp, KD, B3, or B3 Tyr-recombinase site. In some embodiments, the Tyr-integrase site comprises a λ (lambda), HK022, or HP1 Tyr-integrase site. In some embodiments, the serine-resolvase / invertase site comprises a γδ (gamma delta), ParA, Tn3, or Gin serine-resolvase / integrase site. In some embodiments, the serine-integrase site comprises a PhiC31, Bxb1, or R4 serine-integrase site. In some embodiments, the Tyr-recombinase site comprises a Cre Tyr-recombinase site. In some embodiments, the SSRS is a LoxP site. In some embodiments, the LoxP site comprises the nucleic acid sequence set forth in SEQ ID NO: 1 (wild-type LoxP). In some embodiments, the LoxP site comprises a mutant LoxP site. In some embodiments, the mutant LoxP site comprises the nucleic acid sequence set forth in SEQ ID NO: 2 (mutant LoxP). In some embodiments, the mutant LoxP site comprises a nucleic acid selected from the group consisting of SEQ ID NO:3 (Lox511), SEQ ID NO:4 (Lox5171), SEQ ID NO:5 (Lox2272), SEQ ID NO:6 (LoxM2), SEQ ID NO:7 (LoxM3), SEQ ID NO:8 (LoxM7), SEQ ID NO:9 (LoxM11), SEQ ID NO:10 (Lox71), and SEQ ID NO:11 (Lox66). In some embodiments, the mutant LoxP site comprises any LoxP site disclosed herein.In some embodiments, the Tyr-recombinase site comprises an Flp Tyr-recombinase site. In some embodiments, the SSRS is a short flippase recognition target (FRT) site. In some embodiments, the SSRS comprises any FRT site sequence disclosed herein. In some embodiments, the serine-integrase site comprises an att site, e.g., an attP or attB site. In some embodiments, the SSRS comprises any att site disclosed herein.
[0020] In some embodiments of the methods disclosed above, the at least one nucleic acid sequence encoding at least one selectable marker and / or detectable marker is glutamine synthetase (GS) and / or dihydrofolate reductase (DHFR). In some embodiments, the at least one nucleic acid sequence encoding at least one selectable marker and / or detectable marker is a drug resistance gene. In some embodiments, the drug resistance gene is an antibiotic resistance gene. In some embodiments, the antibiotic resistance gene is a puromycin resistance gene. In some embodiments, the puromycin resistance gene is puromycin-N-acetyltransferase. In some embodiments, the at least one nucleic acid sequence encoding at least one selectable marker and / or detectable marker constitutes a protein. In some embodiments, the protein is a fluorescent protein. In some embodiments, the fluorescent protein is mCherry. In some embodiments, the fluorescent protein comprises GFP, ZsGreen1, AcGFP1, EGFP, GFPuv, AcGFP, EBFP, EYFP, ECFP, tdTomato, mCherry, DsRed, AmCyan, ZsGreen, ZsYellow, DsRed2, DsRed-Express, HcRed, AsRed, mOrange, mOrange2, mPlum, mStrawberry, mBanana, YFP, mRaspberry, HcRed1, E2-Crimson, or any combination thereof.
[0021] In some embodiments of the above-disclosed methods, the cells are Chinese hamster ovary (CHO) cells. In some embodiments, the cells are HEK293 or NS0 cells.
[0022] In some embodiments of the methods disclosed above, the nucleic acid encoding the GOI encodes at least one polypeptide. In some embodiments, the at least one polypeptide is an antibody or a fusion protein. In some embodiments, the expression plasmid comprises one, two, or more copies of the GOI, a detectable marker, or a combination thereof.
[0023] In some embodiments, the method disclosed above further comprises determining the expression of GOI, detectable marker, or a combination thereof. In some embodiments, the expression of GOI is determined quantitatively and / or qualitatively. In some embodiments, the expression of GOI is determined by cell sorting, FACS, cell surface staining, Western blot, Northern blot, column chromatography, capillary electrophoresis, microfluidics, UV absorbance, immunohistochemistry, cell size, secreted protein level, transcript level, or any combination thereof.
[0024] In some embodiments of the methods disclosed herein, the landing pad plasmid or expression plasmid is integrated into the genome of the cell at a copy number of 1. In some embodiments, the landing pad plasmid or expression plasmid is integrated into the genome of the cell at a copy number of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30.
[0025] In some embodiments of the methods disclosed herein, (i) the 5' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 18 or a subsequence thereof, and the 3' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 19 or a subsequence thereof; (ii) the 5' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 114 or a subsequence thereof, and the 3' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 115 or a subsequence thereof; or (iii) the 5' homologous recombination site and the 3' homologous recombination site comprise polynucleotide sequences that flank the parental plasmid.
[0026] In some embodiments of the methods disclosed herein, the parental plasmid comprises an open reading frame (ORF) encoding a first GOI, such as an antibody.
[0027] This disclosure describes CG1 / -[P1]-[P2]-[SSRS]-[M]-[SSRS]-[P2]-[P1]- / CG2, CG1 / -[P1]-([P2]-[SSRS]-[M]-[SSRS]-[P2])n-[P1]- / CG2, CG1 / -([P2]-[SSRS]-[M]-[SSRS]-[P2])n- / CG2, CG1 / -[P1]-([P2]-[SSRS]-[M]-[P2])n-[P1]- / CG2, CG1 / -([P2]-[SSRS]-[M]-[P2])n- / CG2, CG1 / -([P1]-([P2]-[SSRS]-[P3]-[SSRS]-[P2])n-[P1])- / CG2, CG1 / -([P2]-[SSRS]-[P3]-[SSRS]-[P2])n- / CG2, CG1 / -([P1]-([P2]-[SSRS]-[P3]-[P2])n-[P1])- / CG2, CG1 / -([P2]-[SSRS]-[P3]-[P2])n- / CG2, CG1 / -[P1]-([P2]-[M]-[SSRS]-[P2])n-[P1]- / CG2, CG1 / -([P2]-[M]-[SSRS]-[P2])n- / CG2, CG1 / -[P1]-[P2]-[P1]- / CG2, CG1 / -[P1]- / CG2, CG1 / -([P1]-([P2])n-[P1])- / CG 2、 -[P1]-[P2]-[P1]-, or -([P1]-([P2])n-[P1])- [Here CG1 and CG2 are parental cell genomic sequences flanking the inserted plasmid; [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [M] is a polynucleotide sequence comprising at least one nucleic acid sequence encoding at least one selectable and / or detectable marker; [SSRS] is a site-specific recombination site (SSRS), n is an integer between 1 and 10, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, A landing pad cell is provided.
[0028] This disclosure describes CG1 / -[P1]-([P2]-[SSRS]-[P3]-[SSRS]-[P2])n-[P1]- / CG2, CG1 / -([P2]-[SSRS]-[P3]-[SSRS]-[P2])n- / CG2, CG1 / -[P1]-([P2]-[SSRS]-[P3]-[P2])n-[P1]- / CG2, CG1 / -([P2]-[SSRS]-[P3]-[P2])n- / CG2, CG1 / -[P1]-([P2]-[P3]-[SSRS]-[P2])n-[P1]- / CG2, or CG1 / -([P2]-[P3]-[SSRS]-[P2])n- / CG2 [Here CG1 and CG2 are parental cell genomic sequences flanking the inserted plasmid; [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [P3] is a polynucleotide sequence derived from a plasmid containing a gene of interest (GOI), [SSRS] is a site-specific recombination site (SSRS), and n is an integer between 1 and 10, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.
[0029] The present disclosure provides cell lines produced by any of the methods disclosed herein. Also provided are kits containing the cells disclosed herein or cells generated according to any of the methods disclosed herein and instructions for their use.
[0030] The present disclosure also provides an isolated cell comprising a polynucleotide sequence comprising a nucleic acid encoding a gene of interest (GOI) integrated at a specific locus in the genome of the cell, wherein the locus is a nucleotide position or polynucleotide subsequence within SEQ ID NO:20 or SEQ ID NO:116.
[0031] Also provided is a method comprising introducing a polynucleotide sequence comprising a nucleic acid encoding a gene of interest (GOI) into a CHO cell and obtaining a CHO cell having the nucleic acid integrated at a specific locus in the genome of the CHO cell, wherein the locus is a nucleotide position or polynucleotide subsequence within SEQ ID NO: 20 or SEQ ID NO: 116. Also provided is a method comprising providing a cell comprising a polynucleotide sequence comprising a nucleic acid encoding a gene of interest (GOI), wherein the polynucleotide sequence comprises a nucleic acid encoding the gene of interest (GOI) operably linked to a promoter, and the nucleic acid is integrated at a specific locus in the genome of the CHO cell, wherein the locus is a nucleotide position or polynucleotide subsequence within SEQ ID NO: 20 or SEQ ID NO: 116. In some embodiments of the methods or isolated cells disclosed herein, (i) the nucleotide subsequence within SEQ ID NO: 20 comprises the sequence set forth in SEQ ID NO: 21, or (ii) the nucleotide subsequence within SEQ ID NO: 116 comprises the sequence set forth in SEQ ID NO: 117. In some embodiments of the methods or isolated cells disclosed herein, (i) the nucleotide subsequence from SEQ ID NO:20 consists of the sequence set forth in SEQ ID NO:21, or (ii) the nucleotide subsequence from SEQ ID NO:116 consists of the sequence set forth in SEQ ID NO:117. In some embodiments of the methods or isolated cells disclosed herein, (i) the nucleotide subsequence from SEQ ID NO:20 is an upstream subsequence (towards the 5' end of SEQ ID NO:20) of the sequence set forth in SEQ ID NO:21, or (ii) the nucleotide subsequence from SEQ ID NO:116 is an upstream subsequence (towards the 5' end of SEQ ID NO:116) of the sequence set forth in SEQ ID NO:117. In some embodiments of the methods or isolated cells disclosed herein, (i) the nucleotide subsequence from SEQ ID NO:20 is a downstream subsequence (towards the 3' end of SEQ ID NO:20) of the sequence set forth in SEQ ID NO:21, or (ii) the nucleotide subsequence from SEQ ID NO:116 is a downstream subsequence (towards the 3' end of SEQ ID NO:116) of the sequence set forth in SEQ ID NO:117.
[0032] In some embodiments, the methods, cells, cell lines, or kits disclosed herein comprise or involve the use of at least two landing pad plasmids or at least two expression plasmids. In some embodiments, the two landing pad plasmids or two expression plasmids are in a configuration selected from the group consisting of head-to-head, tail-to-tail, tail-to-head, and head-to-tail. In some embodiments, each expression plasmid comprises at least a nucleic acid encoding a gene of interest (GOI). In some embodiments, all GOIs are the same. In some embodiments, all GOIs are different. In some embodiments, at least one GOI is different from the rest. In some embodiments, a first GOI comprises an antibody heavy chain (HC) and a second GOI comprises an antibody light chain (LC). In some embodiments, at least one expression plasmid is bicistronic. In some embodiments, a bicistronic expression plasmid encodes a first GOI comprising an antibody HC and a second GOI comprising an antibody LC. In some embodiments, at least one landing pad plasmid is addressable. In some embodiments, each landing pad plasmid contains two Lox sites. In some embodiments, the Lox sites are Lox P and Lox 511. In some embodiments, each landing pad plasmid contains a Lox site and an Frt site. In some embodiments, each landing pad plasmid contains one or two aat sites. In some embodiments, each landing pad plasmid is addressable. In some embodiments, each addressable landing pad plasmid contains a pair of addressable SSRSs that are unique to the landing pad. In some embodiments, the at least one pair of addressable SSRSs is a pair of Lox sites. In some embodiments, the at least one pair of Lox sites is Lox 511 and Lox P. In some embodiments, the at least one pair of Lox sites is Lox m3 and Lox m7.In some embodiments, the first addressable landing pad plasmid comprises a pair of Lox511 and LoxP sites, and the second addressable landing pad plasmid comprises a pair of Loxm3 and Loxm7 sites, hi some embodiments, each addressable landing pad plasmid comprises a non-cross-compatible att site.
[0033] The present disclosure also provides a cell comprising a heterologous polynucleotide sequence comprising a nucleic acid encoding a gene of interest (GOI), a selectable marker, a detectable marker, or a combination thereof integrated at a specific locus in the genome of the cell, wherein the locus is a nucleotide position or polynucleotide subsequence within SEQ ID NO: 20 or SEQ ID NO: 116, or an orthologous sequence thereof. Also provided is a cell comprising a heterologous polynucleotide sequence comprising a nucleic acid encoding a gene of interest (GOI), a selectable marker, a detectable marker, or a combination thereof integrated at a specific locus in the genome of the cell, wherein the locus is a nucleotide position or polynucleotide subsequence within SEQ ID NO: 21 or SEQ ID NO: 117, or an orthologous sequence thereof. Also provided is a cell comprising a heterologous polynucleotide sequence comprising a nucleic acid encoding a gene of interest (GOI), a selectable marker, a detectable marker, or a combination thereof integrated at a specific locus in the genome of the cell, wherein the locus is a polynucleotide subsequence that overlaps with or encompasses SEQ ID NO: 20 or SEQ ID NO: 116, or an orthologous sequence thereof. Also provided are cells comprising a heterologous polynucleotide sequence comprising a nucleic acid encoding a gene of interest (GOI), a selectable marker, a detectable marker, or a combination thereof, integrated at a specific locus in the genome of the cell, the locus being a polynucleotide subsequence that overlaps with or encompasses SEQ ID NO: 21 or SEQ ID NO: 117, or an orthologous sequence thereof. In some embodiments, the cell is a CHO cell. In some embodiments, the orthologous sequence has about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 96%, about 97%, about 98%, or about 99% sequence identity with SEQ ID NO: 20, 21, 116, 117, or a subsequence thereof. In some embodiments, the sequence identity is determined by pairwise alignment using an implementation of the Needleman-Wunsch algorithm. In some embodiments, the cell comprises two landing pad plasmids or two expression plasmids.In some embodiments, the cell comprises three or more landing pad plasmids or three or more expression plasmids. In some embodiments, two landing pad plasmids are addressable. [Brief explanation of the drawings]
[0034] [Figure 1] FIG. 1 is a schematic representation depicting a standard expression cell line generation strategy in which cells are transfected with an expression plasmid that results in its integration at a random location within the cell's genome. [Figure 2] Figure 2 outlines the strategy used to identify two parental cell lines suitable for landing pad cell generation. Parental cell lines 1 and 2 express monoclonal antibodies derived from parental plasmids directed against protein 1 and protein 2, respectively. The arrow and GS represent glutamine synthetase complementation. mAb = monoclonal antibody; LC = light chain; HC = heavy chain; copy number = number of expression plasmids in the cell line (each expression plasmid contains an expression cassette for LC and HC. Copy number was determined by qPCR using GAPDH as an internal control); spPCR = splinkeret PCR (a technique that allows identification of plasmid junction sequences). LC RNA and HC RNA levels were normalized to the levels found for antibody against protein 1 (i.e., antibody against protein 1 = 1.00). [Figure 3] Figure 3 is a simplified depiction of the parental plasmids showing the configuration found in both Cell Line 1 and Cell Line 2. The parental plasmids in both cell lines are in a head-to-tail configuration. The configurations in Cell Line 1 and Cell Line 2 were established by Southern blot analysis and plasmid sequence junction determination, where plasmid-plasmid fusions were detected. The arrow and GS represent glutamine synthetase complementation. [Figure 4]Figures 4A and 4B show two strategies for generating landing pad cell lines containing site-directed recombination sites, such as LoxP. In both strategies, the landing pad plasmid is introduced into cells by homologous recombination, stimulated by restricting the genome of the parent cell line with a site-specific nuclease, such as a CRISPR-associated nuclease (Cas), represented by scissors. In Figure 4A, a parent cell line is identified based on its performance and the number of sites at which the expression plasmid / cassette is found within its genome. In Figure 4B, the parent cell line from Figure 4A is used as the landing pad cell line. In both Figures 4A and 4B, knowledge of the sequence of the cell's genome is required, and sequences homologous to the cell's genome (homology 1, homology 2) are cloned to flank the expression cassette. mCherry represents an open reading frame encoding a fluorescent marker, LoxP sites are sequences used by Cre recombinase, arrow and GS represent glutamine synthetase complementation, and arrow and Puro represent puromycin resistance. [Figure 5]Figures 5A and 5B provide a schematic representation of the universal TI strategy of the present disclosure. The TI technique disclosed herein involves the use of a site-specific endonuclease directed against a parental plasmid sequence in a parent cell line (Figure 5A) or a landing pad cell line (Figure 5B) to stimulate homologous recombination with a second DNA. The parent cell line (Figure 5A) can also serve as a landing pad cell line (Figure 5B). These strategies contrast with techniques that require knowledge and use of genomic sequences (see, e.g., Figures 4A and 4B). The vertical and wavy boxes next to the genomic sequences represent regions of homology between different plasmids. The solid boxes next to each region of homology represent sequences targeted by CRISPR / Cas that are present in the parent cell line (Figure 5A) or the landing pad cell line (Figure 5B) but not in the plasmids recombined into these cell lines. The scissors represent CRISPR / Cas, the mCherry open reading frame encodes a fluorescent marker, the LoxP site is the sequence used by Cre recombinase, the arrow and GS represent glutamine synthetase complementation, and the arrow and Puro represent puromycin resistance. Figure 5C depicts the sequence organization of the expression plasmid (P4) in an expression cell generated according to the methods disclosed herein. These diagrams indicate the location of sequences originating from the parental plasmid (P1), landing pad plasmid (P2), and second GOI plasmid (P3). "Cell genome" indicates flanking genomic sequences. Figure 5D shows a versatile TI strategy using a single SSRS site. Here, a site-specific endonuclease is directed against the parental plasmid sequence in the parental cell line to stimulate homologous recombination. The vertical and wavy boxes next to the genomic sequence represent regions of homology between different plasmids. The black boxes next to each region of homology are sequences that are targeted by a sequence-specific endonuclease, e.g., CRISPR / Cas, that are present in the parental cell line but absent from the plasmids that are recombined into these cell lines.The landing pad cell line contains a single SSRS site, shown here using attB as an example. The GOI plasmid (P3) is inserted into the targeting locus through the single SSRS site. The scissors represent a sequence-specific endonuclease, e.g., CRISPR / Cas. The mCherry open reading frame encodes an exemplary fluorescent marker. The attB and attP sites are sequences used by the integrase. The arrow represents a promoter, and GS represents glutamine synthetase complementation. An is a poly(A) signal sequence. mAb is a monoclonal antibody expression cassette containing a unique promoter and poly(A) signal. Figure 5E shows a TI strategy that uses a cellular genomic sequence for homologous recombination to create a landing pad. Here, a site-specific endonuclease is directed against the cellular genomic sequence to stimulate homologous recombination. The genomic sequence represents the region of homology between the landing pad plasmid and the parent cell. The black boxes next to each homology region represent sequences targeted by CRISPR / Cas that are present in the parent cell line but absent from the plasmid recombined into these cell lines. The landing pad cell line contains a single SSRS site, shown here using attB as an example. Through a single SSRS recombination, the GOI plasmid (P3) is inserted into the targeted locus. The scissors represent CRISPR / Cas, and the mCherry open reading frame encodes a fluorescent marker. The attB and attP sites are sequences used by the integrase. The arrow represents the promoter, and GS represents glutamine synthetase complementation. An is the poly(A) signal sequence. mAb represents a monoclonal antibody expression cassette containing a unique promoter and poly(A). Figure 5F depicts the sequence organization of the expression plasmid (P5) in the expression cell line, generated according to the method described in Figure 5E or using random integration into a new genomic locus. These diagrams show the location of sequences originating from the landing pad plasmid (P2) and the second GOI plasmid (P3). "Cellular genome" indicates the flanking genomic sequences.Plasmid P1 is either completely deleted or not present at the locus, so there is no P1 portion in this expression plasmid construct. [Figure 6] Figures 6A and 6B outline the generation of landing pad cell lines according to the present disclosure. Figure 6A shows the generation of a landing pad cell line expressing a marker (e.g., mCherry) by replacing the monoclonal antibody (mAb)-encoding plasmid in the parent cell line with a portion of the landing pad plasmid's protein (e.g., a linear plasmid containing an open reading frame encoding mCherry and a puromycin resistance gene flanked by LoxP sites). A description of the components in Figure 6A can be found in the description of Figures 5A and 5B above. Figure 6B shows the increased frequency in the generation of mCherry landing pad cell lines stimulated by the presence of a single guide RNA (sgRNA) required for CRISPR / Cas technology. The percentage of clones (approximately 25%) with the desired phenotype, the presence of a single mCherry expression cassette and the absence of a mAb. The landing pad cell lines used for TI were identified by their expression levels (mean fluorescence intensity (MFI)), transcript levels, and the stability of these two parameters as the cells were passaged. [Figure 7]Figures 7A and 7B show the actual application of the proposed targeted integration methodology in the cells of Figures 5A, 5B, 6A, and 6B. In Figure 7A, the mCherry expression cassette was replaced with one expressing an antibody against protein 3 (mAb3) using Cre recombinase. Cells that expressed only mAb3 were single-cell cloned using Berkley Lights (BL) and FACS techniques. Cells were expanded and evaluated for protein expression in AMBR® 15 and AMBR® 250 bioreactor systems and by 24-deep-well fed-batch (24DW FB) bioreactors. Figure 7B shows that the resulting cell populations after selection for GS complementation were screened by FACS for cell surface expression of protein 3 (vertical axis) and mCherry expression (horizontal axis). 5.24% of the cells expressed protein 3 only, 90.06% expressed both proteins, 4.12% expressed mCherry only, and 0.68% expressed neither protein. Cells with only cell surface staining for mAb against protein 3 (mAb3) are the desired cells. Productivities obtained from the screened clones are summarized in the text next to the FACS data. [Figure 8]Figures 8A and 8B outline the universal targeted integration (UTI) technique, which can be implemented using four different strategies (Strategy A, Strategy B, Strategy C, and Strategy D). The universal TI technique disclosed herein involves the use of a site-specific endonuclease directed against a parental plasmid sequence in the parent cell line that is not present in the landing pad plasmid to stimulate homologous recombination with the landing pad plasmid. The advantage of this strategy is that knowledge of the flanking genomic DNA sequence is not required. In this UTI technique, as depicted in Figure 8A, either the parental expression plasmid in the parent cell line is replaced by the landing pad (Strategy A), or the parental expression plasmid is deleted and the landing pad is inserted at an alternative locus in the cell genome (Strategy B). In both cases, a site-specific endonuclease is used to stimulate recombination. Once the landing pad cell line is generated, it is used to generate an expression cell line in which the landing pad is replaced with a second GOI using Cre recombinase. In this UTI technology, as depicted in Figure 8B, a single SSRS site is used to generate expression cell lines in Strategy C and Strategy D. Vertical and wavy boxes represent regions of homology between different plasmids. Solid boxes represent sequences present in the parent expression plasmid but absent from the plasmid recombined into these cell lines, e.g., targeted by CRISPR / Cas. Scissors represent CRISPR / Cas. The mCherry open reading frame encodes a fluorescent marker. LoxP and Lox511 sites are sequences used by Cre recombinase. attB and attP sites are sequences used by integrase. The arrow and GS encode GS complementation. The arrow and Puro encode puromycin resistance. The depictions of these four alternative strategies are exemplary, and the components shown in the figures (e.g., CRISPR / Cas, mCherry, Lox sites, att sites) can be replaced with functional equivalents as disclosed herein. [Figure 9]Figure 9 shows a summary of data from the generation of landing pad cell lines using the strategy shown in Figures 8A and 8B. The photographs demonstrate the increased frequency in the generation of mCherry landing pad cell lines stimulated by the presence of the single guide RNA (sgRNA) required for CRISPR / Cas technology. 25% of clones have the desired phenotype: the presence of the mCherry expression cassette and the absence of the mAb from the parent cell line. The landing pad cell lines used for TI were identified by their mCherry gene copy number, expression level (mean fluorescence intensity, MFI), transcript level, and the stability of these two parameters as the cells were passaged. [Figure 10] Figure 10 summarizes the results of an experiment in which 12 landing pad cell lines were used to construct expressing cell lines using a second GOI plasmid encoding two copies of the light chain and two copies of the heavy chain of the mAb, as well as a plasmid encoding Cre. The percent of expressing cell lines is the percent of mCherry-negative [Red(-)] cells in the bulk culture after selection. [Figure 11] Figures 11A and 11B summarize the results of experiments using five landing pad cell lines that underwent cell line generation. After single-cell cloning, 32 expressing cell lines from each landing pad cell line were randomly selected, expanded, and tested in a 14-day fed-batch assay in 24-deep-well plates (DWPs). This allows for comprehensive characterization of the potential of the landing pad cell lines. The data are summarized in Figure 11A and presented as boxplots in Figure 11B. [Figure 12]Figure 12A shows a head-to-head duo-landing pad configuration. Each landing pad contains two distinct SSRS sites for directional recombination. Recombination inserted one GOI (mAb) into each landing pad locus. The resulting mAb expression plasmid is still in a head-to-head configuration. Figure 12B depicts the duo-landing pad configuration and the effect of Cre recombinase on the duo-landing pad. One landing pad is shown as a diagonal arrow, and the other as a crosshatched arrow. Both landing pads use LoxP and Lox511. The head-to-head and tail-to-tail configurations remain intact in the presence of Cre. In the other two configurations, one of the landing pads can be permanently deleted. The purpose of having more than one landing pad is, for example, to generate bispecific mAbs and increase their titer. When under the control of the same regulatory sequence (e.g., the same promoter), multiple landing pads are likely to have the same activity. In some embodiments, multiple landing pads, for example, three, four, or more, can be present in a 1:1 ratio, or alternative ratios such as 1:2 or 2:1. [Figure 13] Figure 13 shows the results of TI of a second GOI in head-to-head and tail-to-tail dual landing pad configurations. One landing pad is shown as a hatched arrow, and the other as a crosshatched arrow. Both landing pads use LoxP and Lox511. The second GOI is shown as a solid rectangle. In both cases, the expressing cell lines contain two second GOIs. [Figure 14]Figure 14 shows the results of TI of a second GOI in tail-to-head and head-to-tail dual landing pad configurations. One landing pad is shown as a hatched arrow, and the other as a crosshatched arrow. Both landing pads use LoxP and Lox511. The second GOI is shown as a solid rectangle. In both cases, two distinct expressing cell lines were generated, one with one copy of the second GOI and the second with two copies of the second GOI. [Figure 15] Figure 15 shows a duo landing pad configuration containing Frt and LoxP sites, as well as a depiction of the effect of Cre and Flp recombinase on the duo landing pad. One landing pad is shown as a hatched arrow, and the other as a crosshatched arrow. Both landing pads use LoxP and Frt. The head-to-head and tail-to-tail configurations remain intact in the presence of Cre and Flp. In the other two configurations, one of the landing pads can be permanently deleted. This corresponds to what was observed in Figure 12B. [Figure 16] Figure 16 shows a depiction of a duo landing pad configuration using the same aatP sites to flank all landing pads, as well as the result after a second GOI plasmid and Int are transfected into cells. One landing pad is shown as a hatched arrow and the other as a crosshatched arrow. The GOI is shown as a solid rectangle. [Figure 17]Figures 17A and 17B provide a schematic illustrating how the duo landing pad can be used to increase the expression diversity of different subunits that assemble to create a desired composite biologic. The solid and dashed arrows represent the different components required to create a biologic. For illustrative purposes, a composite biologic requires at least one of each arrow. Each GOI plasmid can contain each subunit of the composite biologic in a different configuration, i.e., the composite biologic arrows. A second GOI can be composed of multiple arrows in a different order or each arrow itself. The second GOI plasmid is transfected into the duo landing pad cell line in conjunction with a recombinase. To adjust the level of subunit expression, different combinations of the second GOI plasmid are transfected into the duo landing pad cell line. Different transfections of the second GOI are shown, resulting in gene copy ratios of 1:2, 1:1, and 2:1 between the solid and dashed arrows after TI is complete. For example, in the 1:2 ratio shown in Figure 17A, one second GOI contains one copy of the dashed arrow, while the other second GOI plasmid contains both the solid and dashed arrows in a one-of-two configuration. As shown in Figure 17A, this requires two independent transfections of the duo landing pad cell line. It is clear that this is not an exhaustive list of possible outcomes or inputs. Also, configurations where only a subset, i.e., only one solid line or only one dashed arrow, is incorporated into the landing pad and therefore no combination biologic is produced, are not shown. Figure 17B shows a simplified diagram using addressable landing pads with unique SSRSs. One landing pad is composed of Lox511 and LoxP, while the second contains Lox sites 2272 and M3. The second GOI plasmid would be specifically targeted to one landing pad or the other using the corresponding Lox sites. [Figure 18]Figure 18 demonstrates the utility of having a duo landing pad with addressable landing pads. One landing pad is shown as a diagonal arrow and the other as a crosshatched arrow. The second GOI plasmid is shown as a solid rectangle or a rectangle with vertical lines. In this example, each landing pad is flanked by a unique set of Lox sites that recombine only with each other. This example is illustrative; other recombinases and their target sites may be used. Having addressable landing pads ensures that all four configurations of the duo landing pad have a defined second GOI without the loss of one of the landing pads in the tail-to-head and head-to-tail configurations shown in Figure 12B. [Figure 19] Figure 19 shows a diagram illustrating how having a duo landing pad with a single aatB site in each landing pad eliminates the loss of landing pads. One landing pad is shown as a hatched arrow and the other as a crosshatched arrow. The second GOI plasmid is shown as a solid rectangle. The duo landing pad becomes addressable if the attP sites used are not mutually compatible. [Figure 20] Figure 20 shows a proof-of-concept (POC) of targeted integration using the Duo Landing Pad cell line. The mCherry expression cassette is exchanged for a second GOI mAb expression cassette using Cre recombinase, as outlined in Figure 8b. The resulting cell population after GS complementation selection was screened by FACS for mCherry expression (horizontal axis) and cell surface expression of the mAb (vertical axis). Here, 5.24% of cells express only the mAb, 90.06% express both proteins, 4.12% express only mCherry, and 0.68% express neither protein. [Figure 21]Figure 21 shows that targeted integration of a GOI results in more productive cells relative to random integration. mAbs A and B, in the form of a second GOI plasmid, were integrated into host cells by either random or targeted integration. The landing pad cell line is a direct descendant of the cell line used for random integration. The titer of the cell population used for single-cell cloning was determined. The targeted integration population had a titer approximately three to four times higher than that of the random integration, demonstrating the value of this technique for generating landing pad cell lines that can surpass the industry standard for random integration. [Figure 22] Figure 22 shows an overview of the use of duo landing pad cell lines to generate expression cell lines using a second GOI plasmid containing either 1LC+1HC or 2LC+2HC. Second GOI plasmids containing either 1LC+1HC or 2LC+2HC of mAbs A and B were used in TI cell line construction. Productivity of the top six clones from each group is shown. In both cases, increasing the copy number of LC and HC improved the mean titer by 25%-37% and the median titer by 35%-37%. DETAILED DESCRIPTION OF THE INVENTION
[0035] The present disclosure provides methods for generating landing pad cells, in which a linear plasmid, such as a linear plasmid containing a gene of interest (e.g., one or more open reading frames encoding an antibody), can be inserted into the genome of a host cell without requiring prior knowledge of the host cell's genomic sequence for its targeted insertion. While linear plasmids are often preferred, circular plasmids can also be used to generate landing pad cells.
[0036] The terms "targeted insertion" and "targeted integration" are used interchangeably and refer to a gene targeting method used to direct the direct insertion or integration of a gene or nucleic acid sequence to a specific location on the genome, i.e., direct the gene or nucleic acid sequence to a specific site between two nucleotides in a continuous polynucleotide chain.Targeted insertion can be performed to introduce a small number of nucleotides, or can be performed to introduce an entire gene cassette, for example, including multiple genes, regulatory elements, and / or nucleic acid sequences.The terms "insertion" and "integration" and their grammatical variants are used interchangeably throughout this specification.In some embodiments, targeted integration can be achieved by recombination, for example, site-specific recombination, homologous recombination, or a combination thereof.
[0037] According to these methods, cell lines, such as those historically known to exhibit advantageous properties for the expression of a protein of interest (e.g., high recombinant protein yield, low proteolysis or misfolding, specific glycosylation patterns, or other properties related to post-translational modifications), can be used as parent cell lines to generate landing pad cell lines that can be used to express other genes of interest. Ideally, the parent cell line is a hot cell (i.e., produces high titers of recombinant protein) and contains one or more hot spots (areas in the genome where the introduction of foreign nucleic acid encoding the protein of interest is not problematic and results in high levels of recombinant protein expression). Two hot spots were identified as part of the parent cell selection process disclosed herein.
[0038] A plasmid in a parent cell (parent plasmid) containing an expression cassette integrated into the genome of the parent cell line is partially removed (i.e., without cutting or destroying the parent cell genomic DNA) by excising it (e.g., by homologous recombination) between two locations (e.g., recombination sites) inherent in the parent plasmid, and the excised region is replaced with another DNA sequence (landing pad plasmid) containing two new recombination sites flanking at least one marker (e.g., a selectable and / or screenable marker). This method results in landing pad cells that can be used to insert a nucleic acid sequence (e.g., an expression plasmid or gene-of-interest plasmid) containing a different gene of interest (i.e., a gene of interest different from the gene of interest present in the parent cell) by recombination at the two newly introduced recombination sites. This application discloses different strategies related to this general process.
[0039] These methods are inherently versatile, allowing any particularly advantageous parent cell line to be used as a landing pad cell line based on knowledge of the parent plasmid sequence. Knowledge of the plasmid (parent plasmid) present in the parent cell line, often a commercial plasmid known in the art, facilitates the selection of a suitable recombination site for introducing the landing pad plasmid or a portion thereof into the genome of the parent cell. The newly introduced recombination site in the landing pad plasmid can then be used to integrate a plasmid, such as a linear or circular plasmid containing a gene of interest, or a portion thereof into the genome of the parent cell, resulting in an expressing cell. See, e.g., Figures 4A, 5A, 6A, and 8A. Generalized targeted integration strategies are depicted, for example, in Figure 8A (Strategy A and Strategy B) and Figure 8B (Strategy C and Strategy D). Constructs containing multiple landing pads in different configurations are also provided (see, e.g., Figure 12B), and each pad can be uniquely identified by using a unique combination of SSRSs (see, e.g., Figure 18).
[0040] Identifying parent cell lines that are "hot cell lines" (i.e., have high yields of recombinant protein or another advantageous property) and subsequently identifying the "hot spots" into which the parental plasmids are inserted supports improved methods for making and using landing pad cells that do not require knowledge of the sequence of the parental cell genome, as well as the use of these methods and / or landing pad cells to express suitable alternative biologics, such as monoclonal antibodies.
[0041] Thus, the present disclosure also provides kits containing landing pad cells, landing pad plasmids, and reagents for, e.g., generating landing pad cell lines and / or generating expression cell lines. In addition to providing landing pad cells containing a single landing pad plasmid, the present disclosure provides landing pad cells containing multiple landing pads. In some embodiments, the multiple landing pads in the landing pad cells of the present disclosure may be addressable, e.g., by containing site-specific recombination sites or combinations thereof that uniquely identify each landing pad.
[0042] I. Terminology In order that this disclosure may be more readily understood, certain terms are first defined. As used herein, unless otherwise specified herein, each of the following terms shall have the meaning set forth below. Additional definitions are set forth throughout this application.
[0043] The singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. The terms "a" (or "an"), as well as "one or more" and "at least one," may be used interchangeably herein. In certain embodiments, the term "a" or "an" means "single." In other embodiments, the term "a" or "an" includes "two or more" or "plural."
[0044] Furthermore, when used herein, "and / or" shall be construed as a specific disclosure of each of the two specified features or components, with or without the other. Thus, the term "and / or" used in phrases such as "A and / or B" herein is intended to include "A and B," "A or B," "A" (alone), and "B" (alone). Similarly, the term "and / or" used in phrases such as "A, B, and / or C" is intended to encompass each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).
[0045] The terms "about" or "comprising essentially of" refer to a value or composition that is within an acceptable error range for a particular value or composition, as determined by one of ordinary skill in the art, where the error range depends in part on how the value or composition is measured or determined, i.e., the limitations of the measurement system. For example, "about" or "essentially comprising" can mean within one or more standard deviations, as is customary in the art. Alternatively, "about" or "essentially comprising" can mean a range of up to 10%. Furthermore, particularly with respect to biological systems or processes, these terms can mean up to an order of magnitude, or up to five times, of a value. When a particular value or composition is provided in this application and claims, unless otherwise specified, the meaning of "about" or "essentially comprising" should be assumed to be within an acceptable error range for that particular value or composition.
[0046] Whenever an embodiment is described herein using the word "comprising," it is understood that otherwise similar embodiments described using the terms "consisting of" and / or "consisting essentially of" are also provided.
[0047] As used herein, the term "approximately" when applied to one or more target values refers to a value similar to a specified reference value. In certain embodiments, unless otherwise specified or the context makes clear, the term "approximately" refers to a range of values that are 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less above or below (above or below) the specified reference value (unless such number exceeds 100% of possible values).
[0048] As described herein, any concentration range, percentage range, ratio range, or integer range, unless otherwise stated, should be understood to include every integer value within the stated range, and, where appropriate, fractions thereof (e.g., 1 / 10 and 1 / 100 of an integer).
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this disclosure relates.For example, Concise Dictionary of Biomedicine and Molecular Biology, Juo, Pei-Show, 2nd ed., 2002, CRC Press, The Dictionary of Cell and Molecular Biology, 3rd ed., 1999, Academic Press and Oxford Dictionary of Biochemistry and Molecular Biology, Revised, 2000, Oxford University Press provide those skilled in the art with a general dictionary of many of the terms used in this disclosure.
[0050] Units, prefixes, and symbols are denoted in the accepted form of the International System of Units (SI). The headings provided herein are not intended to limit the various aspects of the disclosure, which may be understood by reference to the specification as a whole. Accordingly, defined terms are more particularly defined by reference to the specification as a whole.
[0051] Abbreviations used herein are defined throughout this disclosure. Various aspects of the disclosure are described in further detail in the following subsections.
[0052] Nucleotides are referred to by their commonly accepted single-letter codes. Unless otherwise specified, nucleic acids are written left to right in a 5' to 3' orientation. Nucleotides are referred to herein by their commonly known single-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Thus, A represents adenine, C represents cytosine, G represents guanine, T represents thymine, and U represents uracil.
[0053] Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Unless otherwise noted, amino acid sequences are written left to right in amino to carboxy orientation.
[0054] The terms "polynucleotide" or "nucleic acid" are used interchangeably herein and refer to polymers of nucleotides of any length, comprising ribonucleotides, deoxyribonucleotides, their analogs, or mixtures thereof. The term refers to the primary structure of the molecule. Thus, the term includes triplex-, double-, and single-stranded deoxyribonucleic acid ("DNA"), as well as triplex-, double-, and single-stranded ribonucleic acid ("RNA"). It also includes modified forms of polynucleotides, for example, by alkylation and / or capping, as well as unmodified forms of polynucleotides. More specifically, the term "polynucleotide" includes polydeoxyribonucleotides (containing 2-deoxy-D-ribose), polyribonucleotides (containing D-ribose) (including mRNA and gRNA, whether spliced or unspliced), any other type of polynucleotide that is an N- or C-glycoside of a purine or pyrimidine base, and other polymers containing a normucleotidic backbone, such as polyamides (e.g., peptide nucleic acids "PNAs") and polymorpholino polymers, and other synthetic sequence-specific nucleic acid polymers, provided that the polymer contains nucleobases in a configuration that allows for base pairing and base stacking as found in DNA and RNA.
[0055] The terms "nucleic acid sequence" and "nucleotide sequence" are used interchangeably and refer to a contiguous nucleic acid sequence. The sequence may be single- or double-stranded DNA or RNA (e.g., gRNA).
[0056] As used herein, the term "subsequence" refers to a subset of contiguous nucleotides in a sequence (either the physical sequence or its symbolic representation).
[0057] The methods disclosed herein can be used, for example, for the production of biologics such as antibodies.
[0058] As used herein, the term "antibody" (Ab) is intended to include, but is not limited to, a glycoprotein immunoglobulin, or antigen-binding portion thereof, that specifically binds to an antigen and comprises at least two heavy (H) chains and two light (L) chains interconnected by disulfide bonds. Each H chain contains a heavy chain variable region (referred to herein as V H The heavy chain constant region comprises three constant domains: C H1 , C H2 and C H3 Each light chain comprises a light chain variable region (referred to herein as V L The light chain constant region contains one constant domain, C L Includes V H and V L The region can be further subdivided into regions of hypervariability called complementarity-determining regions (CDRs), which are interspersed with more conserved regions called framework regions (FRs). H and V L comprises three CDRs and four FRs, arranged from the amino terminus to the carboxy terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, and FR4. The variable regions of the heavy and light chains contain binding domains that interact with antigens. The constant region of the antibody can mediate binding of the immunoglobulin to host tissues or factors, including various immune system cells (e.g., effector cells) and the first component (C1q) of the classical complement system. Thus, for example, the term "anti-PD-1 antibody" includes a complete antibody having two heavy chains and two light chains and specifically binding to PD-1, as well as an antigen-binding portion of the complete antibody. Non-limiting examples of antigen-binding portions are provided elsewhere herein. In some embodiments of the present disclosure, the anti-PD-1 antibody is nivolumab or an antigen-binding portion thereof.
[0059] In some embodiments, the antibody is a bispecific antibody. A "bispecific antibody" is a specific type of "bispecific molecule" or "bispecific binding molecule." The term "bispecific antibody" refers to an antibody that can bind to at least two antigenic determinants (e.g., epitopes) through two different antigen-binding sites. In certain embodiments, a bispecific antibody is capable of simultaneously binding to two antigenic determinants (e.g., epitopes). In some embodiments, a bispecific antibody binds to one antigen (or epitope) in one of its binding arms (a pair of heavy chains / light chains) and to a different antigen (or epitope) in its second binding arm (a different pair of heavy chains / light chains). In some embodiments, a bispecific antibody can have two separate antigen-binding arms (both in specificity and CDR sequences) and is monovalent for each antigen to which it binds. Bispecific antibodies include those generated, for example, by quadroma technology [Milstein & Cuello (1983) Nature 305(5934):537-40], by chemical conjugation of two different monoclonal antibodies [Staerz et al. (1985) Nature 314(6012):628-31], or by knob-into-hole or similar techniques introducing mutations into the Fc region [Holliger et al. (1993) Proc. Natl. Acad. Sci. USA 90(14):6444-6448].
[0060] In recent years, a wide variety of recombinant antibody formats have been created, including trivalent and tetravalent bispecific antibodies. Examples include fusions of IgG antibody formats with single-chain domains (for different formats see, for example, Coloma, MJ, et al, Nature Biotech 15 (1997), 159-163; WO2001 / 077342; Morrison, SL, Nature Biotech 25 (2007), 1233-1234; Holliger, P. et al, Nature Biotech. 23 (2005), 1 126-1 136; Fischer, N., and Leger, O., Pathobiology 74 (2007), 3-14; Shen, J., et al, J. Immunol. Methods 318 (2007), 65-74; Wu, C, et al, Nature Biotech. 25 (2007), 1290-1297). Bispecific antibodies include trivalent or tetravalent bispecific antibodies produced according to the methods disclosed in WO2009 / 080251, WO2009 / 080252, WO2009 / 080253, WO2009 / 080254, WO2010 / 112193, WO2010 / 115589, WO2010 / 136172, WO2010 / 145792, WO2010 / 145793, and WO2011 / 117330, all of which are incorporated herein by reference in their entireties. Those skilled in the art will understand that higher valencies may also be used.
[0061] Recently, a wide variety of recombinant bispecific antibody formats have been created, for example by fusing IgG antibody formats with single chain domains (see Kontermann RE, mAbs 4:2, (2012) 1-16). Bispecific antibodies in which the variable domains VL and VH or the constant domains CL and CH1 are replaced by each other are described in WO2009080251 and WO2009080252.
[0062] To avoid the problem of mispaired by-products, a technique known as "knobs-into-holes" aims to force pairing between two different antibody heavy chains by introducing mutations into the CH3 domain to modify the contact interface. In one chain, a bulky amino acid was replaced with an amino acid with a short side chain to create a "hole." Conversely, an amino acid with a large side chain was introduced into the CH3 domain of the other chain to create a "knob." By coexpressing these two heavy chains (and two identical light chains that must be compatible with both heavy chains), high yields of heterodimer formation ("knobs-into-holes") were observed relative to homodimer formation ("hole-into-holes" or "knobs-into-knobs") (Ridgway JB, Presta LG, Carter P, and WO1996027011). The percentage of heterodimers could be further increased by remodeling the interaction surfaces of the two CH3 domains using phage display techniques and introducing disulfide bridges to stabilize the heterodimers [Merchant AM, et al., Nature Biotech 16 (1998) 677-681; Atwell S, Ridgway JB, Wells JA, Carter P., J Mol Biol 270 (1997) 26-35]. A novel approach using knob-into-hole technology is described, for example, in EP1870459A1. Xie, Z., et al., J Immunol Methods 286 (2005) 95-101, describes a bispecific antibody format using scFvs combined with knob-into-hole technology for the Fc region. See Godar et al. (2018) "Therapeutic bispecific antibody formats: a patent applications review (1994-2017)" Expert. Opin. Ther. Pat. 28(3):251-276, and Brinkmann & Kontermann (2017) "The making of bispecific antibodies" mAbs 9:182-212.both of which are incorporated herein by reference in their entirety. Also, Ridgway et al (1996) Protein Eng 9:617-21, Atwell et al (1997) J. Mol. Biol. 270:26-35, Merchant et al (1998) Nat. Biotechnol. 16:677-681, Moore et al (2011) MAbs 3:546-55, Von Kreudenstein et al (2013) MAbs 5:646-54, Gunasekaran et al (2010) J. Biol. Chem. 285:19637-47, Geuijen et al (2014) J. Clin. Oncology 32:suppl:560, Strop et al (2012) J. Mol. Biol. 420:204-19, Choi et al (2013) Mol. Cancer Ther. 12:2748-59, Choi et al (2015) Mol. Immunol. 65:377-83, Labrijn et al (2013) Proc. Natl. Acad. Sci. USA 110:5145-50, Davis et al (2010) Protein Eng. 23:195-202, Moretti et al (2013) BMC Proceedings 7(Suppl 6):O9, and Leaver-Fey et al (2016) Structure 24:641-51, all of which are incorporated herein by reference in their entireties.Light chain pairing strategies are disclosed, for example, in Schaefer et al. (2011) Proc Natl Acad Sci US A. 108(27):11187-92, Lewis et al. (2014) Nat Biotechnol. 32(2):191-8, Mazor et al. (2015) MAbs. 7(2):377-89, Liu et al. (2015) J Biol Chem. 290(12):7535-62, Dillon et al. (2017) MAbs. 9(2):213-230, and U.S. Patent No. 9,914,785, all of which are incorporated herein by reference in their entireties.
[0063] Immunoglobulins can be derived from any of the widely known isotypes, including, but not limited to, IgA, secretory IgA, IgG, and IgM. IgG subclasses are also well known to those skilled in the art and include, but are not limited to, human IgG1, IgG2, IgG3, and IgG4. "Isotype" refers to the class or subclass of antibody (e.g., IgM or IgG1) encoded by the heavy chain constant region genes. The term "antibody" includes, by way of example, both naturally occurring and non-naturally occurring antibodies, monoclonal and polyclonal antibodies, chimeric and humanized antibodies, human or non-human antibodies, fully synthetic antibodies, and single-chain antibodies. Non-human antibodies can be humanized by recombinant methods to reduce their immunogenicity in humans. Unless otherwise specified and unless the context dictates otherwise, the term "antibody" also includes antigen-binding fragments or portions of any of the aforementioned immunoglobulins, including monovalent and bivalent fragments or portions, and single-chain antibodies.
[0064] An "isolated antibody" refers to an antibody that is substantially free of other antibodies having different antigenic specificities (e.g., an isolated antibody that specifically binds to an antigen such as PD-1 is substantially free of antibodies that specifically bind to antigens other than PD-1). However, an isolated antibody that specifically binds PD-1 may have cross-reactivity with other antigens, such as PD-1 molecules from different species. Moreover, an isolated antibody may be substantially free of other cellular material and / or chemicals.
[0065] The term "monoclonal antibody" (mAb) refers to a non-naturally occurring preparation of antibody molecules of single molecular composition, i.e., antibody molecules essentially identical in primary sequence and displaying a single binding specificity and affinity for a particular epitope. A monoclonal antibody is an example of an isolated antibody. Monoclonal antibodies can be produced by hybridoma, recombinant, transgenic, or other techniques known to those skilled in the art.
[0066] A "human antibody" (HuMAb) refers to an antibody having variable regions in which both the framework and CDR regions are derived from human germline immunoglobulin sequences. Additionally, if the antibody contains a constant region, the constant region also is derived from human germline immunoglobulin sequences. The human antibodies of the present disclosure may include amino acid residues not encoded by human germline immunoglobulin sequences (e.g., mutations introduced by random or site-specific mutagenesis in vitro or by somatic mutation in vivo). However, as used herein, the term "human antibody" is not intended to include antibodies in which CDR sequences derived from the germline of another mammalian species, such as a mouse, have been grafted onto human framework sequences. The terms "human antibody" and "fully human antibody" are used interchangeably.
[0067] A "humanized antibody" refers to an antibody in which some, most, or all of the amino acids outside the CDRs of a non-human antibody have been replaced with the corresponding amino acids from a human immunoglobulin. In one embodiment of a humanized antibody, some, most, or all of the amino acids outside the CDRs have been replaced with amino acids from a human immunoglobulin, while some, most, or all of the amino acids within one or more CDRs remain unchanged. Small additions, deletions, insertions, substitutions, or modifications of amino acids are acceptable as long as they do not eliminate the antibody's ability to bind to a specific antigen. A "humanized antibody" retains antigen specificity similar to that of the original antibody.
[0068] A "chimeric antibody" refers to an antibody whose variable region is derived from one species and whose constant region is derived from another species, e.g., whose variable region is derived from a murine antibody and whose constant region is derived from a human antibody.
[0069] An "anti-antigen antibody" refers to an antibody that specifically binds to an antigen. For example, an anti-PD-1 antibody specifically binds to the PD-1 antigen, and an anti-PD-L1 antibody specifically binds to the PD-L1 antigen.
[0070] An "antigen-binding portion" (also called an "antigen-binding fragment") of an antibody refers to one or more fragments of an antibody that retain the ability to specifically bind to the antigen bound by the whole antibody. It has been shown that the antigen-binding function of an antibody can be performed by fragments of a full-length antibody. Examples of binding fragments encompassed by the term "antigen-binding portion" of an antibody, such as an anti-PD-1 antibody or an anti-PD-L1 antibody, include (i) an Fab fragment (a fragment from papain cleavage), or a V L , V H (ii) a F(ab')2 fragment (a fragment resulting from pepsin cleavage) or a similar bivalent fragment comprising two Fab fragments linked by a disulfide bridge at the hinge region; (iii) a V H and an Fd fragment consisting of the CH1 domain; (iv) a V arm of an antibody L and V HFv fragment consisting of domains, (v) V H (vi) an isolated complementarity-determining region (CDR); and (vii) a combination of two or more isolated CDRs, optionally joined by a synthetic linker. In addition, the two domains of the Fv fragment, V, are also included. L and V H are encoded by separate genes, but recombinant methods can be used to synthesize them into a single protein chain (V L and V H They can be joined by a synthetic linker that allows for the pairing of the domains to form a monovalent molecule (known as a single-chain Fv (scFv); see, e.g., Bird et al. (1988) Science 242:423-426, and Huston et al. (1988) Proc. Natl. Acad. Sci. USA 85:5879-5883). Such single-chain antibodies are also intended to be encompassed by the term "antigen-binding portion" of an antibody. These antibody fragments are obtained using conventional techniques known to those skilled in the art, and the fragments are screened for utility in the same manner as intact antibodies. Antigen-binding portions can be produced by recombinant DNA techniques or by enzymatic or chemical cleavage of intact immunoglobulins.
[0071] In some embodiments, the biologic may be a protein, polypeptide, or polynucleotide. In some embodiments, the biologic is an enzyme, receptor, receptor ligand, protein antibiotic, fusion protein, structural protein, regulatory protein, vaccine, growth factor, hormone, or cytokine. In some embodiments, the biologic may comprise one or more heterologous moieties, such as a moiety for extending the plasma half-life of the biologic, a moiety for facilitating transport across a membrane or blood-brain barrier, a moiety for increasing or decreasing clearance rate, or a moiety for directing the biologic to a specific cell or tissue type (i.e., a targeting moiety).
[0072] As disclosed herein, a polynucleotide, vector, polypeptide, cell, or any composition that is "isolated" is a polynucleotide, vector, polypeptide, cell, or composition that is in a form not found in nature. Isolated polynucleotides, vectors, polypeptides, or compositions include those that have been purified to the extent that they are no longer in the form in which they are found in nature. In some embodiments, an isolated polynucleotide, vector, polypeptide, or composition is substantially pure.
[0073] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acids of any length. The polymers may contain modified amino acids. These terms also encompass amino acid polymers that are modified naturally or by intervention, for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or some other manipulation or modification, such as conjugation with a labeling component. This definition also includes, for example, polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids such as homocysteine, ornithine, p-acetylphenylalanine, D-amino acids, and creatine), as well as other modifications known in the art.
[0074] The term "percent sequence identity" between two polypeptide or polynucleotide sequences refers to the number of identical matching positions shared by the sequences in the comparison window, taking into account additions or deletions (i.e., gaps) that must be introduced for optimal alignment of the two sequences. A matching position is any position where the same nucleotide or amino acid appears in both the target sequence and the reference sequence. Since gaps are neither nucleotides nor amino acids, gaps that appear in the target sequence are not counted. Similarly, gaps that appear in the reference sequence are not counted, since the nucleotides or amino acids of the target sequence are counted, not the nucleotides or amino acids of the reference sequence. When comparing DNA and RNA, thymine (T) and uracil (U) can be considered equivalent.
[0075] The percentage of sequence identity is calculated by determining the number of positions where the same amino acid residue or nucleic acid base occurs in both sequences to determine the number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window, and multiplying the result by 100 to determine the percentage of sequence identity. Sequence comparison and determination of percent sequence identity between two sequences can be achieved using readily available software, both online and downloadable. Suitable software programs are available from various sources for aligning both protein and nucleotide sequences. One suitable program for determining percent sequence identity is bl2seq, which is part of the BLAST suite of programs available from the U.S. government's National Center for Biotechnology Information BLAST website (blast.ncbi.nlm.nih.gov). Bl2seq uses either the BLASTN or BLASTP algorithm to compare two sequences. BLASTN is used to compare nucleic acid sequences, and BLASTP is used to compare amino acid sequences. Other suitable programs are, for example, Needle, Stretcher, Water, or Matcher, which are part of the EMBOSS suite of bioinformatics programs and are available from the European Bioinformatics Institute (EBI) at www.ebi.ac.uk / Tools / psa.
[0076] Different regions in a single polynucleotide or polypeptide target sequence that aligns with a polynucleotide or polypeptide reference sequence can each have a unique percent sequence identity.It is important to note that percent sequence identity values are rounded to one decimal place.For example, 80.11, 80.12, 80.13, and 80.14 are rounded to 80.1, and 80.15, 80.16, 80.17, 80.18, and 80.19 are rounded to 80.2.It is also important to note that length values are always integers.
[0077] In certain embodiments, the percentage identity "%ID" of a first amino acid sequence (or nucleic acid sequence) to a second amino acid sequence (or nucleic acid sequence) is calculated as %ID=100×(Y / Z), where Y is the number of amino acid residues (or nucleic acid bases) scored as identical matches in an alignment of the first and second sequences (when aligned by visual inspection or by a particular sequence alignment program), and Z is the total number of residues in the second sequence. If the length of the first sequence is longer than the second sequence, the percent identity of the first sequence to the second sequence will be higher than the percent identity of the second sequence to the first sequence.
[0078] Those skilled in the art will understand that the generation of sequence alignments for calculating percent sequence identity is not limited to binary sequence comparisons determined solely by primary sequence data. It will also be understood that sequence alignments can be generated by integrating sequence data with data from heterogeneous sources, such as structural data (e.g., crystallographic protein structures), functional data (e.g., mutation locations), or phylogenetic data. A suitable program for integrating heterogeneous data to generate multiple sequence alignments is T-Coffee, available at www.tcoffee.org, or alternatively, from EBI, etc. It will also be understood that the final alignment used to calculate percent sequence identity can be curated automatically or manually.
[0079] The terms "gene," "coding sequence," "encoding nucleic acid," "open reading frame," "ORF," and grammatical variations thereof, are used interchangeably in this disclosure and refer to a nucleic acid (RNA or DNA molecule) comprising a nucleotide sequence that encodes a gene of interest (GOI), which is often a protein, such as a biologic, such as an antibody. The coding sequence may further comprise initiation and termination signals operably linked to regulatory elements comprising a promoter and polyadenylation signal capable of directing expression in the cells of an individual or mammal to which the nucleic acid is administered. The coding sequence may also be codon-optimized.
[0080] As used herein, the term "gene of interest," abbreviated as "GOI," refers to an exogenous protein expressed by the cells disclosed herein. In some embodiments, the GOI is a biologic, such as an antibody or a portion thereof. In some embodiments, the GOI comprises one or more open reading frames encoding one or more recombinant proteins, e.g., operably linked to one or more promoters and / or other regulatory sequences. In some embodiments, the cells disclosed herein can contain a first GOI, which can be replaced by a second GOI. In some embodiments, the first GOI (e.g., the GOI located on the parent plasmid) and the second GOI (e.g., the GOI located on the second GOI plasmid) belong to the same molecular class. For example, if the first GOI was an antibody, the second GOI could also be an antibody because the parent cell line efficiently expressed that type of recombinant protein. In some embodiments, the GOI is a nucleic acid, such as a therapeutic nucleic acid. In some embodiments, the terms GOI and ORF can be used interchangeably, particularly when the GOI is encoded by a single ORF. In some embodiments, the GOI may be encoded by more than one ORF. In some embodiments, the GOI may be a detectable molecule, such as a marker.
[0081] As used herein, "complement" or "complementary" refers to Watson-Crick (e.g., AT / U and CG) or Hoogsteen base pairing between nucleotides or nucleotide analogs of a nucleic acid molecule. "Complementarity" refers to the property shared between two nucleic acid sequences such that when they are aligned antiparallel to each other, the nucleotide bases at each position are complementary.
[0082] The terms "vector," "expression vector," "plasmid," and grammatical variations thereof are used interchangeably in this disclosure and refer to a polynucleotide exogenous to the genome of a host cell that is inserted into a specific location within the genome of the host cell. Generally, a plasmid contains multiple elements, such as recombination sites (e.g., homologous recombination sites and / or site-specific recombination sites), markers (e.g., detectable markers and / or selectable markers), one or more expression cassettes, or any combination thereof. In some embodiments, the plasmid may be a linear plasmid. In other embodiments, the plasmid may be a circular plasmid, e.g., an intact circular plasmid.
[0083] An "expression cassette" comprises a DNA coding sequence operably linked to a promoter. "Operably linked" refers to a juxtaposition of the components so described, in a relationship permitting them to function in their intended manner. For example, a promoter is operably linked to a coding sequence if the promoter affects its transcription or expression.
[0084] As used herein, "host cell" refers to an in vivo or in vitro eukaryotic cell, a prokaryotic cell (e.g., a bacterial or archaeal cell), or a cell from a multicellular organism cultured as a unicellular organism (e.g., a cell line), which eukaryotic or prokaryotic cell can or has become a recombinant host cell by being used to receive heterologous nucleic acid. Thus, the term host cell also includes progeny of the original host cell (i.e., the host cell prior to receiving the heterologous nucleic acid) transformed with heterologous nucleic acid, i.e., a recombinant host cell. It is understood that the progeny of a unicellular cell may not necessarily be completely identical in morphology or genomic or total DNA complement to the original parent due to natural, accidental, or deliberate mutation.
[0085] A "recombinant host cell" or "genetically modified host cell" is a host cell into which heterologous nucleic acid, such as an expression vector, has been introduced. For example, a eukaryotic host cell becomes a recombinant or genetically modified eukaryotic host cell (e.g., a mammalian host cell) by the introduction of exogenous nucleic acid into the eukaryotic host cell.
[0086] As used herein, the terms "hot cell," "hot clone," and "hot cell line" each refer to a cell, clone, or cell line that has advantageous properties, e.g., a higher yield of recombinant protein compared to other cells, clones, or cell lines expressing the same recombinant protein. For example, a hot cell, hot clone, or hot cell line may be capable of expressing increased amounts of recombinant protein, expressing elevated levels of correctly folded recombinant protein, expressing recombinant protein with reduced levels of high molecular weight aggregates, expressing recombinant protein with reduced levels of fragmentation, or any combination thereof, or some other desirable property.
[0087] As used herein, the term "hotspot" refers to a genomic location (locus) into which an exogenous sequence, such as a plasmid containing a polynucleotide sequence encoding a protein for recombinant expression, can be inserted, where (i) transcription of the exogenous sequence is not silenced (e.g., by epigenetic modification), and (ii) transcription of the exogenous sequence occurs at a higher level than the transcription level observed when the exogenous sequence is inserted at another location (e.g., a reference location). In some embodiments, a hotspot does not contain a functional ORF. Thus, in some embodiments, a hotspot does not contain one or more actively transcribed genes. Hotspots lacking actively transcribed genes are particularly advantageous because their partial or complete deletion for inserting a polynucleotide sequence encoding an exogenous gene (gene of interest) does not disrupt the production of endogenous proteins. In some embodiments, the hotspots of the present disclosure are located adjacent to one actively transcribed gene or between two actively transcribed genes, i.e., the hotspot may be flanked by two actively transcribed genes. In some embodiments, inserting a polynucleotide sequence encoding an exogenous gene (gene of interest) into a hotspot of the present disclosure does not affect the expression of one or more actively transcribed genes adjacent to or flanking the hotspot. In some embodiments, inserting a polynucleotide sequence encoding an exogenous gene (gene of interest) into a hotspot of the present disclosure reduces the expression of one or more actively transcribed genes adjacent to or flanking the hotspot by less than about 50%, less than about 45%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, or less than about 10%.
[0088] As used herein, the term "addressable," as applied to the polynucleotide sequences disclosed herein, e.g., the landing pad sequences disclosed herein, refers to a polynucleotide sequence that is uniquely identified by the presence of a unique site-specific recombination site (SSRS) or a combination thereof. For example, a first landing pad having Lox511 and LoxP sites and a second landing pad having Loxm3 and Loxm7 sites are addressable relative to each other. Thus, in some embodiments, a landing pad may be addressable due to the presence of a specific combination of two SSRSs. In other embodiments, a landing pad may be addressable relative to a second landing pad via a single SSRS. For example, a first landing pad may have a first single aatP site and a second landing pad may have a second single aatP site, where these aatP sites are not mutually compatible. In some aspects of the present disclosure, there can be multiple contiguous landing pads, each uniquely addressable by the presence of a unique SSRS or combination thereof that uniquely identifies (addresses) a given landing pad.
[0089] As used herein, the term "addressable SSRS" refers to a unique SSRS or combination thereof that can be specifically targeted for recombination. As used herein, the term "addressable landing pad plasmid" refers to a landing pad plasmid that includes an addressable SSRS or combination thereof that can be specifically targeted for recombination.
[0090] As used herein, the term "incompatible" as applied to a pair of site-specific recombination sites refers to sites that are defective in recombination with alternative SSRSs, i.e., that are unable to recombine or that only partially retain cross-reactivity with alternative SSRSs. For example, two Lox sites, such as LoxP and Lox511, that have reduced recombination ability with each other are considered incompatible. Similarly, two attP-aatB site pairs that have reduced recombination ability with each other are considered incompatible.
[0091] As used herein, the terms "head-to-head," "tail-to-tail," "tail-to-head," and "head-to-tail" refer to the relative orientation of two polynucleotide sequences, such as two landing pads, landing pad plasmids, expression plasmids, or genes of interest, in the genetic constructs disclosed herein. The term "head" refers to the 5' end of a nucleic acid sequence, and the term "tail" refers to the 3' end of a nucleic acid sequence. Thus, a 3'-5'5'-3' configuration is head-to-head because both 5' ends (heads) of the original sequence are adjacent to each other (considering the 5' to 3' ends relative to the construct). Thus, a 5'-3'3'-5' configuration is tail-to-tail, 5'-3'5'-3' is tail-to-head, and 3'-5'3'-5' is head-to-tail.
[0092] II. Landing pad cells The present disclosure provides landing pad cells that can be used for recombinant expression of at least one gene of interest (GOI). In some embodiments, these cell lines include a "landing pad," i.e., one or more specific polynucleotide sequences inserted into the genome of a parent cell, which can be replaced, e.g., by recombination, with another one or more specific polynucleotide sequences that include a nucleotide sequence encoding at least one GOI. In some embodiments, instead of replacing a polynucleotide sequence, e.g., by recombination, one or more specific polynucleotide sequences that include a nucleotide sequence encoding at least one GOI can be inserted at a location within the landing pad, e.g., via an aat site.
[0093] As an overview of one process disclosed herein, a parent cell line (e.g., a "historical" cell line known to efficiently express a particular biologic) is modified by fully or partially replacing an exogenous polynucleotide sequence (i.e., a "parent plasmid") containing a parental GOI or first GOI with a second exogenous polynucleotide sequence (i.e., a "landing pad plasmid or portion thereof"). The resulting cell line, which has integrated the landing pad plasmid or portion thereof rather than the entire parent plasmid, becomes a "landing pad cell." In some embodiments, the landing pad plasmid of the present disclosure includes flanking sequences from the parent plasmid.
[0094] The landing pad plasmid in the landing pad cell can then be replaced (e.g., partially) by recombination with another polynucleotide (a "GOI plasmid") containing a different or second GOI, thereby resulting in an "expressing cell." See, e.g., the process depicted in Figures 5A and 5B and Table 1.
[0095] [Table 1]
[0096] Thus, in some embodiments, the present disclosure provides an expression cell comprising at least one expression plasmid (P4), e.g., a linearized plasmid, integrated into a genomic sequence, wherein each expression plasmid comprises: (i) a polynucleotide sequence derived from an expression plasmid (P4) comprising a nucleic acid encoding a gene of interest (second GOI); (ii) two SSRSs (e.g., when a recombinase system such as Lox is used) or a single SSRS (e.g., when an integrase system such as att is used) flanking the polynucleotide sequence of (i); (iii) a polynucleotide sequence located distal to the polynucleotide of (i) and the SSRS of (ii), wherein both flanking polynucleotide sequences of (iii) are derived from the landing pad plasmid (P2); (iv) providing an expression cell comprising a polynucleotide sequence distally flanking the polynucleotide sequence of (iii), wherein both flanking polynucleotide sequences of (iv) are derived from the parental plasmid (P1).
[0097] The term "site-specific recombination site," abbreviated as "SSRS," as used herein, comprises a nucleotide sequence that can be recognized by a site-specific recombinase and serve as a substrate for a recombination event. In some embodiments, a construct (e.g., a landing pad plasmid or expression plasmid) disclosed herein can comprise two SSRSs, one located upstream and one located downstream relative to the nucleic acid encoding a GOI or marker. In some embodiments, a construct (e.g., a landing pad plasmid or expression plasmid) disclosed herein can comprise a single SSRS located either upstream or downstream relative to the nucleic acid encoding a GOI or marker. In some embodiments, a construct (e.g., a landing pad plasmid or expression plasmid) disclosed herein can comprise three or more SSRSs, all of which are located upstream relative to the nucleic acid encoding a GOI or marker, all of which are located downstream relative to the nucleic acid encoding a GOI or marker, or some of which are located upstream and some of which are located downstream relative to the nucleic acid encoding a GOI or marker.
[0098] In the formulas disclosed herein that include two SSRSs, it should be understood that if recombination occurs using a system that requires a single SSRS (e.g., att) instead of a recombination system that requires two SSRSs (e.g., lox or Frt), one of the two SSRSs in the formula can be omitted and may not be present. When one of the SSRS sites in the above formula is absent, the single SSRS site may be either the upstream SSRS or the downstream SSRS relative to the [M] or [P3] component. In some embodiments, the single SSRS is an att site.
[0099] As used herein, the term "site-specific recombinase" includes enzymes capable of causing recombination between "recombination sites," where the two recombination sites are located within a single nucleic acid molecule or on separate nucleic acid molecules. Examples of "site-specific recombinases" include, but are not limited to, Cre, Flp, and Dre recombinase. In some embodiments, the site-specific recombinase is an integrase, e.g., λ (lambda) integrase. In some embodiments, the site-specific recombinase is a Bxb integrase, e.g., Bxb1 integrase. Bxb1, the integrase encoded by mycobacteriophage Bxb1, is a member of the serine-recombinase family and catalyzes strand exchange between attP and attB (the binding sites for the phage and bacterial host, respectively).
[0100] The present disclosure provides a landing pad cell comprising at least one plasmid, e.g., a linear or circular plasmid, or a combination thereof, integrated into a genomic sequence, wherein each plasmid comprises: (i) a polynucleotide sequence derived from a landing pad plasmid (P2) comprising at least one marker and two site-specific recombination sites (SSRS) flanking the at least one marker (if a recombinase system such as Lox is used, or a single SSRS if an integrase system such as att is used); and (ii) providing a landing pad cell comprising polynucleotide sequences flanking the polynucleotide sequence of (i), wherein both flanking polynucleotide sequences of (ii) are derived from the parental plasmid (P1);
[0101] When a single SSRS is present, a description of its location as "flanking" another element in the formula, e.g., a [P2] or [P3] element (i.e., an element encoding a marker or GOI), should be understood to refer to the location of the SSRS immediately upstream or downstream relative to the flanked element. As an example, in the formula CG1 / -[P1]-[P2]-[SSRS]-[P3]-[P2]-[P1]- / CG2, the [SSRS] would flank [P3], which would encode the GOI.
[0102] The present disclosure provides an expression cell comprising a plasmid, such as a linearized plasmid, inserted into its genomic sequence, wherein the topology of the plasmid is of the formula CG1 / -[P1]-[P2]-[SSRS]-[P3]-[SSRS]-[P2]-[P1]- / CG2 [In the ceremony CG1 and CG2 are parental cell genomic sequences flanking the inserted linear plasmid; [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [P3] is a polynucleotide sequence derived from a second GOI plasmid containing a gene of interest (GOI), [SSRS] is a site-specific recombination site (SSRS)].
[0103] The present disclosure also contemplates landing pad cells comprising multiple plasmids, e.g., landing plasmids or portions thereof. Thus, the present disclosure also provides a landing pad cell comprising at least one plasmid, e.g., one, two, three, or more linear or circular plasmids, integrated into a genomic sequence, each plasmid comprising a polynucleotide sequence derived from a landing pad plasmid (P2) comprising at least one marker and two site-specific recombination sites (SSRS) flanking the at least one marker, and a polynucleotide sequence flanking the polynucleotide sequence of (i), wherein both flanking polynucleotide sequences of (ii) are derived from a parental plasmid (P1).
[0104] Thus, the present disclosure provides an expression cell comprising a plasmid, such as a linearized plasmid, inserted into its genomic sequence, wherein the topology of the plasmid is of the formula CG1 / -([P1]-([P2]-[SSRS]-[P3]-[SSRS]-[P2])n-[P1])- / CG2 [In the ceremony CG1 and CG2 are parental cell genomic sequences flanking the inserted linear plasmid; [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [P3] is a polynucleotide sequence derived from a plasmid containing a gene of interest (GOI), [SSRS] is a site-specific recombination site (SSRS), wherein n is an integer between 1 and 10.
[0105] In some cases, [P3] may comprise a single GOI or multiple GOIs. In some embodiments, either the 5'[SSRS] or the 3'[SSRSA] can be omitted. In some embodiments, the expression cell comprises a plasmid, and the plasmid is an expression plasmid.
[0106] In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, n is 4. In some embodiments, n is 5. In some embodiments, n is 6. In some embodiments, n is 7. In some embodiments, n is 8. In some embodiments, n is 9. In some embodiments, n is 10. In some embodiments, all of the plasmids are identical. In some embodiments, all of the plasmids are different. In some embodiments, at least one of the plasmids is different.
[0107] In some embodiments, CG1 comprises the polynucleotide sequence of SEQ ID NO: 18, or a fragment thereof. In some embodiments, CG2 comprises the polynucleotide sequence of SEQ ID NO: 19, or a fragment thereof.
[0108] In some embodiments, CG1 comprises the polynucleotide sequence of SEQ ID NO: 114, or a fragment thereof. In some embodiments, CG2 comprises the polynucleotide sequence of SEQ ID NO: 115, or a fragment thereof.
[0109] The present disclosure provides a landing pad cell comprising at least one plasmid, e.g., a linearized plasmid, integrated into a genomic sequence, wherein the plasmid comprises: a. a polynucleotide sequence comprising a nucleic acid encoding at least one selectable marker and / or at least one nucleic acid sequence encoding a detectable marker; b. two SSRSs flanking the polynucleotide sequence of (1); c. Two homologous recombination sites located 5' and 3' to the SSRS of (2), which are homologous to the corresponding homologous recombination sites in the parental plasmid; A landing pad cell is provided, comprising:
[0110] In some embodiments, the topology of the plasmid, such as a linear plasmid, in the landing pad cell is represented by the formula CG1 / -[P1]-[P2]-[SSRS]-[M]-[SSRS]-[P2]-[P1]- / CG2 [In the ceremony CG1 and CG2 are parental cell genomic sequences flanking the inserted linear plasmid; [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [M] is a polynucleotide sequence comprising at least one marker, e.g., a screenable marker, a selectable marker, or a combination thereof; [SSRS] corresponds to the site-specific recombination site (SSRS).
[0111] In some embodiments, the topology of the plasmid, such as a linear plasmid, in the landing pad cell is represented by the formula CG1 / -([P1]-([P2]-[SSRS]-[M]-[SSRS]-[P2])n-[P1])- / CG2 where CG1 and CG2 are parent cell genomic sequences flanking the inserted linearized plasmid; [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [M] is a polynucleotide sequence comprising at least one marker, e.g., a screenable marker, a selectable marker, or a combination thereof; [SSRS] is a site-specific recombination site (SSRS), n is an integer between 1 and 10. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, n is 4. In some embodiments, n is 5. In some embodiments, n is 6. In some embodiments, n is 7. In some embodiments, n is 8. In some embodiments, n is 9. In some embodiments, n is 10. In some embodiments, all of the plasmids are identical. In some embodiments, all of the plasmids are different. In some embodiments, at least one of the plasmids is different.
[0112] The present disclosure also provides a landing pad cell comprising a plasmid, such as a linearized plasmid, inserted into its genomic sequence, wherein the topology of the plasmid is as described above. CG1 / -[P1]-[P2]-[P1]- / CG2 [where: CG1 and CG2 are parental cell genomic sequences flanking the inserted plasmid; [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from a landing pad plasmid that includes at least one marker and two site-specific recombination sites (SSRS) flanking the at least one marker.
[0113] The present disclosure also provides a landing pad cell containing a plasmid, such as a linearized plasmid, inserted into its genome sequence, wherein the topology of the plasmid is as described below. CG1 / -[P1 * ]- / CG2 CG3 / -[P2]- / CG4 [where: CG1 and CG2 are parental cell genomic sequences flanking the inserted plasmid at the first hotspot; CG3 and CG4 are parental cell genomic sequences flanking the inserted plasmid at the second hotspot; [P1 * ] is a polynucleotide sequence derived from the parental plasmid containing at least a partial deletion, [P2] is a polynucleotide sequence derived from a landing pad plasmid that includes at least one marker and two site-specific recombination sites (SSRS) flanking the at least one marker].
[0114] The present disclosure also provides a landing pad cell comprising a plasmid, such as a linearized plasmid, inserted into its genomic sequence, wherein the topology of the plasmid is as described above. CG1 / -([P1]-([P2])n-[P1])- / CG2 [where: CG1 and CG2 are parental cell genomic sequences flanking the inserted plasmid; [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from a landing pad plasmid comprising at least one marker and two site-specific recombination sites (SSRS) flanking the at least one marker, where n is an integer between 1 and 10. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, n is 4. In some embodiments, n is 5. In some embodiments, n is 6. In some embodiments, n is 7. In some embodiments, n is 8. In some embodiments, n is 9. In some embodiments, n is 10. In some embodiments, all of the plasmids are identical. In some embodiments, all of the plasmids are different. In some embodiments, at least one of the plasmids is different.
[0115] The present disclosure also provides a landing pad cell comprising a plasmid, such as a linearized plasmid, inserted into its genomic sequence, the topology of the plasmid being as described below. CG1 / -([P1 * ]- / CG2 CG3 / -([P2])n- / CG4 [where: CG1 and CG2 are parental cell genomic sequences flanking the inserted plasmid at the first hotspot; CG3 and CG4 are parental cell genomic sequences flanking the inserted plasmid at the second hotspot; [P1 * ] is a polynucleotide sequence derived from the parental plasmid containing at least a partial deletion, [P2] is a polynucleotide sequence derived from a landing pad plasmid comprising at least one marker and two site-specific recombination sites (SSRS) flanking the at least one marker, where n is an integer between 1 and 10. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, n is 4. In some embodiments, n is 5. In some embodiments, n is 6. In some embodiments, n is 7. In some embodiments, n is 8. In some embodiments, n is 9. In some embodiments, n is 10. In some embodiments, all of the plasmids are identical. In some embodiments, all of the plasmids are different. In some embodiments, at least one of the plasmids is different.
[0116] In some embodiments, CG1 comprises the polynucleotide sequence of SEQ ID NO: 18, or a fragment thereof. In some embodiments, CG2 comprises the polynucleotide sequence of SEQ ID NO: 19, or a fragment thereof.
[0117] In some embodiments, CG3 comprises the polynucleotide sequence of SEQ ID NO: 114, or a fragment thereof. In some embodiments, CG4 comprises the polynucleotide sequence of SEQ ID NO: 115, or a fragment thereof.
[0118] In some embodiments, for example, if a linearized plasmid is inserted into a hotspot that is different from the original hotspot in the parental cell line, the CG1 and CG2 genomic sequences (the parental cell genomic sequences flanking the inserted linearized plasmid) will be replaced by the CG3 and CG4 genomic sequences, respectively, that correspond to the genomic sequences flanking the inserted linearized plasmid at the alternative hotspot.
[0119] The present disclosure also provides a landing pad plasmid for targeted integration into the genome of a host cell, including a plasmid such as a linear plasmid, wherein the topology of the plasmid is of the formula -[P1]-[P2]-[P1]- [In the ceremony [P1] is a polynucleotide sequence derived from the parental plasmid integrated into the host containing the homologous recombination site; [P2] is a polynucleotide sequence comprising at least one marker and two site-specific recombination sites (SSRS) flanking at least one marker].
[0120] The present disclosure also provides a landing pad plasmid for targeted integration into the genome of a host cell, including a plasmid such as a linear plasmid, wherein the topology of the plasmid is of the formula -([P1]-([P2])n-[P1])- [In the ceremony
[0023] In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, n is 4. In some embodiments, n is 5. In some embodiments, n is 6. In some embodiments, n is 7. In some embodiments, n is 8. In some embodiments, n is 9. In some embodiments, n is 10. In some embodiments, all of the plasmids are identical. In some embodiments, all of the plasmids are different. In some embodiments, at least one of the plasmids is different.
[0121] It is understood that the simplified topology of the plasmids disclosed herein (e.g., -[P1]-[P2]-[P1]-) may be described using the terms "description" or "formula" interchangeably.
[0122] The present disclosure also provides a method for generating landing pad cells, comprising: (a) integrating the landing pad plasmid, or a portion thereof, into the genome of the parent cell at a targeted integration site using homologous recombination; The landing pad plasmid is (1) a polynucleotide sequence comprising a nucleic acid encoding at least one selectable marker and / or at least one nucleic acid sequence encoding a detectable marker; (2) two SSRSs flanking the polynucleotide sequence of (1); and (3) two homologous recombination sites located 5' and 3' to the SSRS of (2), which are homologous to the corresponding homologous recombination sites in the parental plasmid; Including, The method provides a method in which a homologous recombination site in the landing pad plasmid recombines with a corresponding homologous recombination site in the parental plasmid, thereby integrating the landing pad plasmid at a location within the parental plasmid that was inserted into the parental cell genomic DNA.
[0123] The targeted integration process disclosed herein involves replacing a polynucleotide subsequence located between two recombination sites in a plasmid with another polynucleotide subsequence located between two corresponding recombination sites in another plasmid.Thus, for example, as illustrated in Figures 4A, 5A, 8A, and 8B, the targeted integration of a landing pad plasmid into a parental plasmid replaces the subsequence of the parental plasmid with the corresponding subsequence from the landing pad plasmid, leaving the remainder of the parental plasmid sequence between the recombination sites in the genome sequence.This targeted integration does not require a complete replacement of one plasmid with another plasmid.
[0124] Similarly, when the second GOI plasmid is recombined with the landing pad plasmid, the partial sequences between the SSRSs (e.g., LoxP sites) will be exchanged, but remnants of the landing pad plasmid will remain between the recombination site and the genomic sequence. In this case, the sequences from the second GOI plasmid will be flanked by sequences originating from the landing pad plasmid, which in turn will be flanked by sequences originating from the parental plasmid.
[0125] As can be seen from these descriptions, references herein to inserting one plasmid into another generally do not require the complete replacement of one plasmid with the other, rather, a plasmid is completely or partially replaced by another plasmid, or an excised plasmid is completely or partially excised.
[0126] In some embodiments, the present disclosure provides a method of generating an expressing cell, comprising: (a) using site-specific recombinase recombination to integrate a second GOI plasmid, such as a linearized plasmid, into the genome of a landing pad cell of the present disclosure at a targeted integration site, wherein the expression plasmid (P4) comprises: (1) a polynucleotide sequence comprising a nucleic acid encoding a GOI; (2) two SSRSs flanking the polynucleotide of (1); Including, The method provides a method in which a site-specific recombinase recombination site of the landing pad plasmid recombines with a corresponding site-specific recombinase recombination site of the GOI plasmid, thereby integrating the GOI plasmid at a location internal to the landing pad plasmid inserted into the genomic DNA of the landing pad cell.
[0127] The present disclosure also provides a method of generating an expressing cell, comprising: (a) using homologous recombination to integrate a second GOI plasmid, such as a linearized plasmid, into the genome of a parent cell containing the parent plasmid at a targeted integration site, wherein the expression plasmid is (1) a polynucleotide sequence comprising a nucleic acid encoding a GOI; (2) two SSRSs flanking the polynucleotide of (1); Including, The method provides a method in which a site-specific recombinase recombination site of a parental plasmid recombines with a corresponding site-specific recombinase recombination site of a GOI plasmid, thereby integrating the GOI plasmid at a location within the parental plasmid that was inserted into the parental cell genomic DNA.
[0128] The present disclosure also provides a method of generating an expressing cell, comprising: using homologous recombination to integrate a second GOI plasmid, such as a linearized plasmid, into the genome of a parent cell containing the parent plasmid at a targeted integration site, wherein the resulting expression plasmid comprises a polynucleotide sequence comprising a nucleic acid encoding the GOI; A method is provided in which a parental plasmid recombines with a GOI plasmid, thereby integrating the GOI plasmid at a location internal to the parental plasmid that is inserted into the parental cell genomic DNA.
[0129] The present disclosure also provides a method of generating an expressing cell, comprising: using homologous recombination to integrate a second GOI plasmid, such as a linearized plasmid, into the genome of a parent cell containing the parent plasmid at a targeted integration site; A method is provided in which the parental plasmid recombines with the flanking genomic sequences, thereby integrating the GOI plasmid into the parental plasmid inserted into the parental cell genomic DNA.
[0130] Although the methods disclosed herein involve two nested nuclease-mediated recombination events, e.g., homologous recombination between a parental plasmid and a landing pad plasmid in a parent cell, and site-specific recombinase recombination between a landing pad plasmid and a second GOI plasmid in a landing pad cell, it will be understood that other combinations of recombination events would be equally applicable, e.g., a first homologous recombination event between P1 and P2, and a second homologous recombination between P2 and P3.
[0131] It should also be understood that the teachings related to the present disclosure regarding the integration of plasmids, such as landing pad plasmids or GOI plasmids, are intended to encompass the insertion of multiple plasmids (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10), which may be the same or different and may differ with respect to their orientation in the final construct (e.g., whether each of the plasmids in the final construct is in a 5'-3' or 3'-5' orientation relative to the original construct and the other plasmids in the final construct).
[0132] In some embodiments, the present disclosure also provides a method of generating an expressing cell, comprising: using homologous recombination to integrate a GOI plasmid, such as a linearized plasmid, into the genome of the parent cell, wherein the expression plasmid comprises a polynucleotide sequence comprising a nucleic acid encoding the GOI; Methods are provided in which a GOI plasmid is integrated by homologous recombination at a location determined to correspond to a hotspot. In some embodiments, at least a portion of the parental plasmid is removed. In some embodiments, the entire parental plasmid is removed.
[0133] In some embodiments, the present disclosure also provides methods for identifying starting cells (parent cell lines) in an efficient manner for generating landing pad cell lines capable of producing high titers. Thus, the present disclosure provides methods for selecting parent cells for generating expressing cells, for example, as disclosed in Examples 1 and 2.
[0134] In some embodiments, the methods disclosed herein comprise removing at least a portion of at least one parental plasmid and introducing one or more landing pad plasmids, or portions thereof, comprising a landing pad into the genome of a cell. In some embodiments, the methods disclosed herein comprise removing only a portion of at least one parental plasmid and introducing one or more landing pad plasmids, or portions thereof, comprising a landing pad into the genome of a cell.
[0135] In some embodiments, the method of selecting a suitable parent cell line for generating a landing cell line of the present disclosure comprises: (i) Selecting a cell line with a high expression titer of the target gene; (ii) Further selection of cells with low copy numbers of the ORF encoding the gene of interest. Includes. In some embodiments, the parental cell has one or two copies of the ORF encoding the gene of interest, hi some embodiments, the parental cell has three or more copies of the ORF encoding the gene of interest.
[0136] In some embodiments, the method of selecting a landing pad cell line comprises screening for the loss of the parental plasmid or a portion thereof and selecting cells with such loss (deletion). In some embodiments, the method of selecting a parental cell line further comprises screening for the presence of a landing pad and selecting cells in which the landing pad is present. In some embodiments, the method further comprises screening the landing pad for characteristics such as the presence or absence of low or high complexity regions, the presence or absence of retrotransposon sequences, the presence or absence of Alu repeats, the presence or absence of long interspersed nuclear elements (LINEs), the presence or absence of islands, the level of cytosine methylation, the level of histone acetylation, the presence or absence of ORFs, and any combination thereof.
[0137] In some embodiments, the cell is a CHO cell. In some embodiments, the hotspot location comprises a sequence selected from SEQ ID NO: 18 or a fragment thereof and SEQ ID NO: 19 or a fragment thereof. In some embodiments, the hotspot location comprises a sequence selected from SEQ ID NO: 114 or a fragment thereof and SEQ ID NO: 115 or a fragment thereof.
[0138] In some embodiments, the GOI plasmid is inserted and integrated by homologous recombination or random integration at a location within the genomic sequence of SEQ ID NO: 18. In some embodiments, the GOI plasmid is inserted and integrated by homologous recombination or random integration at a location within the genomic sequence of SEQ ID NO: 19. In some embodiments, the GOI plasmid is inserted and integrated by homologous recombination or random integration at a location within the genomic sequence of SEQ ID NO: 114. In some embodiments, the GOI plasmid is inserted and integrated by homologous recombination or random integration at a location within the genomic sequence of SEQ ID NO: 115.
[0139] In some embodiments, the GOI plasmid is integrated by homologous recombination at a location within the genomic sequence, wherein the 5' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 18 or a subsequence thereof, and / or the 3' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 19 or a subsequence thereof.
[0140] In some embodiments, the GOI plasmid is integrated by homologous recombination at a location within the genomic sequence, wherein the 5' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 114 or a subsequence thereof, and / or the 3' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 115 or a subsequence thereof.
[0141] In some embodiments, the 5' homologous recombination site is at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 110, at least about 120, at least about 130, at least about 140, at least about 150, at least about 160, at least about 170, at least about 180, at least about 190, at least about 200, at least about 210, at least about 220, at least about 230, at least about 240, at least about 250, at least about 260, at least about 270, at least about 280, at least about 290, at least about 300, at least about 310, at least about 320, at least about 330, at least about 340, at least about 350, at least about 360, at least about 370, at least about 380, at least about 390, at least about 400, at least about 410, at least about 420, at least about 430, at least about 440, at least about 450, at least about 460, at least about 470, at least about 480, at least about 490, at least about 500, at least about 510, at least about 520, at least about 530, at least about 540, at least about 550, at least about 560, at least about 570, at least about 580, at least about 590, at least about 600, at least about 610, at least about 620, at least about 630, 00, at least about 310, at least about 320, at least about 330, at least about 340, at least about 350, at least about 360, at least about 370, at least about 380, at least about 390, at least about 400, at least about 410, at least about 420, at least about 430, at least about 440, at least about 450, at least about 460, at least about 470, at least about 480, at least about 490, at least about 500, at least about 510, at least about 520, at least about 530, at least about 540, at least about 550, at least about 560, at least about 570, at least about 570, at least about 580, at least about 590, or at least about 600 consecutive nucleotides.
[0142] In some embodiments, the 3' homologous recombination site is at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 110, at least about 120, at least about 130, at least about 140, at least about 150, at least about 160, at least about 170, at least about 180, at least about 190, at least about 200, at least about 210, at least about 220, at least about 230, at least about 240, at least about 250, at least about 260, at least about 270, at least about 280, at least about 290, at least about 300, at least about 310, at least about 320, at least about 330, at least about 340, at least about 350, at least about 360, at least about 370, at least about 380, at least about 390, at least about 400, at least about 410, at least about 420, at least about 430, at least about 440, at least about 450, at least about 460, at least about 470, at least about 480, at least about 490, at least about 500, at least about 510, at least about 520, at least about 530, at least about 540, at least about 550, at least about 560, at least about 570, at least about 580, at least about 590, at least about 600, at least about 610, at least about 620, at least about 630, 00, at least about 310, at least about 320, at least about 330, at least about 340, at least about 350, at least about 360, at least about 370, at least about 380, at least about 390, at least about 400, at least about 410, at least about 420, at least about 430, at least about 440, at least about 450, at least about 460, at least about 470, at least about 480, at least about 490, at least about 500, at least about 510, at least about 520, at least about 530, at least about 540, at least about 550, at least about 560, at least about 570, at least about 570, at least about 580, at least about 590, or at least about 600 consecutive nucleotides.
[0143] In some embodiments, the 5' homologous recombination site is at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 110, at least about 120, at least about 130, at least about 140, at least about 150, at least about 160, at least about 170, at least about 180, at least about 190, at least about 200, at least about 210, at least about 220, at least about 230, at least about 240, at least about 250, at least about 260, at least about 270, at least about 280, at least about 290, at least about 300, at least about 310, at least about 320, at least about 330, at least about 340, at least about 350, at least about 360, at least about 370, at least about 380, at least about 390, at least about 400, at least about 410, at least about 420, at least about 430, at least about 440, at least about 450, at least about 460, at least about 470, at least about 480, at least about 490, at least about 500, at least about 510, at least about 520, at least about 530, at least about 540, at least about 550, at least about 560, at least about 570, at least about 580, at least about 590, at least about 600, at least about 610, at least about 620, at least about 630, Also about 220, at least about 230, at least about 240, at least about 250, at least about 260, at least about 270, at least about 280, at least about 290, at least about 300, at least about 310, at least about 320, at least about 330, at least about 340, at least about 350, at least about 360, at least about 370, at least about 380, at least about 390, at least about 400, at least about 410, at least about 420, at least about 430, at least about 440, at least about 450, at least about 460, and the 3' homologous recombination site comprises at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 110, at least about 120, at least about 130, at least about 140, at least about 150, at least about 160, at least about 170, at least about 180, at least about 190, at least about 200, at least about 210, at least about 220, at least about 230, at least about 240, at least about 250, at least about 260, at least about 270, at least about 280, at least about 290, at least about 300, at least about 310, at least about 320, at least about 330, at least about 340, at least about 350, at least about 360, at least about 370, at least about 380, at least about 390, or at least about 400 contiguous nucleotides from SEQ ID NO: 19 or SEQ ID NO: 115. 0, at least about 70, at least about 80, at least about 90, at least about 100, at least about 110, at least about 120, at least about 130, at least about 140, at least about 150, at least about 160, at least about 170, at least about 180, at least about 190, at least about 200, at least about 210, at least about 220, at least about 230, at least about 240, at least about 250, at least about 260, at least about 270, at least about 280, at least about 290, at least about 300, at least about 310,At least about 320, at least about 330, at least about 340, at least about 350, at least about 360, at least about 370, at least about 380, at least about 390, at least about 400, at least about 410, at least about 420, at least about 430, at least about 440, at least about 450, at least about 460, at least about 470, at least about 480, at least about 490, at least about 500, at least about 510, at least about 520, at least about 530, at least about 540, at least about 550, at least about 560, at least about 570, at least about 570, at least about 580, at least about 590, or at least about 600 consecutive nucleotides.
[0144] In some embodiments, the GOI is an antibody. In some embodiments, the GOI comprises the heavy chain (HC) of an antibody. In some embodiments, the GOI comprises the light chain (LC) of an antibody. In some embodiments, the GOI comprises the HC and LC of an antibody. In some embodiments, the GOI comprises the antigen-binding portion of an antibody. In some embodiments, the expression plasmid comprises one, two, or more copies of the GOI. In some embodiments, the expression plasmid comprises one, two, or more expression cassettes. In some embodiments, the expression plasmid is bicistronic. In some embodiments, the expression plasmid is multicistronic.
[0145] In some embodiments, the expression plasmid is integrated into the genome of the expressing cell at a copy number of at least 1. In some embodiments, the expression plasmid is integrated into the genome of the expressing cell at a copy number of 1. In other embodiments, the expression plasmid is integrated into the genome of the expressing cell at a copy number of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, or at least 30 copies. In some embodiments, more than 30 copies are present in the genome of the expressing cell.
[0146] In other embodiments, the expression plasmid is integrated into the genome of the expressing cell in a copy number of about 1 to about 3, about 3 to about 6, about 6 to about 9, about 9 to about 12, about 12 to about 15, about 15 to about 18, about 18 to about 21, about 21 to about 24, about 24 to about 27, about 27 to about 30, about 5 to about 10, about 10 to about 15, about 15 to about 20, about 20 to about 25, about 25 to about 30, about 1 to about 10, about 5 to about 15, about 10 to about 20, about 15 to about 25, about 20 to about 30, about 1 to about 15, about 5 to about 20, about 10 to about 25, about 15 to about 30, about 1 to about 20, about 5 to about 25, about 10 to about 30 copies.
[0147] In some embodiments, the methods disclosed herein include determining the expression of a GOI produced by a host cell following targeted integration of a second GOI plasmid (P3; see, e.g., Figure 5A) in a landing pad cell line to generate an expression plasmid (P4; see, e.g., Figure 5A). In some embodiments, the expression level is determined quantitatively. In other embodiments, the expression is determined qualitatively. Expression of a GOI can be determined using any method known in the art, such as cell sorting, FACS, cell surface staining, Western blot, Northern blot, column chromatography, capillary electrophoresis, microfluidics, UV absorbance, cell size, secreted protein levels, transcript levels, immunohistochemistry, or any combination thereof.
[0148] In the context of the present disclosure, the recombinant expression level of the second GOI can correspond to expression from a single expression cassette or from multiple expression cassettes using expression cells generated according to the methods of the present disclosure, hi some embodiments, expression of the GOI can correspond to multiple cassettes containing the GOI inserted at the same site.
[0149] The present disclosure provides landing pad cells and expression cells produced according to the methods disclosed herein. In some embodiments, the recombinant protein expression level of a second GOI (e.g., a second biologic, such as a second antibody) obtained when using the expression cells produced according to the methods of the present disclosure is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 100%, at least about 110%, at least about 120%, at least about 125%, at least about 130%, at least about 135%, at least about 140%, or at least about 145% of the recombinant protein expression level of a first GOI (e.g., a first biologic, such as a first antibody) observed when the parent cells are cultured under the same conditions. , at least about 150%, at least about 155%, at least about 160%, at least about 165%, at least about 170%, at least about 175%, at least about 180%, at least about 185%, at least about 190%, at least about 195%, at least about 200%, at least about 300%, at least about 400%, at least about 500%, at least about 600%, at least about 700%, at least about 800%, at least about 900%, or at least about 1000%.
[0150] The present disclosure provides landing pad cells and expression cells produced according to the methods disclosed herein. In some embodiments, the recombinant protein expression level of a second GOI (e.g., a second biologic, such as a second antibody) obtained using the expression cells produced according to the methods of the present disclosure is greater than or equal to the first GOI observed when the parent cells are cultured under the same conditions. about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 100%, about 110%, about 120%, about 125%, about 130%, about 135%, about 140%, about 145%, about 150%, about 155%, about 160%, about 165%, about 170%, about 175%, about 180%, about 185%, about 190%, about 195%, about 200%, about 300%, about 400%, about 500%, about 600%, about 700%, about 800%, about 900%, about 1000%, or greater than 1000% of the recombinant protein expression level of the GOI (e.g., a first biologic such as a first antibody).
[0151] The present disclosure provides landing pad cells and expression cells produced according to the methods disclosed herein. In some embodiments, the recombinant protein expression level of a second GOI (e.g., a second biologic, such as a second antibody) obtained using the expression cells produced according to the methods of the present disclosure is greater than or equal to the first GOI observed when the parent cells are cultured under the same conditions. about 50% to about 55%, about 55% to about 60%, about 65% to about 70%, about 70% to about 75%, about 75% to about 80%, about 80% to about 85%, about 85% to about 90%, about 90% to about 100%, about 100% to about 110%, about 110% to about 120%, about 120% to about 125%, about 125% to about 130%, about 130% to about 135%, about 135% to about 140%, about 140% to about 145%, about 145% to about 150%, about 150% to about 155%, about 155% to about 160%, about 160% to about 165%, about 165% to about 170%, about 170% to about 1175%, about 175% to about 180%, about 180% to about 185%, about 185% to about 190%, about 190% to about 195%, about 195% to about 200%, about 200% to about 300%, about 300% to about 400%, about 400% to about 500%, about 500% to about 600%, about 600% to about 700%, about 700% to about 800%, about 800% to about 900%, about 900% to about 1000%, or greater than 1000%.
[0152] In some embodiments of the present disclosure, the cells disclosed herein can be established as cell lines, i.e., cell cultures of cells that are derived from a single cell and therefore have a uniform genetic makeup, which can grow indefinitely in the laboratory under certain conditions, and in the case of expression cell lines, have one or more genes of interest stably integrated into the genome of the cells.
[0153] In some embodiments, the targeted integration site is located within the "Chr3 TI contig" or chromosome 3 targeted integration locus, which is a polynucleotide from chromosome 3 of Cricetulus griseus (Chinese hamster), comprising: (i) a sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, or at least about 96.6% identical to SEQ ID NO:23 (gi|1497155598|ref|NW_020822499.1 26 Mbase contig) at the 5' end of the polynucleotide; and (ii) a sequence at the 3' end of the polynucleotide that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, or at least about 96.6% identical to SEQ ID NO:23 (gi|1497155598|ref|NW_020822499.1 26 Mbase contig). It is defined as a polynucleotide containing a sequence that is at least 96.6% identical to the 3'-terminal 5 kb sequence of a 26 Mbase contig, and the length of this polynucleotide is between 25 Mbases (megabases) and 26.5 Mbases (megabases).
[0154] In some embodiments, the targeted integration site is located within the "Chr5 TI contig" or chromosome 5 targeted integration locus, which is defined as a polynucleotide from chromosome 5 of the Mongolian crucian carp (Chinese hamster), comprising: (i) a sequence at the 5' end of the polynucleotide that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, or at least about 96.6% identical to SEQ ID NO: 119 (the 5' terminal 5kb sequence of the NW_020822577.1 18Mbase contig), and (ii) a sequence at the 3' end of the polynucleotide that is at least 96.6% identical to SEQ ID NO: 120 (the 3' terminal 5kb sequence of the NW_020822577.1 18Mbase contig), wherein the length of the polynucleotide is between 17Mbase (megabase) and 19Mbase (megabase).
[0155] Some of the hotspots in SEQ ID NO:22 (Refseq NW_020822499.1; available at ncbi.nlm.nih.gov / nuccore / NW_020822499.1?report=genbank and ncbi.nlm.nih.gov / nuccore / NW_020822499.1?report=fasta) and SEQ ID NO:118 (Refseq NW_020822577.1; available at ncbi.nlm.nih.gov / nuccore / NW_020822577.1?report=genbank and ncbi.nlm.nih.gov / nuccore / NW_020822577.1?report=fasta) are described in the CHO RNA-seq dataset [Singh et al. Biotechnol J. 2018 Oct;13(10):e1800070, and Lin et al. PLoS Comput Biol 16(12):e1008498, which are incorporated herein by reference in their entireties. There are actively transcribed genes experimentally confirmed by the applicant. For the first hotspot of NW_020822499.1, the closest gene on the 5' side of the deletion where the landing pad is located is Prkg1, 269 kb upstream, and the closest gene on the 3' side of the deletion where the landing pad is located is Mbl2, 43 kb downstream. Between Prkg1 and Mbl2, no other actively transcribed transcripts were identified by the applicant or in the CHO RNA-seq dataset. For the second hotspot of NW_020822577.1, the closest gene on the 5' side of the deletion where the landing pad is located is Ackr1, which is 209 kb upstream, and the closest gene on the 3' side of the deletion where the landing pad is located is Crp, which is 170 kb downstream.No other active transcripts between Ackr1 and Crp were identified by the applicant or in the CHO RNA-seq dataset.In some embodiments, the targeted integration site is located in SEQ ID NO:22 or SEQ ID NO:118 at a position that does not affect the expression of one or more actively transcribed genes.In some embodiments, the actively transcribed gene or genes are located within a hotspot of SEQ ID NO: 22 or a hotspot of SEQ ID NO: 118. In some embodiments, the actively transcribed gene or genes are located near the 5' end of a hotspot of SEQ ID NO: 22 or a hotspot of SEQ ID NO: 118. In some embodiments, the actively transcribed gene or genes are located near a hotspot of SEQ ID NO: 22 or a hotspot of SEQ ID NO: 118. In some embodiments of the present disclosure, an actively transcribed gene is considered to be near the 5' or 3' end of a hotspot disclosed herein when the actively transcribed gene is located at a distance of about 25 kb, 30 kb, 35 kb, 40 kb, 45 kb, 50 kb, 75 kb, 100 kb, 125 kb, 150 kb, 175 kb, 200 kb, 225 kb, 250 kb, 275 kb, 300 kb, 325 kb, 350 kb, 375 kb, 400 kb, 425 kb, 450 kb, 475 kb, or 500 kb from the 5' or 3' end of a hotspot disclosed herein.
[0156] In some embodiments, the targeted integration site is located at a specific location in SEQ ID NO: 22 or the sequence set forth in the Chr3 TI contig. In some embodiments, the targeted integration site is located at a specific location in SEQ ID NO: 22 or the sequence set forth in the Chr3 TI contig within nucleotide positions 1 (5' start position) and 26,290,500 (3' end position) of SEQ ID NO: 22.
[0157] In some embodiments, the targeted integration site is located at a specific location in SEQ ID NO: 118 or the sequence set forth in the Chr5 TI contig. In some embodiments, the targeted integration site is located at a specific location in SEQ ID NO: 118 or the sequence set forth in the Chr5 TI contig within nucleotide positions 1 (5' start position) and 18,231,092 (3' end position) of SEQ ID NO: 118.
[0158] As used herein, the term "specific location" refers to a specific position (e.g., a single base) in a sequence defined in, for example, SEQ ID NO: 22 or the Chr3 TI contig, or in SEQ ID NO: 118 or the Chr5 TI contig, where integration will occur. For example, a specific location at position 100 would mean that integration will occur by insertion between nucleotides 100 and 101. In some embodiments, the term "specific location" refers to a specific range of nucleotides between two positions that will be excised when integration occurs. For example, a specific location between positions 100 and 200 would mean that the original sequence including nucleotides 101 to 199 is deleted and replaced by the integrated sequence.
[0159] In some embodiments, the targeted integration site is located between positions in the sequence set forth in SEQ ID NO:22 (corresponding to the exemplary targeted integration site of SEQ ID NO:21), or the Chr3 TI contig, or SEQ ID NO:118 (corresponding to the exemplary targeted integration site of SEQ ID NO:117), or the Chr5 TI contig. In some embodiments, the enclosed sequence, i.e., the sequence corresponding to the targeted integration site, is replaced by an expression plasmid (e.g., a parent plasmid or a landing pad plasmid described herein). In some embodiments, the expression plasmid (e.g., a parent plasmid or a landing pad plasmid described herein) is integrated on the minus strand corresponding to the sequence set forth in SEQ ID NO:22 or the Chr3 TI contig, or SEQ ID NO:118 or the Chr5 TI contig. Thus, in some embodiments, the underlined sequences upstream (5') and downstream (3') from the boxed sequences correspond to the 3' and 5' junctions, respectively, of the integrated expression plasmid that are integrated on the minus strand corresponding to the sequences set forth in SEQ ID NO: 22 or the Chr3 TI contig or SEQ ID NO: 118 or the Chr5 TI contig.
[0160] In some embodiments, the targeted integration site is between positions 1 and 1,000,000, between positions 1,000,000 and 2,000,000, between positions 2,000,000 and 3,000,000, between positions 3,000,000 and 4,000,000, between positions 4,000,000 and 5,000,000, between positions 5,000,000 and 6,000,000, or between positions 6,000,000 and 7,000,000 of the sequence set forth in SEQ ID NO:22 or the Chr3 TI contig or SEQ ID NO:118 or the Chr5 TI contig. , between 7,000,000th and 8,000,000th, between 8,000,000th and 9,000,000th, between 9,000,000th and 10,000,000th, between 10,000,000th and 11,000,000th, between 11,000,000th and 12,000,000th, between 12,000,000th and 13,000,000th, between 13,000,000th and 14,000,000th Between 14,000,000th and 15,000,000th, Between 15,000,000th and 16,000,000th, Between 16,000,000th and 17,000,000th, Between 17,000,000th and 18,000,000th, Between 18,000,000th and 19,000,000th, Between 19,000,000th and 20,000,000th, Between 20,000,000th and 21,000,000th Between 1,000,000th, between 21,000,000th and 22,000,000th, between 22,000,000th and 23,000,000th, between 23,000,000th and 24,000,000th, between 24,000,000th and 25,000,000th, between 25,000,000th and 26,000,000th, or between 26,000,000th and 26,294,056th.
[0161] In some embodiments, the targeted integration site is the TI contig of SEQ ID NO: 22 or Chr3 or SEQ ID NO: 118 or Chr5 Positions 1 to 100,000, 100,000 to 200,000, 200,000 to 300,000, 300,000 to 400,000, 400,000 to 500,000, 500,000 to 600,000, 600,000 to 700,000, 700,000 to 800,000, 800,000 to 900,000, 900,000 to 1,000,000, 1,000,000 to 1,100,000, 1,100,000 to 1,200,000, 1,200,000th to 1,300,000th, 1,300,000th to 1,400,000th, 1,400,000th to 1,5 00,000, 1,500,000th ~ 1,600,000th, 1,600,000th ~ 1,700,000th, 1,700,00 0th place ~ 1,800,000th place, 1,800,000th place ~ 1,900,000th place, 1,900,000th place ~ 2,000,000th place, 2,000,000th ~ 2,100,000th, 2,100,000th ~ 2,200,000th, 2,200,000th ~ 2,30 0,000th place, 2,300,000th place ~ 2,400,000th place, 2,400,000th place ~ 2,500,000, 2,500,00th place 0th place ~ 2,600,000th place, 2,600,000th place ~ 2,700,000th place, 2,700,000th place ~ 2,800,000th place, 2 , 800,000th place ~ 2,900,000th place, 2,900,000th place ~ 3,000,000th place, 3,000,000th place ~ 3,10 0,000th place, 3,100,000th place ~ 3,200,000th place, 3,200,000th place ~ 3,300,000th place, 3,300,00th place 0th place ~ 3,400,000th place, 3,400,000th place ~ 3,500,000th, 3,500,000th place ~ 3,600,000th place, 3 , 600,000th to 3,700,000th, 3,700,000th to 3,800,000th, 3,800,000th to 3,900th ,000th place, 3,900,000th place ~ 4,000,000th place, 4,000,000th place ~ 4,100,000th place, 4,100,000th place 4,200,000 to 4,200,000, 4,200,000 to 4,300,000, 4,300,000 to 4,400,000, 4,400,000 bits ~ 4,500,000, 4,500,000 bits ~ 4,600,000 bits, 4,600,000 bits ~ 4,700,000 bits, 4,700,000 bits ~ 4,800,000 bits, 4,800,000 bits ~ 4,900,000 bits, 4,900,000 bits ~ 5,000,0 00, 5,000,000 to 5,100,000, 5,100,000 to 5,200,000, 5,200,000 to 5,300,000, 5,300,000 to 5,400,000, 5,400,000 to 5,500,000, 5,500,000 to 5 600,000 bits, 5,600,000 bits ~ 5,700,000 bits, 5,700,000 bits ~ 5,800,000 bits, 5,800,000 bits ~ 5,900,000 bits, 5,900,000 bits ~ 6,000,000 bits, 6,000,000 bits ~ 6,100,000 bits, 6,100 6,000 bits ~ 6,200,000 bits, 6,200,000 bits ~ 6,300,000 bits, 6,300,000 bits ~ 6,400,000 bits, 6,400,000 bits ~ 6,500,000 bits, 6,500,000 bits ~ 6,600,000 bits, 6,600,000 bits ~ 6,700,000 bits 6,700,000 to 6,800,000, 6,800,000 to 6,900,000, 6,900,000 to 7,000,000, 7,000,000 to 7,100,000, 7,100,000 to 7,200,000, 7,200,000 to 7,3 00,000 bits, 7,300,000 bits ~ 7,400,000 bits, 7,400,000 bits ~ 7,500,000 bits, 7,500,000 bits ~ 7,600,000 bits, 7,600,000 bits ~ 7,700,000 bits, 7,700,000 bits ~ 7,800,000 bits, 7,800,000 bits 0 to 7,900,000, 7,900,000 to 8,000,000, 8,000,000 to 8,100,000, 8,100,000 to 8,200,000, 8,200,000 to 8,300,000, 8,300,000 to 8,400,000, 8 400,000 bits ~ 8,500,000, 8,500,000 bits ~ 8,600,000 bits, 8,600,000 bits ~ 8,700,000 bits, 8,700,000 bits ~ 8,800,000 bits, 8,800,000 bits ~ 8,900,000 bits, 8,900,000 bits ~ 9,000000, 9,000,000 to 9,100,000, 9,100,000 to 9,200,000, 9,200,000 to 9,300,000, 9,300,000 to 9,400,000, 9,400,000 to 9,500,000, 9,500,000 ~9,600,000 bits, 9,600,000 bits ~ 9,700,000 bits, 9,700,000 bits ~ 9,800,000 bits, 9,800,000 bits ~ 9,900,000 bits, 9,900,000 bits ~ 10,000,000 bits, 10,000,000 bits ~ 10,100,000 bits 10,100,000 bits ~ 10,200,000 bits, 10,200,000 bits ~ 10,300,000 bits, 10,300,000 bits ~ 10,400,000 bits, 10,400,000 bits ~ 10,500,000 bits, 10,500,000 bits ~ 10,600,000 bits, 10,600,0 ... 0,000 bits ~ 10,700,000 bits, 10,700,000 bits ~ 10,800,000 bits, 10,800,000 bits ~ 10,900,000 bits, 10,900,000 bits ~ 11,000,000 bits, 11,000,000 bits ~ 11,100,000 bits, 11,100,000 Bit ~ 11,200,000, 11,200,000 bits ~ 11,300,000 bits, 11,300,000 bits ~ 11,400,000 bits, 11,400,000 bits ~ 11,500,000 bits, 11,500,000 bits ~ 11,600,000 bits, 11,600,000 bits ~ 11, 700,000 bits, 11,700,000 bits ~ 11,800,000 bits, 11,800,000 bits ~ 11,900,000 bits, 11,900,000 bits ~ 12,000,000 bits, 12,000,000 bits ~ 12,100,000 bits, 12,100,000 bits ~ 12,200,0 00, 12,200,000 to 12,300,000, 12,300,000 to 12,400,000, 12,400,000 to 12,500,000, 12,500,000 to 12,600,000, 12,600,000 to 12,700,000, 12 700,000 bits ~ 12,800,000 bits, 12,800,000 bits ~ 12,900,000 bits, 12,900,000 bits ~ 13,000,000 bits, 13,000,000 bits ~ 13,100,000 bits, 13,100,000 bits ~ 13,200,000 bits, 13,200,000 to 13,300,000, 13,300,000 to 13,400,000, 13,400,000 to 13,500,000, 13,500,000 to 13,600,000, 13,600,000 to 13,700,000, 13,700,000 to 1 3,800,000 bits, 13,800,000~13,900,000 bits, 13,900,000~14,000,000 bits, 14,000,000~14,100,000 bits, 14,100,000~14,200,000 bits, 14,200,000~14,300 14,300,000 bits ~ 14,400,000 bits, 14,400,000 bits ~ 14,500,000 bits, 14,500,000 bits ~ 14,600,000 bits, 14,600,000 bits ~ 14,700,000 bits, 14,700,000 bits ~ 14,800,000 bits 14,800,000 to 14,900,000, 14,900,000 to 15,000,000, 15,000,000 to 15,100,000, 15,100,000 to 15,200,000, 15,200,000 to 15,300,000, 15,300,000 0,000 bits ~ 15,400,000 bits, 15,400,000 bits ~ 15,500,000 bits, 15,500,000 bits ~ 15,600,000 bits, 15,600,000 bits ~ 15,700,000 bits, 15,700,000 bits ~ 15,800,000 bits, 15,800,000 bits ~15,900,000 bits, 15,900,000 bits ~ 16,000,000 bits, 16,000,000 bits ~ 16,100,000 bits, 16,100,000 bits ~ 16,200,000 bits, 16,200,000 bits ~ 16,300,000 bits, 16,300,000 bits ~ 16,4 00,000 bits, 16,400,000 bits ~ 16,500,000 bits, 16,500,000 bits ~ 16,600,000 bits, 16,600,000 bits ~ 16,700,000 bits, 16,700,000 bits ~ 16,800,000 bits, 16,800,000 bits ~ 16,900,000 bits 16,900,000 bits ~ 17,000,000 bits, 17,000,000 bits ~ 17,100,000 bits, 17,100,000 bits ~ 17,200,000 bits, 17,200,000 bits ~ 17,300,000 bits, 17,300,000 bits ~ 17,400,000 bits, 17,400,000 bits ~ 17,500,000, 17,500,000 bits ~ 17,600,000 bits, 17,600,000 bits ~ 17,700,000 bits, 17,700,000 bits ~ 17,800,000 bits, 17,800,000 bits ~ 17,900,000 bits, 17,900,000 bits 00 to 18,000,000, 18,000,000 to 18,100,000, 18,100,000 to 18,200,000, 18,200,000 to 18,300,000, 18,300,000 to 18,400,000, 18,400,000 18,500,000, 18,500,000 to 18,600,000, 18,600,000 to 18,700,000, 18,700,000 to 18,800,000, 18,800,000 to 18,900,000, 18,900,000 to 19,000 0,000, 19,000,000~19,100,000, 19,100,000~19,200,000, 19,200,000~19,300,000, 19,300,000~19,400,000, 19,400,000~19,500,000 19,500,000 to 19,600,000, 19,600,000 to 19,700,000, 19,700,000 to 19,800,000, 19,800,000 to 19,900,000, 19,900,000 to 20,000,000, 20, 000,000 bits ~ 20,100,000 bits, 20,100,000 bits ~ 20,200,000 bits, 20,200,000 bits ~ 20,300,000 bits, 20,300,000 bits ~ 20,400,000 bits, 20,400,000 bits ~ 20,500,000 bits, 20,500,000 bits 0 to 20,600,000, 20,600,000 to 20,700,000, 20,700,000 to 20,800,000, 20,800,000 to 20,900,000, 20,900,000 to 21,000,000, 21,000,000 to 2 1,100,000 bits, 21,100,000 bits ~ 21,200,000 bits, 21,200,000 bits ~ 21,300,000 bits, 21,300,000 bits ~ 21,400,000 bits, 21,400,000 bits ~ 21,500,000 bits, 21,500,000 bits ~ 21,600 bits000, 21,600,000 to 21,700,000, 21,700,000 to 21,800,000, 21,800,000 to 21,900,000, 21,900,000 to 22,000,000, 22,000,000 to 22,100,000, 22, 100,000 bits ~ 22,200,000 bits, 22,200,000 bits ~ 22,300,000 bits, 22,300,000 bits ~ 22,400,000 bits, 22,400,000 bits ~ 22,500,000 bits, 22,500,000 bits ~ 22,600,000 bits, 22,600, 000~22,700,000, 22,700,000~22,800,000, 22,800,000~22,900,000, 22,900,000~23,000,000, 23,000,000~23,100,000, 23,100,000~ 23,200,000 bits, 23,200,000 bits ~ 23,300,000 bits, 23,300,000 bits ~ 23,400,000 bits, 23,400,000 bits ~ 23,500,000 bits, 23,500,000 bits ~ 23,600,000 bits, 23,600,000 bits ~ 23,700 bits 0,000 bits, 23,700,000 bits ~ 23,800,000 bits, 23,800,000 bits ~ 23,900,000 bits, 23,900,000 bits ~ 24,000,000 bits, 24,000,000 bits ~ 24,100,000 bits, 24,100,000 bits ~ 24,200,000 bits 24,200,000 bits ~ 24,300,000 bits, 24,300,000 bits ~ 24,400,000 bits, 24,400,000 bits ~ 24,500,000 bits, 24,500,000 bits ~ 24,600,000 bits, 24,600,000 bits ~ 24,700,000 bits, 24, 700,000 bits ~ 24,800,000 bits, 24,800,000 bits ~ 24,900,000 bits, 24,900,000 bits ~ 25,000,000 bits, 25,000,000 bits ~ 25,100,000 bits, 25,100,000 bits ~ 25,200,000 bits, 25,200,0 00 to 25,300,000, 25,300,000 to 25,400,000, 25,400,000 to 25,500,000, 25,500,000 to 25,600,000, 25,600,000 to 25,700,000, 25,700,000 to 25 800,000 bits, 25,800,000 bits ~ 25,900,000 bits, 25,900,000 bits ~ 26,000,000 bits, 26,000,000 bits ~ 26,100,000 bits, 26,100,000 bits ~ 26,200,000 bits, 26,200,000 bits ~ 26,294 bitsIt is ranked 056th.
[0162] In some embodiments, the targeted integration site is the TI contig SEQ ID NO: 22 or Chr3 or SEQ ID NO: 118 or Chr5 shown above. Positions 1 to 1,000, 1,000 to 2,000, 2,000 to 3,000, 3,000 to 4,000, 4,000 to 5,000, 5,000 to 6,000, 6,000 to 7,000, 7,000 to 8,000, 8,000 to 9,000, 9,000 to 10,000, 10,000 to 11,000, 11,000 to 12,000, 12, 000th to 13,000th, 13,000th to 14,000th, 14,000th to 15,000th, 15,000th to 16,00th 0th place, 16,000th to 17,000th, 17,000th to 18,000th, 18,000th to 19,000th, 19,000th ~20,000, 20,000~21,000, 21,000~22,000, 22,000~23,000, 23,000~24,000, 24,000~25,000, 25,000~26,000, 26,000~27, 000th place, 27,000th ~ 28,000th place, 28,000th ~ 29,000th place, 29,000th ~ 30,000th place, 30,00th place 0th to 31,000th, 31,000th to 32,000th, 32,000th to 33,000th, 33,000th to 34,000th , 34,000th to 35,000th, 35,000th to 36,000th, 36,000th to 37,000th, 37,000th to 3 8,000th, 38,000~39,000, 39,000~40,000, 40,000~41,000, 41, 000th to 42,000th, 42,000th to 43,000th, 43,000th to 44,000th, 44,000th to 45,00th 0th place, 45,000th ~ 46,000th place, 46,000th ~ 47,000th place, 47,000th ~ 48,000th place, 48,000th place ~49,000th, 49,000th ~ 50,000th, 50,000th ~ 51,000th, 51,000th ~ 52,000th, 52,000th ~ 53,000th, 53,000th ~ 54,000th, 54,000th ~ 55,000th, 55,000th ~ 56th,000, 56,000~57,000, 57,000~58,000, 58,000~59,000, 59,000~60,000, 60,000~61,000, 61,000~62,000, 62,000~63,000, 63,000~64,000, 64,000~65,000, 65,000~66,000, 66,000~67,000 67,000-68,000, 68,000-69,000, 69,000-70,000, 70,000-71,000, 71,000-72,000, 72,000-73,000, 73,000-74,000, 74,000-75,000, 75,000-76,000, 76,000-77,000, 77,000-78,000, 7 8,000 to 79,000, 79,000 to 80,000, 80,000 to 81,000, 81,000 to 82,000, 82,000 to 83,000, 83,000 to 84,000, 84,000 to 85,000, 85,000 to 86,000, 86,000 to 87,000, 87,000 to 88,000, 88,000 to 89,000, 89,000 to 89,000 00~90,000, 90,000~91,000, 91,000~92,000, 92,000~93,000, 93,000~94,000, 94,000~95,00 0, 95,000~96,000, 96,000~97,000, 97,000~98,000, 98,000~99,000, 99,000~100,000. ,
[0163] In some embodiments, the targeted integration site is any of the partial ranges of the sequences defined by SEQ ID NO: 22 or the Chr3 TI contig or SEQ ID NO: 118 or the Chr5 TI contig shown above, from positions 1 to 1,000 and from positions 99,000 to 100,000, at positions 1 to 10, 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, 90 to 100, 100 to 110, 110 to 120, 120 to 130, 130 to 140, 140 to 150, 150 to 160, 160 to 170, 170 to 180, 180 to 190, 190 to 200, 200 to 210, 210 to 220, 220 to 230, 230 to 240, 240 to 250, 250 to 260, 260 to 270, 270 to ..... 880 to 890890th to 900th place, 900th to 910th place, 910th to 920th place, 920th to 930th place, 930th to 940th place, 940th to 950th place, 950th to 960th place, 960th to 970th place, 970th to 980th place, 980th to 990th place, 990th to 1000th place, 1st to 100th place, 100th to 200th place, 200th to 300th place, 300th to 400th place, 400th to 500th place, 500th to 600th place, 600th to 700th place, 700th to 800th place, 800th to 900th place, or 900th to 1000th place.
[0164] In some embodiments, the targeted integration site is located at about position 1 to about 10, about position 10 to about 20, about position 20 to about 30, about position 30 to about 40, about position 40 to about 50, about position 50 to about 60, about position 60 to about 70, about position 70 to about 80, about position 80 to about 90, about position 90 to about 100, about position 100 to about 110, about position 110 to about 120, about position 120 to about 130, about position 130 to about 140, about position 140 to about 150, about position 150 to about 160, about position 160, upstream or downstream from the sequence set forth in SEQ ID NO:21 or SEQ ID NO:117. from about 170th place, from about 170th place to about 180th place, from about 180th place to about 190th place, from about 190th place to about 200th place, from about 200th place to about 210th place, from about 210th place to about 220th place, from about 220th place to about 230th place, from about 230th place to about 240th place, from about 240th place to about 250th place, from about 250th place to about 260th place, from about 260th place to about 270th place, from about 270th place to about 280th place, from about 280th place to about 290th place, from about 290th place to about 300th place, from about 300th place to about 310th place, from about 310th place to about 320th place, from about 320th place to about 330th place, from about 330th place to about 340th place, from about 340th place to about 350th place, from about 350th place to about 360th place, approximately 360th to approximately 370th place, approximately 370th to approximately 380th place, approximately 380th to approximately 390th place, approximately 390th to approximately 400th place, approximately 400th to approximately 410th place, approximately 410th to approximately 420th place, approximately 420th to approximately 430th place, approximately 430th to approximately 440th place, approximately 440th to approximately 450th place, approximately 450th to approximately 460th place, approximately 460th to approximately 470th place, approximately 470th to approximately 480th place, approximately 480th to approximately 490th place, approximately 490th to approximately 500th place, approximately 500th to approximately 510th place, approximately 510th to approximately 520th place, approximately 520th to approximately 530th place, approximately 530th to approximately 540th place, approximately 540th to approximately 550th place 0th place, about 550th to about 560th place, about 560th to about 570th place, about 570th to about 580th place, about 580th to about 590th place, about 590th to about 600th place, about 600th to about 610th place, about 610th to about 620th place, about 620th to about 630th place, about 630th to about 640th place, about 640th to about 650th place, about 650th to about 660th place, about 660th to about 670th place, about 670th to about 680th place, about 680th to about 690th place, about 690th to about 700th place, about 700th to about 710th place, about 710th to about 720th place, about 720th to about 730th place, about 730th to about 740th place,Approximately 740 to 750, approximately 750 to 760, approximately 760 to 770, approximately 770 to 780, approximately 780 to 790, approximately 790 to 800, approximately 800 to 810, approximately 810 to 820, approximately 820 to 830, approximately 830 to 840, approximately 840 to 850, approximately 850 to 860, approximately 860 to 870, approximately 870 to 880, approximately 880 to 890, approximately 890 to 900, approximately 900 to 910, approximately 910 to 920, approximately 920 to 930 , approximately 930th to approximately 940th, approximately 940th to approximately 950th, approximately 950th to approximately 960th, approximately 960th to approximately 970th, approximately 970th to approximately 980th, approximately 980th to approximately 990th, approximately 990th to approximately 1000th, approximately 1,000th to approximately 2,000th, approximately 2,000th to approximately 3,000th, approximately 3,000th to approximately 4,000th, approximately 4,000th to approximately 5,000th, approximately 5,000th to approximately 6,000th, approximately 6,000th to approximately 7,000th, approximately 7,000th to approximately 8,000th, approximately 8,000th to approximately 9,000th, approximately 9,000th to approximately 10,000th, Approximately 10,000 to 11,000, approximately 11,000 to 12,000, approximately 12,000 to 13,000, approximately 13,000 to 14,000, approximately 14,000 to 15,000, approximately 15,000 to 16,000, approximately 16,000 to 17,000, approximately 17,000 to 18,000, approximately 18,000 to 19,000, approximately 19,000 to 20,000, approximately 20,000 to 21,000, approximately 21,000 to 22,000, approximately 22,000 to 23,000, Approximately 23,000 to 24,000, approximately 24,000 to 25,000, approximately 25,000 to 26,000, approximately 26,000 to 27,000, approximately 27,000 to 28,000, approximately 28,000 to 29,000, approximately 29,000 to 30,000, approximately 30,000 to 31,000, approximately 31,000 to 32,000, approximately 32,000 to 33,000, approximately 33,000 to 34,000, approximately 34,000 to 35,000, approximately 35,000 to 36,000,Approximately 36,000 to approximately 37,000, approximately 37,000 to approximately 38,000, approximately 38,000 to approximately 39,000, approximately 39,000 to approximately 40,000, approximately 40,000 to approximately 41,000, approximately 41,000 to approximately 42,000, approximately 42,000 to approximately 43,000, approximately 43,000 to approximately 44,000, approximately 44,000 to approximately 45,000, approximately 45,000 to approximately 46,000, approximately 46,000 to approximately 47,000, approximately 47,000 to approximately 48,000, approximately 48,000 to approximately 49,000, Approximately 49,000 to approximately 50,000, approximately 50,000 to approximately 51,000, approximately 51,000 to approximately 52,000, approximately 52,000 to approximately 53,000, approximately 53,000 to approximately 54,000, approximately 54,000 to approximately 55,000, approximately 55,000 to approximately 56,000, approximately 56,000 to approximately 57,000, approximately 57,000 to approximately 58,000, approximately 58,000 to approximately 59,000, approximately 59,000 to approximately 60,000, approximately 60,000 to approximately 61,000, approximately 61,000 to approximately 62,000, Approximately 62,000 to approximately 63,000, approximately 63,000 to approximately 64,000, approximately 64,000 to approximately 65,000, approximately 65,000 to approximately 66,000, approximately 66,000 to approximately 67,000, approximately 67,000 to approximately 68,000, approximately 68,000 to approximately 69,000, approximately 69,000 to approximately 70,000, approximately 70,000 to approximately 71,000, approximately 71,000 to approximately 72,000, approximately 72,000 to approximately 73,000, approximately 73,000 to approximately 74,000, approximately 74,000 to approximately 75,000, Approximately 75,000 to approximately 76,000, approximately 76,000 to approximately 77,000, approximately 77,000 to approximately 78,000, approximately 78,000 to approximately 79,000, approximately 79,000 to approximately 80,000, approximately 80,000 to approximately 81,000, approximately 81,000 to approximately 82,000, approximately 82,000 to approximately 83,000, approximately 83,000 to approximately 84,000, approximately 84,000 to approximately 85,000, approximately 85,000 to approximately 86,000, approximately 86,000 to approximately 87,000, approximately 87,000 to approximately 88,000,Approximately 88,000 to approximately 89,000, approximately 89,000 to approximately 90,000, approximately 90,000 to approximately 91,000, approximately 91,000 to approximately 92,000, approximately 92,000 to approximately 93,000, approximately 93,000 to approximately 94,000, approximately 94,000 to approximately 95,000, approximately 95,000 to approximately 96,000, approximately 96,000 to approximately 97,000, approximately 97,000 to approximately 98,000, approximately 98,000 to approximately 99,000, approximately 99,000 to approximately 100,000, approximately 100,000 to approximately 200,000 Place, approximately 200,000th to approximately 300,000th, approximately 300,000th to approximately 400,000th, approximately 400,000th to approximately 500,000th, approximately 500,000th to approximately 600,000th, approximately 600,000th to approximately 700,000th, approximately 700,000th to approximately 800,000th, approximately 800,000th to approximately 900,000th, approximately 900,000th to approximately 1,000,000th, approximately 1,000,000th to approximately 1,100,000th, approximately 1,100,000th to approximately 1,200,000th, approximately 1,200,000th to approximately 1,300,000th, approximately 1 ,300,000th to about 1,400,000th, about 1,400,000th to about 1,500,000th, about 1,500,000th to about 1,600,000th, about 1,600,000th to about 1,700,000th, about 1,700,000th to about 1,800,000th, about 1,800,000th to about 1,900,000th, about 1,900,000th to about 2,000,000th, about 2,000,000th to about 2,100,000th, about 2,100,000th to about 2,200,000th, about 2,200,000th to about 2,300,000th, about 2,300,000th to approximately 2,400,000th, approximately 2,400,000th to approximately 2,500,000th, approximately 2,500,000th to approximately 2,600,000th, approximately 2,600,000th to approximately 2,700,000th, approximately 2,700,000th to approximately 2,800,000th, approximately 2,800,000th to approximately 2,900,000th, approximately 2,900,000th to approximately 3,000,000th, approximately 3,000,000th to approximately 3,100,000th, approximately 3,100,000th to approximately 3,200,000th, approximately 3,200,000th to approximately 3,300,000th,Approximately 3,300,000th to approximately 3,400,000th, approximately 3,400,000th to approximately 3,500,000th, approximately 3,500,000th to approximately 3,600,000th, approximately 3,600,000th to approximately 3,700,000th, approximately 3,700,000th to approximately 3,800,000th, approximately 3,800,000th to approximately 3,900,000th, approximately 3,900,000th to approximately 4,000,000th, approximately 4,000,000th to approximately 4,100,000th, approximately 4,100,000th to approximately 4,200,000th, approximately 4,200,000th to approximately 4,300,000th, Approximately 4,300,000th to approximately 4,400,000th, approximately 4,400,000th to approximately 4,500,000th, approximately 4,500,000th to approximately 4,600,000th, approximately 4,600,000th to approximately 4,700,000th, approximately 4,700,000th to approximately 4,800,000th, approximately 4,800,000th to approximately 4,900,000th, approximately 4,900,000th to approximately 5,000,000th, approximately 5,000,000th to approximately 5,100,000th, approximately 5,100,000th to approximately 5,200,000th, approximately 5,200,000th to approximately 5,300,000th, Approximately 5,300,000th to approximately 5,400,000th, approximately 5,400,000th to approximately 5,500,000th, approximately 5,500,000th to approximately 5,600,000th, approximately 5,600,000th to approximately 5,700,000th, approximately 5,700,000th to approximately 5,800,000th, approximately 5,800,000th to approximately 5,900,000th, approximately 5,900,000th to approximately 6,000,000th, approximately 6,000,000th to approximately 6,100,000th, approximately 6,100,000th to approximately 6,200,000th, approximately 6,200,000th to approximately 6,300,000th, Approximately 6,300,000th to approximately 6,400,000th, approximately 6,400,000th to approximately 6,500,000th, approximately 6,500,000th to approximately 6,600,000th, approximately 6,600,000th to approximately 6,700,000th, approximately 6,700,000th to approximately 6,800,000th, approximately 6,800,000th to approximately 6,900,000th, approximately 6,900,000th to approximately 7,000,000th, approximately 7,000,000th to approximately 7,100,000th, approximately 7,100,000th to approximately 7,200,000th, approximately 7,200,000th to approximately 7,300,000th,Approximately 7,300,000th to approximately 7,400,000th, approximately 7,400,000th to approximately 7,500,000th, approximately 7,500,000th to approximately 7,600,000th, approximately 7,600,000th to approximately 7,700,000th, approximately 7,700,000th to approximately 7,800,000th, approximately 7,800,000th to approximately 7,900,000th, approximately 7, ,900,000th to about 8,000,000th, about 8,000,000th to about 8,100,000th, about 8,100,000th to about 8,200,000th, about 8,200,000th to about 8,300,000th, about 8,300,000th to about 8,400,000th, about 8,400,000th to about 8,500,000th, about 8,500,000th to about 8,600,000th, about 8,600,000th to about 8,700,000th, about 8,700,000th to about 8,800,000th, about 8,800,000th to about 8,900,000th, about 8 ,900,000th to about 9,000,000th, about 9,000,000th to about 9,100,000th, about 9,100,000th to about 9,200,000th, about 9,200,000th to about 9,300,000th, about 9,300,000th to about 9,400,000th, about 9,400,000th to about 9,500,000th, about 9,500,000th to about 9,600,000th, about 9,600,000th to about 9,700,000th, about 9,700,000th to about 9,800,000th, about 9,800,000th to about 9,900,000th, about 9 ,900,000th to about 10,000,000th, about 10,000,000th to about 10,100,000th, about 10,100,000th to about 10,200,000th, about 10,200,000th to about 10,300,000th, about 10,300,000th to about 10,400,000th, about 10,400,000th to about 10,500,000th, about 10,500,000th to about 10,600,000th, about 10,600,000th to about 10,700,000th, about 10,700,000th to about 10,800,000th, about 10,800,000th 0th place to approximately 10,900,000th place, approximately 10,900,000th place to approximately 11,000,000th place, approximately 11,000,000th place to approximately 11,100,000th place, approximately 11,100,000th place to approximately 11,200,000th place, approximately 11,200,000th place to approximately 11,300,000th place, approximately 11,300,000th place to approximately 11,400,000th place, approximately 11,400,000th place to approximately 11,500,000th place, approximately 11,500,000th place to approximately 11,600,000th place, approximately 11,600,000th place to approximately 11,700,000th place, approximately 11,700,000th place to approximately 11,800,000th place, approximately 11,800,000th to approximately 11,900,000th place, approximately 11,900,000th to approximately 12,000,000th place, approximately 12,000,000th to approximately 12,100,000th place, approximately 12,100,000th to approximately 12,200,000th place, approximately 12,200,000th to approximately 12,300,000th place, approximately 12,300,000th to approximately 12,400,000th place, approximately 12,400,000th to approximately 12,500,000th place, approximately 12,500,000th to approximately 12,600,000th place, approximately 12,600,000th to approximately 12,700,000th place 0th place, approximately 12,700,000th to approximately 12,800,000th place, approximately 12,800,000th to approximately 12,900,000th place, approximately 12,900,000th to approximately 13,000,000th place, approximately 13,000,000th to approximately 13,100,000th place, approximately 13,100,000th to approximately 13,200,000th place, approximately 13,200,000th to approximately 13,300,000th place, approximately 13,300,000th to approximately 13,400,000th place, approximately 13,400,000th to approximately 13,500,000th place, approximately 13,500,000th to approximately 13,600,000th place, approximately 13, 600,000th to approximately 13,700,000th, approximately 13,700,000th to approximately 13,800,000th, approximately 13,800,000th to approximately 13,900,000th, approximately 13,900,000th to approximately 14,000,000th, approximately 14,000,000th to approximately 14,100,000th, approximately 14,100,000th to approximately 14,200,000th, approximately 14,200,000th to approximately 14,300,000th, approximately 14,300,000th to approximately 14,400,000th, approximately 14,400,000th to approximately 14,500,000th, approximately 14,500,000th from about 14,600,000th place, from about 14,600,000th place to about 14,700,000th place, from about 14,700,000th place to about 14,800,000th place, from about 14,800,000th place to about 14,900,000th place, from about 14,900,000th place to about 15,000,000th place, from about 15,000,000th place to about 15,100,000th place, from about 15,100,000th place to about 15,200,000th place, from about 15,200,000th place to about 15,300,000th place, from about 15,300,000th place to about 15,400,000th place, from about 15,400,000th place to about 15,500,000th place, approximately 15,500,000th to approximately 15,600,000th place, approximately 15,600,000th to approximately 15,700,000th place, approximately 15,700,000th to approximately 15,800,000th place, approximately 15,800,000th to approximately 15,900,000th place, approximately 15,900,000th to approximately 16,000,000th place, approximately 16,000,000th to approximately 16,100,000th place, approximately 16,100,000th to approximately 16,200,000th place, approximately 16,200,000th to approximately 16,300,000th place, approximately 16,300,000th to approximately 16,400,000th place 0th place, approximately 16,400,000th to approximately 16,500,000th place, approximately 16,500,000th to approximately 16,600,000th place, approximately 16,600,000th to approximately 16,700,000th place, approximately 16,700,000th to approximately 16,800,000th place, approximately 16,800,000th to approximately 16,900,000th place, approximately 16,900,000th to approximately 17,000,000th place, approximately 17,000,000th to approximately 17,100,000th place, approximately 17,100,000th to approximately 17,200,000th place, approximately 17,200,000th to approximately 17,300,000th place, approximately 17, 300,000th to approximately 17,400,000th, approximately 17,400,000th to approximately 17,500,000th, approximately 17,500,000th to approximately 17,600,000th, approximately 17,600,000th to approximately 17,700,000th, approximately 17,700,000th to approximately 17,800,000th, approximately 17,800,000th to approximately 17,900,000th, approximately 17,900,000th to approximately 18,000,000th, approximately 18,000,000th to approximately 18,100,000th, approximately 18,100,000th to approximately 18,200,000th, approximately 18,200,000th from about 18,300,000th place, from about 18,300,000th place to about 18,400,000th place, from about 18,400,000th place to about 18,500,000th place, from about 18,500,000th place to about 18,600,000th place, from about 18,600,000th place to about 18,700,000th place, from about 18,700,000th place to about 18,800,000th place, from about 18,800,000th place to about 18,900,000th place, from about 18,900,000th place to about 19,000,000th place, from about 19,000,000th place to about 19,100,000th place, from about 19,100,000th place to about 19,200,000th place, approximately 19,200,000th to approximately 19,300,000th place, approximately 19,300,000th to approximately 19,400,000th place, approximately 19,400,000th to approximately 19,500,000th place, approximately 19,500,000th to approximately 19,600,000th place, approximately 19,600,000th to approximately 19,700,000th place, approximately 19,700,000th to approximately 19,800,000th place, approximately 19,800,000th to approximately 19,900,000th place, approximately 19,900,000th to approximately 20,000,000th place, approximately 20,000,000th to approximately 20,100,000th place 0th place, approximately 20,100,000th to approximately 20,200,000th place, approximately 20,200,000th to approximately 20,300,000th place, approximately 20,300,000th to approximately 20,400,000th place, approximately 20,400,000th to approximately 20,500,000th place, approximately 20,500,000th to approximately 20,600,000th place, approximately 20,600,000th to approximately 20,700,000th place, approximately 20,700,000th to approximately 20,800,000th place, approximately 20,800,000th to approximately 20,900,000th place, approximately 20,900,000th to approximately 21,000,000th place, approximately 21, 000,000th to approximately 21,100,000th, approximately 21,100,000th to approximately 21,200,000th, approximately 21,200,000th to approximately 21,300,000th, approximately 21,300,000th to approximately 21,400,000th, approximately 21,400,000th to approximately 21,500,000th, approximately 21,500,000th to approximately 21,600,000th, approximately 21,600,000th to approximately 21,700,000th, approximately 21,700,000th to approximately 21,800,000th, approximately 21,800,000th to approximately 21,900,000th, approximately 21,900,000th from about 22,000,000th place, from about 22,000,000th place to about 22,100,000th place, from about 22,100,000th place to about 22,200,000th place, from about 22,200,000th place to about 22,300,000th place, from about 22,300,000th place to about 22,400,000th place, from about 22,400,000th place to about 22,500,000th place, from about 22,500,000th place to about 22,600,000th place, from about 22,600,000th place to about 22,700,000th place, from about 22,700,000th place to about 22,800,000th place, from about 22,800,000th place to about 22,900,000th place, approximately 22,900,000th to approximately 23,000,000th place, approximately 23,000,000th to approximately 23,100,000th place, approximately 23,100,000th to approximately 23,200,000th place, approximately 23,200,000th to approximately 23,300,000th place, approximately 23,300,000th to approximately 23,400,000th place, approximately 23,400,000th to approximately 23,500,000th place, approximately 23,500,000th to approximately 23,600,000th place, approximately 23,600,000th to Approximately 23,700,000th place, approximately 23,700,000th to approximately 23,800,000th place, approximately 23,800,000th to approximately 23,900,000th place, approximately 23,900,000th to approximately 24,000,000th place, approximately 24,000,000th to approximately 24,100,000th place, approximately 24,100,000th to approximately 24,200,000th place, approximately 24,200,000th to approximately 24,300,000th place, approximately 24,300,000th to approximately 24,400,000th place, approximately 24,400,000th place from about 24,500,000th place, from about 24,500,000th place to about 24,600,000th place, from about 24,600,000th place to about 24,700,000th place, from about 24,700,000th place to about 24,800,000th place, from about 24,800,000th place to about 24,900,000th place, from about 24,900,000th place to about 25,000,000th place, from about 25,000,000th place to about 25,100,000th place, from about 25,100,000th place to about 25,200,000th place, It is located in the following places: 000 to approximately 25,300,000th place, approximately 25,300,000 to approximately 25,400,000th place, approximately 25,400,000 to approximately 25,500,000th place, approximately 25,500,000 to approximately 25,600,000th place, approximately 25,600,000 to approximately 25,700,000th place, approximately 25,700,000 to approximately 25,800,000th place, approximately 25,800,000 to approximately 25,900,000th place, and approximately 25,900,000 to approximately 26,000,000th place.
[0165] In some embodiments, the targeted integration site is at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 110, at least about 120, at least about 130, at least about 140, at least about 150, at least about 160, at least about 170, at least about 180, at least about 190, from the sequence set forth in SEQ ID NO:21 or SEQ ID NO:117. , at least about 200, at least about 210, at least about 220, at least about 230, at least about 240, at least about 250, at least about 260, at least about 270, at least about 280, at least about 290, at least about 300, at least about 310, at least about 320, at least about 330, at least about 340, at least about 350, at least about 360, at least about 370, about 380, at least about 390, at least about 400, at least about 410, at least about 420 , at least about 430, at least about 440, at least about 450, at least about 460, at least about 470, at least about 480, at least about 490, at least about 500, at least about 510, at least about 520, at least about 530, at least about 540, at least about 550, at least about 560, at least about 570, at least about 580, at least about 590, at least about 600, at least about 610, at least about 620, at least about 630, at least about 640, at least about 650, at least about 660, at least about 670, at least about 680, at least about 690, at least about 700, at least about 710, at least about 720, at least about 730, at least about 740, at least about 750, at least about 760, at least about 770, at least about 780, at least about 790, at least about 800, at least about 810, at least about 820, at least about 830, at least about 840, at least about 850, at least about 860, at least about 870,at least about 880, at least about 890, at least about 900, at least about 910, at least about 920, at least about 930, at least about 940, at least about 950, at least about 960, at least about 970, at least about 980, at least about 990, at least about 1,000, at least about 2,000, at least about 3,000, at least about 4,000, at least about 5,000, at least about 6,000, at least about 7,000, at least about 8,000, at least about 9,000 , at least about 10,000, at least about 11,000, at least about 12,000, at least about 13,000, at least about 14,000, at least about 15,000, at least about 16,000, at least about 17,000, at least about 18,000, at least about 19,000, at least about 20,000, at least about 21,000, at least about 22,000, at least about 23,000, at least about 24,000, at least about 25,000, at least about 26,000, at least about 27 ,000, at least about 28,000, at least about 29,000, at least about 30,000, at least about 31,000, at least about 32,000, at least about 33,000, at least about 34,000, at least about 35,000, at least about 36,000, at least about 37,000, at least about 38,000, at least about 39,000, at least about 40,000, at least about 41,000, at least about 42,000, at least about 43,000, at least about 44,000, at least at least about 45,000, at least about 46,000, at least about 47,000, at least about 48,000, at least about 49,000, at least about 50,000, at least about 51,000, at least about 52,000, at least about 53,000, at least about 54,000, at least about 55,000, at least about 56,000, at least about 57,000, at least about 58,000, at least about 59,000, at least about 60,000, at least about 61,000, at least about 62,000,At least about 63,000, at least about 64,000, at least about 65,000, at least about 66,000, at least about 67,000, at least about 68,000, at least about 69,000, at least about 70,000, at least about 71,000, at least about 72,000, at least about 73,000, at least about 74,000, at least about 75,000, at least about 76,000, at least about 77,000, at least about 78,000, at least about 79,000, at least about 80,000 00, at least about 81,000, at least about 82,000, at least about 83,000, at least about 84,000, at least about 85,000, at least about 86,000, at least about 87,000, at least about 88,000, at least about 89,000, at least about 90,000, at least about 91,000, at least about 92,000, at least about 93,000, at least about 94,000, at least about 95,000, at least about 96,000, at least about 97,000, at least about 98,000, at least about 99,000, at least about 100,000, at least about 200,000, at least about 300,000, at least about 400,000, at least about 500,000, at least about 600,000, at least about 700,000, at least about 800,000, at least about 900,000, at least about 1,000,000, at least about 1,100,000, at least about 1,200,000, at least about 1,300,000, at least about 1,400,000, at least about 1,500,000, at least about 1,600,000, at least about 1,700,000, at least about 1,800,000, at least about 1,900,000, at least about 2,000,000, at least about 2,100,000, at least about 2,200,000, at least about 2,300,000, at least about 2,400,000, at least about 2,500,000, at least about 2,600,000, at least about 2,700,000, at least about 2,800,000, at least about 2,900,000,At least about 3,000,000, at least about 3,100,000, at least about 3,200,000, at least about 3,300,000, at least about 3,400,000, at least about 3,500,000, at least about 3,600,000, at least about 3,700,000, at least about 3,800,000, at least about 3,900,000, at least about 4,000,000, at least about 4,100,000, at least about 4,200,000, at least about 4,300,000, at least about 4, 400,000, at least about 4,500,000, at least about 4,600,000, at least about 4,700,000, at least about 4,800,000, at least about 4,900,000, at least about 5,000,000, at least about 5,100,000, at least about 5,200,000, at least about 5,300,000, at least about 5,400,000, at least about 5,500,000, at least about 5,600,000, at least about 5,700,000, at least about 5,800,000, at least about 5,900,000, at least about 6,000,000, at least about 6,100,000, at least about 6,200,000, at least about 6,300,000, at least about 6,400,000, at least about 6,500,000, at least about 6,600,000, at least about 6,700,000, at least about 6,800,000, at least about 6,900,000, at least about 7,000,000, at least about 7,100,000, at least about 7,200,000, at least about 7, 300,000, at least about 7,400,000, at least about 7,500,000, at least about 7,600,000, at least about 7,700,000, at least about 7,800,000, at least about 7,900,000, at least about 8,000,000, at least about 8,100,000, at least about 8,200,000, at least about 8,300,000, at least about 8,400,000, at least about 8,500,000, at least about 8,600,000, at least about 8,700,000,at least about 8,800,000, at least about 8,900,000, at least about 9,000,000, at least about 9,100,000, at least about 9,200,000, at least about 9,300,000, at least about 9,400,000, at least about 9,500,000, at least about 9,600,000, at least about 9,700,000, at least about 9,800,000, at least about 9,900,000, at least about 10,000,000, at least about 10,100,000, at least about 10, 200,000, at least about 10,300,000, at least about 10,400,000, at least about 10,500,000, at least about 10,600,000, at least about 10,700,000, at least about 10,800,000, at least about 10,900,000, at least about 11,000,000, at least about 11,100,000, at least about 11,200,000, at least about 11,300,000, at least about 11,400,000, at least about 11,500,000, at least about 11,600,000, at least about 11,700,000, at least about 11,800,000, at least about 11,900,000, at least about 12,000,000, at least about 12,100,000, at least about 12,200,000, at least about 12,300,000, at least about 12,400,000, at least about 12,500,000, at least about 12,600,000, at least about 12,700,000, at least about 12,800,000, at least about 12,900,000, At least about 13,000,000, at least about 13,100,000, at least about 13,200,000, at least about 13,300,000, at least about 13,400,000, at least about 13,500,000, at least about 13,600,000, at least about 13,700,000, at least about 13,800,000, at least about 13,900,000, at least about 14,000,000, at least about 14,100,000, at least about 14,200,000, at least about 14,300,000,At least about 14,400,000, at least about 14,500,000, at least about 14,600,000, at least about 14,700,000, at least about 14,800,000, at least about 14,900,000, at least about 15,000,000, at least about 15,100,000, at least about 15,200,000, a small number At least about 15,300,000, at least about 15,400,000, at least about 15,500,000, at least about 15,600,000, at least about 15,700,000, at least about 15,800,000, at least about 15,900,000, at least about 16,000,000, at least about 16,100,000, at least about 16,200,000, at least about 16,300,000, at least about 16,400,000, at least about 16,500,000, at least about 16,600,000 at least about 16,700,000, at least about 16,800,000, at least about 16,900,000, at least about 17,000,000, at least about 17,100,000, at least about 17,200,000, at least about 17,300,000, at least about 17,400,000, at least about 17,500,000, at least about 17,600,000, at least about 17,700,000, at least about 17,800,000, at least about 17,900,000, at least about 18,000,000 00, at least about 18,100,000, at least about 18,200,000, at least about 18,300,000, at least about 18,400,000, at least about 18,500,000, at least about 18,600,000, at least about 18,700,000, at least about 18,800,000, at least about 18,900,000, at least about 19,000,000, at least about 19,100,000, at least about 19,200,000, at least about 19,300,000, at least about 19,40 0,000, at least about 19,500,000, at least about 19,600,000, at least about 19,700,000, at least about 19,800,000, at least about 19,900,000, at least about 20,000,000, at least about 20,100,000, at least about 20,200,000, at least about 20,300,000, at least about 20,400,000, at least about 20,500,000, at least about 20,600,000, at least about 20,700,000, at least about 20,800,000, at least about 20,900,000, at least about 21,000,000, at least about 21,100,000, at least about 21,200,000, at least about 21,300,000, at least about 21,400,000, at least about 21,500,000, at least about 21,600,000, at least about 21,700,000, at least about 21,800,000, at least about 21,900,000, at least about 22,000,000, at least about 22,100,000 0, at least about 22,200,000, at least about 22,300,000, at least about 22,400,000, at least about 22,500,000, at least about 22,600,000, at least about 22,700,000, at least about 22,800,000, at least about 22,900,000, at least about 23,000,000, at least about 23,100,000, at least about 23,200,000, at least about 23,300,000, at least about 23,400,000, at least Also about 23,500,000, at least about 23,600,000, at least about 23,700,000, at least about 23,800,000, at least about 23,900,000, at least about 24,000,000, at least about 24,100,000, at least about 24,200,000, at least about 24,300,000, at least about 24,400,000, at least about 24,500,000, at least about 24,600,000, at least about 24,700,000, at least about 24,800,000 At least about 24,900,000, at least about 25,000,000, at least about 25,100,000, at least about 25,200,000, at least about 25,300,000, at least about 25,400,000, at least about 25,500,000, at least about 25,600,000, at least about 25,700,000, at least about 25,800,000, at least about 25,900,000, or at least 26,000,000 nucleobases, downstream or upstream.
[0166] In some embodiments, the targeted integration site is selected from the sequence set forth in SEQ ID NO:21 or SEQ ID NO:117, including about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210, about 220, about 230, about 240, about 250, about 260, about 270, about 280, about 290, about 300, about 310, about 320, about 330, about 340, about 350, about 360, about 370, about 380, about 390, about 400, about 410, about 420, about 430, about 440, about 450, about 460, about 470, about 480, about 490, about 500, about 510, about 520, about 530, about 540, about 550, about 560, about 570, about 580, about 590, about 600, about 610, about 620, about 630, about 640, about 650, about 660, about 670, about 680, about 690, about 700, about 710, about 720, about 730, about 740, about 750, about 760, about 770, about 780, about 790, about 800, about 810, About 360 pieces, about 370 pieces, about 380 pieces, about 390 pieces, about 400 pieces, about 410 pieces, about 420 pieces, about 430 pieces, about 440 pieces, about 450 pieces, about 4 60 pieces, about 470 pieces, about 480 pieces, about 490 pieces, about 500 pieces, about 510 pieces, about 520 pieces, about 530 pieces, about 540 pieces, about 550 pieces, about 560 pieces pieces, about 570 pieces, about 580 pieces, about 590 pieces, about 600 pieces, about 610 pieces, about 620 pieces, about 630 pieces, about 640 pieces, about 650 pieces, about 660 pieces, About 670 pieces, about 680 pieces, about 690 pieces, about 700 pieces, about 710 pieces, about 720 pieces, about 730 pieces, about 740 pieces, about 750 pieces, about 760 pieces, about 77 0 pieces, about 780 pieces, about 790 pieces, about 800 pieces, about 810 pieces, about 820 pieces, about 830 pieces, about 840 pieces, about 850 pieces, about 860 pieces, about 870 pieces , about 880 pieces, about 890 pieces, about 900 pieces, about 910 pieces, about 920 pieces, about 930 pieces, about 940 pieces, about 950 pieces, about 960 pieces, about 970 pieces, about 980 pieces, about 990 pieces, about 1,000 pieces, about 2,000 pieces, about 3,000 pieces, about 4,000 pieces, about 5,000 pieces, about 6,000 pieces, about 7, 000 pieces, about 8,000 pieces, about 9,000 pieces, about 10,000 pieces, about 11,000 pieces, about 12,000 pieces, about 13,000 pieces, about 14,00 0 pieces, about 15,000 pieces, about 16,000 pieces, about 17,000 pieces, about 18,000 pieces, about 19,000 pieces, about 20,000 pieces, about 21,0 00 pieces, about 22,000 pieces, about 23,000 pieces, about 24,000 pieces, about 25,000 pieces, about 26,000 pieces, about 27,000 pieces, about 28,0 00 pieces, about 29,000 pieces, about 30,000 pieces, about 31,000 pieces, about 32,000 pieces, about 33,000 pieces, about 34,000 pieces, about 35, 000 pieces, about 36,000 pieces, about 37,000 pieces, about 38,000 pieces, about 39,000 pieces, about 40,000 pieces, about 41,000 pieces, about 42,000, approximately 43,000, approximately 44,000, approximately 45,000, approximately 46,000, approximately 47,000, approximately 48,000, approximately 49,000, approximately 50,000, approximately 51,000, approximately 52,000, approximately 53,000, approximately 54,000, approximately 55,000, approximately 5 6,000, approximately 57,000, approximately 58,000, approximately 59,000, approximately 60,000, approximately 61,000, approximately 62,000, approximately 63,000, approximately 64,000, approximately 65,000, approximately 66,000, approximately 67,000, approximately 68,000, approximately 69,000, approximately 70,000, approximately 71,000, approximately 72,000, approximately 73,000, approximately 74,000, approximately 75,000, approximately 76,000, approximately 77,000, approximately 78,000, approximately 79,000, approximately 80,000, approximately 81,000, approximately 82,000, approximately 83,000 Approximately 84,000, approximately 85,000, approximately 86,000, approximately 87,000, approximately 88,000, approximately 89,000, approximately 90,000, approximately 91,000, approximately 92,000, approximately 93,000, approximately 94,000, approximately 95,000, approximately 96,000, approximately 97,000 Approximately 98,000, approximately 99,000, approximately 100,000, approximately 200,000, approximately 300,000, approximately 400,000, approximately 500,000, approximately 600,000, approximately 700,000, approximately 800,000, approximately 900,000, approximately 1,000,000, approximately 1,1 00,000, approximately 1,200,000, approximately 1,300,000, approximately 1,400,000, approximately 1,500,000, approximately 1,600,000, approximately 1,700,000, approximately 1,800,000, approximately 1,900,000, approximately 2,000,000, approximately 2,100,000 0, approximately 2,200,000, approximately 2,300,000, approximately 2,400,000, approximately 2,500,000, approximately 2,600,000, approximately 2,700,000, approximately 2,800,000, approximately 2,900,000, approximately 3,000,000, approximately 3,100,000, approximately 3 200,000, approximately 3,300,000, approximately 3,400,000, approximately 3,500,000, approximately 3,600,000, approximately 3,700,000, approximately 3,800,000, approximately 3,900,000, approximately 4,000,000, approximately 4,100,000, approximately 4,200000, approximately 4,300,000, approximately 4,400,000, approximately 4,500,000, approximately 4,600,000, approximately 4,700,000, approximately 4,800,000, approximately 4,900,000, approximately 5,000,000, approximately 5,100,000, approximately 5,200,000 Approximately 5,300,000, approximately 5,400,000, approximately 5,500,000, approximately 5,600,000, approximately 5,700,000, approximately 5,800,000, approximately 5,900,000, approximately 6,000,000, approximately 6,100,000, approximately 6,200,000, approximately 6,3 00,000, approximately 6,400,000, approximately 6,500,000, approximately 6,600,000, approximately 6,700,000, approximately 6,800,000, approximately 6,900,000, approximately 7,000,000, approximately 7,100,000, approximately 7,200,000, approximately 7,300,000 00, approximately 7,400,000, approximately 7,500,000, approximately 7,600,000, approximately 7,700,000, approximately 7,800,000, approximately 7,900,000, approximately 8,000,000, approximately 8,100,000, approximately 8,200,000, approximately 8,300,000, approximately 8,400,000, approximately 8,500,000, approximately 8,600,000, approximately 8,700,000, approximately 8,800,000, approximately 8,900,000, approximately 9,000,000, approximately 9,100,000, approximately 9,200,000, approximately 9,300,000, approximately 9,400,000 0,000, approximately 9,500,000, approximately 9,600,000, approximately 9,700,000, approximately 9,800,000, approximately 9,900,000, approximately 10,000,000, approximately 10,100,000, approximately 10,200,000, approximately 10,300,000, approximately 10,400 0,000, approximately 10,500,000, approximately 10,600,000, approximately 10,700,000, approximately 10,800,000, approximately 10,900,000, approximately 11,000,000, approximately 11,100,000, approximately 11,200,000, approximately 11,300,000 Approximately 11,400,000, approximately 11,500,000, approximately 11,600,000, approximately 11,700,000, approximately 11,800,000, approximately 11,900,000, approximately 12,000,000, approximately 12,100,000, approximately 12,200,000, approximately 12,300,000, approximately 12,400,000, approximately 12,500,000, approximately 12,600,000, approximately 12,700,000, approximately 12,800,000, approximately 12,900,000, approximately 13,000,000, approximately 13,100,000, approximately 13,200,000, approximately 13 300,000, approximately 13,400,000, approximately 13,500,000, approximately 13,600,000, approximately 13,700,000, approximately 13,800,000, approximately 13,900,000, approximately 14,000,000, approximately 14,100,000, approximately 14,200,000 Approximately 14,300,000, approximately 14,400,000, approximately 14,500,000, approximately 14,600,000, approximately 14,700,000, approximately 14,800,000, approximately 14,900,000, approximately 15,000,000, approximately 15,100,000, approximately 15,200 0,000, approximately 15,300,000, approximately 15,400,000, approximately 15,500,000, approximately 15,600,000, approximately 15,700,000, approximately 15,800,000, approximately 15,900,000, approximately 16,000,000, approximately 16,100,000, approximately 16,200,000, approximately 16,300,000, approximately 16,400,000, approximately 16,500,000, approximately 16,600,000, approximately 16,700,000, approximately 16,800,000, approximately 16,900,000, approximately 17,000,000, approximately 17,100,000 00, approximately 17,200,000, approximately 17,300,000, approximately 17,400,000, approximately 17,500,000, approximately 17,600,000, approximately 17,700,000, approximately 17,800,000, approximately 17,900,000, approximately 18,000,000, approximately 18, 100,000, approximately 18,200,000, approximately 18,300,000, approximately 18,400,000, approximately 18,500,000, approximately 18,600,000, approximately 18,700,000, approximately 18,800,000, approximately 18,900,000, approximately 19,000,000 Approximately 19,100,000, approximately 19,200,000, approximately 19,300,000, approximately 19,400,000, approximately 19,500,000, approximately 19,600,000, approximately 19,700,000, approximately 19,800,000, approximately 19,900,000, approximately 20,000000 pieces, about 20,100,000 pieces, about 20,200,000 pieces, about 20,300,000 pieces, about 20,400,000 pieces, about 20,500,000 pieces, about 20,600,000 pieces, about 20,700,000 pieces, about 20, 800,000 pieces, about 20,900,000 pieces, about 21,000,000 pieces, about 21,100,000 pieces, about 21,200,000 pieces, about 21,300,000 pieces, about 21,400,000 pieces, about 21,500,000 pieces, 21,600,000 pieces, 21,700,000 pieces, 21,800,000 pieces, 21,900,000 pieces, 22,000,000 pieces, 22,100,000 pieces, 22,200,000 pieces, 22,300,0 pieces 00 pieces, about 22,400,000 pieces, about 22,500,000 pieces, about 22,600,000 pieces, about 22,700,000 pieces, about 22,800,000 pieces, about 22,900,000 pieces, about 23,000,000 pieces, about 23,10 0,000 pieces, about 23,200,000 pieces, about 23,300,000 pieces, about 23,400,000 pieces, about 23,500,000 pieces, about 23,600,000 pieces, about 23,700,000 pieces, about 23,800,000 pieces, about 2 3,900,000 pieces, about 24,000,000 pieces, about 24,100,000 pieces, about 24,200,000 pieces, about 24,300,000 pieces, about 24,400,000 pieces, about 24,500,000 pieces, about 24,600,000 pieces , about 24,700,000, about 24,800,000, about 24,900,000, about 25,000,000, about 25,100,000, about 25,200,000, about 25,300,000, about 25,400,000, about 25,500,000, about 25,600,000, about 25,700,000, about 25,800,000, about 25,900,000, or 26,000,000 nucleic acid bases, located downstream or upstream.
[0167] In some embodiments, the targeted integration site is located within a high complexity sequence in SEQ ID NO: 22 or the sequence set forth in the Chr3 TI contig. In some embodiments, the targeted integration site comprises a high complexity sequence in SEQ ID NO: 22 or the sequence set forth in the Chr3 TI contig.
[0168] In some embodiments, the targeted integration site is located within a high complexity sequence in SEQ ID NO: 118 or the sequence set forth in the Chr5 TI contig. In some embodiments, the targeted integration site comprises a high complexity sequence in SEQ ID NO: 118 or the sequence set forth in the Chr5 TI contig.
[0169] In some embodiments, the targeted integration site is not located within a low complexity sequence in the sequence defined in SEQ ID NO: 22 or the Chr3 TI contig. In some embodiments, the targeted integration site is not located within a retrotransposon sequence in the sequence defined in SEQ ID NO: 22 or the Chr3 TI contig. In some embodiments, the targeted integration site is not located within an Alu repeat (CHO Alu equivalent) or the like in the sequence defined in SEQ ID NO: 22 or the Chr3 TI contig. In some embodiments, the targeted integration site is not located within a long interspersed nuclear element (LINE) in the sequence defined in SEQ ID NO: 22 or the Chr3 TI contig. In some embodiments, the targeted integration site / locus does not contain a CHO Alu equivalent sequence.
[0170] In some embodiments, the targeted integration site is not located within a low complexity sequence in the sequence defined in SEQ ID NO: 118 or the Chr5 TI contig. In some embodiments, the targeted integration site is not located within a retrotransposon sequence in the sequence defined in SEQ ID NO: 118 or the Chr5 TI contig. In some embodiments, the targeted integration site is not located within an Alu repeat (CHO Alu equivalent) or the like in the sequence defined in SEQ ID NO: 118 or the Chr5 TI contig. In some embodiments, the targeted integration site is not located within a long interspersed nuclear element (LINE) in the sequence defined in SEQ ID NO: 118 or the Chr5 TI contig. In some embodiments, the targeted integration site / locus does not contain a CHO Alu equivalent sequence.
[0171] In some embodiments, the targeted integration site is not composed of multiple low-complexity sequences in the sequence set forth in SEQ ID NO: 22 or the Chr3 TI contig. In some embodiments, the targeted integration site does not comprise a retrotransposon sequence in the sequence set forth in SEQ ID NO: 22 or the Chr3 TI contig. In some embodiments, the targeted integration site does not comprise an Alu repeat in the sequence set forth in SEQ ID NO: 22 or the Chr3 TI contig. In some embodiments, the targeted integration site does not comprise a long interspersed nuclear element (LINE) in the sequence set forth in SEQ ID NO: 22 or the Chr3 TI contig.
[0172] In some embodiments, the targeted integration site is not composed of multiple low-complexity sequences in the sequence set forth in SEQ ID NO: 118 or the Chr5 TI contig. In some embodiments, the targeted integration site does not comprise a retrotransposon sequence in the sequence set forth in SEQ ID NO: 118 or the Chr5 TI contig. In some embodiments, the targeted integration site does not comprise an Alu repeat in the sequence set forth in SEQ ID NO: 118 or the Chr5 TI contig. In some embodiments, the targeted integration site does not comprise a long interspersed nuclear element (LINE) in the sequence set forth in SEQ ID NO: 118 or the Chr5 TI contig.
[0173] In some embodiments, the targeted integration site is not flanked by low complexity sequences in the sequence set forth in SEQ ID NO: 22 or the Chr3 TI contig. In some embodiments, the targeted integration site is not flanked by retrotransposon sequences in the sequence set forth in SEQ ID NO: 22 or the Chr3 TI contig. In some embodiments, the targeted integration site is not flanked by Alu repeats in the sequence set forth in SEQ ID NO: 22 or the Chr3 TI contig. In some embodiments, the targeted integration site is not flanked by long interspersed nuclear elements (LINEs) in the sequence set forth in SEQ ID NO: 22 or the Chr3 TI contig.
[0174] In some embodiments, the targeted integration site is not flanked by low complexity sequences in the sequence set forth in SEQ ID NO: 118 or the Chr5 TI contig. In some embodiments, the targeted integration site is not flanked by retrotransposon sequences in the sequence set forth in SEQ ID NO: 118 or the Chr5 TI contig. In some embodiments, the targeted integration site is not flanked by Alu repeats in the sequence set forth in SEQ ID NO: 118 or the Chr5 TI contig. In some embodiments, the targeted integration site is not flanked by long interspersed nuclear elements (LINEs) in the sequence set forth in SEQ ID NO: 118 or the Chr5 TI contig.
[0175] As used herein, the term "low complexity sequence" refers to a nucleic acid sequence characterized by the presence of repetitive sequences, also known as repetitive elements, repeat units, or repeats. Conversely, the term "high complexity sequence" refers to a nucleic acid sequence characterized by the absence of multiple repetitive sequences. The main types of repetitive sequences are tandem repeats and interspersed repeats, including transposable elements such as retrotransposons.
[0176] Retrotransposons (also known as class I transposable elements or transposons mediated by RNA intermediates) are a type of genetic component (transposon) that copies and pastes itself to different locations in the genome by converting RNA back into DNA through a process of reverse transcription using an RNA transposition intermediate. There are two main types of retrotransposons: long terminal repeats (LTRs) and non-long terminal repeats (non-LTRs). Retrotransposons are classified based on their sequence and method of transposition. Non-LTRs are primarily divided into two types: LINEs (long interspersed nuclear elements) and SINEs (short interspersed nuclear elements). Alus is the most common SINE in primates.
[0177] The Alu family is a family of repetitive elements in primate genomes, including the human genome. Alu elements are short stretches of DNA first characterized by the action of the Arthrobacter luteus (Alu) restriction endonuclease. Alu elements are the most abundant transposable elements, containing over one million copies distributed throughout the human genome. Modern Alu elements are approximately 300 base pairs in length and are therefore classified as short interspersed nuclear elements (SINEs) within the class of repetitive DNA elements. Their typical structure is 5'-part A-A5TACA6-part B-polyA tail-3' (SEQ ID NO: 25), where parts A and B (also known as the "left arm" and "right arm") have similar nucleotide sequences. Two main promoter "boxes" are found in Alu: the 5'A box with the consensus TGGCTCACGCC (SEQ ID NO: 26) and the 3'B box with the consensus GWTCGAGAC (IUPAC nucleic acid notation).
[0178] In the context of this disclosure, reference to an Alu element as applied to the Mongolian crucian carp sequences disclosed herein refers to the CHO Alu equivalent, i.e., the Alu-like element present in the Mongolian crucian carp genome as described by Haynes et al. (1981) Molecular and Cellular Biology 1(7):573-583. Haynes et al. described a consensus sequence for the major interspersed deoxyribonucleic acid repeat in the genome of Chinese hamster ovary cells (CHO cells) that is highly homologous to human Alu sequences and mouse B1 interspersed repeat sequences. Because the CHO consensus sequence shows significant homology with the human Alu sequence, it is referred to as the CHO Alu equivalent sequence. A conserved structure surrounding CHO Alu equivalent family members can be recognized. This is similar to that surrounding the human Alu and mouse B1 sequences and is represented as follows: direct repeat, CHO-Alu-A-rich sequence-direct repeat. The consensus sequence for the CHO Alu equivalent sequence is disclosed in Figure 1 of Haynes et al., which is incorporated herein by reference in its entirety.
[0179] Long interspersed nuclear elements (LINEs), also known as long interspersed nucleotide elements or long interspersed elements, are a group of non-LTR (long terminal repeat) retrotransposons that are widespread in the genomes of many eukaryotes. They comprise approximately 21.1% of the human genome. LINEs constitute a family of transposons, each approximately 7,000 base pairs long. The only abundant LINE in humans is LINE1. The human genome contains an estimated 100,000 truncated LINE-1 elements and 4,000 full-length LINE-1 elements. Due to the accumulation of random mutations, the sequences of many LINEs have degenerated to the point that they are no longer transcribed or translated.
[0180] In some embodiments, the targeted integration site does not comprise a CpG island in SEQ ID NO: 22 or the sequence set forth in the Chr3 TI contig. In some embodiments, the targeted integration site is not present within a CpG island in SEQ ID NO: 22 or the sequence set forth in the Chr3 TI contig. In some embodiments, the targeted integration site is not flanked by a CpG island in SEQ ID NO: 22 or the sequence set forth in the Chr3 TI contig.
[0181] In some embodiments, the targeted integration site does not comprise a CpG island in SEQ ID NO: 118 or the sequence set forth in the Chr5 TI contig. In some embodiments, the targeted integration site is not present within a CpG island in SEQ ID NO: 118 or the sequence set forth in the Chr5 TI contig. In some embodiments, the targeted integration site is not flanked by a CpG island in SEQ ID NO: 118 or the sequence set forth in the Chr5 TI contig.
[0182] A CpG island (or CG island) is a region containing a high frequency of CpG sites. Although there are limited objective definitions for a CpG island, the usual official definition is a region of at least 200 bp, with a GC percentage greater than 50% and a CpG observed-to-expected ratio greater than 60%. The "CpG observed-to-expected ratio" is calculated by dividing the observed value by the number of CpGs and the expected value by (number of Cs × number of Gs) / length of the sequence, or [(number of Cs + number of Gs) / 2]. 2 / It can be derived by calculating the length of the sequence.
[0183] In mammalian genomes, CpG islands are typically 300-3,000 base pairs in length and are found in or near the promoters of approximately 40% of mammalian genes. Over 60% of human genes and almost all housekeeping genes have their promoters embedded in CpG islands.
[0184] Based on extensive surveys of the complete sequences of human chromosomes 21 and 22, DNA regions larger than 500 bp are likely to be "true" CpG islands associated with the 5' region of a gene if they have a GC content greater than 55% and a CpG observed-to-expected ratio of 65%.
[0185] CpG islands are characterized by a CpG dinucleotide content of at least 60% of what is statistically expected (approximately 4-6%), while the rest of the genome has a significantly lower CpG frequency (approximately 1%), a phenomenon called CpG suppression. Unlike CpG sites within the coding region of a gene, CpG sites within promoter CpG islands are, in most cases, unmethylated when the gene is expressed. Most of the methylation differences between tissues, or between normal and cancer samples, occur at a short distance from the CpG island (at the "CpG island shore"), rather than within the island itself.
[0186] Because histone acetylation and cytosine demethylation enhance transcription, the targeted integration site is located, for example, within a locus that contains above-average levels of acetylated histones and / or above-average levels of unmethylated cytosines.
[0187] Therefore, in some embodiments, targeted integration site is located in, comprises, or is flanked by the subsequence of the sequence set forth in SEQ ID NO: 22 or Chr3 TI contig or SEQ ID NO: 118 or Chr5 TI contig, characterized by above-average level of unmethylated cytosine.As used herein, above-average level of unmethylated cytosine is considered based on the number of unmethylated cytosine per kilobase, for example, over a specific polynucleotide length.Therefore, the percentage of unmethylated cytosine can be calculated, for example, for the sequence set forth in SEQ ID NO: 22 or SEQ ID NO: 118, to obtain the average level of unmethylated cytosine. Subsequences in SEQ ID NO:22 or SEQ ID NO:118 can then be scored according to whether the percentage of unmethylated cytosines within the subsequence (e.g., 10 nt, 100 nt, 1000 nt, 10000 nt or more) is above or below the average number of unmethylated cytosines calculated for the entire sequence set forth in SEQ ID NO:22 or SEQ ID NO:118.
[0188] Similarly, in some embodiments, the targeted integration site is located within, includes, or is flanked by a subsequence in the sequence set forth in SEQ ID NO: 22 or Chr3 TI contig, or SEQ ID NO: 118 or Chr5 TI contig, characterized in that it is associated with a histone having an above-average level of acetylation. As used herein, an above-average level of histone acetylation is considered based on the number of acetylated histones over a specific polynucleotide length, for example, per kilobase. Thus, the percentage of acetylated histones can be calculated, for example, for the sequence set forth in SEQ ID NO: 22 or SEQ ID NO: 118, to obtain the average level of acetylated histones. Then, the subsequence in SEQ ID NO: 22 or SEQ ID NO: 118 can be scored according to whether the percentage of acetylated histones in the subsequence (e.g., 10nt, 100nt, 1000nt, 10000nt, or more) is higher or lower than the average number of acetylated histones calculated for the entire sequence set forth in SEQ ID NO: 22 or SEQ ID NO: 118.
[0189] Methods for identifying methylation markers and transcription or open chromatin boundaries are described, for example, in Sharmin et al. (2016) BMC Cancer 16:88, Wang et al. (2012) Nucleic Acids Res. 40:511-29, Papin et al. (2020) J. Mol. Biol. doi:10.1016 / j.jmb.2020.09.018, Li et al. (2013) BMC Genomics 14:553, Butcher & Beck (2015) Methods 72:21-8, Chen et al. (2020) Epigenetics 22:1-22, Keller et al. (2016) Mol. Biol. Evol. 33:1019-28, and Symmons et al. (2014) Genome Res. 24:390-400, or Mifsud et al. (2015) Nat. Genet. 47:598-606, Collings & Anderson (2017) Epigenetics and Chromatin 10 doi.org / 10.1186 / s13072-017-0125-5, all of which are incorporated herein by reference in their entireties.
[0190] In some embodiments, the targeted integration site is located within, comprises, or is flanked by a subsequence in SEQ ID NO: 22 or the sequence set forth in the Chr3 TI contig, characterized as being a region where early initiation of replication occurs.
[0191] In some embodiments, the targeted integration site is located within, comprises, or is flanked by a subsequence in the sequence set forth in SEQ ID NO: 118 or the Chr5 TI contig, characterized as being a region where early initiation of replication occurs.
[0192] The early start of replication is associated with open chromatin and areas of transcription. Methods for identifying replication origins and their association with chromatin state and transcription are provided, for example, in Smith & Aladjem (2014) J. Mol. Biol. 426:3330-41, Dellino et al. (2013) Genome Res. 23:1-11, Boos & Ferreira (2019) Genes 10:199, Boulos et al. (2015) FEBS Lett. 489:2944-57, or Gomez & Brockdorff (2004) Proc. Natl. Acad. Sci. USA 101:6923-6928, all of which are incorporated herein by reference in their entirety. Based on these methods, the replication origins in the sequence set forth in SEQ ID NO: 22 or the Chr3 TI contig, or SEQ ID NO: 118 or the Ch5 TI contig, can be classified and ranked as early, middle, and late replication initiation regions. In some embodiments, the targeted integration site is located within, contains, or is flanked by a subsequence in the sequence set forth in SEQ ID NO: 22 or the Chr3 TI contig, or SEQ ID NO: 118 or the Ch5 TI contig, that is within the top 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the replication initiation regions.
[0193] In some embodiments, the targeted integration site comprises the sequence set forth in SEQ ID NO: 21, or a portion thereof, located at positions 20,002 to 20,019 of SEQ ID NO: 20. In some embodiments, the targeted integration site is located within or comprises a subsequence of the sequence of SEQ ID NO: 21 within SEQ ID NO: 20.
[0194] In some embodiments, the targeted integration site comprises the sequence set forth in SEQ ID NO: 117, or a portion thereof. In some embodiments, the targeted integration site is located within or comprises a subsequence of the sequence of SEQ ID NO: 117 within SEQ ID NO: 116.
[0195] In some embodiments, the targeted integration site is located upstream from SEQ ID NO:21 or SEQ ID NO:117.
[0196] In some embodiments, the targeted integration site is located downstream from SEQ ID NO:21 or SEQ ID NO:117.
[0197] In some embodiments, the targeted integration site is selected from the group consisting of about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 250, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, about 2000, about 2500, about 3000, about 3500, about 4000, about 5000, about 6000, about 7500, about 8000, about 8500, about 9000, about 9500, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, about 2000, about 2500, about 3000, about The sequence is located about 450, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, about 2000, about 2100, about 2200, about 2300, about 2400, about 2500, about 2600, about 2700, about 2800, about 2900, or about 3000 nt upstream.
[0198] In some embodiments, the targeted integration site is selected from the group consisting of about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 250, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, about 2000, about 2500, about 3000, about 3500, about 4000, about 5000, about 6000, about 7500, about 8000, about 8500, about 9000, about 9500, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, about 2000, about 2500, about 3000, about about 450, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, about 2000, about 2100, about 2200, about 2300, about 2400, about 2500, about 2600, about 2700, about 2800, about 2900, or about 3000 nt downstream.
[0199] In some embodiments, the targeted integration site is located within the sequence set forth in SEQ ID NO: 20, or a fragment thereof, e.g., a sequence orthologous to SEQ ID NO: 21. In some embodiments, the targeted integration site is located within the sequence set forth in SEQ ID NO: 20, or a fragment thereof, e.g., a sequence paralogous to SEQ ID NO: 21.
[0200] In some embodiments, the targeted integration site is located within the sequence set forth in SEQ ID NO: 116, or a fragment thereof, e.g., a sequence orthologous to SEQ ID NO: 117. In some embodiments, the targeted integration site is located within the sequence set forth in SEQ ID NO: 116, or a fragment thereof, e.g., a sequence paralogous to SEQ ID NO: 117.
[0201] In some embodiments of the present disclosure, the targeted integration site is located within SEQ ID NO: 20. SEQ ID NO: 20 is a partial sequence of SEQ ID NO: 22 (a 26 Mbase sequence from chromosome 3 of the Mongolian cress, Chinese hamster), which includes 20 Kb on each side of the integration site of SEQ ID NO: 21.
[0202] In some embodiments of the present disclosure, the targeted integration site is located within SEQ ID NO: 116. SEQ ID NO: 116 is a partial sequence of SEQ ID NO: 118 (an 18 Mbase sequence from chromosome Chr5 of the Mongolian cress, Chinese hamster), which includes 20 Kb on each side of the integration site of SEQ ID NO: 117.
[0203] In some embodiments, the disclosure provides an isolated cell comprising a polynucleotide sequence (exogenous nucleic acid) comprising a nucleic acid encoding a gene of interest (GOI) integrated into a specific locus in the genome of the cell, wherein the locus has at least about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity to SEQ ID NO:20, SEQ ID NO:116, or a subsequence thereof.
[0204] While the sequences disclosed herein are derived from the Mongolian crescendo rat, it should be understood that the present disclosure also encompasses orthologous sequences from other species, e.g., human, mouse, rabbit, rat, pig, or dog. Thus, reference to any of the sequences set forth in SEQ ID NOS: 14-24 and 110-120 also encompasses variant sequences having at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to their parent or reference sequences (i.e., any of the sequences set forth in SEQ ID NOS: 14-24 and 110-120, or fragments or subsequences thereof), as determined, for example, by pairwise alignment using an implementation of the Needleman-Wunsch algorithm. As used herein, the term "orthologous" refers to multiple polynucleotides that have similar nucleic acid sequences due to separation by a speciation event, i.e., they represent homologous sequences in different organisms due to ancestral relationships, and therefore perform similar functions in the different organisms. Thus, sequences (or subsequences) orthologous to the sequences (or subsequences) disclosed herein are considered to be functionally equivalent to the known sequences from the Mongolian crescendo mouse disclosed herein, i.e., can similarly be used as specific loci for targeted integration.
[0205] In some embodiments, the disclosure provides an isolated cell comprising a polynucleotide sequence (exogenous nucleic acid) comprising a nucleic acid encoding a gene of interest (GOI) integrated into a specific locus in the genome of the cell, wherein the locus has at least about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity between SEQ ID NO:20 and SEQ ID NO:116 or a subsequence thereof.
[0206] In some embodiments, the subsequence is about 18, about 20, about 25, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, about 130, about 140, about 150, about 175, about 200, about 225, about 250, about 275, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 700, about 800, about 900, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, or about 2000 nucleotides in length, and the subsequence comprises the sequence set forth in SEQ ID NO:21 or SEQ ID NO:117.
[0207] In some embodiments, the subsequence comprises about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, about 130, about 140, about 150, about 175, about 200, about 225, about 250, about 275, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 700, about 800, about 900, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, or about 2000 nucleotides upstream to the sequence set forth in SEQ ID NO:21 or SEQ ID NO:117.
[0208] In some embodiments, the subsequence comprises about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, about 130, about 140, about 150, about 175, about 200, about 225, about 250, about 275, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 700, about 800, about 900, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, or about 2000 nucleotides downstream relative to the sequence set forth in SEQ ID NO:21 or SEQ ID NO:117.
[0209] In some embodiments, the present disclosure provides an isolated cell comprising a polynucleotide sequence (exogenous nucleic acid) comprising a nucleic acid encoding a gene of interest (GOI) integrated into a specific locus in the genome of the cell, wherein the locus (e.g., a targeted integration site of the present disclosure) is a nucleotide position or nucleotide sequence within SEQ ID NO: 20 or SEQ ID NO: 116. The present disclosure also provides a method comprising introducing a polynucleotide sequence (exogenous nucleic acid) comprising a nucleic acid encoding a gene of interest (GOI) into a mammalian cell, such as a CHO cell, and obtaining a mammalian cell, such as a CHO cell, wherein the exogenous nucleic acid has integrated into a specific locus in the genome of the CHO cell, wherein the locus is a nucleotide position or nucleotide sequence within SEQ ID NO: 20 or SEQ ID NO: 116. Also provided is a method comprising: (a) providing a cell comprising a polynucleotide sequence (exogenous nucleic acid) comprising a nucleic acid encoding a gene of interest (GOI), wherein the polynucleotide sequence comprises a nucleic acid encoding a gene of interest (GOI) operably linked to a promoter, and the locus is a nucleotide position or nucleotide sequence within SEQ ID NO: 20 and SEQ ID NO: 116. In some embodiments, the locus overlaps with SEQ ID NO:20 or SEQ ID NO:16.
[0210] In some embodiments, a polynucleotide sequence comprising a nucleic acid encoding a gene of interest (GOI) is integrated into a specific site at either position of SEQ ID NO:20 or SEQ ID NO:116.
[0211] In some embodiments, the specific site is a sequence of nucleotides having numbers 1 to 1000, 1001 to 2000, 2001 to 3000, 3001 to 4000, 4001 to 5000, 5001 to 6000, 6001 to 7000, 7001 to 8000, 8001 to 9000, 9001 to 10000, 10001 to 11000, 110 01-12000, 12001-13000, 13001-14000, 14001-15000, 15001-16000, 16001-17000, 17001-18000, 18001-19000, 190001-20000, 20001-21000, 21001-22000, 22 001-23000, 23001-24000, 24001-25000, 25001-26000, 26001-27000, 27001-28000, 28001-29000, 29001-30000, 30001-31000, 31001-32000, 32001-33000, 33 and at a position in SEQ ID NO: 20 selected from the group consisting of nucleotides spanning positions 001 to 34000, 34001 to 35000, 35001 to 36000, 36001 to 37000, 37001 to 38000, 38001 to 39000, 39001 to 40000, or 40001 to 40020.
[0212] In some embodiments, the specific site is a sequence of amino acids 1 to 1000, 1001 to 2000, 2001 to 3000, 3001 to 4000, 4001 to 5000, 5001 to 6000, 6001 to 7000, 7001 to 8000, 8001 to 9000, 9001 to 10000, 10001 to 11000, 11001 to 12000, 12001 to 13000, 13001 to 14000, 140 01-15000, 15001-16000, 16001-17000, 17001-18000, 18001-19000, 190001-20000, 20001-21000, 21001-22000, 22001-23000, 23001-24000, 24001-25000, 25001-26000, 26001-27000, 27001-28000, 2 8001-29000, 29001-30000, 30001-31000, 31001-32000, 32001-33000, 33001-34000, 34001-35000, 35001-36000, 36001-37000, 37001-38000, 38001-39000, 39001-40000, 40001-41000, 410001-42000 , 42001~43000, 43001~44000, 44001~45000, 45001~46000, 46001~47000, 47001~48000, 49001~50000, 50001~51000, 51001~52000, 52001~53000, 53001~54000, 54001~55000, 55001~56000, 56001~5700. 57001~58000, 58001~59000, 59001~60000, 60001~61000, 61001~62000, 62001~63000, 63001~64000, 64001~65000, 65001~66000, 66001~67000, 67001~ 68000, 68001-69000, 69001-70000, 70001-71000, 71001-72000, 72001-7300, 73001-74000, 74001-75000, 75001-76000, 76001-77000, 77001-78000,78001-79000, 79001-80000, 80001-81000, 81001-82000, 82001-83000, 83001-84000, 84001-85000, 85001-86000, 86001-87000, 87001-8 8000, 88001-89000, 89001-90000, 90001-91000, 91001-92000, 92001-93000, 93001-94000, 94001-95000, 95001-96000, 96001-97000, 9 7001~98000, 98001~99000, 99001~100000, 100001~101000, 101001~102000, 102001~103000, 103001~104000, 104001~105000, 105001~106 No. 000, No. 106001-107000, No. 107001-108000, No. 108001-109000, No. 109001-110000, No. 110001-111000, No. 111001-112000, No. 112001-113000, No. 113001-114000 , 114001~115000, 115001~116000, 116001~117000, 117001~118000, 118001~119000, 119001~120000, 120001~121000, 121001~122000, 122 001~123000, 123001~124000, 124001~125000, 125001~126000, 126001~127000, 127001~128000, 128001~129000, 129001~130000, 130001~ 131000, 131001-132000, 132001-133000, 133001-134000, 134001-135000, 135001-136000, 136001-137000, 137001-138000, 138001-1390 00, 139001~140000, 140001~141000, 141001~142000, 142001~143000, 143001~144000, 144001~145000, 145001~146000, 146001~147000,147001-148000, 148001-149000, 149001-150000, 150001-151000, 151001-152000, 152001-153000, 153001-154000, 154001-155000, 155001-156000, 156001-157000, 157 001~158000, 158001~159000, 159001~160000, 160001~161000, 161001~162000, 162001~163000, 163001~164000, 164001~165000, 165001~166000, 166001~167000, 167001 ~168000, 168001~169000, 169001~170000, 170001~171000, 171001~172000, 172001~173000, 173001~174000, 174001~175000, 175001~176000, 176001~177000, 177001~17 116 is located at a position in SEQ ID NO: 116 selected from the group consisting of nucleotides spanning positions 8000, 178001 to 179000, 179001 to 180000, 180001 to 181000, 181001 to 182000, 182001 to 183000, 183001 to 184000, or 184001 to 185000.
[0213] In some embodiments, the specific site is a sequence numbered 19,000 to 21,000, 18,000 to 22,000, 17,000 to 23,000, 16,000 to 24,000, 15,000 to 25,000, 14,000 to 26,000, 13,000 to 27,000, 12,000 to 28,000, 11,000 to 29,000, 10,000 to 30,000, 9,000 to 10,000, or 11,000 to 12,000. and at a position in SEQ ID NO: 20 selected from the group consisting of nucleotides spanning positions 31000, 8000-32000, 7000-33000, 6000-34000, 5000-35000, 4000-36000, 3000-37000, 2000-38000, 1000-39000, or 1-40020.
[0214] In some embodiments, the specific site is 19000-19100, 19100-19200, 19200-19300, 19300-19400, 19400-19500, 19500-19600, 19600-19700, 19700-19800, 19800-19900, 19900-20000, 20000-20100 20, at a position in SEQ ID NO: 20 selected from the group consisting of nucleotides spanning positions 20100-20200, 20200-20300, 20300-20400, 20400-20500, 20500-20600, 20600-20700, 20700-20800, 20800-20900, or 20900-21000.
[0215] In some embodiments, a polynucleotide sequence comprising a nucleic acid encoding a gene of interest (GOI) is integrated into a specific site at any position of SEQ ID NO: 20, the specific site being any of positions 20000-20020, 19990-20030, 19980-20040, 19970-20050, 19960-20060, 19950-20070, 19940-20080, 19930-20090, 19920-2010 0, 19910~20110, 19900~20120, 19890~20120, 19880~20130, 19870~20140, 19860~20150, 19850~20160, 19840~20170, 19830~20180, 19820~20190, 19810~20200, 19800~20210, 19790~20220, 19780~20230, 19770~202 30, 19760~20240, 19750~20250, 19740~20260, 19730~20270, 19720~20280, 19710~20290, 19700~20300, 19690~20310, 19680~20320, 19670~20330, 19660~20340, 19650~20350, 19640~20360, 19630~20370, 19620~20 and a position in SEQ ID NO: 20 consisting of nucleotides spanning positions 380, 19610-20390, 19600-20400, 19590-20410, 19580-20420, 19570-20430, 19560-20440, 19550-20450, 19540-20460, 19530-20470, 19520-20480, 19510-20490, or 19500-20500.
[0216] In some embodiments, a polynucleotide sequence comprising a nucleic acid encoding a gene of interest (GOI) is integrated into a specific site (hotspot) at a position within SEQ ID NO: 20 (the genomic sequence comprising hotspot 1) or SEQ ID NO: 116 (the genomic sequence comprising hotspot 2), or that overlaps with SEQ ID NO: 20 or SEQ ID NO: 116.
[0217] In some embodiments, the specific site at a position in SEQ ID NO: 20 is selected from the group consisting of nucleotide positions or subsequences ranging from positions 20,002 to 20,019 (corresponding to the 18-mer sequence set forth in SEQ ID NO: 21).
[0218] In some embodiments, the specific site is located at nucleotide positions 19900, 19901, 19902, 19903, 19904, 19905, 19906, 19907, 19908, 19909, 19910, 19911, 19912, 19913, 19914, 19915, 19916, 19917, 19918, 19919, 19920, 19921, 19922, 19923, 19924, 19925, 19926, 19927, 19928, 1992 ...30, 19931, 19932, 19933, 19934, 19935, 19936, 19937, 19938, 19939, 19940, 19941, 19942, 19943, 19944, 19945, 19946, 19947, 19948, 19949, 19950, 9927, 19928, 19929, 19930, 19931, 19932, 19933, 19934, 19935, 19936, 19937, 19938, 19939, 19949, 19941, 19942, 19943, 19944, 19945, 19946, 19947, 19948, 19949, 19950, 19951, 19952, 19953, 19954, 19955, 19956, 19956, 19958, 19959, 19960, 19961, 19962, 19963, 19964, 19965, 19966, 19967, 19968 , 19969, 19970, 19971, 19971, 19972, 19973, 19974, 19975, 19976, 19977, 19978, 19979, 19980, 19981, 19982, 19983, 19984, 19985, 19986, 19987, 199 88, 19989, 19990, 19991, 19992, 19993, 19994, 19995, 19996, 19997, 19998, 19999, 20000, 20001, 20002, 20003, 20004, 20005, 20006, 20007, 20008, 20 009, 20010, 20011, 20012, 20013, 20014, 20015, 20016, 20017, 20018, 20019, 20020, 20021, 20022, 20023, 20024, 20025, 20026, 20027, 20028, 20029, 2 0030, 20031, 20032, 20033, 20034, 20035, 20036, 20037, 20038, 20039, 20040, 20041, 20042, 20043, 20044, 20045, 20046, 20047, 20048, 20049, 20050,20051, 20052, 20053, 20054, 20055, 20056, 20057, 20058, 20059, 20060, 20061, 20062, 20063, 20064, 20065, 20066, 20067, 20068, 20069, 20070, 20071, 20072, 20073, 20074, 20075, 20076, 20077, 20078 20090, 20091, 20092, 20093, 20094, 20095, 20096, 20097, 20098, 20099, or 20100.
[0219] The present disclosure also provides a method that allows for the generation of landing pad and expression cell lines and the identification of additional hotspots within the genome of the parental cell line without any prior knowledge of the genomic sequence surrounding the parental plasmid. This versatile TI technique uses a site-specific endonuclease directed against a parental plasmid sequence in the parental cell line that is not present in the landing pad plasmid. An advantage of this approach is that knowledge of the flanking genomic DNA sequence is not required. For example, Figure 4A illustrates the need to know the genomic sequence targeted by CRISPR / Cas, as indicated by the solid box next to the scissors representing CRISPR / Cas. In contrast, Figures 8A and 8B show that the sequence targeted by CRISPR / Cas is internal to the parental plasmid. The boxes with vertical and wavy lines represent regions of homology between different plasmids.
[0220] According to these new strategies, a parent cell line with a high expression titer (e.g., 3-4 g / L of antibody) and a low copy number (e.g., 2) is first selected, as shown, for example, in Figure 2 and related disclosure. Once such a cell line, or "hot cell line," is identified, the hot cell line can be used according to two different strategies. In both strategies, the landing pad plasmid encodes a marker, such as a fluorescent marker like b1mCherry, and expresses a selectable marker, such as puromycin resistance, that is different from the parental plasmid present in the parent cell line, and the polynucleotide sequence encoding the marker is flanked by heterologous site-specific recombination sites (SSRS). Exemplary SSRSs shown in Figures 12A, 12B, 13, and 14 are two Lox sites (LoxP and Lox511) that are targets of Cre recombinase. However, as disclosed below, these strategies may also be implemented using alternative SSRSs, such as Lox, Frt, att, or combinations thereof. For example, the combination of Lox and Frt is depicted in FIG. 15, the use of att sites (binding sites) is shown in FIG. 19, etc.
[0221] In the presence of a site-specific endonuclease (e.g., CRISPR / Cas) and a landing pad plasmid, the first GOI (e.g., mAb expression cassette) in the parental cell line is either replaced with a landing pad, denoted as mCherry flanked by Lox sites (Strategy A), or deleted, with the landing pad plasmid integrated into an alternative locus within the genome of the hot cell line (Strategy B). Thus, in Strategy A, the landing pad plasmid is inserted into a hot spot that supports high expression, the same hot spot used in the parental cell line. In Strategy B, the first GOI (e.g., mAb expression cassette) in the parental cell line is deleted, and the landing pad plasmid is inserted at an alternative location within the genome of the parental cell line. Because the parental cell line is a hot cell, the identification of additional hot spots results in a landing pad cell line that can generate expressing cell lines with desirable properties, such as high titer. See Figure 8A.
[0222] The present disclosure provides a method for identifying a landing pad cell line, comprising: (1) generating a parent cell population that does not contain the first GOI by removing the first GOI from a plasmid integrated into the genomic sequence of the parent cell (e.g., hot cell); (2) generating a library of candidate cells by integrating a landing pad plasmid containing at least one marker (e.g., Cherry) into the parent cell population of (1) at an alternative genomic locus; and (3) screening a library of candidate cells containing at least one copy of the landing pad plasmid integrated at at least one alternative genomic locus, wherein candidate cell lines are selected if they meet desired properties, such as (a) a cell titer above a predetermined threshold level, (b) a plasmid copy number of a predetermined value, (c) an RNA expression level above a predetermined threshold level, or (d) a specific plasmid configuration if multiple plasmid copies are present. The present invention provides a method comprising:
[0223] In some embodiments, only cells containing a landing pad plasmid at the newly identified hotspot are selected. In some embodiments, cells containing two or more landing pad plasmids at the newly identified hotspot are selected. In some embodiments, the parent cell is a historical cell line, e.g., a cell line characterized by high titer expression of a GOI, such as an antibody or antigen-binding portion thereof. In some embodiments, the library of candidate cells is a library generated by random integration of landing pad sequences at multiple locations within the genome of a parent cell modified by deletion / excision / removal of an expression cassette encoding a protein of interest, such as an antibody or antigen-binding portion thereof. In some embodiments, this method selects hot cells containing at least one landing pad plasmid integrated at the new hotspot. In some embodiments, the parent cell line is a CHO cell line.
[0224] The present disclosure provides a method of generating a landing pad cell, comprising integrating a landing pad plasmid into the genome of a parent cell (e.g., a CHO hot cell) at a targeted integration site using homologous recombination (e.g., using CRISPR / Cas), wherein the sequence targeted for homologous recombination is located in the parent plasmid, i.e., the sequence targeted for homologous recombination is not a genomic sequence, and the homologous recombination site of the landing pad plasmid recombines with the corresponding homologous recombination site of the parent plasmid, thereby integrating the landing pad plasmid at a location internal to the parent plasmid where it is inserted into the parent cell genomic DNA.
[0225] In some embodiments, each landing pad plasmid comprises (i) a polynucleotide sequence comprising a nucleic acid encoding at least one selectable marker and / or at least one nucleic acid sequence encoding a detectable marker, (ii) two site-specific recombination sites (SSRS) flanking the polynucleotide sequence of (i), and (iii) two homologous recombination sites located 5' and 3' to the SSRS of (ii) that are homologous to corresponding homologous recombination sites in the parental plasmid.
[0226] The present disclosure also provides methods for generating expression cells, comprising integrating a GOI plasmid (e.g., a plasmid encoding an antibody or antigen-binding portion thereof) into the genome of a landing pad cell (e.g., a CHO hot cell) disclosed above using site-specific recombinase recombination (e.g., using the Cre / Lox system), wherein the site-specific recombination sites of the landing pad plasmid recombine with corresponding site-specific recombination sites of the GOI plasmid, thereby integrating the GOI. In some embodiments, the resulting expression plasmid comprises (i) a polynucleotide sequence comprising a nucleic acid encoding at least one GOI, and (ii) two SSRSs flanking the polynucleotide of (i).
[0227] Also provided are methods for generating expression cells, comprising: (a) integrating a landing pad plasmid into the genome of a parent cell (e.g., a parent hot cell) at a targeted integration site using homologous recombination, wherein the sequence targeted for homologous recombination is located in the parent plasmid and the homologous recombination site in the landing pad plasmid recombines with a corresponding homologous recombination site of the parent plasmid within the landing pad at a different genomic locus, thereby integrating the landing pad plasmid at a location internal to the landing pad at a different genomic locus in the parent cell genomic DNA; and (b) integrating a GOI plasmid into the genome of the landing pad cell using site-specific recombinase recombination, wherein the site-specific recombination site of the landing pad plasmid recombines with a corresponding site-specific recombination site of the GOI plasmid, thereby integrating the GOI plasmid at a location internal to the landing pad plasmid in the landing pad cell. In some embodiments of this method, each landing pad plasmid comprises (i) at least one polynucleotide sequence comprising a nucleic acid encoding at least one selectable marker and / or at least one nucleic acid sequence encoding a detectable marker, (ii) two SSRSs flanking the polynucleotide sequence of (i), and (iii) two homologous recombination sites located 5' and 3' to the SSRS of (ii) that are homologous to corresponding homologous recombination sites in the parental plasmid. In some embodiments of this method, the resulting expression plasmid comprises (i) a polynucleotide sequence comprising a nucleic acid encoding at least one GOI, and (ii) two SSRSs flanking the polynucleotide of (i).
[0228] Also provided is a method for generating a landing pad cell, comprising: (a) removing a parental plasmid or a portion thereof from a first hotspot location in a parental cell line; and (b) using homologous recombination to integrate the landing pad plasmid into a second hotspot location within the genome of the parental cell at a targeted integration site, wherein the sequence targeted for homologous recombination is present in the parental plasmid, and the homologous recombination site of the landing pad plasmid recombines with a corresponding homologous recombination site of the parental plasmid, thereby integrating the landing pad plasmid at a location within the parental plasmid inserted into the parental cell genomic DNA. In some embodiments of this method, the landing pad plasmid comprises: (i) a polynucleotide sequence comprising at least one nucleic acid sequence encoding a selectable marker and / or at least one nucleic acid sequence encoding a detectable marker; (ii) two site-specific recombination sites (SSRS) flanking the polynucleotide sequence of (i); and (iii) two homologous recombination sites located 5' and 3' to the SSRS of (ii), which are homologous to corresponding homologous recombination sites in the parental plasmid. Step (a) results in the generation of a cell population derived from a parent cell line (e.g., a hot cell line) that does not contain the first GOI (e.g., an antibody that was highly expressed in the parent cell line). In step (b), insertion of the landing pad plasmid into the genome of the cell population of step (a) results in the generation of a cell population containing land cell pads integrated at multiple locations, which can then be screened to identify new hot cells and their corresponding hotspots.
[0229] Also provided is a method of generating an expressing cell, comprising: (a) removing a parental plasmid from a first hotspot location in a parental cell line; (b) using homologous recombination to integrate a landing pad plasmid into a second hotspot location in the genome of the parental cell at a targeted integration site, wherein the sequence targeted for homologous recombination is that which was present in the parental plasmid, and wherein each landing pad plasmid comprises, for example, (i) a polynucleotide sequence comprising at least one nucleic acid encoding a selectable marker and / or at least one nucleic acid sequence encoding a detectable marker; (ii) two site-specific recombination sites (SSRS) flanking the polynucleotide sequence of (i); and (iii) two homologous recombination sites located 5' and 3' to the SSRS of (ii), which are homologous to corresponding homologous recombination sites in the parental plasmid; (c) using site-specific recombinase recombination to integrate the GOI plasmid into the genome of the landing pad cell, wherein the expression plasmid comprises, for example, (i) a polynucleotide sequence comprising a nucleic acid encoding a GOI, and (ii) two SSRSs flanking the polynucleotide of (1), and the site-specific recombination sites of the landing pad plasmid recombine with corresponding site-specific recombination sites of the GOI plasmid, thereby integrating the GOI plasmid at a location internal to the landing pad plasmid in the landing pad cell.
[0230] In some embodiments, the landing pad cells are CG / -[P1]-([P2]-[SSRS]-[M]-[SSRS]-[P2])n-[P1]- / CG, CG / -([P2]-[SSRS]-[M]-[SSRS]-[P2])n- / CG, CG / -[P1]-([P2]-[SSRS]-[M]-[P2])n-[P1]- / CG, CG / -([P2]-[SSRS]-[M]-[P2])n- / CG, CG / -[P1]-([P2]-[M]-[SSRS]-[P2])n-[P1]- / CG, or CG / -([P2]-[M]-[SSRS]-[P2])n- / CG [wherein CG is a parental cell genomic sequence flanking the inserted plasmid, [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from a landing pad plasmid, [M] is a polynucleotide sequence comprising at least one marker, [SSRS] is a site-specific recombination site (SSRS), and n is an integer between 1 and 10. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, n is 4. In some embodiments, n is 5. In some embodiments, n is 6. In some embodiments, n is 7. In some embodiments, n is 8. In some embodiments, n is 9. In some embodiments, n is 10.
[0231] Note that in any of the formulas in this disclosure, the labels [P1], [P2], and [SSRS] are merely descriptors of the origin or type of components to represent the topology of the construct. The nucleic acid sequences of each component [P1] and [P2] are different, i.e., the nucleic acid sequence of the first [P1] is different from the nucleic acid sequence of the second [P1], but they share a common origin, i.e., the parental plasmid. Similarly, the nucleic acid sequence of the first [P2] is different from the nucleic acid sequence of the second [P2], but they share a common origin, i.e., the landing pad plasmid. In some embodiments, the CG sequence in the landing pad cell is different from the CG sequence in the parental cell line, i.e., the plasmid is located in a hotspot that is different from the original hotspot in the parental cell line.
[0232] The [SSRS] components can be, for example, Cre / Lox sites, each of which can have a different sequence. However, in some embodiments, in any of the formulas containing an [SSRS] pair presented throughout this disclosure, one of the shown [SSRS]s can be omitted. When integration is performed using, for example, serine integrase, a single [SSRS] is required. Thus, in certain such embodiments, a single att site, e.g., an attP site, can be present in place of an [SSRS] pair.
[0233] In some embodiments, the topology of the plasmid incorporated into the expression cell is as described above. CG / -[P1]-([P2]-[SSRS]-[P3]-[SSRS]-[P2])n-[P1]- / CG, CG / -([P2]-[SSRS]-[P3]-[SSRS]-[P2])n- / CG, CG / -[P1]-([P2]-[SSRS]-[P3]-[P2])n-[P1]- / CG, CG / -([P2]-[SSRS]-[P3]-[P2])n- / CG, CG / -[P1]-([P2]-[P3]-[SSRS]-[P2])n-[P1]- / CG, or CG / -([P2]-[P3]-[SSRS]-[P2])n- / CG [where CG corresponds to the parental cell genomic sequence flanking the inserted plasmid, [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [P3] is a polynucleotide sequence derived from the plasmid containing the gene of interest (GOI), and [SSRS] is a site-specific recombination site (SSRS), and n is an integer between 1 and 10. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3. In some embodiments, n is 4. In some embodiments, n is 5. In some embodiments, n is 6. In some embodiments, n is 7. In some embodiments, n is 8. In some embodiments, n is 9. In some embodiments, n is 10. In some embodiments, the CG sequence in the landing pad cell is different from the CG sequence in the parental cell line, i.e., the plasmid is located in a hotspot that is different from the original hotspot in the parental cell line.
[0234] In some embodiments, homologous recombination is mediated by CRISPR / Cas system, TALEN system or ZFN system, which will be described in detail below.In some embodiments, homologous recombination system, for example, CRISPR / Cas system, further comprises single guide RNA (sgRNA).Depending on the homologous recombination used, additional components may be required, as will be described in detail below.
[0235] In some embodiments, the site-specific recombinase recombination site (SSRS) is a Tyr-recombinase site, a Tyr-integrase site, a serine-resolvase / invertase site, a serine-integrase site, or a combination thereof. In some embodiments, the Tyr-recombinase site comprises a Cre, Dre, Flp, KD, B3, or B3 Tyr-recombinase site. In some embodiments, the Tyr-integrase site comprises a λ (lambda), HK022, or HP1 Tyr-integrase site. In some embodiments, the serine-resolvase / invertase site comprises a γδ (gamma delta), ParA, Tn3, or Gin serine-resolvase / integrase site. In some embodiments, the serine-integrase site comprises a PhiC31, Bxb1, or R4 serine-integrase site. In some embodiments, the Tyr-recombinase site comprises a Cre Tyr-recombinase site. In some embodiments, the SSRS is a LoxP site. In some embodiments, the LoxP site comprises the nucleic acid sequence set forth in SEQ ID NO: 1 (wild-type LoxP). In some embodiments, the LoxP site comprises a mutant LoxP site. In some embodiments, the mutant LoxP site comprises the nucleic acid sequence set forth in SEQ ID NO: 2 (mutant LoxP). In some embodiments, the mutant LoxP site comprises a nucleic acid selected from the group consisting of, for example, SEQ ID NO: 3 (Lox511), SEQ ID NO: 4 (Lox5171), SEQ ID NO: 5 (Lox2272), SEQ ID NO: 6 (LoxM2), SEQ ID NO: 7 (LoxM3), SEQ ID NO: 8 (LoxM7), SEQ ID NO: 9 (LoxM11), SEQ ID NO: 10 (Lox71), and SEQ ID NO: 11 (Lox66). In some embodiments, the Tyr-recombinase site comprises a Flp Tyr-recombinase site. In some embodiments, the SSRS is a short flippase recognition target (FRT) site. In some embodiments, the serine-integrase site comprises an att site, eg, an attP or attB site.
[0236] In some embodiments, each SSRS of a pair of SSRSs in a plasmid disclosed herein may belong to a different class. For example, the first SSRS may be, for example, a Tyr-recombinase site, and the second SSRS may be, for example, a Ser-integrase site. In some embodiments, the SSRS pair comprises two sites selected from wild-type LoxP, mutant LoxP, Lox511, Lox5171, Lox2272, LoxM2, LoxM3, LoxM7, LoxM11, Lox71, Lox66, or any combination thereof. In some embodiments, the SSRS pair comprises a LoxP site and a Lox511 site. In some embodiments, the SSRS pair comprises a LoxP site and an Frt site. In some embodiments, the SSRS pair comprises two aat sites, for example, two attP sites. In some embodiments, the SSRS pair comprises two aat sites, for example, two attR sites. In some embodiments, the SSRS pair comprises a Lox 2272 site and a Lox M3 site. In some embodiments, the SSRS pair comprises a Lox m3 site and a Lox m7 site.
[0237] In some embodiments, the plasmids disclosed herein comprise at least one single selection marker. In some embodiments, the plasmids disclosed herein comprise a single selection marker. In some embodiments, the plasmids disclosed herein comprise two or more single selection markers, for example, two selection markers. In some embodiments, at least one selection marker is glutamine synthetase (GS). In some embodiments, at least one selection marker is dihydrofolate reductase (DHFR). In some embodiments, at least one selection marker comprises a glutamine synthetase (GS) marker and a dihydrofolate reductase (DHFR) marker. There are several selection markers suitable for generating stably transfected Chinese hamster ovary (CHO) cell lines. Due to different modes of action, each selection marker has its own optimal selection stringency in different host cells to achieve high productivity. See Yeo et al. (2017) Biotechnol J 12(12), the entire contents of which are incorporated herein by reference.
[0238] In some embodiments, at least one selectable marker is a drug resistance gene, for example, an antibiotic resistance gene. In some embodiments, the antibiotic resistance gene is selected from the group consisting of an actinomycin D resistance gene, a bleomycin resistance gene, a chloramphenicol resistance gene, a G418 resistance gene, a hydromycin resistance gene, a mitomycin C resistance gene, a mycophenolic acid resistance gene, a puromycin resistance gene, and any combination thereof. In some embodiments, the antibiotic resistance gene is a puromycin resistance gene. In some embodiments, the puromycin resistance gene is puromycin-N-acetyltransferase.
[0239] In some embodiments, the at least one detectable marker comprises a protein such as a fluorescent protein. In some embodiments, the fluorescent protein is mCherry. In some embodiments, the fluorescent protein is selected from the group consisting of GFP, ZsGreen1, AcGFP1, EGFP, GFPuv, AcGFP, EBFP, EYFP, ECFP, tdTomato, mCherry, DsRed, AmCyan, ZsGreen, ZsYellow, DsRed2, DsRed-Express, HcRed, AsRed, mOrange, mOrange2, mPlum, mStrawberry, mBanana, YFP, mRaspberry, HcRed1, E2-Crimson, and any combination thereof.
[0240] In some embodiments, the parental cell is selected from the group consisting of a Chinese hamster ovary (CHO) cell, an HEK293 cell, and an NS0 cell, or a derivative or equivalent thereof. In some embodiments, the CHO cell is a CHO DG44 cell or a CHO K1 cell.
[0241] In some embodiments, the GOI encodes at least one polypeptide, e.g., an antibody or a fusion protein. In some embodiments, the antibody specifically binds to an immune checkpoint protein, such as T-cell immunoglobulin and mucin domain-containing protein 3 (TIM3), a Tau protein, e.g., an N-terminal fragment of tau (eTau), or PD-1 of PD-L1. In some embodiments, the antibody is nivolumab. In some embodiments, the GOI is the heavy chain (HC) of an antibody. In some embodiments, the GOI is the light chain (LC) of an antibody. In some embodiments, the GOI comprises the HC and LC of an antibody (e.g., in a bicistronic construct). In some embodiments, the GOI is a bispecific antibody or a portion thereof, e.g., the HC or LC of a bispecific antibody, or any combination thereof. In some embodiments, the expression plasmid comprises one copy, two copies, or more than two copies of the GOI.
[0242] In some embodiments, the method disclosed herein comprises determining the expression of GOI or marker disclosed herein.In some embodiments, the expression of GOI or marker is determined quantitatively and / or qualitatively.In some embodiments, the expression of GOI or marker is determined by, for example, cell sorting, FACS, cell surface staining, Western blot, Northern blot, column chromatography, capillary electrophoresis, microfluidics, UV absorbance, immunohistochemistry, cell size, secreted protein level, transcript level, or any combination thereof.
[0243] In some embodiments, the landing pad plasmid (second GOI plasmid) or expression plasmid (P4) is integrated into the genome of the cell at a copy number of 1. In some embodiments, the landing pad plasmid (second GOI plasmid) or expression plasmid (P4) is integrated into the genome of the cell at a copy number of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30.
[0244] In some embodiments, the 5' homologous recombination site of the plasmids disclosed herein comprises the polynucleotide sequence of SEQ ID NO: 18 or a subsequence thereof, and the 3' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 19 or a subsequence thereof.
[0245] In some embodiments, the 5' homologous recombination site of the plasmids disclosed herein comprises the polynucleotide sequence of SEQ ID NO: 114 or a subsequence thereof, and the 3' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 115 or a subsequence thereof.
[0246] In some embodiments, an isolated cell or isolated cell population of the present disclosure comprises a polynucleotide sequence comprising a nucleic acid encoding a gene of interest (GOI) integrated into a specific locus in the genome of the cell, the locus comprising a nucleotide subsequence selected from SEQ ID NO:20 and SEQ ID NO:116. In some embodiments, a method disclosed herein comprises introducing a polynucleotide sequence comprising a nucleic acid encoding at least one gene of interest (GOI) into a cell, such as a CHO cell, or another suitable cell line, and obtaining a cell, such as a CHO cell, wherein the exogenous nucleic acid has been integrated into a specific locus in the genome of the cell, the locus comprising a nucleotide subsequence selected from SEQ ID NO:20 and SEQ ID NO:116. In some embodiments, a method disclosed herein comprises (a) providing a cell comprising a polynucleotide sequence comprising a nucleic acid encoding a gene of interest (GOI), the polynucleotide sequence comprising a nucleic acid encoding a gene of interest (GOI) operably linked to a promoter, the locus comprising a nucleotide subsequence selected from SEQ ID NO:20 and SEQ ID NO:116. In some embodiments, the nucleotide subsequence selected from SEQ ID NO:20 comprises the sequence set forth in SEQ ID NO:21. In some embodiments, the nucleotide subsequence selected from SEQ ID NO:116 comprises the sequence set forth in SEQ ID NO:117. In some embodiments, the nucleotide subsequence selected from SEQ ID NO:20 consists of the sequence set forth in SEQ ID NO:21. In some embodiments, the nucleotide subsequence selected from SEQ ID NO:116 consists of the sequence set forth in SEQ ID NO:117. In some embodiments, the nucleotide subsequence selected from SEQ ID NO:20 is an upstream subsequence (towards the 5' end of SEQ ID NO:20) relative to the sequence set forth in SEQ ID NO:21. In some embodiments, the nucleotide subsequence selected from SEQ ID NO:20 is a downstream subsequence (towards the 3' end of SEQ ID NO:20) relative to the sequence set forth in SEQ ID NO:21. In some embodiments, the nucleotide subsequence selected from SEQ ID NO:116 is an upstream subsequence (towards the 5' end of SEQ ID NO:116) relative to the sequence set forth in SEQ ID NO:117.In some embodiments, the nucleotide subsequence selected from SEQ ID NO:116 is a subsequence downstream (towards the 3' end of SEQ ID NO:116) relative to the sequence set forth in SEQ ID NO:117.
[0247] The present disclosure provides landing pad cell lines containing a single landing pad plasmid. However, landing pad cell lines containing two or more landing pad plasmids offer an opportunity to further refine the expression of multi-subunit biologics, such as bispecific monoclonal antibodies (mAbs). Therefore, the cell screening method disclosed herein can be used to identify landing pad cell lines containing two landing pad plasmids at the same locus, i.e., duo-landing pad cells. This ensures equal expression from both landing pad plasmids, since they reside at the same genomic locus.
[0248] The duo landing pads of the present disclosure can integrate in four different orientations: head-to-head, tail-to-tail, tail-to-head, and head-to-tail. Because the head-to-head and tail-to-tail configurations are functionally indistinguishable from each other, they are generally used when single-site-directed recombinases such as Cre / Lox or Flp / Frt are used. Unlike tail-to-head and head-to-tail configurations, which can result in the deletion of one of the landing pads in the presence of Cre / Lox, head-to-head and tail-to-tail configurations simply flip, resulting in the same starting configuration.
[0249] When a second GOI plasmid is used with each of the four duo-landing pad configurations (head-to-head, tail-to-tail, tail-to-head, and head-to-tail), the head-to-head and tail-to-tail configurations can each generate two cell lines in which the sequence between the two recombination sites flanking the plasmid junction may be inverted, but the two cell lines are otherwise identical. When the head-to-tail or tail-to-head configuration is used with a second GOI plasmid, a cell line containing two second GOI plasmids is produced. However, if sufficient Cre activity is present, one of the second GOI plasmids can be excised, resulting in a second GOI plasmid cell line containing a single second GOI plasmid.
[0250] If the landing pad uses an Frt recognition site for Flp instead of a Lox site, such as Lox511, and both Cre / Lox and Flp are used, the same results are obtained: deletions occur in the tail-to-head and head-to-tail orientations, and inversions occur in the head-to-head and tail-to-tail orientations. However, in the tail-to-tail and head-to-head configurations, recombination of a second GOI plasmid into the duo landing pad using attP / attB with integrase does not result in inversions, whereas in the tail-to-head and head-to-tail configurations, deletion of one of the landing pads can still occur. If each of the landing pads has a single attP site, a single integration of the second GOI plasmid containing a single attB site occurs in any of the four duo landing pad configurations, resulting in no deletions.
[0251] As used herein, the term "single landing pad" refers to a landing pad containing a single landing pad plasmid or a second GOI plasmid. As used herein, the term "duo landing pad" refers to a landing pad containing two landing pad plasmids or a second GOI plasmid.
[0252] The use of a duo landing pad offers an alternative method for producing biologics containing different GOIs, such as antibodies containing heavy and light chains. In one embodiment, the present disclosure provides methods and compositions in which a second GOI plasmid contains multiple expression cassettes encoding, for example, antibody heavy and light chains. In another embodiment, each expression cassette may be in a different second GOI plasmid, with both second GOI plasmids located within the duo landing pad.
[0253] The use of duo landing pad cell lines has advantages over landing pad cell lines containing a single landing pad (i.e., a landing pad containing a single second GOI plasmid). In the case of single landing pad cell lines, because the cell line only accepts a single second GOI, all expression cassettes required to generate a multi-component biologic must be placed in a single second GOI plasmid. This is not the case with duo landing pad cell lines. Duo landing pad cell lines offer the opportunity to design for increased levels of expression diversity and create expression cell lines with superior properties. Diversity can be generated in multiple ways using different configurations of the second GOI plasmid. In one example, the second GOI plasmid contains all expression cassettes required to generate a combination biologic in a unique configuration. In a second example, the second GOI plasmid may contain a subset of the expression cassettes that need to be present in the same cell to generate the expression cell line. A third example is a combination of the previous two examples, where one or more second GOI plasmids have all of the expression cassettes in a unique configuration required to make a combination biologic, together with a set of second GOI plasmids containing a subset of all of the expression cassettes in a unique configuration.
[0254] It is understood that the same method for generating duo-landing pads disclosed herein can be used to generate cell lines containing higher combinations of landing pad plasmids. For example, the method disclosed herein for identifying landing pad cell lines containing two landing pad plasmids in a hotspot can be used to select landing pad cell lines with three, four, or more landing pad plasmids. Landing pad cell lines and expression cells with hotspots containing three or more landing pad plasmids can be used, for example, to produce biologics containing three or more different subunits.
[0255] In some embodiments, a duo landing pad configuration can include both landing pad plasmids with the same recombinase or Int recognition sequence, but each landing pad plasmid can have a unique recombination "address," i.e., each landing pad plasmid can be addressable. For recombinases such as Cre and Flp, four unique recognition sequences can be used. Thus, each landing pad plasmid will have a unique pair of recognition sites. In some embodiments, four incompatible Lox sites can be used. Langer, SJ, Ghafoori, AP, Byrd, M. and Leinwand, L. (2002) A genetic screen identifies novel non-compatible loxP sites. Nucleic Acids Res., 30, 3067-3077, Missirlis, PI, Smailus, DE and Holt, RA (2006) A high-throughput screen identifying sequence and promiscuity characteristics of the loxP spacer region in See Cre-mediated recombination. BMC Genomics, 7, 73. and Siegel, RW, Jain, R. and Bradbury, A. (2001) Using an in vivo phagemid system to identify non-compatible loxP sequences. FEBS Lett., 505, 467-473.
[0256] Additional strategies include replacing two Lox sites with two incompatible Frt sites and using Cre with Frt [see Lauth, M., Spreafico, F., Dethleffsen, K. and Meyer, M. (2002) Stable and efficient cassette exchange under non-selective conditions by combined use of two site-specific recombinases. Nucleic Acids Res., 30, e115], and using integrase with two to four incompatible aat sites [Jusiak, B., Jagtap, K., Gaidukov, L., Duportet, X., Bandara, K., Chu, J., Zhang, L., Weiss, R. and Lu, TK (2019) Comparison of Integrases Identifies Bxb1-GA Mutant as the Most Efficient Site-Specific Integrase System in Mammalian Cells. ACS Synth Biol, 8, 16-24], the use of two or more integrases, such as BxB1 and phiC3 [see Smith, MC, Brown, WR, McEwan, AR and Rowley, PA (2010) Site-specific recombination by phiC31 integrase and other large serine recombinases. Biochem. Soc. Trans., 38, 388-394], and combinations thereof. The use of a single att site in each landing pad is sufficient for insertion of the second GOI plasmid into each landing pad. In this case, the second GOI plasmid must be circular, since a linear plasmid would essentially restrict the chromosome.It is also clear that a landing pad can contain multiple att sites, each containing a unique address.
[0257] The duo landing pad configuration, using landing pads with unique addresses, can also be used to generate a more defined diversity of expressing cell lines compared to when they are not addressable, and a higher diversity relative to landing pad cell lines containing a single landing pad.
[0258] An additional use of addressable landing pads is the option to express two independent biologics, each with a specific independent function: one of the biologics may support the expression of the second biologic by the expressing cell line, or the first biologic may cause a specific post-translational modification of the second biologic, or may modify some other component of the expressing cell line.
[0259] In some embodiments, the methods, cells, cell lines, or kits disclosed herein comprise at least two landing pad plasmids or at least two expression plasmids in tandem. In other words, in some embodiments, CG / -([P1]-([P2]-[SSRS]-[M]-[SSRS]-[P2])n-[P1])- / CG, CG / -([P2]-[SSRS]-[M]-[SSRS]-[P2])n- / CG, CG / -([P1]-([P2]-[SSRS]-[M]-[P2])n-[P1])- / CG, CG / -([P1]-([P2]-[M]-[SSRS]-[P2])n-[P1])- / CG, CG / -([P2]-[SSRS]-[M]-[P2])n- / CG, CG / -([P2]-[M]-[SSRS]-[P2])n- / CG, CG / -([P1]-([P2]-[SSRS]-[P3]-[SSRS]-[P2])n-[P1])- / CG, CG / -([P2]-[SSRS]-[P3]-[SSRS]-[P2])n- / CG, CG / -([P1]-([P2]-[SSRS]-[P3]-[P2])n-[P1])- / CG, CG / -([P1]-([P2]-[P3]-[SSRS]-[P2])n-[P1])- / CG, CG / -([P2]-[SSRS]-[P3]-[P2])n- / CG, or CG / -([P2]-[P3]-[SSRS]-[P2])n- / CG, Or in any other formula containing a value of n disclosed herein, the value of n can be 2 or greater. In some specific embodiments, n is 2. Thus, in some embodiments, at least two landing pad plasmids or at least two expression plasmids arranged in tandem are present in the constructs disclosed herein. In some embodiments, n is an integer such as 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some embodiments, n is greater than 10, for example, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30.
[0260] In some embodiments, the two landing pad plasmids or two expression plasmids are in a configuration selected from the group consisting of head-to-head, tail-to-tail, tail-to-head, and head-to-tail. In some embodiments, each expression plasmid contains at least a nucleic acid encoding a gene of interest (GOI). In some embodiments, all GOIs are the same. In some embodiments, all GOIs are different. In some embodiments, at least one GOI is different from the rest. In some embodiments, a first GOI is an antibody HC and a second GOI is an antibody LC. In some embodiments, at least one expression plasmid is bicistronic or polycistronic. In some embodiments, a bicistronic expression plasmid encodes a first GOI comprising an antibody HC and a second GOI comprising an antibody LC.
[0261] In some embodiments, each landing pad plasmid in the duo landing pad is addressable. In some embodiments, each addressable landing pad plasmid contains a pair of SSRSs, which can be unique or incompatible. In some embodiments, the landing pad plasmid contains two Lox sites. In some embodiments, the Lox sites are Lox P and Lox 511. In some embodiments, each landing pad plasmid contains a Lox site and an Frt site. In some embodiments, each landing pad plasmid contains one or two aat sites, for example, two aatP sites.
[0262] In some embodiments, each landing pad plasmid is addressable. In some embodiments, each addressable landing pad plasmid contains a pair of addressable SSRSs that are unique to the landing pad. In some embodiments, at least one pair of addressable SSRSs is a pair of Lox sites. In some embodiments, at least one pair of Lox sites is Lox 511 and Lox P. In some embodiments, at least one pair of Lox sites is Lox m3 and Lox m7.
[0263] In some embodiments, the methods, cell lines, cells, or kits of the disclosure comprise a first addressable landing pad plasmid comprising a pair of Lox sites, Lox 511 and Lox P, and a second addressable landing pad plasmid comprising a pair of Lox sites, Lox m3 and Lox m7. In some embodiments, each addressable landing pad plasmid comprises non-reciprocally compatible attP sites.
[0264] In some embodiments, the LoxP site is selected from the group consisting of SEQ ID NOs: 1-11 and 28-82, and any combination thereof. In some embodiments, the Frt site is selected from the group consisting of SEQ ID NOs: 12 and 83-91, and any combination thereof. In some embodiments, the addressable pad disclosed herein may comprise an SSRS, or a combination thereof, selected from the group consisting of SEQ ID NOs: 1-13 and 28-109, and any combination thereof.
[0265] In some embodiments, the att sites are selected from the group consisting of SEQ ID NOs: 92-109 and any combination thereof. In some embodiments, the pair of att sites comprises an attB site of SEQ ID NO: 92 and an attP site of SEQ ID NO: 93. In some embodiments, the pair of att sites comprises an attB site of SEQ ID NO: 94 and an attP site of SEQ ID NO: 95. In some embodiments, the pair of att sites comprises an attB site of SEQ ID NO: 96 and an attP site of SEQ ID NO: 97. In some embodiments, the pair of att sites comprises an attB site of SEQ ID NO: 98 and an attP site of SEQ ID NO: 99. In some embodiments, the pair of att sites comprises an attB site of SEQ ID NO: 100 and an attP site of SEQ ID NO: 101. In some embodiments, the pair of att sites comprises an attB site of SEQ ID NO: 102 and an attP site of SEQ ID NO: 103. In some embodiments, the pair of att sites comprises an attB site of SEQ ID NO: 104 and an attP site of SEQ ID NO: 105. In some embodiments, the pair of att sites comprises an attB site of SEQ ID NO: 106 and an attP site of SEQ ID NO: 107. In some embodiments, the pair of att sites comprises an attB site of SEQ ID NO:108 and an attP site of SEQ ID NO:109.
[0266] nuclease As used herein, the term "nuclease" refers to an enzyme that has the catalytic activity of DNA cleavage.
[0267] In some embodiments, the nuclease agent can promote homologous recombination between two plasmids, such as the linear plasmids disclosed herein, for example, a parent plasmid and a landing pad plasmid. In some embodiments, the plasmid (parent plasmid, P1) and the landing pad plasmid (P2) integrated into the genome of the parent cell line contain regions of homology, and in the parent plasmid integrated into the parent cell line, a sequence targeted by a nuclease such as a CRISPR / Cas nuclease is present next to each homologous region, but is not present in the landing pad plasmid recombined into the parent cell line.
[0268] The size of the recognition site for nuclease-mediated homologous recombination can vary, e.g., at least about 4, at least about 6, at least about 8, at least about 10, at least about 12, at least about 14, at least about 16, at least about 18, at least about 19, at least about 20, at least about 21, at least about 22, at least about 23, at least about 24, at least about 25, at least about 26, at least about 27, at least about 28, at least about 29, at least about 30, at least about 31, at least about 32, at least about 33, at least about 34, at least about 35, at least about 36, at least about 37, at least about 38, at least about 39, at least about 40, at least about 41, at least about 42, at least about 43, at least about 44, at least about 45, at least about 46, at least about 47, at least about 48, at least about 49, at least about 50, at least about 51, at least about 52, at least about 53, at least about 54, at least about 55, at least about 56, at least about 57, at least about 58, at least about 59, at least about 60, at least about 61, at least about 62, at least about 63, at least about 64, at least about 65, at least about 66, at least about 67, at least about 68, at least about 69, at least about 70, at least about 71, at least about 72, at least about 73, at least about 74, at least about 75, at least about 76, at least about 77, at least about 78, at least about 79, at least about 4, at least about 35, at least about 36, at least about 37, at least about 38, at least about 39, at least about 40, at least about 41, at least about 42, at least about 43, at least about 44, at least about 45, at least about 46, at least about 47, at least about 48, at least about 49, at least about 50, at least about 51, at least about 52, at least about 53, at least about 54, at least about 55, at least about 56, at least about 57, at least about 58, at least about 59, at least about 60, at least about 61, At least about 62, at least about 63, at least about 64, at least about 65, at least about 66, at least about 67, at least about 68, at least about 69, at least about 70, at least about 80, at least about 90, at least about 100, at least about 110, at least about 120, at least about 130, at least about 140, at least about 150, at least about 160, at least about 170, at least about 180, at least about 190, at least about 200, at least about 210, at least about 220, at least about 230, at least about 240, at least about at least about 250, at least about 260, at least about 270, at least about 280, at least about 290, at least about 300, at least about 310, at least about 320, at least about 330, at least about 340, at least about 350, at least about 360, at least about 370, at least about 380, at least about 390, at least about 400, at least about 410, at least about 420, at least about 430, at least about 440, at least about 450, at least about 460, at least about 470, at least about 480, at least about 490,at least about 500, at least about 510, at least about 520, at least about 530, at least about 540, at least about 550, at least about 560, at least about 570, at least about 580, at least about 590, at least about 600, at least about 610, at least about 620, at least about 630, at least about 640, at least about 650, at least about 660, at least about 670, at least about 680, at least about 690, at least about 700, at least about 710, at least about 720, at least about 730, at least about 740, at least about 750, at least about 760, at least about 770, at least about 780, at least about 790, at least about 800, at least about 810, at least about 820, at least about 830, at least about 840, at least about 850, at least about 860, at least about 870, at least about 880, and / or a recognition site that is at least about 890, at least about 900, at least about 910, at least about 920, at least about 930, at least about 940, at least about 950, at least about 960, at least about 970, at least about 980, at least about 990, at least about 1000, at least about 1010, at least about 1020, at least about 1030, at least about 1040, at least about 1050, at least about 1060, at least about 1070, at least about 1080, at least about 1090, at least about 1100, at least about 1110, at least about 1120, at least about 1130, at least about 1140, at least about 1150, at least about 1160, at least about 1170, at least about 1180, at least about 1190, at least about 1200, at least about 2010, at least about 2020, or more nucleotides in length.
[0269] The size of the recognition site for nuclease-mediated homologous recombination can vary, e.g., about 4, about 6, about 8, about 10, about 12, about 14, about 16, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58 , about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 80, about 90, about 100, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210, about 220, about 230, about 240, about 250, about 260, about 270, about 280, about 290, about 300, about 310, about 320, about 330, about 340, about 350, about 360, about 370, about 380, about 390, about 400, about 410, about 420, about 43 0, about 440, about 450, about 460, about 470, about 480, about 490, about 500, about 510, about 520, about 530, about 540, about 550, about 560, about 570, about 580, about 590, about 600, about 610, about 620, about 630, about 640, about 650, about 660, about 670, about 680, about 690, about 700, about 710, about 720, about 730, about 740, about 750, about 760, about 770, about 780, about 790, about 800, about 810, about 820, about 830, about 840, about 850, about 860, about 870, about 88 The present invention also includes recognition sites that are 0, about 890, about 900, about 910, about 920, about 930, about 940, about 950, about 960, about 970, about 980, about 990, about 1000, about 1010, about 1020, about 1030, about 1040, about 1050, about 1060, about 1070, about 1080, about 1090, about 1100, about 1110, about 1120, about 1130, about 1140, about 1150, about 1160, about 1170, about 1180, about 1190, about 1200, about 2010, about 2020, or more nucleotides in length.
[0270] The size of the recognition sites for nuclease-mediated homologous recombination can vary, for example, from about 4 to about 10, from about 10 to about 20, from about 20 to about 30, from about 30 to about 40, from about 40 to about 50, from about 50 to about 60, from about 60 to about 70, from about 70 to about 80, from about 80 to about 90, from about 90 to about 100, from about 100 to about 125, from about 125 to about 150, from about 150 to about 175, from about 175 to about 200, from about 200 to about 300, and the like. to about 225, about 225 to about 250, about 250 to about 275, about 275 to about 300, about 300 to about 325, about 325 to about 350, about 350 to about 375, about 375 to about 400, about 400 to about 425, about 425 to about 450, about 450 to about 475, about 475 to about 500, about 500 to about 525, about 525 to about 550, about 550 to about 575, about 575 to about 600, about 600 to about 625, about 625 to about 650, about 650 to about 675, about 675 to about 700, about 700 to about 725, about 725 to about 750, about 750 to about 775, about 775 to about 800, about 800 to about 825, about 825 to about 850, about 850 to about 875, about 875 to about 900, about 900 to about 925, about 925 to about 950, about 950 to about 975, about 975 to about 1000, about 1 The present invention also includes recognition sites that are 000 to about 1100, about 1100 to about 1200, about 1200 to about 1300, about 1300 to about 1400, about 1400 to about 1500, about 1500 to about 1600, about 1600 to about 1700, about 1700 to about 1800, about 1800 to about 1900, about 1900 to about 2000, or about 2000 to about 2100, or more, nucleotides in length.
[0271] In one embodiment, each monomer of a nuclease agent recognizes a recognition site of at least 9 nucleotides. In other embodiments, the recognition site is about 9 to about 12 nucleotides in length, about 12 to about 15 nucleotides in length, about 15 to about 18 nucleotides in length, or about 18 to about 21 nucleotides in length, and any combination of such subranges (e.g., 9 to 18 nucleotides). The recognition site can be a batch structure, i.e., reading a sequence on one strand can be the same as reading it in the opposite direction on the complementary strand. It is recognized that a given nuclease agent can bind to and cleave the recognition site, or alternatively, the nuclease agent can bind to a sequence different from the recognition site. Furthermore, the term recognition site includes both the nuclease agent binding site and the nick / cleavage site, regardless of whether the nick / cleavage site is within or outside the nuclease agent binding site. In another variation, cleavage by a nuclease agent may occur at nucleotide positions immediately opposite each other, resulting in a blunt-ended cut, or in other cases, the cuts may be offset, resulting in a single-stranded overhang, also called a "sticky end," which may be either a 5' overhang or a 3' overhang.
[0272] In some embodiments, one of the sequences in the parent plasmid that is absent in the landing pad plasmid is SEQ ID NO:14 and the other is SEQ ID NO:15.
[0273] In some embodiments, one of the sequences in the parent plasmid that is absent in the landing pad plasmid is SEQ ID NO:14.
[0274] In some embodiments, one of the sequences in the parent plasmid that is absent in the landing pad plasmid is SEQ ID NO:15.
[0275] The methods and compositions disclosed herein may use any nuclease agent that induces a nick or double-strand break at the desired recognition site. Naturally occurring or native nuclease agents can be used as long as the nuclease agent induces a nick or double-strand break at the desired recognition site. Alternatively, modified or engineered nuclease agents can be used. "Engineered nuclease agents" include nucleases engineered (modified or derived) from natural forms to specifically recognize a desired recognition site and induce a nick or double-strand break there. Thus, engineered nuclease agents may be derived from natural, naturally occurring nuclease agents or may be artificially created or synthesized. The modification of a nuclease agent may be as little as one amino acid in a protein cleavage agent or one nucleotide in a nucleic acid cleavage agent. In some embodiments, an engineered nuclease induces a nick or double-strand break at a recognition site that is not the sequence that would be recognized by a natural (unengineered or unmodified) nuclease agent. Creating a nick or double-strand break in the recognition site or other DNA may be referred to herein as "cutting" or "cleaving" the recognition site or other DNA.
[0276] Homologous recombination system In some embodiments of the present disclosure, the homologous recombination is mediated by a CRISPR / Cas system, a TALEN system, a ZFN system, a meganuclease, or a restriction endonuclease.
[0277] CRISPR / Cas In some embodiments, the nuclease agent used for homologous recombination in the various methods and compositions disclosed herein may comprise a CRISPR / Cas system. It should be noted that the depiction of CRISPR / Cas in the figures as the "default" homologous recombination system is merely exemplary, and the process diagrammed in the figures may be performed using alternative homologous recombination systems, such as TALEN systems, ZFN systems, meganucleases, or restriction endonucleases. Such CRISPR / Cas systems may, for example, use Cas9 nuclease, which in some cases is codon-optimized for the desired cell type in which it is expressed. Such systems may also use guide RNAs (gRNAs) comprising two separate molecules. An exemplary bimolecular gRNA comprises a crRNA-like ("CRISPR RNA" or "targeter RNA" or "crRNA" or "crRNA repeat") molecule and a corresponding tracrRNA-like ("trans-acting CRISPR RNA" or "activator RNA" or "tracrRNA" or "scaffold") molecule.
[0278] The crRNA contains both the DNA-targeting segment (single-stranded) of the gRNA and a stretch of nucleotides that forms one half of the double-stranded RNA (dsRNA) duplex of the protein-binding segment of the gRNA. The corresponding tracrRNA (activator RNA) contains a stretch of nucleotides that forms the other half of the dsRNA duplex of the protein-binding segment of the gRNA. Thus, the stretch of nucleotides in the crRNA is complementary to and hybridizes with the stretch of nucleotides in the tracrRNA to form the dsRNA duplex of the protein-binding domain of the gRNA. Thus, each crRNA can be said to have a corresponding tracrRNA. The crRNA also provides a single-stranded DNA-targeting segment. Thus, the gRNA contains a sequence that hybridizes with the target sequence and the tracrRNA. Thus, the crRNA and tracrRNA (as a corresponding pair) hybridize to form the gRNA. When used for intracellular modification, the exact sequence and / or length of a given crRNA or tracrRNA molecule can be designed to be specific to the species in which the RNA molecule is used.
[0279] The naturally occurring genes encoding the three elements (Cas9, tracrRNA, and crRNA) are typically organized as an operon. Naturally occurring CRISPR RNAs vary depending on the Cas9 system and organism, but often contain a targeting segment of 21 to 72 nucleotides in length flanked by two direct repeats (DRs) of 21 to 46 nucleotides in length (see, for example, WO2014 / 131833). In the case of S. pyogenes, the DRs are 36 nucleotides in length and the targeting segment is 30 nucleotides in length. The 3'-located DRs are complementary to and hybridize with the corresponding tracrRNA, allowing the tracrRNA to bind to the Cas9 protein.
[0280] Alternatively, this system also uses a fused crRNA-tracrRNA construct (i.e., a single transcript) that functions with a codon-optimized Cas9. This single RNA is often referred to as the guide RNA or gRNA. Within the gRNA, the crRNA portion is identified as the "target sequence" for a given recognition site, and the tracrRNA is often referred to as the "scaffold." Briefly, a short DNA fragment containing the target sequence is inserted into a guide RNA expression plasmid. The gRNA expression plasmid contains the target sequence (approximately 20 nucleotides in some embodiments), a form of tracrRNA sequence (scaffold), and a suitable promoter active in the cell and the necessary elements for proper processing in eukaryotic cells. Many of these systems rely on custom-designed complementary oligonucleotides that are annealed to form double-stranded DNA and then cloned into the gRNA expression plasmid.
[0281] Then gRNA expression cassette and Cas9 expression cassette are introduced into cell.See for example, Mali P et al. (2013) Science 2013 Feb.15;339(6121):823-6, Jinek M et al. Science 2012 Aug.17;337(6096):816-21, Hwang WY et al. Nat Biotechnol 2013 March;31(3):227-9, Jiang W et al. Nat Biotechnol 2013 March;31(3):233-9 and Cong L et al. Science 2013 Feb.15;339(6121):819-23.Each of these is incorporated herein by reference. See also, e.g., WO / 2013 / 176772A1, WO / 2014 / 065596A1, WO / 2014 / 089290A1, WO / 2014 / 093622A2, WO / 2014 / 099750A2, and WO / 2013142578A1, each of which is incorporated herein by reference.
[0282] In some embodiments, the Cas9 nuclease may be provided in the form of a protein. In some embodiments, the Cas9 protein may be provided in the form of a complex with a gRNA. In other embodiments, the Cas9 nuclease may be provided in the form of a nucleic acid encoding the protein. The nucleic acid encoding the Cas9 nuclease may be RNA [e.g., messenger RNA (mRNA)] or DNA. In some embodiments, the gRNA may be provided in the form of RNA. In other embodiments, the gRNA may be provided in the form of DNA encoding the RNA. In some embodiments, the gRNA may be provided in the form of separate crRNA and tracrRNA molecules, or separate DNA molecules encoding the crRNA and tracrRNA, respectively.
[0283] In one embodiment, the method for generating a landing pad cell disclosed herein further comprises introducing into the cell: (a) a first expression construct comprising a first promoter operably linked to a first nucleic acid sequence encoding a CRISPR-associated (Cas) protein; and (b) a second expression construct comprising a second promoter operably linked to a genomic target sequence linked to a guide RNA (gRNA), wherein the genomic target sequence is flanked by protospacer adjacent motifs. The genomic target sequence may be flanked at its 3' end by a protospacer adjacent motif (PAM) sequence.
[0284] In some embodiments, the gRNA comprises a third nucleic acid sequence encoding a CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) RNA (crRNA) and a transactivating CRISPR RNA (tracrRNA). In one embodiment, the Cas protein is a type I Cas protein. In one embodiment, the Cas protein is a type II Cas protein. In one embodiment, the type II Cas protein is Cas9. In one embodiment, the type II Cas, for example, Cas9, is a human codon-optimized Cas.
[0285] In certain embodiments, the Cas protein is a "nickase" that can create a single-strand break (i.e., a "nick") at a target site without cutting both strands of double-stranded DNA (dsDNA). Cas9, for example, contains two nuclease domains, a RuvC-like nuclease domain and an HNH-like nuclease domain, which are involved in cleaving opposing DNA strands. Mutations in either of these domains can create a nickase. Examples of nickase-creating mutations can be found, for example, in WO / 2013 / 176772A1 and WO / 2013 / 142578A1, each of which is incorporated herein by reference.
[0286] In certain embodiments, two separate Cas proteins (e.g., nickases) specific to target sites on each strand of dsDNA can create overhang sequences in another nucleic acid, or overhang sequences complementary to separate regions on the same nucleic acid. The overhang ends created by contacting a nucleic acid with two nickases specific to target sites on both strands of dsDNA can be either 5' or 3' overhang ends. For example, to create an overhang sequence, a first nickase can create a single-strand break in the first strand of dsDNA, and a second nickase can create a single-strand break in the second strand of dsDNA. The target sites of each nickase that create a single-strand break can be selected so that the created overhang end sequence is complementary to the overhang end sequence on a different nucleic acid molecule. The complementary overhang ends of two different nucleic acid molecules can be annealed by the methods disclosed herein. In some embodiments, the target site for the nickase on the first strand is different from the target site for the nickase on the second strand.
[0287] In some embodiments, the first nucleic acid comprises a mutation that destroys at least one amino acid residue in the nuclease active site in the Cas protein, and the mutant Cas protein generates a break in only one strand of the target DNA region, and the mutation reduces non-homologous recombination in the target DNA region. In one embodiment, the first nucleic acid encoding the Cas protein further comprises a nuclear localization signal (NLS). In one embodiment, the nuclear localization signal is an SV40 nuclear localization signal.
[0288] TALEN In some embodiments, the nuclease agent used for homologous recombination in various methods and compositions disclosed herein can comprise TALEN system.Thus, in one embodiment, the nuclease agent is transcription activator-like effector nuclease (TALEN).TAL effector nuclease is a class of sequence-specific nucleases that can be used to create double-strand breaks in specific target sequences in the genome of prokaryotes or eukaryotes.TAL effector nuclease is produced by fusing natural or engineered transcription activator-like (TAL) effector or its functional part with the catalytic domain of endonuclease, for example, FokI.
[0289] The unique modular TAL effector DNA-binding domain a...
Claims
1. A method for selecting parent cells suitable for generating landing pad cell lines, comprising: (i) screening and selecting cell lines with high expression titers of the gene of interest (GOI) located on the parental plasmid integrated into the genomic sequence of the cell line; (ii) further screening the cells of (i) to select cells with a low copy number of the parental plasmid comprising the nucleic acid encoding the GOI, wherein the copy number is 1 or 2; (iii) deleting the parental plasmid containing the GOI or a portion thereof; and (iv) introducing a landing pad plasmid or a portion thereof containing a landing pad into a cell; A method comprising:
2. 2. The method of claim 1, wherein the parental plasmid contains two site-specific recombination sites (SSRS), one SSRS, or no SSRS. (v) screening for loss of the parental plasmid or a portion thereof in the parent cell line and selecting cells with such loss (deletion); and (vi) further screening the cells of (v) for the presence of a landing pad and selecting cells in which a landing pad is present.
3. The method of claim 1 or 2, further comprising:
4. The landing pad in the landing pad cell, (a) the presence or absence of regions of low or high complexity; (b) the presence or absence of retrotransposon sequences; (c) the presence or absence of Alu repeats; (d) the presence or absence of long-chain interspersed nuclear elements (LINEs); (e) the presence or absence of CpG islands; (f) the level of cytosine methylation; (g) the level of histone acetylation; (h) the presence or absence of active transcription, and (i) any combination thereof 4. The method of claim 3, further comprising screening for a feature in the polynucleotide sequence selected from the group consisting of:
5. 2. The method of claim 1, wherein a landing pad plasmid or a portion thereof comprising a landing pad is inserted at the site of the deletion of step (iii) of claim 1.
6. 2. The method of claim 1, wherein the landing pad plasmid or a portion thereof comprising the landing pad is inserted at a site that is not the site of the deletion of step (iii) of claim 1.
7. 1. A method for generating landing pad cells, comprising: using homologous recombination to integrate the landing pad plasmid into the genome of the parent cell at a targeted integration site, wherein the sequence targeted for homologous recombination is located in the parent plasmid; Each landing pad plasmid is (1) a polynucleotide sequence comprising a nucleic acid encoding at least one selectable marker and / or at least one nucleic acid sequence encoding a detectable marker; (2) two site-specific recombination sites (SSRS) flanking the polynucleotide sequence of (1); and (3) two homologous recombination sites located 5' and 3' to the SSRS of (2), which are homologous to corresponding homologous recombination sites in the parental plasmid; Including, A method in which a homologous recombination site in the landing pad plasmid recombines with a corresponding homologous recombination site in the parental plasmid, thereby integrating the landing pad plasmid at a location within the parental plasmid that was inserted into the parental cell genomic DNA.
8. 10. The method of claim 1 or 7, wherein the parental plasmids are located at two or more genomic loci.
9. 1. A method for identifying landing pad cells, comprising: (1) removing at least a portion of the first GOI from the parental plasmid integrated into the genomic sequence of the parental cell; (2) integrating the landing pad plasmid at an alternative genomic locus; (3) screening a library of candidate cells containing at least one copy of the landing pad plasmid integrated at at least one alternative genomic locus, wherein the candidate cell lines exhibit the following characteristics: (a) cell titer above a predetermined threshold level; (b) the copy number of the landing pad plasmid or landing pad is a predetermined value; (c) an RNA expression level above a predetermined threshold level; (d) multiple plasmid copies, if present, have a particular plasmid organization; (e) deletion of at least a portion of the first GOI from the parental plasmid; and (f) the presence of at least one landing pad containing a functional SSRS screening, which is assessed for one or more of A method comprising:
10. 10. The method of claim 9, wherein the library of candidate cells is a library generated by random integration of landing pad sequences at multiple locations within the genome of the parent cell.
11. The method of claim 9 or 10, wherein hot cells containing landing pad sequences integrated into hot spots are selected.
12. 10. The method of claim 1, 7, or 9, wherein the parent cell line is a CHO cell line.
13. 10. A method of generating an expression cell, comprising integrating a second GOI plasmid into the genome of a landing pad cell of claim 1, 7, or 9 using site-specific recombinase recombination, wherein the resulting expression plasmid comprises: (1) a polynucleotide sequence comprising a nucleic acid encoding a second GOI; and (2) two SSRSs flanking the polynucleotide of (1); Including, A method in which a site-specific recombination site in the landing pad plasmid recombines with a corresponding site-specific recombination site in a second GOI plasmid, thereby integrating the GOI plasmid at a location internal to the landing pad plasmid in the landing pad cell.
14. 1. A method for generating an expressing cell, comprising: (a) using homologous recombination to integrate a landing pad plasmid or a portion thereof into the genome of a parent cell at a targeted integration site, wherein the sequence targeted for homologous recombination is located in the parent plasmid; Each landing pad plasmid or a portion thereof is (1a) a polynucleotide sequence comprising a nucleic acid encoding at least one selectable marker and / or at least one nucleic acid sequence encoding a detectable marker; (2a) two SSRSs flanking the polynucleotide sequence of (1a); and (3a) two homologous recombination sites located 5' and 3' to the SSRS of (2a), which are homologous to corresponding homologous recombination sites in the parental plasmid; Including, a homologous recombination site of the landing pad plasmid or a portion thereof recombines with a corresponding homologous recombination site of the parental plasmid within the landing pad at a different genomic locus, thereby integrating the landing pad plasmid or a portion thereof at a location internal to the landing pad at a different genomic locus in the parental cell genomic DNA; and (b) integrating a second GOI plasmid into the genome of the landing pad cell using site-specific recombinase recombination, wherein the expression plasmid comprises: (1b) a polynucleotide sequence comprising a nucleic acid encoding a GOI; and (2b) two SSRSs flanking the polynucleotide of (1b); and Including, The site-specific recombination site of the landing pad plasmid recombines with the corresponding site-specific recombination site of the GOI plasmid, thereby integrating the GOI plasmid at a location within the landing pad plasmid in the landing pad cell. A method comprising:
15. 1. A method for generating landing pad cells, comprising: (a) removing at least a portion of the parental plasmid from a first hotspot location in the parental cell line; and (b) integrating the landing pad plasmid into a second hotspot location within the genome of the parent cell at a targeted integration site using homologous recombination or random integration, wherein the sequence targeted for homologous recombination or random integration was present in the landing pad plasmid; Each landing pad plasmid is (1) a polynucleotide sequence comprising a nucleic acid encoding at least one selectable marker and / or at least one nucleic acid sequence encoding a detectable marker; (2) two site-specific recombination sites (SSRS) flanking the polynucleotide sequence of (1); and (3) two homologous recombination sites located 5' and 3' to the SSRS of (2), which are homologous to corresponding homologous recombination sites in the parent cell line genome; to include, to incorporate A method comprising:
16. 1. A method for generating an expressing cell, comprising: (a) removing the parental plasmid or a portion thereof from a first hotspot location in the parental cell line; (b) using homologous recombination to integrate the landing pad plasmid into a second hotspot location within the genome of the parent cell at a targeted integration site, wherein the sequence targeted for homologous recombination was present in the parent cell line; Each landing pad plasmid is (1b) a polynucleotide sequence comprising a nucleic acid encoding at least one selectable marker and / or at least one nucleic acid sequence encoding a detectable marker; (2b) two site-specific recombination sites (SSRS) flanking the polynucleotide sequence of (1b); and (3b) two homologous recombination sites located 5' and 3' to the SSRS of (2b), which are homologous to corresponding homologous recombination sites in the parent cell line; Including, The homologous recombination sites of the landing pad plasmid recombine with corresponding homologous recombination sites of the parent cell line, thereby integrating the landing pad plasmid at an internal location in the parent cell genomic DNA; and (c) integrating the GOI plasmid into the genome of the landing pad cell using site-specific recombinase recombination, wherein the expression plasmid comprises: (1c) a polynucleotide sequence comprising a nucleic acid encoding a first GOI; and (2c) two SSRSs flanking the polynucleotide of (1c); and Including, The site-specific recombination site of the landing pad plasmid recombines with the corresponding site-specific recombination site of the GOI plasmid, thereby integrating the GOI plasmid at a location within the landing pad plasmid in the landing pad cell. A method comprising:
17. Landing pad cells described CG 1 / -[P1]-[P2]-[SSRS]-[M]-[SSRS]-[P2]-[P1]- / CG 2 、 CG 1 / -[P1]-([P2]-[SSRS]-[M]-[SSRS]-[P2])n-[P1]- / CG 2 、 CG 1 / -([P2]-[SSRS]-[M]-[SSRS]-[P2])n- / CG 2 、 CG 1 / -[P1]-([P2]-[SSRS]-[M]-[P2])n-[P1]- / CG 2 、 CG 1 / -([P2]-[SSRS]-[M]-[P2])n- / CG 2 、 CG 1 / -([P1]-([P2]-[SSRS]-[P3]-[SSRS]-[P2])n-[P1])- / CG 2 、 CG 1 / -([P2]-[SSRS]-[P3]-[SSRS]-[P2])n- / CG 2 、 CG 1 / -([P1]-([P2]-[SSRS]-[P3]-[P2])n-[P1])- / CG 2 、 CG 1 / -([P2]-[SSRS]-[P3]-[P2])n- / CG 2 、 CG 1 / -[P1]-([P2]-[M]-[SSRS]-[P2])n-[P1]- / CG 2 、 CG 1 / -([P2]-[M]-[SSRS]-[P2])n- / CG 2 、 CG 1 / -[P1]-[P2]-[P1]- / CG 2 、 CG 1 / -[P1]- / CG 2 、 CG 1 / -([P1]-([P2])n-[P1])- / CG 2、 -[P1]-[P2]-[P1]-, or -([P1]-([P2])n-[P1])- [where CG 1 and C.G. 2 is the parent cell genomic sequence flanking the inserted plasmid, [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [M] is a polynucleotide sequence comprising at least one nucleic acid sequence encoding at least one selectable and / or detectable marker; [SSRS] is a site-specific recombination site (SSRS), and n is an integer between 1 and 10, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10; 17. The method of claim 1, 7, 9, 14, 15, or 16.
18. The topology of the plasmid to be integrated into the expression cells is described CG 1 / -[P1]-([P2]-[SSRS]-[P3]-[SSRS]-[P2])n-[P1]- / CG 2 、 CG 1 / -([P2]-[SSRS]-[P3]-[SSRS]-[P2])n- / CG 2 、 CG 1 / -[P1]-([P2]-[SSRS]-[P3]-[P2])n-[P1]- / CG 2 、 CG 1 / -([P2]-[SSRS]-[P3]-[P2])n- / CG 2 、 CG 1 / -[P1]-([P2]-[P3]-[=SSRS]-[P2])n-[P1]- / CG 2、 CG 1 / -([P2]-[P3]-[SSRS]-[P2])n- / CG 2 [where CG 1 and C.G. 2 is the parent cell genomic sequence flanking the inserted plasmid, [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [P3] is a polynucleotide sequence derived from a plasmid containing a gene of interest (GOI), [SSRS] is a site-specific recombination site (SSRS), n is an integer between 1 and 10, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10.
19. The topology of the plasmid to be integrated into the expression cells is described CG 1 / -[P1]-([P2]-[SSRS]-[P3]-[SSRS]-[P2])n-[P1]- / CG 2 、 CG 1 / -([P2]-[SSRS]-[P3]-[SSRS]-[P2])n- / CG 2 、 CG 1 / -[P1]-([P2]-[SSRS]-[P3]-[P2])n-[P1]- / CG 2 、 CG 1 / -([P2]-[SSRS]-[P3]-[P2])n- / CG 2 、 CG 1 / -[P1]-([P2]-[P3]-[=SSRS]-[P2])n-[P1]- / CG 2、 CG 1 / -([P2]-[P3]-[SSRS]-[P2])n- / CG 2 [where CG 1 and C.G. 2 is the parent cell genomic sequence flanking the inserted plasmid, [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [P3] is a polynucleotide sequence derived from a plasmid containing a gene of interest (GOI), [SSRS] is a site-specific recombination site (SSRS), n is an integer between 1 and 10, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10.
20. The method of claim 7, 14, 15, or 16, wherein the homologous recombination is mediated by a CRISPR / Cas system, a TALEN system, or a ZFN system.
21. 17. The method of claim 2, 7, 14, 15, or 16, wherein the site-specific recombinase recombination site (SSRS) is a Tyr-recombinase site, a Tyr-integrase site, a serine-resolvase / invertase site, or a serine-integrase site.
22. 22. The method of claim 21, wherein the SSRS is a LoxP site.
23. 17. The method of claim 1, 9, 14, or 16, wherein the nucleic acid encoding the GOI encodes at least one polypeptide.
24. 24. The method of claim 23, wherein at least one polypeptide is an antibody or a fusion protein.
25. (i) the 5' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 18 or a subsequence thereof, and the 3' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 19 or a subsequence thereof; or (ii) the 5' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 114 or a subsequence thereof, and the 3' homologous recombination site comprises the polynucleotide sequence of SEQ ID NO: 115 or a subsequence thereof; or (iii) the 5' homologous recombination site and the 3' homologous recombination site comprise polynucleotide sequences flanking the parental plasmid; 17. The method of claim 7, 9, 14, or 16.
26. Description CG 1 / -[P1]-[P2]-[SSRS]-[M]-[SSRS]-[P2]-[P1]- / CG 2 、 CG 1 / -[P1]-([P2]-[SSRS]-[M]-[SSRS]-[P2])n-[P1]- / CG 2 、 CG 1 / -([P2]-[SSRS]-[M]-[SSRS]-[P2])n- / CG 2 、 CG 1 / -[P1]-([P2]-[SSRS]-[M]-[P2])n-[P1]- / CG 2 、 CG 1 / -([P2]-[SSRS]-[M]-[P2])n- / CG 2 、 CG 1 / -([P1]-([P2]-[SSRS]-[P3]-[SSRS]-[P2])n-[P1])- / CG 2 、 CG 1 / -([P2]-[SSRS]-[P3]-[SSRS]-[P2])n- / CG 2 、 CG 1 / -([P1]-([P2]-[SSRS]-[P3]-[P2])n-[P1])- / CG 2 、 CG 1 / -([P2]-[SSRS]-[P3]-[P2])n- / CG 2 、 CG 1 / -[P1]-([P2]-[M]-[SSRS]-[P2])n-[P1]- / CG 2 、 CG 1 / -([P2]-[M]-[SSRS]-[P2])n- / CG 2 、 CG 1 / -[P1]-[P2]-[P1]- / CG 2 、 CG 1 / -[P1]- / CG 2 、 CG 1 / -([P1]-([P2])n-[P1])- / CG 2、 -[P1]-[P2]-[P1]-, or -([P1]-([P2])n-[P1])- [where CG 1 and C.G. 2 is the parent cell genomic sequence flanking the inserted plasmid, [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [M] is a polynucleotide sequence comprising at least one nucleic acid sequence encoding at least one selectable and / or detectable marker; [SSRS] is a site-specific recombination site (SSRS), and n is an integer between 1 and 10, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.
27. description, CG 1 / -[P1]-([P2]-[SSRS]-[P3]-[SSRS]-[P2])n-[P1]- / CG 2 、 CG 1 / -([P2]-[SSRS]-[P3]-[SSRS]-[P2])n- / CG 2 、 CG 1 / -[P1]-([P2]-[SSRS]-[P3]-[P2])n-[P1]- / CG 2 、 CG 1 / -([P2]-[SSRS]-[P3]-[P2])n- / CG 2 、 CG 1 / -[P1]-([P2]-[P3]-[SSRS]-[P2])n-[P1]- / CG 2 ,or, CG 1 / -([P2]-[P3]-[SSRS]-[P2])n- / CG 2 [where CG 1 and C.G. 2 is the parent cell genomic sequence flanking the inserted plasmid, [P1] is a polynucleotide sequence derived from the parental plasmid, [P2] is a polynucleotide sequence derived from the landing pad plasmid, [P3] is a polynucleotide sequence derived from a plasmid containing a gene of interest (GOI), [SSRS] is a site-specific recombination site (SSRS), and n is an integer between 1 and 10, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.
28. 17. The method of claim 1, 7, 9, 14, 15, or 16, comprising at least two landing pad plasmids or at least two expression plasmids.
29. 29. The method of claim 28, wherein the two landing pad plasmids or the two expression plasmids are in a configuration selected from the group consisting of head-to-head, tail-to-tail, tail-to-head, and head-to-tail.
30. 28. The cell of claim 26 or 27, comprising at least two landing pad plasmids or at least two expression plasmids.
31. 31. The cell of claim 30, wherein the two landing pad plasmids or the two expression plasmids are in a configuration selected from the group consisting of head-to-head, tail-to-tail, tail-to-head, and head-to-tail.