LOCI with increased transgene expression and uses thereof

TRIP technology addresses the challenge of CHO cell line development by identifying transcriptional hotspots for targeted integration, achieving significantly higher mRNA and titer levels than traditional methods.

WO2026096425A1PCT designated stage Publication Date: 2026-05-07AMGEN INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
AMGEN INC
Filing Date
2025-10-28
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Current CHO cell line development for therapeutic protein production is laborious and lengthy due to cell-to-cell heterogeneity in transgene expression and dynamic epigenetic silencing, making it challenging to identify stable clones with high and consistent expression.

Method used

The method involves integrating transgenes into specific genomic loci using targeted integration (TI) at coordinates identified by the Thousands of Reporters Integrated in Parallel (TRIP) technology, which utilizes barcoded reporters to identify transcriptional hotspots in CHO cells, allowing for high and stable expression without the need for clonal isolation.

Benefits of technology

TRIP technology identifies loci with up to 9.4-fold increase in mRNA levels and 5.6-fold increase in fed-batch titers, surpassing traditional methods by providing consistent and high transgene expression, even with single copy integration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000013_0001
    Figure IMGF000013_0001
  • Figure IMGF000014_0001
    Figure IMGF000014_0001
  • Figure IMGF000015_0001
    Figure IMGF000015_0001
Patent Text Reader

Abstract

The present disclosure relates to methods of inserting and expressing a transgene, and uses thereof.
Need to check novelty before this filing date? Find Prior Art

Description

10988-W001-SEC Electronically Filed:. October 28, 2025LOCI WITH INCREASED TRANSGENE EXPRESSION AND USES THEREOFCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application No.63 / 713,315, filed October 29, 2024.DESCRIPTION OF THE TEXT FILE SUBMITTED ELECTRONICALLY

[0002] The present application contains a Sequence Listing, which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. The computer readable format copy of the Sequence Listing, which was created on October 28, 2025, is named 10988-W001-SEC_ST26.xml and is 5,428 bytes in size.BACKGROUND OF THE INVENTION

[0003] Chinese Hamster Ovary (CHO) cells are commonly used for producing therapeutic proteins in the biopharmaceutical industry (Walsh (2018) Biopharmaceutical benchmarks 2018, Nat Biotechnol 36, 1136-1145; and Wurm, F. M. (2004) Production of recombinant protein therapeutics in cultivated mammalian cells, Nat Biotechnol 22, 1393- 1398). Targeted integration (TI) of therapeutic protein-encoding transgenes into pre-determined high and stably expressing genomic loci can simplify the cell line development processes for biologies production. Establishing a successful TI system requires identifying genomic loci that allow high expression of the integrated transgenes.

[0004] The CHO cell line is a commonly used mammalian cell line in the biopharmaceutical industry for commercial production of therapeutic proteins. Classical stable CHO mammalian expression systems producing therapeutic proteins are dependent on random or transposase-mediated genomic integration of transgenes coding for the therapeutic protein and a selectable marker. The stable pools arising from such approaches show cell-to-cell heterogeneity in expression of the transgenes. This heterogeneity in expression is mainly due to variability in where the transgenes integrate (also known as the “position-effect) and cell-to-cell variation in transgene copy number. Cell line development (CLD) processes that identify high- producing clones can overcome this issue; however, such high-producing clones can show a decrease in transgene expression over time due to potential dynamic epigenetic silencing or10988-W001-SEC copy number changes. Thus, CLD efforts for identifying stable clones are laborious and lengthy, often requiring an extensive screen of large numbers of single cell clones.SUMMARY OF THE INVENTION

[0005] The present disclosure provides a method of expressing a transgene, said method comprising integrating the transgene into a locus, wherein the locus has coordinates selected from the group consisting of: NC_048599.1 :42048280-44048302 (positive strand); NW_023276806.1 :86353406-88353428 (negative strand); NC_048595.1 : 14745469-16745491 (negative strand); and NC_048596.1 :52895116-54895138 (positive strand); wherein the coordinates are based on CriGri-PICRH-1.0 genome assembly. In an embodiment, the coordinates are: NC_048599.1 :43048280-43048302 (positive strand);NW_023276806.1 :87353406-87353428 (negative strand); NC_048595.1 : 15745469-15745491 (negative strand); or NC_048596.1 :53895116-53895138 (positive strand); wherein the coordinates are based on CriGri-PICRH-1.0 genome assembly. In an embodiment, the coordinates are NC_048599.1 :42048280-44048302 (positive strand). In an embodiment, the coordinates are NW_023276806.1 :86353406-88353428 (negative strand). In an embodiment, the coordinates are NC_048595.1 : 14745469-16745491 (negative strand). In an embodiment, the coordinates are NC_048596.1 :52895116-54895138 (positive strand). In an embodiment, the coordinates are NC_048599.1 :43048280-43048302 (positive strand). In an embodiment, the coordinates are NW_023276806.1 :87353406-87353428 (negative strand). In an embodiment, the coordinates are NC_048595.1 : 15745469-15745491 (negative strand). In an embodiment, the coordinates are NC 048596.1 :53895116-53895138 (positive strand). In an embodiment, the locus is in a CHO cell. In an embodiment, the locus is a CHO cell locus. In an embodiment, the locus is a CHO homologous sequence in a cell that is not a CHO cell. In an embodiment, the locus is a CHO homologous sequence in a HEK 293 cell.

[0006] The present disclosure provides a cell comprising a transgene integrated into a locus, wherein the locus has coordinates selected from the group consisting of:NC_048599.1 :42048280-44048302 (positive strand); NW_023276806.1 :86353406-88353428 (negative strand); NC_048595.1 : 14745469-16745491 (negative strand); andNC_048596.1 :52895116-54895138 (positive strand); wherein the coordinates are based on CriGri-PICRH-1.0 genome assembly. In an embodiment, the cell comprises a transgene integrated into a locus, wherein the locus has coordinates selected from the group consisting of: NC_048599.1 :43048280-43048302 (positive strand); NW_023276806.1 :87353406-8735342810988-W001-SEC(negative strand); NC_048595.1:15745469-15745491 (negative strand); andNC_048596.1 :53895116-53895138 (positive strand); wherein the coordinates are based on CriGri-PICRH-1.0 genome assembly. In an embodiment, the locus is in a CHO cell. In an embodiment, the locus is a CHO cell locus. In an embodiment, the locus is a CHO homologous sequence in a cell that is not a CHO cell. In an embodiment, the locus is a CHO homologous sequence in a HEK 293 cell.

[0007] The present disclosure provides a cell comprising a transgene integrated into a locus, wherein the locus comprises an nucleic acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 1. The present disclosure provides a cell comprising a transgene integrated into a locus, wherein the locus comprises an nucleic acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 2. The present disclosure provides a cell comprising a transgene integrated into a locus, wherein the locus comprises an nucleic acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 3. The present disclosure provides a cell comprising a transgene integrated into a locus, wherein the locus comprises an nucleic acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 4. The present disclosure provides a cell comprising a transgene integrated into a locus, wherein the locus comprises an nucleic acid sequence comprising SEQ ID NO: 1. The present disclosure provides a cell comprising a transgene integrated into a locus, wherein the locus comprises an nucleic acid sequence comprising SEQ ID NO: 2. The present disclosure provides a cell comprising a transgene integrated into a locus, wherein the locus comprises an nucleic acid sequence comprising SEQ ID NO: 3. The present disclosure provides a cell comprising a transgene integrated into a locus, wherein the locus comprises an nucleic acid sequence comprising SEQ ID NO: 4. The present disclosure provides a cell comprising a transgene integrated into a locus, wherein the locus comprises an nucleic acid sequence consisting of SEQ ID NO: 1. The present disclosure provides a cell comprising a transgene integrated into a locus, wherein the locus comprises an nucleic acid sequence consisting of SEQ ID NO: 2. The present disclosure provides a cell comprising a transgene integrated into a locus, wherein the locus comprises an nucleic acid sequence consisting of SEQ ID NO: 3. The present disclosure provides a cell comprising a transgene integrated into a locus, wherein the10988-W001-SEC locus comprises an nucleic acid sequence consisting of SEQ ID NO: 4. In an embodiment, the locus is in a HEK 293 cell. In an embodiment, the locus is a HEK 293 cell locus.

[0008] In an embodiment, one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, or fifteen copies of the transgene is expressed in the locus of the present disclosure. In an embodiment, one copy of the transgene is expressed in the locus. In an embodiment, ten copies of the transgene is expressed in the locus.

[0009] The present disclosure provides a transgene expressed by a method of the present disclosure or a cell of the present disclosure.

[0010] In an embodiment, the cell is a CHO cell.

[0011] The present disclosure provides a transgene expressed by a method or cell of the present disclosure.

[0012] In an embodiment, the transgene is integrated into the locus by homologous recombination, gene editing or recombinase-mediated recombination.

[0013] In an embodiment, the transgene the transgene is integrated into the locus by CRISPR or TALEN.

[0014] In an embodiment, the transgene is integrated into the locus by recombinase- mediated recombination with BxPl or CRE-Lox.BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1A, Figure IB, Figure 1C, and Figure ID. TRIP Results in CHOCells. (A) Cloning of the barcoded reporter library. EGFP coding sequence was PCR amplified with a reverse primer containing 16 nt random barcode sequences. Both forward and reverse primers also included Type IIS enzyme BsmBI sites and the necessary overhangs for cloning. The resulting amplicon was cloned into the destination vector via Golden Gate reaction using BsmBI enzyme. In the final library, GS (glutamine synthetase) selection marker and the barcoded EGFP cassette were flanked by inverted terminal repeats (ITRs). pA: polyadenylation signal. (B) TRIP overview. Barcoded reporter library was cotransfected with a piggyBac (PB) transposase plasmid. TRIP cell pools were created after GS selection, and these pools were split into two technical replicates ten days before genomic DNA (gDNA) and total RNA collection. Complementary DNA (cDNA) was synthesized from the RNA. Targeted next-generation sequencing (NGS) of the reporter barcodes in the gDNA and cDNA was used for calculating normalized expression of the barcode from each TRIP locus. Barcode integration sites were mapped by restriction10988-W001-SEC digestion (using DpnII) of the gDNA and self-ligation, followed by inverse PCR (iPCR) of the circularized fragments and subsequent NGS. (C) Distribution of the normalized and log2 -transformed IR expression values (i.e. log2(IR expression)) and the correlation between two technical replicates are shown. The diagonal dashed line shows the identity line, and the diagonal solid line is the linear fit. R2: R-squared value. N indicates the total number of unique barcodes from the two TRIP pools. (D) Distribution of IR expression Z-scores of mapped TRIP loci is shown. Standard normal distribution with mean 0 and standard deviation (s.d.) 1 is overlayed as a comparison to the empirical distribution of Z-scores. TRIP loci were categorized into 5 groups of “high”, high shoulder”, “medium”, “low shoulder” and “low”, based on their corresponding IR expression Z-scores (See “Methods” section”). For “low”, “medium” and “high” groups, percentage of the TRIP loci that fall into these categories are shown at the bottom. The number of loci that fall into these categories are indicated in parenthesis.

[0016] Figure 2. Studying epigenetic contributors to high TRIP reporter expression. (A) Box plot showing distribution of ATAC-seq and CUT&RUN reads around TRIP loci (125 bp bidirectional base-pair extended from the IR integration site) in “low”, “medium” and “high” TRIP groups. Two-sample Wilcoxon test was applied. *p < 0.05, ** p < 0.01, *** p < 0.001, **** p < 0.0001. Data is shown in loglO scale and pseudo-count of 1 was added for visualization purposes.

[0017] Figure 3. Studying the effect of proximity to epigenetic feature peaks on TRIP reporter expression. (A) Distribution of the distance of TRIP loci to the nearest epigenomic peak for all 5 features analyzed. Two-sample Wilcoxon test was applied. *p < 0.05, ** p < 0.01, *** p < 0.001, **** p < 0.0001. Data is shown in loglO scale. Pseudocount of 1 was added for visualization purposes.

[0018] Figure 4A and Figure 4B. ChromHMM analysis in CHO cells. (A)ChromHMM chromatin states. Rows show the ChromHMM states (seven in total), columns show epigenetic features used in the model. Darker color indicates higher epigenetic feature frequency in each ChromHMM state. (B) Distribution of the distance of TRIP loci within each indicated TRIP group to the nearest “weak promoter” and “active promoter” ChromHMM state. Two-sample Wilcoxon test was applied. *p < 0.05, ** p < 0.01, *** p < 0.001, **** p < 0.0001. Data is shown in loglO scale. Pseudo-count of 1 was added for visualization purposes.10988-W001-SEC

[0019] Figure 5A and 5B. Validation of the top 4 TRIP loci and their comparison to a mid-level expressing TRIP locus. (A) Experimental design for screening candidate TRIP loci by CRISPR knock-in of a test Fc molecule. (B) Donor DNA copy number (determined by quantifying PuroR sequence) in selected clones. ACTB gene copy number was used for normalization. Clones 1-29, 2-10, 2-13, 3-32 and Ctrl-23 had more than one copy of the donor DNA; thus, these clones were excluded from any further analysis. Clones 2-02 and Ctrl-07 did not grow well and were also excluded from any further analysis. Error bars represent mean ± s.d. (n= 3 replicates). The table at the bottom lists the single knock-in clones that were chosen for expression analysis.

[0020] Figure 6A, Figure 6B, Figure 6C, Figure 6D, Figure 6E, Figure 6F, Figure 6G. Expression analysis of single knock-in clones. (A) mRNA quantification of Fc expression by ddPCR from single knock-in clones for each selected locus. Z-score associated with each TRIP locus is shown at the top. The mRNA expression level from each clone is shown relative to the mean expression from TRIP control clones. (B) Day- 13 titers from the 4 ml fed-batch culture of single knock-in clones for each selected locus. (C) Day- 13 Qp from the 4 ml fed-batch culture of single knock-in clones for each selected locus. For panels (A-C), the dashed line is the average estimate observed for the control clones. For data shown in panels (A-C), corresponding clone numbers are presented at the bottom. (D- G) 30 ml fed-batch culture for select clones (1-19, 2-09 and 2-17) and comparison of the titers to the PiggyBac (PB) protocol. (D) Titer, (E) Qp, (F) viable cell density and (G) viability over the course of the fed-batch process are shown. All experiments were performed in triplicate.DETAILED DESCRIPTION

[0021] Herein, we demonstrate that the Thousands of Reporters Integrated in Parallel (TRIP) technology can identify such transcriptional hotspots in CHO cells, and resulting loci that allow for increased transgene expression. TRIP simplifies screening for transcriptional hotspots since it utilizes randomly integrated barcoded reporters and uses each barcode as a unique identifier to track the genomic location and activity of the corresponding reporter by next-generation sequencing. Transcriptional hotspots identified by TRIP resulted in up to a 9.4-fold increase in mRNA levels and a 5.6-fold increase in fed-batch titers of a test molecule compared to a medium-expressing control locus. Moreover, single copy expression from one of the identified transcriptional hotspots resulted in up to a 1.6-fold higher titer10988-W001-SEC than the piggyBac-mediated stable expression of the same molecule. Reporter expression levels from TRIP loci showed positive correlations with active chromatin marks; however, proximity to active marks was not consistently deterministic. These results suggest TRIP is a powerful functional screening method for identifying transcriptional hotspots in CHO cells without the need for more complex epigenomic analyses.

[0022] To improve stable mammalian expression, targeted integration (Tl)-based expression systems may be used (Lee et al. (2019) Mitigating Clonal Variation in Recombinant Mammalian Cell Lines, Trends Biotechnol 37, 931-942; Huhtinen et al. (2023) Selection of biophysically favorable antibody variants using a modified Flp-In CHO mammalian display platform, Front Bioeng Biotechnol 11, 1170081; Zhou et al (2010) Generation of stable cell lines by site-specific integration of transgenes into engineered Chinese hamster ovary strains using an FLP-FRT system, J Biotechnol 147, 122-129; Zhang et al (2015) Recombinase- mediated cassette exchange (RMCE) for monoclonal antibody expression in the commercially relevant CHOK1SV cell line, Biotechnol Prog 31, 1645-1656). TI utilizes site-specific recombinases to integrate transgenes into specific loci. TI systems require creating mammalian expression hosts with landing pads (LPs), which contain site-specific recombination sequences, inserted (integrated) into pre-determined loci. Plasmids with the cognate recombination sequences and the desired transgene can then be integrated into the LPs when the necessary site-specific recombinase is expressed. To ensure high transgene expression via TI, it is important to identify genomic regions that are transcriptional hotspots, which allow high expression of the integrated transgenes and can be utilized as LP loci.

[0023] One approach for identifying transcriptional hotspots in the CHO genome relied on sequence homology to known safe harbor sites in other species such as the Hippl 19 and Rosa2610 loci. Alternative strategies focused on examining various CHO -omics data such as analysis of transcriptomics, or epigenomic analysis of active histone marks. Recently, 3D genome organization was also integrated into the analysis of the CHO epigenome. A detailed description of the epigenomic profiling methods can be found elsewhere. In one example study, differences in the 3D chromatin structure (assayed by high-throughput chromosome conformation capture, or Hi-C) and the transcriptome (assayed by RNA-seq) were discovered between a mAb-producing CHO-K1 cell line and its parental host, which helped identify epigenetically and transcriptionally conserved domains in the CHO genome as candidate hotspot regions (Hilliard et al (2021) Systematic identification of safe harbor regions in the CHO genome through a comprehensive epigenome analysis, Biotechnol Bioeng 118, 659-675).10988-W001-SEC

[0024] Based on the previously published histone mark (assayed by chromatin immunoprecipitation sequencing, or ChlP-seq) and DNA methylation data (assayed by bisulfite sequencing) in other CHO cell lines from the literature, the hotspot regions reported by Hillard et al. exhibited the expected epigenetic characteristics. However, the transgene integration site in the mAb-producing CHO-K1 cell line, as well as several other CHO hotspots from the literature, were not within the transcriptionally active domains predicted by Hillard et al., highlighting potential additional layers of transgene regulation that need to be considered for defining active domains. In another example, researchers created a detailed integrative map of the 3D architecture and chromatin accessibility of the CHO genome by using a combination of Hi-C, ATAC-seq (assay for transposase-accessible chromatin using sequencing), and histone modification ChlP-seq data in CHO cells to improve identification of novel transcriptional hotspots. Based on their reported model, the previously reported Ferll4 LP locus (Zhang et al., (2015) Recombinase-mediated cassette exchange (RMCE) for monoclonal antibody expression in the commercially relevant CHOK1SV cell line, Biotechnol Prog 31, 1645-1656) was found to be within an active compartment, interacting with a number of promoters; however, the de novo predictive power of this approach remains to be tested.

[0025] Transcriptional regulation in mammalian cells is a complex process that involves the action of enhancers and other regulatory elements, as well as 3D regulation of the genome via genome compartmentalization, spatial positioning of genes within the nucleus, and interaction of genes with nuclear landmarks. Thus, establishing an accurate epigenetic predictive model for identifying transcriptional hotspots requires a multi-dimensional analysis. Moreover, although epigenomic methods can define chromatin domains that are permissive to transcription, it remains challenging to identify, at base-pair resolution, a shortlist of candidate hotspots.

[0026] In addition to the predictive methods described above, some previous work used more empirical approaches for discovering transcriptional hotspots by directly measuring reporter expression from various genomic loci. These assays involved random integration of a reporter gene throughout the genome to identify the “position effect” in gene expression. Such functional screening of genomic loci for transgene expression levels can increase the chances of identifying an amenable high transcriptional activity locus. However, previous strategies to carry out these functional screens were laborious and relatively low throughput due to the requirement to isolate single cell clones with correct integration of the reporter gene.10988-W001-SEC

[0027] Herein, a method previously reported by Akhtar et al. ((2013) Chromatin position effects assayed by thousands of reporters integrated in parallel, Cell 154, 914-927; and (2014) Using TRIP for genome-wide position effect analysis in cultured cells, Nat Protoc 9, 1255-1281) called TRIP (Thousands of Reporters Integrated in Parallel) was used to measure normalized reporter activity from various genomic loci in CHO cells, ultimately identifying LP loci with high transgene expression. TRIP is based on a barcoded reporter library, which is integrated throughout the genome via piggyBac (PB) transposase. Each barcode serves as a tracker for identifying integrated reporter (IR) location and transcriptional activity. As a result of using barcodes, the TRIP method does not require clonal isolation, and the analysis can be performed using the cell pool after PB-mediated integration of the plasmid library. Skipping the time-consuming clonal isolation step increases the throughput of the screen.

[0028] A locus with up to a 9.4-fold increase in mRNA level and a 5.6-fold increase in fed-batch titer compared to a medium-expressing TRIP control locus in a CHO cell host was identified. Surprisingly, single copy expression from this locus resulted in up to a 1.6-fold higher titer than the PB-based stable expression, which typically has multicopy integration and expression of the transgene. We further found that, while active epigenomic markers such as ATAC-seq, H3 Ac, and H3K4me3 positively correlated with TRIP loci reporter expression levels, they were not always deterministic and additional epigenetic factors may need to be assayed. Thus, a functional assay such as TRIP can be a more direct and effective approach for identifying transcriptional hotspots as compared to complex predictive epigenetic analyses.

[0029] In the context of transgene expression, multiple copies of a transgene can possibly be integrated into the genome to increase the overall expression of the transgene. Additional copies can occur in the same single genomic locus or it can be integrated at two or more loci.

[0030] As used herein, a “transgene” refers to a gene (DNA) that has been integrated into a genome, and which is not naturally present in the genome. A transgene can be integrated into the genome at a locus by methods described herein and / or known in the art (see e.g. Lanigan et al. Genes. 2020; 11(3):291). Integration generally involves a double stranded break in DNA, followed by insertion of the transgene at the site of the DNA break, and joining (e.g. ligating) the ends of the transgene with the DNA at the site of the break. Transgenes can be expressed transiently or stably. For stable expression of transgenes, the transgene can be integrated into the genome via homologous recombination, gene editing or recombinase- mediated recombination.10988-W001-SEC

[0031] Any and all transgenes are part of the methods of the present disclosure.

[0032] A “locus” (plural, “loci”) refers to a physical location on a chromosome. The physical location of a locus can be determined, for example, based on a genome assembly.

[0033] Genome assemblies are known in the art. CriGri-PICRH-1.0 genome assembly is the reference CHO genome sequence used in this study for all genomic and epigenomic analyses.

[0034] A RefSeq identifier is a unique identifier for a sequence in the NCBI RefSeq database. RefSeq stands for Reference Sequence, and the database is a collection of sequences that includes genomic DNA, proteins, and transcripts.

[0035] A genomic coordinate refers to the location of a gene or genetic element on a chromosome and includes chromosome name, start position, end position, and chromosome strand.

[0036] The nucleic acid sequences of the exemplified loci are provided in Table 5.

[0037] A homologous sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% homologous to one of the sequences from SEQ ID 1-4. Preferably, a homologous sequence is at least 90% homologous to one of the sequences from SEQ ID 1-4.

[0038] Sequences of the loci described here were identified from CHO cells, however homologous sequences can potentially be found in cell lines derived from Cricetulus griseus or other homologous species. A "homologous sequence" in the context of nucleic acid sequences, refers to a sequence that is substantially homologous to a reference nucleic acid sequence. In some embodiments, two sequences are considered to be homologous if at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more of their corresponding nucleotides are identical over a relevant stretch of residues. In some embodiments, the relevant stretch is a complete (i.e., full) sequence.

[0039] A DNA strand is one of the two polymers of nucleotides that make up the DNA molecule, where one end is called the 5 '-end and the other end is called the 3 '-end. Forward or positive DNA strand is the strand with its 5' end at the tip of the short chromosome arm (p). Reverse or negative strand is reverse complementary to the positive strand.EXAMPLESEXAMPLE 1: Identifying Transcriptional Hotspots in CHO Cells Using TRIP10988-W001-SEC

[0040] To perform TRIP for the discovery of novel transcriptional hotspots in CHO cells, a barcoded enhanced green fluorescent protein (EGFP) reporter library which also contained a glutamine synthetase (GS) expression cassette for positive selection was created (Figure 1 A). Barcode cloning was verified by Sanger sequencing. The barcoded reporters were integrated throughout the genome of a CHO GS knockout cell line by PB transposition, followed by selection in growth media lacking glutamine (Figure IB). After selection, TRIP pools were created by sub-culturing 50,000 cells per pool. Two TRIP pools were analyzed, and each pool was split into two technical replicates for downstream analysis (See Example 5 for details).

[0041] The barcode associated with each IR serves as a unique identifier for tracking reporter location and transcriptional activity. The location of each integrated barcode was determined using an inverse PCR (iPCR)-based approach (Figure IB). Ideally, each barcode should be present at only one genomic location for accurate assignment of the barcode expression value to a single genomic location. The transcription level of each integrated barcode was quantified by next-generation sequencing (NGS) of the barcodes in a cDNA preparation. The count for each barcode in the cDNA library was normalized using its count in genomic DNA (gDNA) (Figure IB). Barcodes were filtered to only include uniquely mapped barcodes in the final list.In total, 1,335 unique barcodes were mapped across two cell pools, allowing interrogation of reporter activity from 1,335 loci. It is important to note that the number of loci screened can be further increased when additional cell pools are analyzed. A log2 -transformed normalized expression value was calculated for each barcoded IR within a technical replicate and observed a significant correlation between replicates (R2 = 0.76, F Test P < 2.2 * 10-16; Figure 1C), indicating minimal technical variance. Subsequently, the expression values were averaged across replicates and the resulting values showed a long-tailed standard normal distribution. Averaged IR expression values were then standardized to robust Z-scores, which were used to rank activity of the mapped TRIP locus associated with each IR (Figure ID). 134 loci located in the top decile (Z > 1.289, right-tailed P < 0.098) were classified as “high” expressing group, 134 loci located in the bottom decile (Z < -1.373, left-tailed P < 0.085) were classified as “low” expressing group, and 667 loci with IR expression Z-scores falling within the interquartile range (Z > -0.686, Z < 0.633) were considered “medium” expressing group. Highly active TRIP loci were detected across CHO chromosomes, albeit without significant chromosomal enrichment.10988-W001-SECEXAMPLE 2: ATAC-Seq, H3Ac, and H3K4me3 Signal Levels Positively Correlate with TRIP Loci Reporter Expression

[0042] To understand the epigenetic determinants of high transgene activity in CHO cells, we performed ATAC-seq and CUT&RUN (Cleavage Under Targets & Release Using Nuclease) chromatin profiling. ATAC-seq identifies open chromatin regions in the genome using a hyperactive Tn5 transposase. Purified nuclei of cells are treated with the Tn5 transposase loaded with NGS adapters, which simultaneously fragments open chromatin regions and tags them with the NGS adapters (also known as “tagmentation”). The tagged DNA fragments can then be identified by NGS to reveal open chromatin regions in the genome. CUT&RUN is an alternative to ChlP-seq for analyzing protein-DNA interactions. This technique can be used for mapping binding sites of chromatin-associated proteins or sites of enrichment of various histone post-translational modifications (histone marks) that can affect gene expression. In this in situ protocol, cells are permeabilized and the protein of interest is labeled with an antibody, which is subsequently targeted by a micrococcal nuclease (MNase)- protein A (pA) fusion protein. MNase digestion of the DNA regions that are interacting with the protein of interest releases the cleaved DNA fragments out of the nucleus. These DNA fragments can be purified and sequenced by NGS for chromatin profiling.

[0043] Histone marks H3 Ac (histone H3 acetylation), H3K4me3 (histone H3 lysine 4 tri-methylation), H3K9me3 (histone H3 lysine 9 tri-methylation) and H3K27me3 (histone H3 lysine 27 tri-methylation) were assessed using CUT&RUN in CHO cells. High ATAC-seq, H3 Ac, and H3K4me3 signals are often associated with transcriptional activation, whereas H3K9me3 and H3K27me3 are mainly associated with transcriptional repression (Table 1). Highly accessible chromatin regions showing high ATAC-seq signal are often associated with cis-regulatory elements, including promoters and enhancers. Typically, high levels of H3K4me3 and H3 Ac are associated with active promoters, whereas high H3 Ac and low H3K4me3 are associated with active enhancers.Table 1. Chromatin features and their effects on transcription.10988-W001-SEC

[0044] To compare ATAC-seq and CUT&RUN results to the TRIP reporter expression levels, CHO cells from the day of transfection were tested with the TRIP reporter library. To understand the correlation between various chromatin marks and TRIP reporter expression levels, we quantified the signal (read pile-up) for each chromatin feature in a 250 bp region centered at each IR integration site. ATAC-Seq, H3 Ac, and H3K4me3 signal levels positively correlated with TRIP IR expression levels (Figure 2A). Conversely, H3K9me3 and H3K27me3 levels showed no correlation. Median signal around the “high” TRIP group loci compared to the “low” TRIP group loci was 6.1-fold higher for ATAC-seq, 1.7-fold higher for H3 Ac and 5.1 -fold higher for H3K4me3. However, given the large range of values for the reads at TRIP locus for each of the analyzed epigenetic features, it is difficult to assign a TRIP locus to a particular TRIP group solely based on any single feature’s signal level.

[0045] The distance between each TRIP locus and the closest ATAC-seq or CUT&RUN peak was analyzed. Higher expressing TRIP loci demonstrated increased proximity to ATAC-Seq, H3 Ac, and H3K4me3 peaks (Figure 3 A). Median distance to the closest ATAC-seq peak was 0.2 kb for “high” TRIP loci and 1.4 kb for “low” TRIP loci. Interestingly, the median distance to the closest H3 Ac peak was zero kb for “high” TRIP loci, highlighting that many of these loci overlapped an H3 Ac peak. The corresponding distance was 0.8 kb for the “low” TRIP loci. Finally, the median distance to the closest H3K4me3 peak was 6.6 kb for “high” TRIP loci and 45.9 kb for “low” TRIP loci, showing a higher fold-difference between the “high” and “low” TRIP groups as compared to the other epigenetic features.

[0046] The frequency of overlap between epigenetic feature peaks and TRIP loci was quantified. 38.8%, 64.2%, and 43.3% of loci in the “high” TRIP group overlapped with ATAC- seq, H3Ac and H3K4me3 peaks, respectively (Table 2). For loci in the “low” TRIP group, the corresponding percentages of overlap were 25.6%, 38.8%, and 22.3%, respectively. While it was remarkable to see a fold-increase of 1.5 for ATAC-seq, 1.7 for H3Ac and 1.9 for10988-W001-SECH3K4me3 peaks for percentage of overlap with loci in the “high” TRIP group compared to the “low” group, a significant number of loci in the “low” TRIP group still overlapped with these active marks. Thus, absent knowledge of the TRIP group, proximity to any one of these marks alone would not be deterministic, but rather a modest predictor.Table 2. Fraction of loci overlapping with ATAC-seq or CUT&RUN peaks for each TRIP group.EXAMPLE 3: TRIP Loci Reporter Expression Levels Positively Correlate with Proximity to Active ChromHMM Chromatin States

[0047] To better understand the combinatorial effects of various chromatin marks on TRIP IR expression levels, the histone marks assayed by CUT&RUN were integrated with ATAC-seq data to derive a ChromHMM model.28 ChromHMM utilizes a hidden Markov model (HMM) across the observed data to define different chromatin “states”, which are represented by different correlation strengths with the provided histone mark & ATAC-seq signals. Observation of these marks combined with a priori knowledge allows for finer-grained interpretation into distinct functional chromatin states. Using ChromHMM, a 7-state model was created (Figure 4A). The ChromHMM states are: Repressed PolyComb (H3K27me3 signal),10988-W001-SECNo Signal, Heterochromatin (H3K9me3 signal), Poised Chromatin (ATAC-seq signal), Weak Promoter (H3K4me3 and ATAC-seq signals), Active Promoter (H3K4me3, H3Ac and ATAC- seq signals) and Repressed Promoter (H3K4me3 signal).

[0048] When the percentage of loci that overlapped with active ChromHMM states was analyzed, 27.6% of “high” TRIP loci overlapped with the “weak promoter” state and 18.7% overlapped with the “active promoter” state; whereas 14.9% of “low” loci overlapped with the “weak promoter” state and 11.9% overlapped with the “active promoter” state (Table 3).Table 3. Percentage of loci overlapping with each ChromHMM state for each TRIP group.

[0049] The distance between TRIP loci and the “weak promoter” or “active promoter” states was analyzed, hypothesizing that proximity to such putative regulatory domains, instead of a perfect overlap, might be a better indicator of high IR expression. It is important to note that the main difference between the “active promoter” and “weak promoter” states is the10988-W001-SEC strength of the H3Ac signal. In some of the cases within the dataset, the “weak promoter” state can be immediately next to an “active promoter” state. In those cases, the “weak promoter” state can be interpreted as a transition state from the “active promoter” state, instead of being considered as a separate, independent domain.

[0050] Loci in the “high” TRIP group showed closer proximity to both “weak promoter” and “active promoter” states as compared to loci in the “low” group (Figure 4B). Interestingly, 89.6% and 62.7% of the loci in the “high” group were within 30 kb distance from the “weak promoter” and “active promoter” states, respectively. For the loci in the “low” group, the corresponding percentages were 65.4% and 43.4%. Moreover, the percentages of loci in the “high” group within 1 kb of the “weak promoter” and “active promoter” states were 50.7% and 29.9%, respectively, whereas the corresponding percentages for loci in the “low” group were 26.9% and 18.6%, respectively. These results suggest that proximity to “weak promoter” or “active promoter” states can be an indicator of high expression IRs, although not consistently deterministic.EXAMPLE 4: Transcriptional Hotspots Identified by TRIP were Validated by Measuring Titers and Transgene mRNA Levels from Clonal Cells After CRISPR Knock-In of a Test Molecule

[0051] To validate high reporter expression from the loci in the “high” expressing TRIP group, an IgGl crystallizable fragment (Fc) molecule was knocked-in to select loci (Figure 5A) and both mRNA expression and the Fc titer from clones with single copy insertion was measured. For this purpose, four loci from the “high” TRIP group were selected as potential transcriptional hotspots, and these loci were numbered based on their TRIP IR expression Z- scores, with locus 1 having the highest score. For Fc knock-in, an sgRNA target site within a -500 bp region around each TRIP locus IR integration site was selected.

[0052] Table 4 and Table 5 provide information about the tested loci (Table 4) and knock-in information (Table 5). All coordinates are based on CriGri-PICRH-1.0 genome assembly. Table 6 shows the RefSeq identifier of each CHO chromosome.Table 4, Location of loci.10988-W001-SECTable 5, Knock-in information of loci.Table 6, RefSeq identifiers of CHO chromosomes (CriGri-PICRH-1 0).10988-W001-SEC

[0053] The selected TRIP loci demonstrated different epigenetic patterns in a 50 kb region centered at each corresponding sgRNA target site (Table 7). Locus 1 had no active ChromHMM states nearby, with only a few H3 Ac peaks, even though it had the highest TRIP IR expression Z-score. It is important to note that there may also be additional factors contributing to gene regulation, which have not assayed here, such as other histone modifications or 3D genome organization. Unlike locus 1, locus 2 was proximal to “active promoter” and “weak promoter” states, as well as ATAC-seq and H3 Ac peaks. Locus 3 and 4 were included in our analysis because although they both were in the “high” TRIP group with a comparable score to locus 2, and were proximal to peaks for active marks and chromatin states, both locus 3 and 4 also had repressive marks nearby (H3K27me3 and H3K9me3 peaks, respectively). A TRIP locus from the “medium” group was also included in the screen as a negative control (referred to as “TRIP control”). Interestingly, the TRIP control locus was proximal to active ChromHMM states and peaks for active chromatin marks. However, the control locus was also proximal to repressive H3K27me3 peaks, which might be contributing to transgene repression.Table 7, Presence of epigenetic feature peaks and ChromHMM states in a 50 kb region surrounding each knock-in site.10988-W001-SEC

[0054] After co-transfection of CRISPR ribonucleoprotein (RNP) complex and a linear donor DNA with 50 bp short homology arms, cells were selected with puromycin, and clonal isolation was performed. Clones were first verified for correct insertion by genotyping PCR and Sanger sequencing of the knock-in allele. Copy number of the insert was verified by ddPCR and only clones with single copy insertion were selected for further analysis (Figure 5B). All TRIP loci from the “high” group had higher Fc mRNA expression levels than the TRIP control locus (Figure 6A), with one of the locus 2 clones having up to a 9.4-fold increase over the average mRNA expression of the control clones. Some clonal variability in mRNA levels was observed, but the average mRNA level from locus 2 clones showed a 7.6-fold increase over the average of control clones. Moreover, mRNA expression levels for the locus 1 and 2 clones were higher than the locus 3 and 4 clones. Average mRNA expression from the locus 2 clones was 3.3-fold higher than the average of locus 3 clones and 2-fold higher than the average of locus 4 clones. The repressive marks near locus 3 and 4 might explain the lower transcriptional activity of these loci compared to locus 1 and 2.

[0055] To analyze expression at the protein level, single knock-in clones for each locus were tested in 4 ml fed-batch production. Titer (Figure 6B) and specific productivity (Qp, Figure 6C) from each clone were measured. The normalized results are summarized in Table 8. Locus 2 yielded the highest titer, again with some clone-to-clone variation, and both locus 2 clones had titers significantly higher than the average titer of the TRIP control clones, with up10988-W001-SEC to a 5.6-fold increase. Locus 3 and 4 yielded lower titers than locus 2, which was consistent with their lower TRIP IR expression Z-scores and Fc mRNA levels. Interestingly, even though the mRNA level from locus 1 was high and comparable to at least one of the locus 2 clones, the titer of the locus 1 clone 1-19 was lower than the titers of the locus 2 clones, with clone 2-09 having a 3.2-fold higher titer than the 1-19 clone. This observation might reflect a fed-batch process-related effect.Table 8, Expression analysis of single knock-in clones.

[0056] Next, titers from the clones with single copy insertion into locus 1 and 2 were compared to the titer from a cell pool after PB-mediated insertion of the same DNA insert, which is a conventional method for creating recombinant CHO cell pools. For accurate10988-W001-SEC comparison of the single knock-in clones to the cell pool after PB-mediated insertion, CHO host cells were co-transfected with a PB plasmid and a plasmid that contains the same insert sequence as in the CRISPR knock-in donor DNA, but this time flanked by inverted terminal repeats (ITRs). Based on the fed-batch titers (Figure 6D) and Qp analysis (Figure 6E), clone 2- 09 with single copy knock-in into locus 2 again had the highest titer with 559 mg / L. 2-09 and 2-17 clones had 1.6-fold (p < 0.05) and 1.3-fold (ns) higher titers than the PB control, respectively. Clone 1-19 from knock-in into locus 1 again resulted in low titer with 113 mg / L. For all fed-batch experiments, the same feeding protocol, determined based on the average consumption of glucose, was used for all the cells tested. Thus, titers from the top performing clones can potentially be further improved after clone-specific optimization of the feeding protocol. Indeed, glucose consumption from clone 1-19 was lower compared to clones 2-09 and 2-17, indicating potential metabolic differences. Moreover, the tested clones showed differences in their growth patterns (Figure 6F and 6G). Thus, the unexpected lower titer from the locus 1 clone could be a result of cell health- and / or metabolism-related differences, which may have been exacerbated during fed-batch cultivation, and these differences could be due to potential genetic and / or epigenetic changes in this clone. These genetic / epigenetic changes might have been introduced during the clonal isolation process or due to potential CRISPR off- target effects, which would need further investigation.

[0057] Overall, these results demonstrate that TRIP was successful at identifying transcriptional hotspots in the CHO genome, leading to increased mRNA and titer levels of a tested molecule. To identify high-productivity cell lines, screening multiple clones can further improve overall titers because although single site integration can reduce heterogeneity in transgene epigenetic regulation, clonal heterogeneity due to cell health- and metabolism-related differences can also impact expression.EXAMPLE 5: MethodsPlasmid Cloning

[0058] For TRIP, we first cloned a destination vector which could be used for Golden Gate cloning of a barcoded reporter library. For this purpose, a plasmid containing a GS cassette with SR alpha (SRa) promoter, followed by a chimeric CMV promoter, was cloned by Golden Gate reaction using Bsal enzyme. All these elements were cloned between inverted terminal repeats (ITRs). Two BsmBI sites were included downstream of the CMV promoter10988-W001-SEC and upstream of a polyA signal (pA) to allow cloning of a reporter gene-coding sequence by Golden Gate reaction using the BsmBI sites.

[0059] To generate the barcoded EGFP library, 16 bp random barcodes were added 3’ of EGFP, downstream of the stop codon, by PCR amplification using a reverse primer with 16- nt degenerate sequences at the 5’ end (Figure 1 A). The amplicon contained a DpnII site immediately upstream of the barcode, which is necessary for the mapping step in TRIP. EGFP- barcode amplicon was cloned downstream of the chimeric CMV promoter in the SRa-GS- chimeric CMV plasmid by Golden Gate reaction using BsmBI enzyme. The complexity of the library was around 160,000 as estimated by colony counting. No-insert control resulted in no bacterial colonies. Quality control for the barcoded EGFP library was performed by Sanger sequencing at and around the barcode region, demonstrating equal representation of all four nucleotides at each position along the barcode.

[0060] For knock-in experiments, a test Fc coding sequence, followed by a T2A linker and a puromycin resistance cassette (PuroR), was cloned under the chimeric CMV promoter. All these elements were flanked by ITRs to enable use of the same plasmid in PB control experiments. When the Fc-T2A-PuroR plasmid was used for PCR amplification of the knock-in donor DNA (described below), primers were designed to avoid ITR sequences in the final amplicon.TRIP Transfection, Library Preparation and Next-Generation Sequencing (NGS)

[0061] CHO Cells were co-transfected with the barcoded reporter library and a PB plasmid using Lipofectamine LTX (Thermo Fisher Scientific). On the day of transfection, an aliquot of the host cells was also collected for ATAC-seq and CUT&RUN experiments. 48 hours post-transfection, cells were placed into selection media matching the growth media but without glutamine. No methionine sulfoximine (MSX) was added to keep IR copy number per cell within the range recommended by Akhtar et al, which is at least 1 copy / cell and ideally around 10 copies / cell.23 Controlling the IR copy number per cell is important to ensure cell library complexity.

[0062] After 8 days of positive selection, TRIP pools were created by sub-culturing 50,000 cells per pool. After an additional 11 days, TRIP pools were split into 2 technical replicates and grown for 10 more days. Once the cells reached high density, cells were collected for gDNA & RNA purification. A total of 2 TRIP cell pools were created from two independent transfections. The results from these pools were combined into a final list for statistical analysis.10988-W001-SEC

[0063] NGS library preparation for TRIP was performed as previously described.23 Briefly, barcoded regions from gDNA and cDNA preparations were amplified by two rounds of nested PCRs (PCR-1 and PCR-2), resulting in normalization and expression libraries, respectively. Different than the previously described protocol, PCR-1 was performed using primers “PB-cDNA-forward” and “PB-cDNA-Rev-New”; and for PCR-2, we used primers from xGen UDI Primer Plate 2, 8nt, 96rxn (IDT) for indexing and incorporation of adapters.

[0064] For mapping library preparation, gDNA from a TRIP pool was first digested with the DpnII enzyme, which cuts immediately upstream of the barcode and the DpnII sites in gDNA downstream of the 3’ ITR. Digested DNA pieces were then circularized by intramolecular ligation, bringing the barcode and gDNA sequences together, which enabled amplification of the barcode and the genomic DNA sequences using inverse PCR (iPCR) with primers annealing to the constant sequences in these circular DNAs (Figure IB). Mapping libraries were created by three rounds of PCRs (iPCR-1, iPCR-2 and iPCR-3). Different than the previously described protocol, iPCR-1 was performed using “PB-outer-F-2” and “cDNA- ampl-R” primers; iPCR-2 was performed using “PB-cDNA-forward” and “InvPCR-F-TruSeq- new” primers; and for iPCR-3, we used primers from xGen UDI Primer Plate 2, 8nt, 96rxn (IDT) for indexing and incorporation of adapters.

[0065] For all NGS, we utilized Illumina NextSeq 2000 P3 reagents 300 cycles kit (2x150 bp paired end sequencing).TRIP Data Analysis

[0066] To analyze mapping, normalization (gDNA) and expression (cDNA) libraries, we used the publicly available Perl-based pipeline from previous work23 (https: / / trip.nki.nl / ). For mapping, Chinese hamster PICRH assembly was used. For IR expression analysis, barcodes extracted from the gDNA library reads were filtered to exclude barcodes with less than 5 reads, resulting in a list of genuine barcodes. The barcodes were further filtered to only include the uniquely mapped barcodes. A pseudocount of 1 was added to the count of each barcode from the cDNA library. Barcode counts from the gDNA and cDNA libraries were normalized using the sum of the barcode counts within each of the corresponding sequencing library to account for differences in sequencing depth.31

[0067] A normalized IR expression value was then calculated as previously described using the normalized cDNA and gDNA counts of the associated barcode in each technical replicate. Normalized cDNA counts were divided by the normalized gDNA counts for each barcode, followed by log2 transformation (i.e. IR expression = log2(normalized cDNA counts / 10988-W001-SEC normalized gDNA counts)). This expression value accounts for differences in IR copy number in gDNA for each pool and replicate. A log2 transformation is applied to normalize the distribution of the copy number adjusted expression values. To calculate a final expression value for each barcoded IR in each TRIP pool, IR expression values from the two technical replicates were averaged.

[0068] Since the average IR expression values showed normal, but long-tailed, distribution across TRIP pools, we calculated robust Z-scores for the average IR expression values. Robust Z-score centers each pool to the median and scales the values to an estimate of the standard deviation, which is based on the median absolute deviation, and is therefore robust to the observed long tails and extreme values.

[0069] Barcoded IRs from both TRIP pools, and hence their associated TRIP loci, were ranked by the IR expression Z-score in a combined list and categorized into TRIP expression groups based on percentile ranking. Barcodes in the bottom 10% were classified as “low” expressing, those in the top 10% were classified as “high” expressing, and those in the interquartile range (the center 50%) were classified as “medium” expressing loci. Barcodes within the 10th to 25th percentile, and 75th to 90th percentile were classified as “low shoulder” and “high shoulder” respectively.ATAC-seq Analysis

[0070] ATAC-seq was performed in duplicates by MedGenome, Inc. using an ATAC- seq kit (ActiveMotif) with 100,000 cells as input. ATAC-seq data was processed using version 2.0 of the nf-core35 ATAC-seq pipeline. The pipeline’s default parameters were used, with reads aligned to the PICRH-1.0 genome, narrow peak calling using MACS2,37 and a macs gsize parameter of 2.3E9. BigWig and narrowPeak files were visualized using IGV38 to examine read signal and identified peaks, respectively.

[0071] To examine ATAC-seq reads around TRIP loci, subread featureCounts was used to extract aligned reads at the 250 bp region around each TRIP locus. Average of replicate 1 and 2 reads were calculated. Two-sample Wilcoxon test was applied to examine differences between TRIP groups.

[0072] To assess intersections and distances between ATAC-seq peaks and TRIP loci coordinates, we employed bedtools version 2.22.1, specifically using the “intersect” and “closest” functions. Two-sample Wilcoxon test was applied to examine differences between TRIP groups. All statistical analyses and final visualization of the data were performed in R using a custom script.10988-W001-SECCUT&RUN Analysis

[0073] CUT&RUN was performed in duplicates by MedGenome, Inc. using a CUT&RUN kit (Epicypher), 100,000 cells and 0.5 pg primary Ab per reaction. The antibodies used were rabbit anti-H3ac (Active Motif; Cat. No.: 860 39139), rabbit anti-H3K4me3 (Active Motif; Cat. No.: 39159), rabbit anti-H3K9me3 (Abeam; Cat. No.: ab8898) and rabbit anti- 143 K27me3 (Cell Signaling Technologies, Cat. No: 9733 S).

[0074] CUT&RUN data was processed using version 3.2 of the nf-core CUT&RUN pipeline. Default parameters were largely used, but counts per million (CPM), rather than spike-in, was chosen for the normalization method upon observation of highly variable spike-in alignment rates across control & treatment samples. All reads were aligned to the PICRH-1.0 genome, and peaks were called for each mark using SEACR. Tracks for each histone mark were visualized in IGV.

[0075] To examine CUT&RUN reads around TRIP loci, subread featureCounts was used to extract aligned reads at the 250 bp region around each TRIP locus per histone mark assessed, and the rest of the analysis was performed in a similar manner as described for ATAC-seq. To assess intersections and distances between CUT&RUN peaks and TRIP loci coordinates, we utilized bedtools in a similar manner as described for ATAC-seq.ChromHMM Modeling

[0076] Aligned histone mark and ATAC-seq reads were used as inputs to ChromHMM in order to build a model of chromatin states. We primarily used default ChromHMM processing parameters with 200 bp resolution and settled on a 7-state model after testing models using between 5-8 states. Distance of each TRIP locus to either “active promoter” or “weak promoter” was calculated as described above for ATAC-seq and CUT&RUN peaks. sgRNA Design

[0077] To design 20 nt guide sequences of sgRNAs for each knock-in, -500 bp regions around TRIP loci coordinates were submitted to CRISPOR tool. sgRNAs in low complexity and repetitive regions were avoided and the sgRNA with the highest specificity score was selected for each knock-in.PCR for Creating Linear Donor DNA

[0078] Linear donor DNAs with short HAs were created via PCR as described before, 46 using the Fc-T2A-PuroR plasmid as the template. Primers for donor DNA PCR were designed to include 50 nt 5’ extension sequences homologous to the target site (a.k.a. HAs). HAs were designed to exclude PAM sequence and also the first 3-4 nt upstream of the10988-W001-SEC expected CRISPR cut site. Primers were ordered from Integrated DNA Technologies (IDT). The linear donor DNA created by donor DNA PCR contained Fc-T2A-PuroR cassette, flanked by 50 bp HAs. PCRs were performed using Q5 High Fidelity DNA Polymerase (New England Biolabs (NEB)), using manufacturer’s protocol with the following changes to the PCR cycle: 98°C for 3 min; followed by 30 cycles of 98°C for 30 s, 69°C for 15 s and 72°C for 3 min. The final extension was 72°C for 6 min. Donor DNA PCRs were purified with QIAquick PCR Purification Kit (Qiagen).CRISPR Knock-in Transfections

[0079] To perform transfections for knock-in of the Fc-T2A-PuroR donor (Figure 5A), nucleofection of the CHO host cell line with CRISPR RNP and donor DNA was performed using SF Cell Line 4D Nucleofector Kit (Lonza), following manufacturer’s protocol. 1 week after nucleofection, 20 pg / ml puromycin selection was started. After 1 week selection, clonally derived subclones were generated by depositing single cells from starting pools into 96-well plates by BD FACSAria™ II (BD Biosciences). Light microscopy imaging was performed periodically to track the formation of a single colony in each well. Colonies were scaled up to suspension culture and screened by full length and left junction genotyping PCR. Clones with expected PCR results were further verified for single copy knock-in via ddPCR. Sequence of the insert was verified by sequencing. gDNA Extraction and ddPCR Copy Number Analysis

[0080] gDNA was extracted using QuickExtract DNA Extraction Solution (Epicentre). Briefly, -100,000 cells were collected and 50 pl QuickExtract solution was added per cell pellet. After vortexing, the mixture was incubated at 65°C for 15 min, followed by 98°C for 10 min. For setting up the ddPCR reaction, 2X ddPCR Supermix for Probes (no dUTP) (Bio-Rad) was used and the reaction was setup according to manufacturer’s instructions, using 5U of Hindlll enzyme for DNA fragmentation. Beta-actin (ACTB) was used as a reference gene (copy number 4). PuroR cassette was used as the target. Primers and probes were ordered from IDT.

[0081] Droplets were generated by QX200 AutoDG Automated Droplet Generator (Bio-Rad), followed by thermal cycling using the cycle conditions recommended by the 2X ddPCR Supermix for Probes (no dUTP) protocol. Droplet reading was performed using QX200 Droplet Reader. Data were analyzed using QuantaSoft Software (Bio-Rad).RT-ddPCR10988-W001-SEC

[0082] RNA from the clones were isolated using Qiagen RNeasy Plus Mini kit (96-well format). For two-step RT-ddPCR, SuperScript III First-Strand Synthesis SuperMix for qRT- PCR (Invitrogen) was used for cDNA synthesis, following manufacturer’s protocol. ddPCR was performed as described above but without any DNA digestion step. Moreover, target gene of choice was the Fc coding sequence. ACTB was used as a reference gene.Fed-batch Analysis

[0083] Cells were inoculated at a density of 1 x 106 cells / mL into proprietary chemically defined production media, without any selection reagent, to initiate the fed-batch production. Fed-batch production assays were performed with the culture volume of either 30 mL in 125 mL shake flasks or 4 mL in 24 deep well-plates. Cells were maintained in humidified incubators at 36°C, 5% CO2 with shaking at either 150 rpm (125 mL shake flask) or 220 rpm (24 deep well-plate).

[0084] The production was carried out for 13 or 14 days by feeding the cells with a proprietary chemically defined feed media every two to three days. The viability and viable cell density (VCD) of the culture were measured using Vi-CELL BLU (Beckman Coulter). The Bioprofile Flex2 from Nova Biomedical was used for measurement of the metabolites. Glucose levels on feed days were adjusted based on the average glucose in cell culture supernatant of the evaluated clones. Secreted Fc concentration on feed days was quantified using ProA probes on a ForteBio Octet QKe system following the manufacturer's protocol. Specific productivity (Qp; pg / cell / day) for each time point was calculated as described previously, 47 using the below formula: Qp=[Titer / IVCD]x 1000.

[0085] IVCD stands for Integral Viable Cell Density. Viable cell densities (VCDs) were reported in le6 cells / ml.10988-W001-SECSEQUENCES10988-W001-SECREFERENCES

[0001] Walsh, G. (2018) Biopharmaceutical benchmarks 2018, Nat Biotechnol 36, 1136-1145.

[0002] Wurm, F. M. (2004) Production of recombinant protein therapeutics in cultivated mammalian cells, Nat Biotechnol 22, 1393-1398.

[0003] Gray, D. (2001) Overview of protein expression by mammalian cells, Curr Protoc Protein Sci Chapter 5, Unit5 9.

[0004] Lee, J. S., Kildegaard, H. F., Lewis, N. E., and Lee, G. M. (2019) Mitigating Clonal Variation in Recombinant Mammalian Cell Lines, Trends Biotechnol 37, 931-942.

[0005] Kim, M., O'Callaghan, P. M., Droms, K. A., and James, D. C. (2011) A mechanistic understanding of production instability in CHO cell lines expressing recombinant monoclonal antibodies, Biotechnol Bioeng 108, 2434-2446.

[0006] Huhtinen, O., Salbo, R., Lamminmaki, U., and Prince, S. (2023) Selection of biophysically favorable antibody variants using a modified Flp-In CHO mammalian display platform, Front Bioeng Biotechnol 11, 1170081.

[0007] Zhou, H., Liu, Z. G., Sun, Z. W., Huang, Y., and Yu, W. Y. (2010) Generation of stable cell lines by site-specific integration of transgenes into engineered Chinese hamster ovary strains using an FLP-FRT system, J Biotechnol 147, 122-129.

[0008] Zhang, L., Inniss, M. C., Han, S., Moffat, M., Jones, H., Zhang, B., Cox, W. L., Rance, J. R., and Young, R. J. (2015) Recombinase-mediated cassette exchange (RMCE) for monoclonal antibody expression in the commercially relevant CHOK1SV cell line, Biotechnol Prog 31, 1645-1656.

[0009] Chi, X., Zheng, Q., Jiang, R., Chen-Tsai, R. Y., and Kong, L. J. (2019) A system for site-specific integration of transgenes in mammalian cells, PLoS One 14, e0219842.

[0010] Gaidukov, L., Wroblewska, L., Teague, B., Nelson, T., Zhang, X., Liu, Y., Jagtap, K., Mamo, S., Tseng, W. A., Lowe, A., Das, J., Bandara, K., Baijuraj, S., Summers, N. M., Lu, T. K., Zhang, L., and Weiss, R. (2018) A multi-landing pad DNA integration platform for mammalian cell engineering, Nucleic Acids Res 46, 4072-4086.

[0011] Pristovsek, N., Nallapareddy, S., Grav, L. M., Hefzi, H., Lewis, N. E., Rugbjerg, P., Hansen, H. G., Lee, G. M., Andersen, M. R., and Kildegaard, H. F. (2019) Systematic Evaluation of Site-Specific Recombinant Gene Expression for Programmable Mammalian Cell Engineering, ACS Synth Biol 8, 758-774.

[0012] Hertel, O., Neuss, A., Busche, T., Brandt, D., Kalinowski, J., Bahnemann, J., and Noll, T. (2022) Enhancing stability of recombinant CHO cells by CRISPR / Cas9-mediated10988-W001-SEC site-specific integration into regions with distinct histone modifications, Front Bioeng Biotechnol 10, 1010719.

[0013] Dhiman, H., Campbell, M., Melcher, M., Smith, K. D., and Borth, N. (2020) Predicting favorable landing pads for targeted integrations in Chinese hamster ovary cell lines by learning stability characteristics from random transgene integrations, Comput Struct Biotechnol J 18, 3632-3648.

[0014] Mehrmohamadi, M., Sepehri, M. H., Nazer, N., and Norouzi, M. R. (2021) A Comparative Overview of Epigenomic Profiling Methods, Front Cell Dev Biol 9, 714687.

[0015] Hilliard, W., and Lee, K. H. (2021) Systematic identification of safe harbor regions in the CHO genome through a comprehensive epigenome analysis, Biotechnol Bioeng 118, 659-675.

[0016] Feichtinger, J., Hernandez, I., Fischer, C., Hanscho, M., Auer, N., Hackl, M., Jadhav, V., Baumann, M., Krempl, P. M., Schmidl, C., Farlik, M., Schuster, M., Merkel, A., Sommer, A., Heath, S., Rico, D., Bock, C., Thallinger, G. G., and Borth, N. (2016) Comprehensive genome and epigenome characterization of CHO cells in response to evolutionary pressures and over time, Biotechnol Bioeng 113, 2241-2253.

[0017] Bevan, S., Schoenfelder, S., Young, R. J., Zhang, L., Andrews, S., Fraser, P., and O'Callaghan, P. M. (2021) High-resolution three-dimensional chromatin profiling of the Chinese hamster ovary cell genome, Biotechnol Bioeng 118, 784-796.

[0018] Dekker, J., Alber, F., Aufmkolk, S., Beliveau, B. J., Bruneau, B. G., Belmont, A. S., Bintu, L., Boettiger, A., Calandrelli, R., Disteche, C. M., Gilbert, D. M., Gregor, T., Hansen,A. S., Huang, B., Huangfu, D., Kalhor, R., Leslie, C. S., Li, W., Li, Y., Ma, J., Noble, W. S., Park, P. J., Phillips-Cremins, J. E., Pollard, K. S., Rafelski, S. M., Ren, B., Ruan, Y., Shav-Tal, Y., Shen, Y., Shendure, J., Shu, X., Strambio-De-Castillia, C., Vertii, A., Zhang, H., and Zhong, S. (2023) Spatial and temporal organization of the genome: Current state and future aims of the 4D nucleome project, Mol Cell 83, 2624-2640.

[0019] Consortium, E. P., Moore, J. E., Purcaro, M. J., Pratt, H. E., Epstein, C. B., Shoresh, N., Adrian, J., Kawli, T., Davis, C. A., Dobin, A., Kaul, R., Halow, J., Van Nostrand, E. L., Freese, P., Gorkin, D. U., Shen, Y., He, Y., Mackiewicz, M., Pauli-Behn, F., Williams,B. A., Mortazavi, A., Keller, C. A., Zhang, X. O., Elhajjajy, S. I., Huey, J., Dickel, D. E., Snetkova, V., Wei, X., Wang, X., Rivera-Mulia, J. C., Rozowsky, J., Zhang, J., Chhetri, S. B., Zhang, J., Victorsen, A., White, K. P., Visel, A., Yeo, G. W., Burge, C. B., Lecuyer, E., Gilbert, D. M., Dekker, J., Rinn, J., Mendenhall, E. M., Ecker, J. R., Kellis, M., Klein, R. J.,10988-W001-SECNoble, W. S., Kundaje, A., Guigo, R., Famham, P. J., Cherry, J. M., Myers, R. M., Ren, B., Graveley, B. R., Gerstein, M. B., Pennacchio, L. A., Snyder, M. P., Bernstein, B. E., Wold, B., Hardison, R. C., Gingeras, T. R., Stamatoyannopoulos, J. A., and Weng, Z. (2020) Expanded encyclopaedias of DNA elements in the human and mouse genomes, Nature 583, 699-710.

[0020] Dekker, J., Belmont, A. S., Guttman, M., Leshyk, V. O., Lis, J. T., Lomvardas, S., Mimy, L. A., O'Shea, C. C., Park, P. J., Ren, B., Politz, J. C. R., Shendure, J., Zhong, S., and Network, D. N. (2017) The 4D nucleome project, Nature 549, 219-226.

[0021] O'Brien, S. A., Lee, K., Fu, H. Y., Lee, Z., Le, T. S., Stach, C. S., McCann, M. G., Zhang, A. Q., Smanski, M. J., Somia, N. V., and Hu, W. S. (2018) Single Copy Transgene Integration in a Transcriptionally Active Site for Recombinant Protein Synthesis, Biotechnol J 13, el800226.

[0022] Akhtar, W., de Jong, J., Pindyurin, A. V., Pagie, L., Meuleman, W., de Ridder, J., Bems, A., Wessels, L. F., van Lohuizen, M., and van Steensel, B. (2013) Chromatin position effects assayed by thousands of reporters integrated in parallel, Cell 154, 914-927.

[0023] Akhtar, W., Pindyurin, A. V., de Jong, J., Pagie, L., Ten Hoeve, J., Bems, A., Wessels, L. F., van Steensel, B., and van Lohuizen, M. (2014) Using TRIP for genome-wide position effect analysis in cultured cells, Nat Protoc 9, 1255-1281.

[0024] Buenrostro, J. D., Giresi, P. G., Zaba, L. C., Chang, H. Y., and Greenleaf, W. J. (2013) Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position, Nat Methods 10, 1213-1218.

[0025] Skene, P. J., Henikoff, J. G., and Henikoff, S. (2018) Targeted in situ genomewide profiling with high efficiency for low cell numbers, Nat Protoc 13, 1006-1019.

[0026] Skene, P. J., and Henikoff, S. (2017) An efficient targeted nuclease strategy for high-resolution mapping of DNA binding sites, Elife 6.

[0027] Heintzman, N. D., Stuart, R. K., Hon, G., Fu, Y., Ching, C. W., Hawkins, R. D., Barrera, L. O., Van Calcar, S., Qu, C., Ching, K. A., Wang, W., Weng, Z., Green, R. D., Crawford, G. E., and Ren, B. (2007) Distinct and predictive chromatin signatures of transcriptional promoters and enhancers in the human genome, Nat Genet 39, 311-318.

[0028] Ernst, J., and Kellis, M. (2017) Chromatin-state discovery and genome annotation with ChromHMM, Nat Protoc 12, 2478-2492.

[0029] Hilliard, W., and Lee, K. H. (2023) A compendium of stable hotspots in the CHO genome, Biotechnol Bioeng 120, 2133-2143.10988-W001-SEC

[0030] Hilliard, W., MacDonald, M. L., and Lee, K. H. (2020) Chromosome-scale scaffolds for the Chinese hamster reference genome assembly to facilitate the study of the CHO epigenome, Biotechnol Bioeng 117, 2331-2339.

[0031] Leemans, C., van der Zwalm, M. C. H., Brueckner, L., Comoglio, F., van Schaik, T., Pagie, L., van Arensbergen, J., and van Steensel, B. (2019) Promoter-Intrinsic and Local Chromatin Features Determine Gene Repression in LADs, Cell 177, 852-864 e814.

[0032] Park, C., Kim, H., and Wang, M. (2022) Investigation of finite-sample properties of robust location and scale estimators, Communications in Statistics - Simulation and Computation 51, 2619-2645.

[0033] Rousseeuw, P. J., and Croux, C. (1993) Alternatives to the Median Absolute Deviation, Journal of the American Statistical Association 88, 1273-1283.

[0034] Anand, L., and Rodriguez Lopez, C. M. (2022) ChromoMap: an R package for interactive visualization of multi-omics data and annotation of chromosomes, BMC Bioinformatics 23, 33.

[0035] Ewels, P. A., Peltzer, A., Fillinger, S., Patel, H., Alneberg, J., Wilm, A., Garcia, M. U., Di Tommaso, P., and Nahnsen, S. (2020) The nf-core framework for community- curated bioinformatics pipelines, Nat Biotechnol 38, 276-278.

[0036] Harshil Patel, Jose Espinosa-Carrasco, Bjorn Langer, Phil Ewels, nf-core bot, Maxime U Garcia, Robert Syme, Alexander Peltzer, Adam Talbot, Drew Behrens, Gisela Gabemet, Mingda Jin, Matthias Hortenhuber, Jonatan Gonzalez Rodriguez, Kevin Menden, and An, O. (2023) nf-core / atacseq: [2.1.2] - 2022-08-07 (2.1.2), Zenodo.

[0037] Zhang, Y., Liu, T., Meyer, C. A., Eeckhoute, J., Johnson, D. S., Bernstein, B. E., Nusbaum, C., Myers, R. M., Brown, M., Li, W., and Liu, X. S. (2008) Model-based analysis of ChlP-Seq (MACS), Genome Biol 9, R137.

[0038] Robinson, J. T., Thorvaldsdottir, H., Winckler, W., Guttman, M., Lander, E. S., Getz, G., and Mesirov, J. P. (2011) Integrative genomics viewer, Nat Biotechnol 29, 24-26.

[0039] Ramirez, F., Ryan, D. P., Gruning, B., Bhardwaj, V., Kilpert, F., Richter, A. S., Heyne, S., Dundar, F., and Manke, T. (2016) deepTools2: a next generation web server for deep-sequencing data analysis, Nucleic Acids Res 44, W160-165.

[0040] Liao, Y., Smyth, G. K., and Shi, W. (2013) featureCounts: an efficient general purpose program for assigning sequence reads to genomic features, Bioinformatics (Oxford, England) 30, 923-930.10988-W001-SEC

[0041] Quinlan, A. R., and Hall, I. M. (2010) BEDTools: a flexible suite of utilities for comparing genomic features, Bioinformatics 26, 841-842.

[0042] Cheshire, C., charlotte, w., Ronkko, T., bot, n.-c., Patel, H., tamara, h., Ladd, D., Thiery, A., Fields, C., Deu-Pons, J., Ewels, P., Moller, S., and Menden, K. (2023) nf- core / cutandrun: nf-core / cutandrun v3.2.1 Mercury Turkey, Zenodo.

[0043] Meers, M. P., Tenenbaum, D., and Henikoff, S. (2019) Peak calling by Sparse Enrichment Analysis for CUT&RLTN chromatin profiling, Epigenetics & Chromatin 12, 42.

[0044] Ernst, J., and Kellis, M. (2017) Chromatin-state discovery and genome annotation with ChromHMM, Nature Protocols 12, 2478-2492.

[0045] Concordet, J. P., and Haeussler, M. (2018) CRISPOR: intuitive guide selection for CRISPR / Cas9 genome editing experiments and screens, Nucleic Acids Res 46, W242- W245.

[0046] Tasan, I., Sustackova, G., Zhang, L., Kim, J., Sivaguru, M., HamediRad, M., Wang, Y., Genova, J., Ma, J., Belmont, A. S., and Zhao, H. (2018) CRISPR / Cas9-mediated knock-in of an optimized TetO repeat for live cell imaging of endogenous loci, Nucleic Acids Res 46, elOO.

[0047] Diep, J., Le, H., Le, K., Zasadzinska, E., Tat, J., Yam, P., Zastrow, R., Gomez, N., and Stevens, J. (2021) Microfluidic chip-based single-cell cloning to accelerate biologic production timelines, Biotechnol Prog 37, e3192.

Claims

10988-W001-SECCLAIMSWhat is claimed:

1. A method of expressing a transgene, said method comprising integrating the transgene into a CHO cell locus, wherein the locus has coordinates selected from the group consisting of: a. NC_048599.1 :42048280-44048302 (positive strand); b. NW_023276806.1 :86353406-88353428 (negative strand); c. NC_048595.1 : 14745469-16745491 (negative strand); and d. NC_048596.1 :52895116-54895138 (positive strand); wherein the coordinates are based on CriGri-PICRH-1.0 genome assembly.

2. The method of Claim 1, said method comprising integrating the transgene into a locus, wherein the locus has coordinates selected from the group consisting of: a. NC_048599.1 :43048280-43048302 (positive strand); b. NW_023276806.1 :87353406-87353428 (negative strand); c. NC_048595.1 : 15745469-15745491 (negative strand); and d. NC_048596.1 :53895116-53895138 (positive strand); wherein the coordinates are based on CriGri-PICRH-1.0 genome assembly.

3. The method of Claim 1, wherein the coordinates are NC_048599.1 :42048280- 44048302 (positive strand).

4. The method of Claim 1, wherein the coordinates are NW_023276806.1 : 86353406- 88353428 (negative strand).

5. The method of Claim 1, wherein the coordinates are NC_048595.1 : 14745469- 16745491 (negative strand).

6. The method of Claim 1, wherein the coordinates are NC 048596.1 :52895116- 54895138 (positive strand).

7. The method of Claim 2, wherein the coordinates are NC_048599.1 :43048280- 43048302 (positive strand).

8. The method of Claim 2, wherein the coordinates are NW_023276806.1 :87353406- 87353428 (negative strand).

9. The method of Claim 2, wherein the coordinates are NC_048595.1 : 15745469- 15745491 (negative strand).10988-W001-SEC10. The method of Claim 2, wherein the coordinates are NC 048596.1 :53895116- 53895138 (positive strand).

11. The method of any one of Claims 1-10, wherein one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, or fifteen copy or copies of the transgene is expressed in the locus.

12. The method of any one of Claims 1-11, wherein one copy of the transgene is expressed in the locus.

13. The method of any one of Claims 1-11, wherein ten copies of the transgene is expressed in the locus.

14. A cell comprising a transgene integrated into a locus, wherein the locus has coordinates selected from the group consisting of: a. NC_048599.1 :42048280-44048302 (positive strand); b. NW_023276806.1 :86353406-88353428 (negative strand); c. NC_048595.1 : 14745469-16745491 (negative strand); and d. NC_048596.1 :52895116-54895138 (positive strand); wherein the coordinates are based on CriGri-PICRH-1.0 genome assembly.

15. A cell comprising a transgene integrated into a locus, wherein the locus has coordinates selected from the group consisting of: a. NC_048599.1 :43048280-43048302 (positive strand); b. NW_023276806.1 :87353406-87353428 (negative strand); c. NC_048595.1 : 15745469-15745491 (negative strand); and d. NC_048596.1 :53895116-53895138 (positive strand); wherein the coordinates are based on CriGri-PICRH-1.0 genome assembly.

16. The cell of Claim 14 or Claim 15, which is a CHO cell.

17. A transgene expressed by a method of any one of Claims 1-13, or the cell of Claim 14 or Claim 15.

18. The method of any one of Claims 1-13, or the cell of Claim 14 or Claim 15, wherein the transgene is integrated into the locus by homologous recombination, gene editing or recombinase-mediated recombination.

19. The method of any one of Claims 1-13, or the cell of Claim 14 or Claim 15, wherein the transgene is integrated into the locus by CRISPR or TALEN.

20. The method of any one of Claims 1-13, or the cell of Claim 14 or Claim 15, wherein the transgene is integrated into the locus by recombinase-mediated recombination with BxPl or CRE-Lox.