Methods for high-throughput identification of genomic harbors for transgene expression in biological samples

The use of self-recording CRISPR guide RNAs addresses the inefficiencies in transgene insertion by quantifying mutation fractions, identifying optimal genomic harbors for stable transgene expression, enhancing gene therapy and biomanufacturing efficiency.

WO2025245448A1PCT designated stage Publication Date: 2025-11-27JOHNS HOPKINS UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/030774
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-24
Filing Date
2025-05-23
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Current methods for identifying suitable genomic locations for transgene insertion are inefficient and often result in unstable or silenced expression, lacking a systematic approach to characterize insertion sites across different cell types and states.

Method used

A high-throughput method using self-recording CRISPR guide RNAs (hgRNAs) that integrate pseudo-randomly into the genome, recording their expression via Cas9-induced mutagenesis, allowing for the quantification of mutation fractions to identify optimal harbors for transgene expression through sequencing.

Benefits of technology

Enables the rapid identification of safe and stable genomic locations for transgene expression, facilitating targeted gene therapy and biomanufacturing by reducing the need for extensive screening and ensuring consistent transgene activity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025030774_27112025_PF_FP_ABST
    Figure US2025030774_27112025_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are methods of high-throughput identification to identify sites for transgene integration in the genome of biological samples. The method integrates a self-recording nucleic-acid cassette across the genome in a massively parallel manner, and the recording is activated by introducing Cas9 protein into the system. Quality of the insertion site is determined by analyzing the number of mutations by recording process in a polyclonal expansion group of cells in a biological sample.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS FOR HIGH-THROUGHPUT IDENTIFICATION OF GENOMIC HARBORS FOR TRANSGENE EXPRESSION IN BIOLOGICAL SAMPLES

[0002] CROSS-REFERENCE TO RELATED APPLICATION

[0003] This application claims priority to U.S. Provisional Patent Application No. 63 / 651,692, filed May 24, 2024. The content of this application is incorporated herein by reference in its entirety.

[0004] SEQUENCE LISTING

[0005] This application contains a Sequence Listing that has been submitted electronically as an XML file named “44807-0482WOl_ST26_SL.” The XML file, created on May 7, 2025, is 13,204 bytes in size. The material in the XML file is hereby incorporated by reference in its entirety.

[0006] STATEMENT REGARDING FEDERAL FUNDING

[0007] This invention was made with government support under DA056556 awarded by the U.S. National Institutes of Health and the U.S. Department of Health and Human Services. The government has certain rights in the invention.

[0008] TECHNICAL FIELD

[0009] The present disclosure relates to the area of high throughput identification of locations for insertions of transgene in a genome.

[0010] BACKGROUND

[0011] Many gene-editing and transgenic technologies and molecular circuits used in synthetic biology approaches require insertion of a transgene into the genome. However, different loci show variable behavior upon transgene insertion. In some loci, inserted transgenes are silenced broadly or in tissue-specific fashion. In some others, inserted transgenes can express broadly in most tissues, such as the ROSA26 locus and the AAVS1 locus. See e.g., Irion et al., “Identification and targeting of the ROSA26 locus in human embryonic stem cells,” Nature Biotechnology volume 25, pagesl477-1482 (2007) and Dutheil et al., “Adeno-associated virus site-specifi cally integrates into a muscle-specific DNA region,” Proc Natl Acad Sci U S A. 2000 Apr 25; 97(9): 4862-4866, each of which in incorporated by reference in its entirety. Because the behavior of an insertion in a given locus in a given cell type cannot be determined a priori, there is a need to develop platform technology that allows identification of insertion sites for transgene expression in the genome of a cell.

[0012] The present disclosure introduces a platform for characterizing the behavior of inserted transgenes. This platform depends on small nuclear ribonucleic acids (snRNAs) that record their own expression. For example, CRISPR / Cas depends on an snRNA called the single guide RNA (sgRNA) which confers specificity to the system by identifying the target DNA locus. Our platform of recording the behavior of inserted locus uses a homing guide RNA (hgRNA) scaffold that directs the Cas9-hgRNA complex to target the DNA locus of the hgRNA itself. See Kalhor et al., Nat Methods. 2017 Feb; 14(2): 195-200, which is incorporated by reference in its entirety. After insertion, the hgRNA expression and related molecular events (e.g., recording of molecular events into genomic DNA) can be tracked. See e.g., Perli et al., Science. 2016 Sep 9;353(6304):aag0511, which provided a memory system for storing analog biological information in the form of accumulating DNA mutations in human cells. Perli et al. is incorporated by reference in its entirety. As an example, a protocol that generates barcodes based on synthetically-induced mutations in two MARC1 mice cell lines to enable lineage analysis has been described previously. See Leeper et al., Nat Protoc. 2021 Apr; 16(4): 2088-2108, which is incorporated by reference in its entirety.

[0013] SUMMARY OF THE DISCLOSURE

[0014] Disclosed herein are methods for high-throughput screening of harbors (i.e., locations) in the genome (i.e., on a genome-wide scale). The present method is based on a self-recording CRISPR guide RNAs (hgRNA) construct that is -400 base pairs in length and contains a U6 promoter driving the expression of an hgRNA. The hgRNA pseudo-randomly inserts itself into the genome, which self-records its expression at the locus via Cas9-induced mutagenesis. This method distills the intricate relationship between transgene expression and local chromatin properties into a discrete unit: mutation of the locus. The resulting DNA-stored information on transgene expression level can be quantified as the mutation fraction within a clonal population, leveraging high-throughput sequencing to simultaneously read out the genomic location and the mutation. As such, the present method overcomes the throughput limitations of conventional assays by coupling transgene locus and its expression in a single assay.

[0015] More specifically, when the hgRNA is expressed in the presence of Cas9 nuclease, it targets its own locus (i.e., the construct) for double-strand breaks, leading to mutations. The overall amount of mutations in an hgRNA locus across a cell population represents the expression level of hgRNA. Therefore, the hgRNA construct records its expression level in its DNA sequence in form of mutations.

[0016] Overall, the present approach identifies suitable harbors for expressing transgenes in target cells, which are commercially valuable in cell therapy and biomanufacturing spaces. For example, the methods disclosed herein can be used to identify safe and open harbors for modifying CAR-T cells. It can be used to identify safe and open harbors for modifying pluripotent stem cells. It can also be used for identifying harbors in CHO cells or other biomanufacturing cell lines that maximize the production of the biologic of interest without causing instability in the cell line. It can also be used to identify harbors that are open in specific cell states and cell type, which can be used to regulate transgene expression or limit transgene expression to desired cells, cell states, times, or locations.

[0017] In some instances, the methods disclosed herein identify cell-specific locations for transgene insertion. The locations in the genome identified as possible harbor sites can be protected with the insertion of a switch. See e.g., Castanon et al, “CRISPR-mediated biocontainment,” 2020 (doi.org / 10.1101 / 2020.02.03.922146), which is incorporated by reference in its entirety. In some instances, the switch is a kill switch. The switch can also be an activation switch or a cell-state transition switch. One skilled in the art can design the switch based on identification of the genomic location for insertion. One can also place the switch in a location that one does not want transgene expression. It is appreciated that the switch can be designed to be cell-specific. For example, if one wishes to ensure that a gene is hepatocyte-specific but does not express in stem cells, they have to use a hepatocyte-specific promoter that is silent in stem cells. In some instances, a switch or a transgene can be inserted with a universal promoter in a cell-specific harbor that is known to be closed in the another cell, thereby achieving cell specific expression without cell specific promoter. In some instances, cell-specific harbors are used in place of cell-specific promoters to accomplish cell-specific expression. In some instances, cell- specific harbors are used in conjunction with cell-specific promoters to accomplish highly specific expression.

[0018] Finding insertion sites in the genome for transgene expression can be difficult. Current loci were largely identified by trial and error and only a limited number of them exist (e.g., Rosa26 in mice or AAV1T1 in humans). The present disclosure overcomes the problem of finding an appropriate genomic location — or harbors — in cell types of interest for expressing transgenes. Genetic modification of human cells requires such harbors that are both safe (do not cause cancer) and allow for effective transgene expression. Cell and regenerative therapies require such loci.

[0019] Additionally, biomanufacturing of biologies (e.g., monoclonal antibodies, hormones, and gene therapy delivery vectors) relies on integrating transgenes of interest in cell line genomes. Because most integrations are either unstable or get silenced, a large number of clones must be screened at a heavy cost to isolate stable ones. This disclosure can be used to identify the stable harbors in producer lines (e.g., CHO) for producing transgenes of interest. It can further be used to generate designer producer lines (e.g., CHO) that have engineered harbors ready for integration of transgene-expressing constructs. It will be appreciated that the methods of this disclosure can be developed into a screening platform or used to develop customized cell lines with superior performance.

[0020] Thus, disclosed herein is a method for high-throughput identification of one or more insertion sites for transgene expression in a genome of a biological sample. In some instances, the method includes: (a) integrating a self-recording nucleic acid sequence into one or more genomic loci of the biological sample, wherein the integrating induces one or more changes at the one or more genomic loci; (b) inducing the self-recording nucleic acid sequence to record its activity by generating changes in the self-recording nucleic acid sequence; and (c) determining the activity of the self-recording nucleic acid in each integrated genomic locus by (i) identifying the location of the self-recording nucleic acid in the genome and (ii) measuring the number of changes that are accumulated in the self-recording nucleic acid sequence, thereby characterizing the integrated genomic locus or loci as harbor or harbors for transgene expression.

[0021] In another aspect, disclosed herein is a method of identifying a location for random insertion of a nucleic acid into a genome of a biological sample. In some instances, the method includes: (a) integrating a self-recording nucleic acid sequence into one or more genomic loci of the biological sample, wherein the integrating induces one or more changes at the one or more genomic loci; (b) inducing the self-recording nucleic acid sequence to record its activity by generating changes in the self-recording nucleic acid sequence; and (c) determining the location of the random insertion of the nucleic acid (i) identifying the location of the self-recording nucleic acid in the genome and (ii) measuring the number of changes that are accumulated in the self-recording nucleic acid sequence.

[0022] In some instances, the methods also include introducing a transgenic nucleic acid sequence into the location of the genome. In some instances, the transgenic nucleic acid sequence comprises a sequence that encodes for all or part of an antibody or an antigen-binding fragment thereof. In some instances, the biological sample is Chinese hamster ovary cells.

[0023] In some instances, the method includes identifying a cell-type specific harbor that allows transgene expression in a subset of cell types but not others. In some instances, the method includes identifying a cell-state specific harbor that allows transgene expression in a subset of cell states but not others.

[0024] In yet another aspect, disclosed is a method for introducing a self-recording nucleic acid sequence into one or more random locations of a biological sample. In some instances, the method includes (a) contacting the biological sample with a gene transfer system, thereby introducing the self-recording nucleic acid sequence into the genome of the biological sample; (b) inducing the self-recording nucleic acid sequence to record its activity by generating changes in the self-recording nucleic acid sequence; and (c) determining the location of an insertion of the self-recording nucleic acid in a genome of the biological sample by (i) identifying the location of the self-recording nucleic acid in the genome and (ii) measuring the number of changes that are accumulated in the self-recording nucleic acid sequence.

[0025] In another aspect, the disclosure includes a method of identifying putative genomic safe harbor locations within a genome in a biological sample. In some instances, the method includes introducing into multiple genomic sites of the genome in the biological sample a plurality of self- recording nucleic acid sequences, wherein each self-recording nucleic acid sequence of the plurality is capable of self-recording its own activity; allowing the self-recording nucleic acid sequences of the plurality to record their activity at their genomic sites; and measuring the recorded activity at each site to identify putative genomic safe harbor locations. In some instances, the methods also include introducing a transgenic nucleic acid sequence into the location of the genome. In some instances, the transgenic nucleic acid sequence comprises a sequence that encodes for all or part of an antibody or an antigen-binding fragment thereof. In some instances, the biological sample is Chinese hamster ovary cells. In some instances, the location is modified by the insertion of a nucleic acid switch. In some instances, the nucleic acid switch is a kill switch, an activation switch, or a cell-state transition switch. In some instances, the location for insertion of the nucleic acid switch is cell-type specific. In some instances, the location for insertion of the nucleic acid switch is specific for T- cells. In some instances, the transgenic nucleic acid sequence comprises a universal primer. In some instances, the nucleic acid switch is an inducible nucleic acid switch. In some instances, the nucleic acid switch is a capable of inducing nucleic acid breaks at one or more locations in the genome.

[0026] In some instances, the self-recording nucleic acid sequence is a homing CRISPR guide RNA. In some instances, the activity is RNA expression. In some instances, the inducing utilizes a S. pyogenes Cas9 nuclease. In some instances, the changes are mutations to the self-recording nucleic acid sequence. In some instances, the changes are changes to the methylation status at the location. In some instances, the changes are histone modifications at the location. In some instances, the histone modification comprises histone acetylation, histone deacetylation, histone methylation, histone ubiquitination, and histone citrullination.

[0027] In some instances, the methods disclosed in this application include isolating the genomic DNA from the biological sample after induction; treating the biological sample with an endonuclease that cuts the genomic DNA at sites outside but flanking the self-recording nucleic acid sequence, thereby generating fragments of genomic DNA that comprise all or part of the self-recording nucleic acid sequence and a part of a sequence of genomic DNA endogenous to the genome of the biological sample; ligating adaptors to the fragments; optionally amplifying the fragments; and sequencing (i) all or part of the self-recording nucleic acid sequence and (i) a part of a sequence of genomic DNA endogenous to the genome of the biological sample. In some instances, all or part of the self-recording nucleic acid sequence comprises a mutation induced by a nuclease upon expression of the self-recording nucleic acid sequence.

[0028] Also disclosed herein is a method of creating a transgenic biological sample. In some instances, the method includes (a) introducing into the biological sample a self-recording nucleic acid sequence; and (b) identifying that the self-recording nucleic acid sequence has randomly incorporated into the genome of the biological sample, thereby creating the transgenic biological sample.

[0029] In some instances, the biological sample comprises a primary cell, a cell in culture, or a cell from a cell line. In some instances, the biological sample is from a tissue sample. In some instances, the self-recording nucleic acid sequence is part of a sequence of a vector comprising one or more of a scaffolding sequence, a protospacer adjacent motif (PAM) sequence, and spacer sequence, and any combination thereof. In some instances, the vector comprises one or more selection markers. In some instances, the vector comprises one or more detectable markers. In some instances, the methods includes detecting expression of part of the vector inserted into the genome using the one or more detectable markers. In some instances, the self-recording nucleic acid sequence forms a homing complex using a Cas9 protein at the location. In some instances, the Cas9:homing complex targets the location of the genome for a double-strand break at the location. In some instances, the methods include repairing the double-stranded break, thereby inducing the one or more mutations. In some instances, the methods also include integrating the self-recording system to the biological sample using a transposase. In some instances, the transposon is a piggyBac transposase, a sleeping beauty transposase, a CRISPR transposase. In some instances, the methods also include treating the genome with a restriction enzyme, thereby generating a digested self-recording nucleic acid fragment that comprises the one or more mutations and a part of a sequence of the genome of the biological sample.

[0030] In some instances, the methods include attaching an adaptor sequence(s) to the 3’ and / or 5’ ends of the digested self-recording nucleic acid fragment. In some instances, the attaching comprises ligating the adaptor sequences(s). In some instances, the methods include amplifying the digested self-recording nucleic acid fragment. In some instances, the determining step comprises sequencing the digested self-recording nucleic acid fragment.

[0031] Also provided herein is a method of identifying a location for a random insertion of a nucleic acid into a genome of a biological sample. In some instances, the method includes (a) providing the biological sample; (b) adding to the biological sample a vector comprising a self- recording nucleic acid sequence and a transposon; (c) randomly integrating the self-recording nucleic acid sequence into one or more genomic loci of the genome of the biological sample, wherein the integrating induces one or more mutations at one or more genomic loci; (d) inducing the self-recording nucleic acid sequence to record its activity by generating changes in the self- recording nucleic acid sequence; (e) adding a restriction enzyme to the biological sample, thereby generating a digested self-recording nucleic acid fragment; (f) ligating adaptors to the digested self-recording nucleic acid fragment; (g) amplifying the digested self-recording nucleic acid fragment; and (h) sequencing the digested self-recording nucleic acid fragment, or a complement thereof, thereby identifying a location of random insertion of a nucleic acid into a genome of a biological sample.

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Although methods and materials similar or equivalent to those described herein can be used to practice the disclosure, suitable methods and materials are described below.

[0033] Where values are described in terms of ranges, it should be understood that the description includes the disclosure of all possible sub-ranges within such ranges, as well as specific numerical values that fall within such ranges irrespective of whether a specific numerical value or specific sub-range is expressly stated.

[0034] The terms “each,” when used in reference to a collection of items, is intended to identify an individual item in the collection but does not necessarily refer to every item in the collection, unless expressly stated otherwise, or unless the context of the usage clearly indicates otherwise.

[0035] All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.

[0036] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the disclosure will be apparent from the description and drawings, and from the claims.

[0037] BRIEF DESCRIPTION OF DRAWINGS

[0038] FIG. 1 shows a flow chart of the methods of the disclosure.

[0039] FIG. 2A shows a vector diagram showing a self-recording nucleic acid sequence in a vector having a selection marker. FIG. 2B shows a vector diagram showing a self-recording nucleic acid sequence in a vector without a selection marker.

[0040] FIG. 3 shows a schematic of amplifying single clones using polyclonal expansion followed by Cas9 treatment to identify the mutational fraction of a clone in a biological sample.

[0041] FIG. 4A shows a ranking of hgRNA loci in a mouse genome based on the expression of hgRNAs in the MARC1 mouse model. The in vivo expression of hgRNAs in adult mice was captured using RNA-seq. This figure displays RNA read counts of hgRNAs on a symmetric log scale (y-axis) plotted against their expression level rank.

[0042] FIG. 4B shows a comparison of hgRNA ranks in RNA expression level between two animals demonstrating a high degree of concordance, with a Spearman's rank correlation coefficient of 0.94.

[0043] FIG. 5 shows a schematic of human chromosomes depicting a comprehensive display of insertion locations and their hgRNA mutation fractions along with ATAC-seq scores.

[0044] FIG. 6 shows a histogram of mutation fractions for 1,380 hgRNA insertion locations, each with more than 40 UMIs from date shown in FIG. 5. This figure identifies the number of closed harbor sites, intermediate harbor sites, and open harbor sites in K562 cells.

[0045] FIG. 7A shows a scatter plot of mutation fraction (y-axis) versus mean ATAC-seq signal within a ±500 bp range of the insertion site (x-axis). The plot includes an OLS regression line with the corresponding slope (0.11) and Pearson R value (0.10).

[0046] FIG. 7B shows a scatter plot of mutation fraction (y-axis) versus logio transformed mean ATAC-seq signal within a ±500 bp range of the insertion site (x-axis). The plot includes an OLS regression line with the corresponding slope (12.58) and Pearson R value (0.23).

[0047] FIG. 8 shows a graph providing correlation of pairs of hgRNAs in adjacent locations. Thick solid line: Represents the mean Pearson R of sets of adjacent pairs within each window; dashed line: represents the number of adjacent pairs within each window; dotted line (roughly horizontal to the x-axis): represents the correlation of 100 randomly selected hgRNA pairs.

[0048] FIGs. 9A-9C show genome-wide mapping of transgene expression in marmoset embryonic stem cells (ESCs). High-throughput transgene expression assay in marmoset genome identified 2,181 loci with more than 50 unique molecular identifiers (UMIs). FIG. 9A shows chromosomal ideogram binned at 1 Mb, with white-to-black color scale for insertion count and blue-to-red color scale indicating the mean hgRNA mutation fraction of the inserts per bin. FIG. 9B is a histogram showing the bimodal distribution of mutation fractions across all loci. FIG. 9A shows a box-and- whisker plots of mutation fraction by chromosome; extreme outliers represent hyperactive sites that emerge as leading genomic safe-harbor candidates.

[0049] FIGs. 10A-10C show functional context of high-confidence insertion sites. FIG. 10A shows genomic distribution of insertions. About 42 % map to unannotated intergenic regions. FIG. 10B shows the remainder intersect protein-coding genes (51.2 %), long non-coding RNAs (6.2 %), or other biotypes (< 1 %). FIG. IOC shows mutation fraction profiles for intragenic versus intergenic loci. Although the overall MF distributions are similar, the most transcriptionally permissive “hyper-active” sites (outliers) are enriched in intergenic space.

[0050] DETAILED DESCRIPTION

[0051] A. Background

[0052] Expression of exogenous or synthetic genes in a cell often requires their insertion into the genome. However, the site of genomic insertion — also referred throughout this application as a “harbor” — determines how well the inserted constructs express: some harbors silence the construct while others allow the construct to express at differing levels. See e.g., Chen et al., Nat Struct Mol Biol. 2017 Jan;24(l):47-54, which is incorporated by reference in its entirety. Moreover, harbors perform differently in different cell types. An active harbor in one cell type can be inactive in another. There has been no method to systematically characterize the potential of various loci in the genome for characterizing inserted construct expression.

[0053] As will be evident in the disclosure, hgRNA expression levels can be different based on the harbor in which they integrate. Self-recording guide RNAs (also called homing guide RNAs or hgRNAs) (which include a “self-recording (or homing) nucleic acid sequence”) target their own loci to create highly diverse and unique random mutations. Using these self-recording nucleic acid sequences, one can collect a large volume of data compared to other means of identification, such as cell sorting or fluorescent expression. This is because high-throughput sequencing of each insertion site in the genome can be performed, and has the capability to identify hundreds, thousands, or even millions of insertion sites in a single genome.

[0054] This disclosure provides multiple ways of recording the insertion of the hgRNA (or self- recording) nucleic acid into the genome. Recording, as used herein, refers to a process by which the self-recording nucleic acid modifies itself. In some instances, the modification is permanent. In some instances, the modification is on a timeline long enough to complete an experiment. For instance, clonal expansion can be performed, in which multiple (e.g., 2, 4, 6, 8, 10, 20, 50, 100, 500, 1000, or more) rounds of cell division occur. In these instances, the self-recording mutation will be maintained.

[0055] In one instance, upon insertion, one or more mutations are introducing a mutation into the self-recording nucleic acid sequence itself. By identifying these sequences (e.g., via high- throughput sequencing), one can identify the location — and frequency — of the mutation.

[0056] In another aspect, the recording can be performed by inducing epigenetic changes. In some instances, the change induces a different methylation status compared to the previous methylation status (i.e., before the experiment). This methylation status can be measured using methods known in the art, including ATAC-seq (e.g., methyl-ATAC-seq), sodium bisulfite conversion and sequencing, differential enzymatic cleavage of DNA, and affinity capture of methylated DNA.

[0057] In another aspect, the epigenetic changes that are recorded include one or more histone modifications. The one or more modifications can also include one or multiple events of acetylation, deacetylation, methylation, ubiquitination, citrullination, and phosphorylation of specific amino acids within the histone protein. In some instances, the histone modification can be measured using Chromatin Immunoprecipitation (ChIP), ChlP-seq, mass spectrometry, or other methods known in the art.

[0058] As a final example of self-recording, one can introduce modified nucleotides into the site of mutation. For example, nucleotides can be modified using non-natural nucleotides such as locked-nucleic acids, or using nucleotides that are tagged with a detectable moiety such as a fluorescent moiety. For instance, labeled ddNTPs can be used.

[0059] Some harbors only allow low or no expression while others allow high expression. Based on this observation, Applicant has developed a method to randomly integrate hgRNA constructs across the genome in a population of cells of interest using transposases. A Cas9 protein is expressed to induce mutations in these constructs. Genomic DNA is then extracted from the cells and simultaneously sequence the mutation level of hgRNA constructs together with their genomic position. This approach allows us to quantify the expression level of hgRNA construct in a large number of genomic harbors. This expression level is a representative of each genomic locations harbor potential in the cell type analyzed. Thus, provided herein are methods for high-throughput identification of one or more insertion sites for transgene expression in a genome of a biological sample. Also included in this disclosure are methods of identifying a location of random insertion of a nucleic acid into a genome of a biological sample. In addition, some methods disclosed herein include introducing a self- recording nucleic acid sequence into one or more random locations of a biological sample. In some instances, the methods include the steps of contacting the biological sample with a gene transfer system, thereby introducing the self-recording nucleic acid sequence into the genome of the biological sample, and determining the location of an insertion of the self-recording nucleic acid in a genome of the biological sample by (i) identifying the location of the insertion in the genome and (ii) measuring the number of one or more mutations accumulated in the self-recording nucleic acid sequence.

[0060] Additional methods and compositions are further described.

[0061] B. Methods and Uses

[0062] Referring to FIG. 1, the methods disclosed herein are treated with one or more hgRNA constructs. A transposase such as a piggyBac transposon can be used. A protein such as Cas9 can be introduced to activate recording (e.g., induce mutations of the hgRNA transcript into the locus). Mutations are accumulated in the genome, and after a given time for recording, genomic DNA can be harvested from the cell population. The genomic DNA can be digested with a chosen restriction enzyme, such as DpnII. It is appreciated that many restriction enzymes can be used, with the understanding of selecting one that (1) targets the hgRNA sequence and (2) has limited sites in the human genome. Ideally, the enzyme does not recognize downstream U6 promoter up to the next recognition site in the genome where hgRNA is inserted. Accordingly, it would result in a digested genome sequence that includes both part of the hgRNA sequence (or a complement thereof) and part of the endogenous genome sequence, which allows a user to identify the location of the insertion. Ligation adaptors are added to the digested genomic DNA fragments and sequencing can be performed.

[0063] In some instances, provided is a method for high-throughput identification of one or more insertion sites for transgene expression in a genome of a biological sample. In some instances, the methods include integrating a self-recording nucleic acid sequence into one or more genomic loci of the biological sample, wherein the integrating induces one or more changes at the one or more genomic loci; inducing the self-recording nucleic acid sequence to record its activity by generating changes in the self-recording nucleic acid sequence; and determining the activity of the self-recording nucleic acid in each integrated genomic locus by (i) identifying the location of the self-recording nucleic acid in the genome and (ii) measuring the number of changes that are accumulated in the self-recording nucleic acid sequence, thereby characterizing the integrated genomic locus or loci as harbor or harbors for transgene expression.

[0064] In some instances, the activity includes expression of RNA. That is, the step of recording the activity includes recording RNA expression of the self-recording nucleic acid sequence.

[0065] The methods described herein can also be used to identify a location of random insertion of a nucleic acid into a genome of a biological sample. In other aspects, the methods include introducing a self-recording nucleic acid sequence into one or more random locations of a biological sample.

[0066] In some instances, the methods described herein include providing a biological sample; adding to the biological sample a vector comprising a self-recording nucleic acid sequence and a transposon; randomly integrating the self-recording nucleic acid sequence into one or more genomic loci of the genome of the biological sample, wherein the integrating induces one or more mutations at one or more genomic loci; detecting expression of the self-recording nucleic acid sequence; adding a restriction enzyme to the biological sample, thereby generating a digested self-recording nucleic acid fragment; ligating adaptors to the digested self-recording nucleic acid fragment; amplifying the digested self-recording nucleic acid fragment; and sequencing the digested self-recording nucleic acid fragment, or a complement thereof, thereby identifying a location of random insertion of a nucleic acid into a genome of a biological sample. In some instances, all or part of the self-recording nucleic acid sequence comprises a mutation induced by a transposon upon expression of the self-recording nucleic acid sequence.

[0067] In some instances, integrating a self-recording nucleic acid sequence into one or more genomic loci of the biological sample, wherein the integrating induces one or more mutations at the one or more genomic loci. In some instances, the self-recording nucleic acid sequence forms a piggyBac: homing complex using a Cas9 protein at the location. In some instances, the Cas9:homing complex targets the location of the genome for a double-strand break at the location. In some instances, the methods include repairing the double-stranded break, thereby inducing the one or more mutations. In some instances, the methods include delivering a transposon to the biological sample. In some instances, transposon is a piggyBac transposon. In some instances, transposon is codelivered to the biological sample with the self-recording nucleic acid sequence. In some instances, co-delivered transposon facilitates random insertion of the self-recording nucleic acid sequence into the genome of the biological sample.

[0068] In some instances, the methods include determining a location in the genome of the self- recording nucleic acid insertion of the nucleic acid by (i) identifying the location of the mutation in the genome and (ii) measuring the number of the one or more mutations accumulated in the self-recording nucleic acid sequence, thereby identifying the genomic harbor for transgene expression.

[0069] In some instances, after insertion of the hgRNA sequence, the cell can be expanded prior to Cas9 expression. As shown in FIG. 3, a cell with an hgRNA nuclei acid is inserted into one location (called location A in FIG. 3). In another cell, insertion occurs at a second location (e.g., location B). After polyclonal expansion of each cell, Cas9 can be introduced. After Cas9 treatment, which induces mutations, the cells can be analyzed for expression of an induced mutation. Given that the same sequence has been introduced, the percentage of mutations introduced into the population can be measured. This number results in a “mutation fraction” of insertion at that location can be determined. A higher mutation fraction at a specific area indicates that this locus in the genome is an open harbor, and that one or more transgenes can be inserted at this region more readily than loci with lower mutation fraction.

[0070] C. Vectors, Transposons, and Systems

[0071] The present disclosure provides a genome-targeting nucleic acid (e.g., vector) that can direct the activities of an associated polypeptide (e.g., a site-directed polypeptide) to a specific target sequence within a target nucleic acid. The genome-targeting nucleic acid can be an RNA. A genome-targeting RNA is referred to as a “self-recording” nucleic acid (and the like), a “guide RNA” or “hgRNA” herein. In some instances, the self-recording nucleic acid sequence is a homing CRISPR guide RNA. A guide RNA comprises at least a spacer sequence that hybridizes to a target nucleic acid sequence of interest, and a CRISPR repeat sequence. In Type II systems, the gRNA also comprises a second RNA called the tracrRNA sequence. In the Type II guide RNA (gRNA), the CRISPR repeat sequence and tracrRNA sequence hybridize to each other to form a duplex. In the Type V guide RNA (gRNA), the crRNA forms a duplex. In both systems, the duplex binds a site-directed polypeptide, such that the guide RNA and site-direct polypeptide form a complex. In some embodiments, the genome-targeting nucleic acid provides target specificity to the complex by virtue of its association with the site-directed polypeptide. The genome-targeting nucleic acid thus directs the activity of the site-directed polypeptide.

[0072] In some embodiments, the genome-targeting nucleic acid is a double-molecule guide RNA. In some embodiments, the genome-targeting nucleic acid is a single-molecule guide RNA.

[0073] In some instances, the vector comprises one or more scaffolding sequences, protospacer adjacent motif (PAM) sequences, spacer sequences, and any combination thereof. The vector can also include one or more selection markers that confers ampicillin or puromycin resistance.

[0074] Any vector that will facilitate insertion of hgRNA or a transgene can be used. In some instances, viral vectors such as retroviruses or lentiviruses can be used for introduction of a transgene into a cell. In some instances, the viral vectors used for CRISPR / Cas delivery include adeno-associated viruses (AAVs), adenoviral vectors (AdVs), and lentiviral vectors (LVs). AAVs are small, non-enveloped single-stranded DNA viruses that are not pathogenic to humans. These members of the Parvoviridae family have attracted the attention of researchers as gene delivery systems. An AdV is a double-strand DNA virus that can transduce a broad spectrum of both dividing and non-dividing cells. These vectors can carry a genetic cargo up to 37 kb, and following transduction, they create an episomal DNA adjacent to the host DNA instead of integrating into the genome. The HIV-l-derived LVs are single-stranded RNA (ssRNA) viruses primarily used for integrating the desired transgene into dividing and non-dividing cells. Their delivery capacity for genetic cargo is around 9 kb, encompassed by a lipid-enriched capsid.

[0075] In some instances, the vector comprises one or more detectable markers such as a radioisotope, a fluorophore, a chemiluminescent compound, a bioluminescent compound, or a combination thereof. In various embodiments, the fluorescent label is selected from the group consisting of Atto dyes, Alexafluor dyes, quantum dots, Hydroxycoumarin, Aminocouramin, methoxy Coumarin (Methoxycourmarin), Cascade Blue, Pacific Blue, Pacific Orange, Lucifer Yellow, NBD, R-phycoerythrin (PE), PE-Cy5 conjugate, PE-Cy7 conjugate, Red 613, PerCP, TruRed , FluorX, Fluorescent Yellow, BODIPY-FL, Cy2, Cy3, Cy3B, Cy3.5, Cy5, Cy5.5, Cy7, TRITC, X-Rose Red, Lissamine Rose Red B, Texas Red, Allophycocyanin (APC ), APC-Cy7 conjugate, Indo-1, Fluo-3, Fluo-4, DCFH, DHR, SNARF, GFP (Y66H mutation), GFP (Y66F mutation), EBFP, EBFP2, Azurite ), GFPuv, T-Sapphire, Cerulean, mCFP, mTurquoise2, ECFP, CyPet, GFP (Y66W mutation), mKeima-Red, TagCFP, AmCyanl, mTFPl, GFP (S65A mutation), Midorishi Cyan, wild Type GFP, GFP (S65C mutation), TurboGFP, TagGFP, GFP (S65L mutation), Emerald, GFP (S65T mutation), EGFP, Azami Green, ZsGreenl, TagYFP, EYFP, Topaz, Venus, mCitrine, YPet, TurboYFP, ZsYellowl, Kusabira Orange, mOrange, Allophycocyanin (APC), mKO, TurboRFP, tdTomato, TagRFP, DsRed monomer, DsRed2 (“RFP”), mStrawberry, TurboFP602, AsRed2, mRFPl, J-Red, R-Phycoerythrin ( RPE), B- phycoerythrin (BPE), mCherry, HcRedl, Katusha, P3, dinophycoxanthin-chlorophyll protein complex (PerCP), mKate (TagFP635), TurboFP635, mPlum and mRaspberry. In some instances, the detectable marker is GFP.

[0076] Additional elements can be included in the vector. These additional elements include, without limitation, promoters (e.g., a EF-l-alpha promoter) and enhancers. In some embodiments, a vector comprises one or more pol III promoter (e.g., 1, 2, 3, 4, 5, or more pol III promoters), one or more pol II promoters (e.g., 1,2,3, 4, 5, or more pol II promoters), one or more pol I promoters (e.g., 1,2, 3, 4, 5, or more pol I promoters), or combinations thereof. Examples of pol III promoters include, but are not limited to, U6 and Hl promoters. Examples of pol II promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer), (see, e.g., Boshart et al, Cell, 41 :521-530 (1985), which is incorporated by reference in its entirety), the SV40 promoter, the dihydrofolate reductase promoter, the - actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EFla promoter. Also encompassed by the term "regulatory element" are enhancer elements, such as WPRE; CMV enhancers; the R-U51segment in LTR of HTLV-I; SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit P-globin. It will be appreciated by those skilled in the art that the design of the expression vector can depend on such factors as the choice of the host cell to be transformed, the level of expression desired, etc.

[0077] Representative vectors are shown in FIGs. 2A and 2B.

[0078] In some instances, the vectors includes one of the following sequences: (SEQ ID NO:1)

[0079] 1 actcttcctt tttcaatatt attgaagcat ttatcagggt tattgtctca tgagcggata 61 catatttgaa tgtatttaga aaaataaaca aataggggtt ccgcgcacat ttccccgaaa 121 agtgccacct aaattgtaag cgttaatatt ttgttaaaat tcgcgttaaa ttttgttaa 181 atcagctcat tttttaacca ataggccgaa atcggcaaaa tcccttataa atcaaaagaa 241 tagaccgaga tagggttgag tgttgttcca gtttggaaca agagtccact attaaagaac 301 gtggactcca acgtcaaagg gcgaaaaacc gtctatcagg gcgatggccc actacgtgaa 361 ccatcaccct aatcaagttt tttggggtcg aggtgccgta aagcactaaa tcggaaccct 421 aaagggagcc cccgatttag agcttgacgg ggaaagccgg cgaacgtggc gagaaaggaa 481 gggaagaaag cgaaaggagc gggcgctagg gcgctggcaa gtgtagcggt cacgctgcgc 541 gtaaccacca cacccgccgc gcttaatgcg ccgctacagg gcgcgtccca ttcgccattc 601 aggctgcgca actgttggga agggcgatcg gtgcgggcct cttcgctatt acgccagctg 661 gcgaaagggg gatgtgctgc aaggcgatta agttgggtaa cgccagggtt ttcccagtca 721 cgacgttgta aaacgacggc cagtgagcgc gcctcgttca ttcacgtttt tgaacccgtg 781 gaggacgggc agactcgcgg tgcaaatgtg ttttacagcg tgatggagca gatgaagatg 841 ctcgacacgc tgcagaacac gcagctagat taaccctaga aagataatca tattgtgacg 901 tacgttaaag ataatcatgt gtaaaattga cgcatgtgtt ttatcggtct gtatatcgag 961 gtttatttat taatttgaat agatattaag ttttattata tttacactta catactaata 1021 ataaattcaa caaacaattt atttatgttt atttatttat taaaaaaaac aaaaactcaa 1081 aatttcttct ataaagtaac aaaactttta tgagggacag ccccccccca aagcccccag 1141 ggatgtaatt acgtccctcc cccgctaggg ggcagcagcg agccgcccgg ggctccgctc 1201 cggtccggcg ctccccccgc atccccgagc cggcagcgtg cggggacagc ccgggcacgg 1261 ggaaggtggc acgggatcgc tttcctctga acgcttctcg ctgctctttg agcctgcaga 1321 cacctggggg gatacgggga aaaggctagc cgataacttc gtataatgta tgctatacga 1381 agttatgata tcttgtcttc gttgggagtg aattagccct tccagtcccc ttttcttttg 1441 ttcgatgcat ggggtcgtgc gctcctttcg gtcgggcgct gcgggtcgtg gggcgggcgt 1501 caggcaccgg gcttgcgggt catgcaccag gtgcgcggtc cttcgggcac ctcgacgtcg 1561 gcggtgacgg tgaagccgag ccgctcgtag aaggggaggt tgcggggcgc ggaggtctcc 1621 aggaaggcgg gcaccccggc gcgctcggcc gcctccactc cggggagcac gacggcgctg 1681 cccagaccct tgccctggtg gtcgggcgag acgccgacgg tggccaggaa ccacgcgggc 1741 tccttgggcc ggtgcggcgc caggaggcct tccatctgtt gctgcgcggc cagccgggaa 1801 ccgctcaact cggccatgcg cgggccgatc tcggcgaaca ccgcccccgc ttcgacgctc 1861 tccggcgtgg tccagaccgc caccgcggcg ccgtcgtccg cgacccacac cttgccgatg 1921 tcgagcccga cgcgcgtgag gaagagttct tgcagctcgg tgacccgctc gatgtggcgg 1981 tccggatcga cggtgtggcg cgtggcgggg tagtcggcga acgcggcggc gagggtgcgt

[0080] 2041 acggccctgg ggacgtcgtc gcgggtggcg aggcgcaccg tgggcttgta ctcggtcatg

[0081] 2101 gttgtggcca tattatcatc gtgtttttca aaggaaaacc acgtccccgt ggttcggggg

[0082] 2161 gcctagacgt ttttttaacc tcgactaaac acatgtaaag catgtgcacc gaggccccag

[0083] 2221 atcagatccc atacaatggg gtaccttctg ggcatccttc agccccttgt tgaatacgct

[0084] 2281 tgaggagagc catttgactc tttccacaac tatccaactc acaacgtggc actggggttg

[0085] 2341 tgccgccttt gcaggtgtat cttatacacg tggcttttgg ccgcagaggc acctgtcgcc

[0086] 2401 aggtgggggg ttccgctgcc tgcaaagggt cgctacagac gttgtttgtc ttcaagaagc

[0087] 2461 ttccagagga actgcttcct tcacgacatt caacagacct tgcattcctt tggcgagagg

[0088] 2521 ggaaagaccc ctaggaatgc tcgtcaagaa gacagggcca ggtttccggg ccctcacatt

[0089] 2581 gccaaaagac ggcaatatgg tggaaaataa catatagaca aacgcacacc ggccttattc

[0090] 2641 caagcggctt cggccagtaa cgttaggggg gggggaggga gagggggggg gggcggaatt

[0091] 2701 ccgcgggccc gtcgacgcgg ccgctttact tgtacagctc gtccatgccg agagtgatcc

[0092] 2761 cggcggcggt cacgaactcc agcaggacca tgtgatcgcg cttctcgttg gggtctttgc

[0093] 2821 tcagggcgga ctgggtgctc aggtagtggt tgtcgggcag cagcacgggg ccgtcgccga

[0094] 2881 tgggggtgtt ctgctggtag tggtcggcga gctgcacgct gccgtcctcg atgttgtggc

[0095] 2941 ggatcttgaa gttcaccttg atgccgttct tctgcttgtc ggccatgata tagacgttgt

[0096] 3001 ggctgttgta gttgtactcc agcttgtgcc ccaggatgtt gccgtcctcc ttgaagtcga

[0097] 3061 tgcccttcag ctcgatgcgg ttcaccaggg tgtcgccctc gaacttcacc tcggcgcggg

[0098] 3121 tcttgtagtt gccgtcgtcc ttgaagaaga tggtgcgctc ctggacgtag ccttcgggca

[0099] 3181 tggcggactt gaagaagtcg tgctgcttca tgtggtcggg gtagcggctg aagcactgca

[0100] 3241 cgccgtaggt cagggtggtc acgagggtgg gccagggcac gggcagcttg ccggtggtgc

[0101] 3301 agatgaactt cagggtcagc ttgccgtagg tggcatcgcc ctcgccctcg ccggacacgc

[0102] 3361 tgaacttgtg gccgtttacg tcgccgtcca gctcgaccag gatgggcacc accccggtga

[0103] 3421 acagctcctc gcccttgctc accatggtgg cgaccggtgg atcccccggg ctgcaggaat

[0104] 3481 tcgatatcaa gcttacctag ccagcttggg tctccctata gtgagtcgta ttagtaccaa

[0105] 3541 gctaattcct cacgacacct gaaatggaag aaaaaaactt tgaaccactg tctgaggctt

[0106] 3601 gagaatgaac caagatccaa actcaaaaag ggcaaattcc aaggagaatt acatcaagtg

[0107] 3661 ccaagctggc ctaacttcag tctccaccca ctcagtgtgg ggaaactcca tcgcataaaa

[0108] 3721 cccctccccc caacctaaag acgacgtact ccaaaagctc gagaactaat cgaggtgcct

[0109] 3781 ggacggcgcc cggtactccg tggagtcaca tgaagcgacg gctgaggacg gaaaggccct 3841 tttcctttgt gtgggtgact cacccgcccg ctctcccgag cgccgcgtcc tccattttga

[0110] 3901 gctccctgca gcagggccgg gaagcggcca tctttccgct cacgcaactg gtgccgaccg

[0111] 3961 ggccagcctt gccgcccagg gcggggcgat acacggcggc gcgaggccag gcaccagagc

[0112] 4021 aggccggcca gcttgagact acccccgtcc gattctcggt ggccgcgctc gcaggccccg

[0113] 4081 cctcgccgaa catgtgcgct gggacgcacg ggccccgtcg ccgcccgcgg ccccaaaaac

[0114] 4141 cgaaatacca gtgtgcagat cttggcccgc atttacaaga ctatcttgcc agaaaaaaag

[0115] 4201 cgtcgcagca ggtcatcaaa aattttaaat ggctagagac ttatcgaaag cagcgagaca

[0116] 4261 ggcgcgaagg tgccaccaga ttcgcacgcg gcggccccag cgcccaggcc aggcctcaac

[0117] 4321 tcaagcacga ggcgaagggg ctccttaagc gcaaggcctc gaactctccc acccacttcc 4381 aacccgaagc tcgggatcaa gaatcacgta ctgcagccag gtggaagtaa ttcaaggcac 4441 gcaagggcca taacccgtaa agaggccagg cccgcgggaa ccacacacgg cacttacctg 4501 tgttctggcg gcaaacccgt tgcgaaaaag aacgttcacg gcgactactg cacttatata

[0118] 4561 cggttctccc ccaccctcgg gaaaaaggcg gagccagtac acgacatcac tttcccagtt

[0119] 4621 taccccgcgc caccttctct aggcaccggt tcaattgccg acccctcccc ccaacttctc

[0120] 4681 ggggactgtg ggcgatgtgc gctctgccca ctgacgggca ccggagccaa ttcccactcc

[0121] 4741 tttcaagacc tagaaggtcc attagctgca aagattcctc tctgtttaaa actttatcca

[0122] 4801 tctttgcaaa gcttatcgat tacctccacg gccactagtt taatacgact cactatagcg

[0123] 4861 agggcctatt tcccatgatt ccttcatatt tgcatatacg atacaaggct gttagagaga

[0124] 4921 taattggaat taatttgact gtaaacacaa agatattagt acaaaatacg tgacgtagaa

[0125] 4981 agtaataatt tcttgggtag tttgcagttt taaaattatg ttttaaaatg gactatcata

[0126] 5041 tgcttaccgt aacttgaaag tatttcgatt tcttggcttt atatatcttg tggaaaggac

[0127] 5101 gaaacaccgg tcccctccac cccacagtgg ggttagagct agaaatagca agttaaccta

[0128] 5161 aggctagtcc gttatcaact tgaaaaagtg gcaccgagtc ggtgcttttt ttctatagtg

[0129] 5221 tcacctaaat catgcgtcaa ttttacgcag actatctttc tagggttaag ctcttccgct

[0130] 5281 tcctcgctca ctgactcgct gcgctcggtc gttcggctgc ggcgagcggt atcagctcac

[0131] 5341 tcaaaggcgg taatacggtt atccacagaa tcaggggata acgcaggaaa gaacatgtga

[0132] 5401 gcaaaaggcc agcaaaaggc caggaaccgt aaaaaggccg cgtgctggc gttttccat 5461 aggctccgcc cccctgacga gcatcacaaa aatcgacgct caagtcagag gtggcgaaac 5521 ccgacaggac tataaagata ccaggcgttt ccccctggaa gctccctcgt gcgctctcct 5581 gttccgaccc tgccgcttac cggatacctg tccgcctttc tcccttcggg aagcgtggcg 5641 ctttctcata gctcacgctg taggtatctc agttcggtgt aggtcgttcg ctccaagctg 5701 ggctgtgtgc acgaaccccc cgttcagccc gaccgctgcg ccttatccgg taactatcgt 5761 cttgagtcca acccggtaag acacgactta tcgccactgg cagcagccac tggtaacagg 5821 attagcagag cgaggtatgt aggcggtgct acagagttct tgaagtggtg gcctaactac 5881 ggctacacta gaaggacagt atttggtatc tgcgctctgc tgaagccagt taccttcgga 5941 aaaagagttg gtagctcttg atccggcaaa caaaccaccg ctggtagcgg tggttttttt 6001 gtttgcaagc agcagattac gcgcagaaaa aaaggatctc aagaagatcc tttgatcttt 6061 tctacggggt ctgacgctca gtggaacgaa aactcacgtt aagggatttt ggtcatgaga 6121 ttatcaaaaa ggatcttcac ctagatcctt ttaaattaaa aatgaagttt taaatcaatc 6181 taaagtatat atgagtaaac ttggtctgac agttaccaat gcttaatcag tgaggcacct 6241 atctcagcga tctgtctatt tcgttcatcc atagttgcct gactccccgt cgtgtagata 6301 actacgatac gggagggctt accatctggc cccagtgctg caatgatacc gcgagaccca 6361 cgctcaccgg ctccagattt atcagcaata aaccagccag ccggaagggc cgagcgcaga 6421 agtggtcctg caactttatc cgcctccatc cagtctatta attgttgccg ggaagctaga 6481 gtaagtagtt cgccagttaa tagtttgcgc aacgttgttg ccattgctac aggcatcgtg 6541 gtgtcacgct cgtcgtttgg tatggcttca ttcagctccg gttcccaacg atcaaggcga 6601 gttacatgat cccccatgtt gtgcaaaaaa gcggttagct ccttcggtcc tccgatcgtt 6661 gtcagaagta agttggccgc agtgttatca ctcatggtta tggcagcact gcataattct 6721 cttactgtca tgccatccgt aagatgcttt tctgtgactg gtgagtactc aaccaagtca 6781 ttctgagaat agtgtatgcg gcgaccgagt tgctcttgcc cggcgtcaat acgggataat 6841 accgcgccac atagcagaac tttaaaagtg ctcatcattg gaaaacgttc ttcggggcga 6901 aaactctcaa ggatcttacc gctgttgaga tccagttcga tgtaacccac tcgtgcaccc 6961 aactgatctt cagcatcttt tactttcacc agcgtttctg ggtgagcaaa aacaggaagg 7021 caaaatgccg caaaaaaggg aataagggcg acacggaaat gttgaatact cat

[0133] (SEQ ID NO:2)

[0134] 1 ctgataccgc tcgccgcagc cgaacgaccg agcgcagcga gtcagtgagc gaggaagcgg 61 aagagcttaa ccctagaaag atagtctgcg taaaattgac gcatgattta ggtgacacta 121 tagaaaaaaa gcaccgactc ggtgccactt tttcaagttg ataacggact agccttaggt 181 taacttgcta tttctagctc taaccccact gtggggtgga ggggaccggt gtttcgtcct 241 ttccacaaga tatataaagc caagaaatcg aaatactttc aagttacggt aagcatatga 301 tagtccattt taaaacataa ttttaaaact gcaaactacc caagaaatta ttactttcta 361 cgtcacgtat tttgtactaa tatctttgtg tttacagtca aattaattcc aattatctct

[0135] 421 ctaacagcct tgtatcgtat atgcaaatat gaaggaatca tgggaaatag gccctcgcta 481 tagtgagtcg tattaaacta gtggccgtgg aggtaatcga taagctttgc aaagatggat 541 aaagttttaa acagagagga atctttgcag ctaatggacc ttctaggtct tgaaaggagt 601 gggaattgac gcgtcctgca ggacagaccg ataaaacaca tgcgtcaatt ttacacatga 661 ttatctttaa cgtacgtcac aatatgatta tctttctagg gttaatctag ctgcgtgttc 721 tgcagcgtgt cgagcatctt catctgctcc atcacgctgt aaaacacatt tgcaccgcga 781 gtctgcccgt cctccacggg ttcaaaaacg tgaatgaacg aggcgcgctc actggccgtc 841 gttttacaac gtcgtgactg ggaaaaccct ggcgttaccc aacttaatcg ccttgcagca 901 catccccctt tcgccagctg gcgtaatagc gaagaggccc gcaccgatcg cccttcccaa 961 cagttgcgca gcctgaatgg cgaatgggac gcgccctgta gcggcgcatt aagcgcggcg 1021 ggtgtggtgg ttacgcgcag cgtgaccgct acacttgcca gcgccctagc gcccgctcct 1081 ttcgctttct tcccttcctt tctcgccacg ttcgccggct ttccccgtca agctctaaat 1141 cgggggctcc ctttagggtt ccgatttagt gctttacggc acctcgaccc caaaaaactt 1201 gattagggtg atggttcacg tagtgggcca tcgccctgat agacggtttt tcgccctttg 1261 acgttggagt ccacgttctt taatagtgga ctcttgttcc aaactggaac aacactcaac 1321 cctatctcgg tctattcttt tgatttataa gggattttgc cgatttcggc ctattggtta 1381 aaaaatgagc tgatttaaca aaaatttaac gcgaatttta acaaaatatt aacgcttaca 1441 atttaggtgg cacttttcgg ggaaatgtgc gcggaacccc tatttgttta tttttctaaa 1501 tacattcaaa tatgtatccg ctcatgagac aataaccctg ataaatgctt caataatatt 1561 gaaaaaggaa gagtatgagt attcaacatt tccgtgtcgc ccttattccc ttttttgcgg 1621 cattttgcct tcctgttttt gctcacccag aaacgctggt gaaagtaaaa gatgctgaag 1681 atcagttggg tgcacgagtg ggttacatcg aactggatct caacagcggt aagatccttg 1741 agagttttcg ccccgaagaa cgttttccaa tgatgagcac ttttaaagtt ctgctatgtg 1801 gcgcggtatt atcccgtatt gacgccgggc aagagcaact cggtcgccgc atacactatt 1861 ctcagaatga cttggttgag tactcaccag tcacagaaaa gcatcttacg gatggcatga 1921 cagtaagaga attatgcagt gctgccataa ccatgagtga taacactgcg gccaacttac 1981 ttctgacaac gatcggagga ccgaaggagc taaccgcttt tttgcacaac atgggggatc 2041 atgtaactcg ccttgatcgt tgggaaccgg agctgaatga agccatacca aacgacgagc 2101 gtgacaccac gatgcctgta gcaatggcaa caacgttgcg caaactatta actggcgaac 2161 tacttactct agcttcccgg caacaattaa tagactggat ggaggcggat aaagttgcag 2221 gaccacttct gcgctcggcc cttccggctg gctggtttat tgctgataaa tctggagccg

[0136] 2281 gtgagcgtgg gtctcgcggt atcattgcag cactggggcc agatggtaag ccctcccgta 2341 tcgtagttat ctacacgacg gggagtcagg caactatgga tgaacgaaat agacagatcg 2401 ctgagatagg tgcctcactg attaagcatt ggtaactgtc agaccaagtt tactcatata 2461 tactttagat tgatttaaaa cttcattttt aatttaaaag gatctaggtg aagatccttt 2521 ttgataatct catgaccaaa atcccttaac gtgagttttc gttccactga gcgtcagacc 2581 ccgtagaaaa gatcaaagga tcttcttgag atcctttttt tctgcgcgta atctgctgct 2641 tgcaaacaaa aaaaccaccg ctaccagcgg tggtttgttt gccggatcaa gagctaccaa 2701 ctctttttcc gaaggtaact ggcttcagca gagcgcagat accaaatact gtccttctag 2761 tgtagccgta gttaggccac cacttcaaga actctgtagc accgcctaca tacctcgctc 2821 tgctaatcct gttaccagtg gctgctgcca gtggcgataa gtcgtgtctt accgggttgg 2881 actcaagacg atagttaccg gataaggcgc agcggtcggg ctgaacgggg ggttcgtgca 2941 cacagcccag cttggagcga acgacctaca ccgaactgag atacctacag cgtgagctat 3001 gagaaagcgc cacgcttccc gaagggagaa aggcggacag gtatccggta agcggcaggg 3061 tcggaacagg agagcgcacg agggagcttc cagggggaaa cgcctggtat ctttatagtc 3121 ctgtcgggtt tcgccacctc tgacttgagc gtcgattttt gtgatgctcg tcaggggggc 3181 ggagcctatg gaaaaacgcc agcaacgcgg cctttttacg gttcctggcc ttttgctggc 3241 cttttgctca catgttcttt cctgcgttat cccctgattc tgtggataac cgtattaccg 3301 cctttgagtg ag

[0137] In some instances, the hgRNA sequence in the vector comprises a spacer sequence. A spacer sequence is a sequence (e.g., a 20 base pair sequence) that defines the target sequence (e.g., a DNA target sequences, such as a genomic target sequence) of a target nucleic acid of interest. The “target sequence” is adjacent to a PAM sequence and is the sequence modified by an RNA-guided nuclease (e.g., Cas9). The “target nucleic acid” is a double-stranded molecule: one strand comprises the target sequence and is referred to as the “PAM strand,” and the other complementary strand is referred to as the “non-PAM strand.” One of skill in the art recognizes that the gRNA spacer sequence hybridizes to the reverse complement of the target sequence, which is located in the non-PAM strand of the target nucleic acid of interest. Thus, the gRNA spacer sequence is the RNA equivalent of the target sequence. In some instances, such as in a CRISPR / Cas system herein, the spacer sequence is designed to hybridize to a region of the target nucleic acid that is located 5' of a PAM of the Cas9 enzyme used in the system. The spacer may perfectly match the target sequence or may have mismatches. Each Cas9 enzyme has a particular PAM sequence that it recognizes in a target DNA. For example, S. pyogenes recognizes in a target nucleic acid a PAM that comprises the sequence 5 -NRG-3', where R comprises either A or G, where N is any nucleotide and N is immediately 3' of the target nucleic acid sequence targeted by the spacer sequence. In some instances, the methods of inducing described herein utilize a S. pyogenes Cas9 nuclease.

[0138] In some instances, the vector includes a minimum CRISPR repeat sequence that comprises nucleotides that can hybridize to a minimum tracrRNA sequence in a cell. The minimum CRISPR repeat sequence and a minimum tracrRNA sequence form a duplex, i.e., a base-paired double- stranded structure. Together, the minimum CRISPR repeat sequence and the minimum tracrRNA sequence bind to the site-directed polypeptide.

[0139] In some instances, the disclosure features a gene transfer system comprising a transposon according to any one of the above aspects; and a piggyBac (PB) transposase. PB insertion has been described in Yoshida et al., Sci Rep. 2017 Mar 2:7:43613, which is incorporated by reference in its entirety. In one embodiment, the piggyBac transposase is from the family Noctuidae. In a related embodiment, the piggyBac transposase is from the species Trichoplusia ni. The PB transposon system employs a genetically engineered transposase enzyme to insert a gene into a cell's genome. It is built upon the natural PB transposable element (transposon), enabling the back and forth movement of genes between chromosomes and genetic vectors such as plasmids through a "cut and paste" mechanism. PiggyBac vectors are one of the most active and flexible class 2 transposon systems available for the stable transfection of mammalian cells (Wilson et al., Mol. Ther. 15: 139-145 (2007); Wu et al., Proc. Natl. Acad. Sci. U.S.A. 103: 15008-15013 (2006)). Transposase catalyzes the excision of the transposon from one DNA source (i.e., a delivered plasmid) and allows its subsequent re-integration into another DNA source (i.e., the host cell genome).

[0140] In some embodiments, the transposon (i.e., agent) can be any programmable nuclease. In some embodiments, the agent is a natural homing meganuclease. In some embodiments, the agent is a TALEN-based agent, a ZFN-based agent, or a CRISPR-based agent, or any biologically active fragment, fusion, derivative or combination thereof. In some embodiments, the agent is a deaminase or a nucleic acid encoding a deaminase. In some embodiments, a cell is engineered to stably and / or transiently express a TALEN-based agent, a ZFN-based agent, and / or a CRISPR-based agent.

[0141] In some instances, Cas proteins devoid of nucleolytic activity (dead Cas proteins; dCas) can be used to deliver functional cargo to programmed sites in the genome. CRISPR-Cas systems have been described previously, such as in Xu and Li “CRISPR-Cas systems: Overview, innovations and applications in human disease research and gene therapy,” Comput Struct Biotechnol J. 2020; 18: 2401-2415, which is incorporated by reference in its entirety.

[0142] D. Insertion of a Transgene at the Genomic Harbor

[0143] The methods used herein can identify one or more harbors (i.e., locations) in a genome for insertion of a transgene. In some instances, the transposase gene (e.g., hgRNA) is removed from the transposon and replaced by transgenes of interest; the transposase is then usually delivered to the cell, typically by a separate plasmid. Thus, the disclosure further provides an efficient method for producing a transgenic biological sample (e.g., a cell), including the step of applying the inventive gene transfer system to the biological sample.

[0144] Transgenic DNA has not been efficiently inserted into chromosomes. Only about one in a million of the foreign DNA molecules is inserted into the cellular genome, generally several cleavage cycles into development. Consequently, most transgenic animals are mosaic. As a result, animals raised from embryos into which transgenic DNA has been delivered must be cultured until gametes can be assayed for the presence of inserted foreign DNA. Many transgenic animals fail to express the transgene due to position effects. A simple, reliable procedure that directs early insertion of exogenous DNA into the chromosomes of animals at the one-cell stage is needed. The present system helps to fill this need.

[0145] In some instances, the gene transfer system of this invention can readily be used to produce transgenic animals that carry a particular marker or express a particular protein in one or more cells at certain locations of a genome in a biological sample. Generally, methods for producing transgenic animals are known in the art and incorporation of the inventive gene transfer system into these techniques does not require undue experimentation, e.g., there are a variety of methods for producing transgenic cells for research or for protein production including, but not limited to Hackett et al. (1993, supra). Other methods for producing transgenic animals are described in the art (e ., M. Markkula et al. Rev. Reprod., 1, 97-106 (1996); R. T. Wall et al., J. Dairy Sci., 80, 2213-2224 (1997)), J. C. Dalton, et al. (Adv. Exp. Med. Biol., 411, 419-428 (1997)) and H. Lubon et al. (Transfus. Med. Rev., 10, 131-143 (1996)).

[0146] In another embodiment, the present invention features a transgenic animal produced by the methods described herein, preferably by using the gene transfer system presently described. For example, transgenic animals may preferably contain a nucleic acid sequence inserted into the genome of the animal by the gene transfer system, thereby enabling the transgenic animal to produce its gene product, e.g., a protein. Promoters can be used that promote expression in milk, urine, blood or eggs and these promoters include, but are not limited to, casein promoter, the mouse urinary protein promoter, beta-globin promoter and the ovalbumin promoter respectively. Recombinant growth hormone, recombinant insulin, and a variety of other recombinant proteins have been produced using other methods for producing protein in a cell. Nucleic acids encoding these or other proteins can be inserted into the transposon of this invention and transfected into a cell. Efficient transfection of the inventive transposon as defined above into the DNA of a cell occurs when mammalian piggyBac transposase protein is present. Where the cell is part of a tissue or part of a transgenic animal, large amounts of recombinant protein can be obtained.

[0147] The minTRs and IDs are crucial for the effective integration of the transposon into the host genome and together (known as terminal domains) consist of more than 700 base pairs each (Zhuang et al., Acta. Biochim. Biophys. Sin (Shanghai) 42:426-431 (2010)). The 5' terminal domain also serves as a native promoter for transposase expression. As part of the transposition, the terminal domains are integrated into the host cell genome, exclusively at TTAA integration site, alongside the delivered transgene of interest (Elick et al., Genetica 98:33-41 (1996); Fraser et al., Insect Mol. Biol. 5: 141-151 (1996)). Therefore, like integrated viruses, they deliver a significant amount of extra DNA to the target cell genome. Although the terminal domains are required for successful transposition, once integrated into the host cell genome, they perform no useful function. In fact, they may increase the risk of insertional mutagenesis (Meir et al., BMC Biotechnol 2011; 11:28 (2011)), due to any apparent or potential promoter or enhancer activity that the terminal domains might exert on host cell oncogenes (Cadinanos et al., Nucleic Acids Res. 35:e87 (2007); Shi et al., BMC Biotechnol. 7:5 (2007). Neither the 5' nor the 3' piggyBac minTRs contain known active promoters or enhancers (Handler et al., Proc. Natl. Acad. Sci. USA 95:7520-7525 (1998); Shi et al., BMC Biotechnol. 7:5 (2007)). In some instances, the transgene produces a biologic such as a monoclonal antibody or hormone. In some instances, the biologic is a mammalian antibody. Because most integrations in this field are either unstable or get silenced, a large number of clones must be screened at a heavy cost to isolate stable ones. This invention can be used to identify the stable harbors in producer lines (e.g., CHO) for producing transgenes of interest. It can further be used to generate designer producer lines (e.g., CHO) that have engineered harbors ready for intergration of transgene-expressing constructs. In some instances, the antibody can include IgG, IgA, IgM, or IgE antibodies, antibody fragments such as Fc regions, antibody Fab regions, antibody heavy chains, antibody light chains, antibody CDRs, nanobodies, chimeric antibodies and other IgG domains; T cell receptors (TCRs).

[0148] In some instances, the methods disclosed herein can be used to identify safe and open harbors for modifying CAR-T cells. CARs are hybrid transmembrane receptors comprising an extracellular tumor-antigen binding moiety, typically in the form of a single-chain variable fragment (scFv), linked to signaling endodomains allowing T cell activation and effector functions upon target engagement. In some instances, viral vectors such as retroviruses or lenti viruses can be used for introduction of the transgene into a cell. While first-generation CARs include the endodomain of CD3 zeta only (for signal 1 of T cell activation), second- and third- generation CARs further comprise one or more costimulatory endodomains (signal 2), respectively, usually derived from CD28 or 4-1BB molecules. In some instances, disclosed herein is a genetically engineered cell comprising in its genome a modified human T cell receptor gene, wherein the cell has modified cell surface expression of the endogenous T cell receptor.

[0149] E. Biological Samples

[0150] In one embodiment, the biological sample is a cell is obtained from an animal. In another embodiment, the cell is from a vertebrate or an invertebrate. In a further embodiment, the vertebrate is a mammal.

[0151] In some instances, the biological sample is a cultured cell sample. In some instances, the biological sample comprises a primary cell, a cell in culture, or a cell from a cell line. In some instances, the biological sample is from a tissue sample. In some instances, biological samples can be derived from a homogeneous culture or population of the subjects or organisms mentioned herein or alternatively from a collection of several different organisms, for example, in a community or ecosystem. According to the disclosure, the biological sample includes prokaryotic cells (e.g., E. coli) or progressive cells (e.g., dendritic cells, B cells, CHO cells, COS cells, K562 cells, HEK293 cells, HELA cells, yeast cells and insect cells) according to the present disclosure. In some instances, the cultured cell comprises CHO cells. In some instances, the cultured cell comprises K562 cells.

[0152] In some instances, the cultured cells are derived from model organisms. For example, cultured cells can include marmoset embryonic stem cells (CJ367 cells).

[0153] In some instances, a biological sample is typically obtained from the subject for analysis using any of a variety of techniques including, but not limited to, biopsy, surgery, and laser capture microscopy (LCM), and generally includes cells and / or other biological material from the subject. In some embodiments, the biological sample is a tissue sample. In some embodiments, the biological sample (e.g., tissue sample) is a tissue microarray (TMA). A tissue microarray contains multiple representative tissue samples - which can be from different tissues or organisms - assembled on a single histologic slide. The TMA can therefore allow for high throughput analysis of multiple specimens at the same time. Tissue microarrays may be paraffin blocks produced by extracting cylindrical tissue cores from different paraffin donor blocks and re-embedding these tissue cores into a single recipient (microarray) block at defined array coordinates.

[0154] The biological sample as used herein can be any suitable biological sample described herein or known in the art. In some embodiments, the biological sample is a tissue sample. In some embodiments, the tissue sample is a solid tissue sample. In some embodiments, the biological sample is a tissue section (e.g., a fixed tissue section). In some embodiments, the tissue is flash-frozen and sectioned. Any suitable method described herein or known in the art can be used to flash-freeze and section the tissue sample. In some embodiments, the biological sample, e.g., the tissue, is flash-frozen using liquid nitrogen before sectioning. In some embodiments, the biological sample, e.g., a tissue sample, is flash-frozen using nitrogen (e.g., liquid nitrogen), isopentane, or hexane.

[0155] The biological sample can be from a mammal. In some instances, the biological sample is from a human, mouse, or rat. In addition to the subjects described above, the biological sample can be obtained from non-mammalian organisms (e.g., a plant, an insect, an arachnid, a nematode (e.g., Caenorhabditis elegans), a fungus, an amphibian, or a fish (e.g., zebrafish)). A biological sample can be obtained from a prokaryote such as a bacterium, e.g., Escherichia coli, Staphylococci or Mycoplasma pneumoniae,' an archaeon; a virus such as Hepatitis C virus or human immunodeficiency virus; or a viroid. Subjects from which biological samples can be obtained can be healthy or asymptomatic individuals, individuals that have or are suspected of having a disease (e.g., cancer) or a pre-disposition to a disease, and / or individuals that are in need of therapy or suspected of needing therapy.

[0156] Biological samples can include one or more diseased cells. A diseased cell can have altered metabolic properties, gene expression, protein expression, and / or morphologic features. Examples of diseases include inflammatory disorders, metabolic disorders, nervous system disorders, and cancer. Cancer cells can be derived from solid tumors, hematological malignancies, cell lines, or obtained as circulating tumor cells.

[0157] In any of the foregoing, the biological sample can further be stained, imaged, and / or destained. For example, in some embodiments, a fresh frozen tissue sample or fixed frozen tissue sample is stained (e.g., via eosin and / or hematoxylin), imaged, destained (e.g., via HCl), or a combination thereof. In some embodiments, when a fresh frozen tissue sample is fixed in methanol, the sample is treated with isopropanol prior to being stained (e.g., via eosin and / or hematoxylin), imaged, destained (e.g., via HC1), or a combination thereof. In some embodiments when a fixed frozen tissue sample is treated with a sucrose gradient, the sample can be rehydrated using an ethanol gradient before being stained, (e.g., via eosin and / or hematoxylin), imaged, destained (e.g., via HC1), decrosslinked (e.g., via TE buffer or citrate buffer), or a combination thereof. In some embodiments, the biological sample can undergo further fixation (e.g., while mounted on a substrate), stained, imaged, and / or destained. For example, a fixed frozen biological sample may be subject to an additional fixing step (e.g., using PF A) before optional ethanol rehydration, staining, imaging, and / or destaining.

[0158] The tissue sample can be obtained from any suitable location in a tissue or organ of a subject, e.g., a human subject. In some instances, the sample is a mouse sample. In some instances, the sample is a human sample. In some embodiments, the sample can be derived from skin, brain, breast, lung, liver, kidney, prostate, tonsil, thymus, testes, bone, lymph node, ovary, eye, heart, or spleen. In some instances, the sample is a human or mouse breast tissue sample. In some instances, the sample is a human or mouse brain tissue sample. In some instances, the sample is a human or mouse lung tissue sample. In some instances, the sample is a human or mouse tonsil tissue sample. In some instances, the sample is a human or mouse liver tissue sample. In some instances, the sample is a human or mouse bone, skin, kidney, thymus, testes, or prostate tissue sample. In some embodiments, the tissue sample is derived from normal or diseased tissue. In some embodiments, the sample is an embryo sample. The embryo sample can be a non-human embryo sample. In some instances, the sample is a mouse embryo sample.

[0159] EXAMPLES

[0160] The disclosure is further described in the following examples, which do not limit the scope of the disclosure described in the claims.

[0161] Example 1 — Detection of mutated hgRNA expression level corresponds to its genomic location in the nwuse genome.

[0162] The present example examines the correlation between the level of mutations in hgRNA and insertion of the same into different harbors in the mouse genome.

[0163] Mouse cells from transgenic mice (i.e., treated with an hgRNA construct (e.g., as shown in FIGs. 2A)) were used for this experiment. Generation of transgenic mice were described in Kalhor et al., Science, 2018 Aug 31;361(6405) and Leeper et al., Nat Protoc. 2021 Apr;16(4):2088-2108, each of which is incorporated by reference in its entirety. Cas9 was introduced into the mouse to activate recording. Cas9-hgRNA complex induced mutations in the hgRNA spacer, in turn, recording the expression of the hgRNA transcript into the locus. In particular, the hgRNA transcript expressed by the locus formed a complex with Cas9 protein and targeted the locus for double-strand break. As the non-homologous end joining (NHEJ) repair system repaired the cut, it introduced (e.g., recorded) mutations in the hgRNA locus. After, genomic DNA was harvested from the mouse cells using a chosen restriction enzyme (DpnII) that does not recognize downstream U6 promoter up to the next recognition site in the genome where hgRNA was inserted. The genomic DNA was purified, were provided adaptors via ligation at the 5’ and 3’ ends, and were sequenced.

[0164] The measured RNA expression level from 25 hgRNAs inserted into different harbors in the mouse genome is shown in FIG. 4A. As shown in FIG. 4A, RNA expression varies based on location and covers a wide range. When we compared the expression rankings among the 16 hgRNAs shared between two MARC1 mice, we observed a high degree of concordance, with a Spearman's rank correlation coefficient of 0.94 as shown in FIG. 4B.

[0165] Example 2 — hgRNA expression level corresponds to its genomic location in K562 cells.

[0166] In an experiment similar to Example 1, a human immortalized myelogenous leukemia cell line (K562 cells) was cultured in standard media, and an hgRNA construct (e.g., as shown in FIGs. 2A) was delivered to the cells.

[0167] We implemented our assay in K562 cells, obtaining recorded mutation fractions for approximately 10,000 loci, thus uncovering a wide spectrum of expression levels. We observed limited correlation between our transgene mutation fraction and chromatin accessibility, as measured by ATAC-seq, suggesting that information captured by our assay is not fully represented by chromatin openness. Therefore, our approach complements existing genome assays in predicting transgene activity across different genomic locations.

[0168] Cas9 was introduced into the K562 cells to activate recording. Cas9-hgRNA complex induced mutations in the hgRNA spacer, in turn, recording the expression of the hgRNA transcript into the locus. In particular, the hgRNA transcript expressed by the locus formed a complex with Cas9 protein and targeted the locus for double-strand break. As the non- homologous end joining (NHEJ) repair system repaired the cut, it introduced (e.g., recorded) mutations in the hgRNA locus. After, genomic DNA was harvested from the K562 cells using a chosen restriction enzyme (DpnII) that does not recognize downstream U6 promoter up to the next recognition site in the genome where hgRNA was inserted. The genomic DNA was purified, were provided adaptors via ligation at the 5’ and 3’ ends, and were sequenced.

[0169] The measured RNA insertion level inserted into different harbors in the human genome is shown in FIG. 5. The outer layer of FIG. 5 shows an ideogram displaying the cytobands of the hg38 human genome. The next (second) layer, which is labeled with “hgRNA,” shows a series of dots. Each dot represents the mutation fraction at the identified insertion location. The next (third) layer, labeled with “ATAC,” provides dots that represent the ATAC-seq signal at the identified insertion location. The signal value is calculated as the mean of ATAC signals within a ±500 bp range of the insertion site. The mean value is then logio normalized for display, i.e. ATAC score = logl0( x + 1) where x = mean(ATAC signal ±500 bp). Finally, the most interior (fourth) layer, labeled with “Insertion,” shows a histogram depicting the number of unique insertion locations. Each bin represents a 1 Mbp span.

[0170] As shown in FIG. 5, there was genome-wide distribution of hgRNA insertion sites. 1,380 hgRNA insertion locations with more than 40 UMIs were identified. These locations are distributed across the genome (excluding the Y chromosome, as K562 cells are derived from a female individual), demonstrating the genome-wide application of the devised method.

[0171] In total, over 13,000 unique locations were identified in the human genome of K562 cells. At each site, the number of hgRNA mutations were measured. FIG. 6 shows mutation fractions for 1,380 hgRNA insertion locations, each with more than 40 UMIs. The devised method captures a wide spectrum of mutation percentages, and allows us to identify areas of the genome that are “open harbor candidates,” (i.e., where mutations most frequently occur compared to the rest of the genome), and “intermediate harbor candidates,” where mutation percentage was slightly lower. These data demonstrate that these locations are possible open harbor candidates for insertion of transgenes.

[0172] We next were interested to examine the relationship between mutation fraction and ATAC-seq scores at 1,380 insertion sites, each with at least 40 UMIs. As shown in FIGs. 7A and 7B, there was very little correlation between the mutation fraction (y-axis) versus mean ATAC- seq score (FIG. 7A; Pearson R value: 0.10). Further, adapting the x-axis to represent the logio(l + mean ATAC-seq score) value, showed little correlation (FIG. 7B; Pearson R value: 0.23). These data indicate that there is very little correlation between ATAC-seq signals and areas of hgRNA integration and mutation, suggesting that these methods can provide a different way to determine insertion location sites.

[0173] Finally, we examined whether adjacent locations of two different insertion sites correlated with one another. Said another way, we hypothesized that the closer that the locations of two insertion sites on a chromosome were, the more likely the insertion sites would occur at that window. As shown in FIG. 8, there indeed appeared to be a correlation of location to mutation pair. The Pearson R scores were calculated for sets of hgRNA pairs where the distance between pairs falls within defined ranges. To mitigate bias from axis arrangement, the order of pairs was randomized and repeated 10 times to obtain the mean Pearson R value. The lower end of the window size was fixed at 500 bp, while the upper end increased from 2,000 bp to 800,000 bp in 2,000 bp increments. Thus, the smallest data point on the x-axis represents the mean Pearson R of pairs with distances between 500 and 2,000 bp, the next point represents distances between 500 and 4,000 bp, and so on, up to the 500-800,000 bp window. FIG. 8 and these data indicate that there are windows along a chromosome where random insertion more likely would occur compared to other areas of a chromosome.

[0174] Taken together, these data indicate that one could measure hgRNA directly and connect these data to harbor locations in a high-throughput fashion, since the mutations in hgRNAs record their expression in their genomic sequence. This genome-wide transgene expression profde assay enhances understanding of the relationship between transgenes and their genomic locations. Additionally, it can aid in the identification of genomic safe harbors — regions of the genome that support stable and consistent transgene expression.

[0175] Example 3 — Identification of safe harbor locations in the marmoset genome.

[0176] We next utilized a high-throughput method for identifying genomic harbors — an assay that quantifies transgene expression based on the mutation fraction (MF) at each insertion site — and applied it to marmoset embryonic stem cells (CJ367 cells) to systematically screen for safe harbor candidates in the marmoset genome. There is a growing need in the biomedical research community for reliable transgenic marmoset models, but progress has been limited by the absence of well-characterized genomic safe harbor sites suitable for targeted transgene insertion. To address this bottleneck, we leveraged our screening platform to identify transcriptionally permissive loci with minimal dismption to host gene function.

[0177] Using this approach, we applied our assay to the CJ367 ESC line and identified 4,230 insertion sites across the marmoset genome and calculated the MF associated with each locus. To ensure robust MF estimates, we restricted downstream analysis to the 2,181 loci supported by > 50 unique molecular identifiers (UMIs), which are distributed across all chromosomes (FIG. 9A). MF values exhibited a clear bimodal distribution that separates transcriptionally silent (“inactive”) from expression-positive (“active”) loci (FIG. 9B). Within the active cohort we detected a subset of “hyper-active” sites that drive exceptionally high transgene expression, marking them as leading safe-harbor candidates (FIG. 9C).

[0178] Additionally, functional annotation revealed that the most highly expressed insertions were preferentially located in intergenic regions, nominating these sites as strong candidates for genomic safe harbors. Leveraging the latest marmoset genome annotation, we classified each high-confidence insertion (n = 2,181) as intragenic or intergenic. As shown in FIG. 10A, forty-two percent reside in unannotated intergenic regions, while the remaining 58 % intersect annotated genes. Within the intragenic group, 88.9 % fall within protein-coding genes and 10.8 % overlap long non-coding RNAs; other biotypes each account for less than 0.1 % (FIG. 10B). Although the overall MF distributions of intragenic and intergenic loci are similar (FIG.

[0179] 10C), hyper-active sites are more favorably enriched in intergenic regions (9), underscoring their potential as genomic safe-harbor loci.

[0180] Altogether, this study demonstrates the utility of our method in uncovering safe harbor sites in emerging model organisms such as the marmoset, enabling more precise and reliable transgenesis in this species.

Claims

1. WHAT IS CLAIMED IS:

1. A method for high-throughput identification of one or more insertion sites for transgene expression in a genome of a biological sample, the method comprising:(a) integrating a self-recording nucleic acid sequence into one or more genomic loci of the biological sample, wherein the integrating induces one or more changes at the one or more genomic loci;(b) inducing the self-recording nucleic acid sequence to record its activity by generating changes in the self-recording nucleic acid sequence; and(c) determining the activity of the self-recording nucleic acid in each integrated genomic locus by (i) identifying a location of the self-recording nucleic acid in the genome and (ii) measuring a number of changes that are accumulated in the self-recording nucleic acid sequence, thereby characterizing the integrated genomic locus or loci as harbor or harbors for transgene expression.

2. A method of identifying a location for random insertion of a nucleic acid into a genome of a biological sample, the method comprising:(a) integrating a self-recording nucleic acid sequence into one or more genomic loci of the biological sample, wherein the integrating induces one or more changes at the one or more genomic loci;(b) inducing the self-recording nucleic acid sequence to record its activity by generating changes in the self-recording nucleic acid sequence; and(c) determining the location of the random insertion of the nucleic acid (i) identifying the location of the self-recording nucleic acid in the genome and (ii) measuring a number of changes that are accumulated in the self-recording nucleic acid sequence.

3. A method for introducing a self-recording nucleic acid sequence into one or more random locations of a biological sample, the method comprising:(a) contacting the biological sample with a gene transfer system, thereby introducing the self-recording nucleic acid sequence into a genome of the biological sample;(b) inducing the self-recording nucleic acid sequence to record its activity by generating changes in the self-recording nucleic acid sequence; and(c) determining the location of an insertion of the self-recording nucleic acid in a genome of the biological sample by (i) identifying the location of the self-recording nucleic acid in the genome and (ii) measuring a number of changes that are accumulated in the self-recording nucleic acid sequence.

4. A method of identifying putative genomic safe harbor locations within a genome in a biological sample, the method comprising:(a) introducing into multiple genomic sites of the genome in the biological sample a plurality of self-recording nucleic acid sequences, wherein each self-recording nucleic acid sequence of the plurality is capable of self-recording its own activity;(b) allowing the self-recording nucleic acid sequences of the plurality to record their activity at their genomic sites; and(c) measuring the recorded activity at each site to identify putative genomic safe harbor locations.

5. The method of any one of claims 1-4, further comprising introducing a transgenic nucleic acid sequence into the location of the genome.

6. The method of claim 5, wherein the transgenic nucleic acid sequence comprises a sequence that encodes for all or part of an antibody or an antigen-binding fragment thereof.

7. The method of claim 6, wherein the biological sample is Chinese hamster ovary cells.

8. The method of any one of claims 1-7, wherein the location is modified by the insertion of a nucleic acid switch.

9. The method of claim 8, wherein the nucleic acid switch is a kill switch, an activation switch, or a cell-state transition switch.

10. The method of claim 8 or 9, wherein the location for insertion of the nucleic acid switch is cell-type specific.

11. The method of any one of claims 8-10, wherein the location for insertion of the nucleic acid switch is specific for T-cells.

12. The method of any one of claims 5-11, wherein the transgenic nucleic acid sequence comprises a universal primer.

13. The method of any one of claims 8-12, wherein the nucleic acid switch is an inducible nucleic acid switch.

14. The method of any one of claims 8-12, wherein the nucleic acid switch is a capable of inducing nucleic acid breaks at one or more locations in the genome.

15. The method of any one of claims 1-14, wherein the self-recording nucleic acid sequence is a homing CRISPR guide RNA.

16. The method of any one of claims 1-15, wherein the activity is RNA expression.

17. The method of any one of claims 1-16, wherein the inducing utilizes a S. pyogenes Cas9 nuclease.

18. The method of any one of claims 1-17, wherein the changes are mutations to the self- recording nucleic acid sequence.

19. The method of any one of claims 1-17, wherein the changes comprise changes to a methylation status at the location.

20. The method of any one of claims 1-17, wherein the changes are histone modifications at the location.21 . The method of claim 20, wherein the histone modification comprises histone acetylation, histone deacetylation, histone methylation, histone ubiquitination, and histone citrullination.

22. The method of any one of claims 1-21, further comprising: isolating genomic DNA from the biological sample after induction; treating the biological sample with an endonuclease that cuts the genomic DNA at sites outside but flanking the self-recording nucleic acid sequence, thereby generating fragments of genomic DNA that comprise all or part of the self-recording nucleic acid sequence and a part of a sequence of genomic DNA endogenous to the genome of the biological sample; ligating adaptors to the fragments; optionally amplifying the fragments; and sequencing (i) all or part of the self-recording nucleic acid sequence and (i) a part of a sequence of genomic DNA endogenous to the genome of the biological sample.

23. The method of claim 22, wherein all or part of the self-recording nucleic acid sequence comprises a mutation induced by a nuclease upon expression of the self-recording nucleic acid sequence.

24. A method of creating a transgenic biological sample, the method comprising:(a) introducing into the biological sample a self-recording nucleic acid sequence; and(b) identifying that the self-recording nucleic acid sequence has randomly incorporated into a genome of the biological sample, thereby creating the transgenic biological sample.

25. The method of any one of claims 1-24, wherein the biological sample comprises a primary cell, a cell in culture, or a cell from a cell line.

26. The method of any one of claims 1-25, wherein the biological sample is from a tissue sample.

27. The method of any one of claims 1-26, wherein the biological sample comprises dendritic cells, B cells, CHO cells, COS cells, K562 cells, HEK293 cells, HELA cells, yeast cells, insect cells, or marmoset embryonic stem cells.

28. The method of any one of claims 1-27, wherein the self-recording nucleic acid sequence is part of a sequence of a vector comprising one or more of a scaffolding sequence, a protospacer adjacent motif (PAM) sequence, and spacer sequence, and any combination thereof.

29. The method of claim 28, wherein the vector comprises one or more selection markers, and / or wherein the vector comprises one or more detectable markers.

30. The method of claim 29, further comprising detecting expression of part of the vector inserted into the genome using the one or more detectable markers.

31. The method of any one of claims 1-30, wherein the self-recording nucleic acid sequence forms a homing complex using a Cas9 protein at the location.

32. The method of claim 31, wherein the Cas9:homing complex targets the location of the genome for a double-strand break at the location.

33. The method of claim 32, further comprising repairing the double-stranded break, thereby inducing one or more mutations.

34. The method of any one of claims 1-33, further comprising integrating the self-recording nucleic acid sequence to the biological sample using a transposase.

35. The method of claim 34, wherein the transposase is a piggyBac transposase, a sleeping beauty transposase, or a CRISPR transposase.

36. The method of any one of claims 1-35, further comprising treating the genome with a restriction enzyme, thereby generating a digested self-recording nucleic acid fragment thatcomprises one or more mutations and a part of a sequence of the genome of the biological sample.

37. The method of claim 36, further comprising attaching an adaptor sequence(s) to the 3’ and / or 5’ ends of the digested self-recording nucleic acid fragment.

38. The method of claim 37, wherein the attaching comprises ligating the adaptor sequences(s).

39. The method of any one of claims 36-38, further comprising amplifying the digested self- recording nucleic acid fragment.

40. The method of any one of claims 36-39, wherein the determining step comprises sequencing the digested self-recording nucleic acid fragment.

41. A method of identifying a location for a random insertion of a nucleic acid into a genome of a biological sample, the method comprising:(a) providing the biological sample;(b) adding to the biological sample a vector comprising a self-recording nucleic acid sequence and a transposon;(c) randomly integrating the self-recording nucleic acid sequence into one or more genomic loci of the genome of the biological sample, wherein the integrating induces one or more mutations at one or more genomic loci;(d) inducing the self-recording nucleic acid sequence to record its activity by generating changes in the self-recording nucleic acid sequence;(e) adding a restriction enzyme to the biological sample, thereby generating a digested self-recording nucleic acid fragment;(f) ligating adaptors to the digested self-recording nucleic acid fragment;(g) amplifying the digested self-recording nucleic acid fragment; and(h) sequencing the digested self-recording nucleic acid fragment, or a complement thereof, thereby identifying a location of random insertion of a nucleic acid into a genome of a biological sample.

Citation Information

Patent Citations

  • Targeted genomic modification with partially single-stranded donor molecules

    US20160145644A1

  • Self-targeting genome editing system

    US20180291372A1