Targeted sequencing of genomic loci
The method addresses imprecise transgene integration by using tagmentation and long-read sequencing to efficiently characterize genomic loci, facilitating rapid and accurate identification of transgene sites in transgenic organisms.
Patent Information
- Application Number
- PCT/US2025/044261
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-30
- Filing Date
- 2025-08-29
- Publication Date
- 2026-03-05
AI Technical Summary
Current methods for identifying transgene integration sites in transgenic animals are imprecise, often result in off-target insertions, and are cumbersome, especially for large genomes like human cell lines or mouse models, lacking efficient techniques to amplify long stretches of DNA across unknown regions.
A method combining tagmentation using a transposase, enrichment, and sequencing to characterize genomic loci, employing a hyper-active transposase to fragment DNA, append transposon DNA, and use locus-specific and universal primers for PCR amplification, followed by long-read sequencing.
Enables rapid and efficient identification of transgene integration sites with minimal prior knowledge, overcoming limitations of existing techniques by providing long-range amplification and accurate sequencing of complex genomic structures.
Smart Images

Figure US2025044261_05032026_PF_FP_ABST
Abstract
Description
TARGETED SEQUENCING OF GENOMIC LOCIBACKGROUND
[0001] Transgenic animal models are crucial for the study of gene function and disease, and are widely utilized in basic biological research, agriculture and pharma industries. Since the current methods for generating transgenic animals result in the random integration of the transgene under study, the phenotype may be compromised due to disruption of known genes or regulatory regions. Though more modern techniques for transgene integration have been developed such as CRISPR-targeted homologous recombination, which are expected to be much more precise in their targeting, there are still off-target insertions. Identifying the location of these off-target integration events are of utmost importance to the field.SUMMARY
[0002] There are many situations in molecular biology (in addition to transgenic modifications) where you might want to amplify a region in the genome, but you know only one end of that region which you are investigating. If you know a lot about the region you can design two primers to perform PCR, but sometimes it can be hard to find compatible primer pairs that amplify efficiently together, and other times you simply don’t know precisely where the other end is. If you have a transgene insertion, for instance, but don’t know where that integration took place it can be very hard to determine that location by conventional means. Even if you think you know where the insertion may have occurred, designing primers that will bind on both sides of the site can be tricky. Many of these insertion events also tend to cut out stretches of DNA as they integrate and might remove one or more of the primer binding sites.
[0003] Amplification problems can be compounded by the fact that sometimes you might need to amplify long stretches of DNA to cross any insertion junctions and PCR amplification across greater distances gets increasingly harder to accomplish. It also can be difficult to do so reproducibly in diverse individuals due to polymorphisms in either of the two priming sites in different individuals of a population. This is more likely to be an issue the more nucleotides involved, i.e. with more primer sequences used in the assay. Until recently it was difficult to get long-range information about a region without sequencing the entire genome of an individual or by amplifying smaller portions of the target region at a time and Sanger sequencing the products.
[0004] Ways to amplify across regions of unknown makeup in the past involved employing a locus-specific primer paired with a universal priming site, introduced either by using degenerate primers to target random places close to the locus priming site, or by adding adapters to ends ofthe DNA. These adapters, are attached by ligation to ends created either by restriction enzyme digestion (fixed-end) or mechanical shearing (sheared-end), can then be targeted for primer binding and amplification. These approaches have always been limited by the length of reads possible on short read sequencers or hampered by the difficulties of long-range PCR. Being able to amplify loci without knowing both ends of the molecule(s) you’re investigating opens so many avenues for research especially when you consider that you can now sequence much longer fragments so easily.
[0005] Knowing the arrangement of transgene integrations is critical for the analysis of transgenic organisms because the transgene’s effect may be influenced by this arrangement. The insertion position in the genome, copy number and orientation of any tandem repeats or concatemers of the vector can all affect the efficacy of transgene expression. Insertion in the incorrect location or close-by to an oncogene can cause mutations which will have deleterious effects on the organism under study. It can be important that researchers can ascertain the location of integration events.
[0006] Since most techniques for transgene insertion are random, finding the exact location of an integration in a model organism or analyzing a distinct part of the genome has been largely intractable, until now. Investigating mutations induced by selective modification of the genome (CRISPR, etc.), or focusing on a specific part of the genome that may be important for human disease, is also a long-standing need that has only been partially met by existing technologies. Targeting a specific part of the genome with long reads has proven to be difficult. You can amplify across a suspected locus but if the insertion is too large it is very unlikely to be amplifiable with current PCR techniques.
[0007] The methods described herein address the aforementioned limitations of prior methods. Recognized herein is that a synergistic combination of tagmentation of genomic material using a transposase, enrichment, sequencing, and bioinformatic analysis can be used to characterize a genomic location of interest (e.g., without knowing the genomic location prior to performing the method).
[0008] Many protocols do exist to investigate the location of a transgene vector integration, or other induced events in the genome, but these approaches tend to be overly expensive and / or cumbersome for the average researcher, especially when working with the larger genomes of human cell lines or mouse models. What is needed is a rapid assay, which can quickly identify transgene insertion sites or verify sequence changes at locations targeted by new genome modification technologies, with limited to no knowledge of the location of interest. The present invention fits that bill.
[0009] In an aspect, provided herein is a method for characterizing a genomic locus. Themethod can comprise providing a quantity of genomic DNA having a locus of interest therein; contacting the genomic DNA with a transposase, which transposase fragments the genomic DNA and appends transposon DNA on a terminus of the fragments, thereby creating tagmented DNA fragments; performing a PCR reaction utilizing (i) the tagmented DNA fragments; (ii) a first primer that anneals to a portion of the transposon DNA; and (iii) a second primer that anneals to a portion of the locus of interest; and sequencing a product of the PCR reaction.
[0010] In some embodiments, the locus of interest is a transgenic segment.
[0011] In some embodiments, the method further comprises, prior to sequencing the product of the PCR reaction, enriching the product of the PCR reaction for sequences having at least a minimum number of base pairs.
[0012] In some embodiments, the method further comprises aligning the sequences of the product of the PCR reaction to characterize the genomic DNA adjacent to the locus of interest.
[0013] In some embodiments, the transposase is a hyper-active transposase.
[0014] In some embodiments, the transposase has at least about 75% amino acids in common with a MuA or TN5 transposase.
[0015] In some embodiments, the transposase is MuA transposase.
[0016] In some embodiments, the transposon DNA is double stranded.
[0017] In some embodiments, the transposon DNA comprises an adapter.
[0018] In some embodiments, the tagmented DNA fragments are double stranded.
[0019] In some embodiments, the tagmented DNA fragments have a single strand nick in proximity to a location at which the transposon DNA is appended to the genomic DNA.
[0020] In some embodiments, the first primer anneals to the adapter.
[0021] In some embodiments, the second primer anneals to a portion of the locus of interest that is in proximity to a terminus of the locus of interest.
[0022] In some embodiments, the sequencing is performed using a method that results in sequence contigs having at least about 200 base pairs.
[0023] In some embodiments, the amplicons are passed through a nanopore, thereby generating a plurality of signals corresponding to an identity of a series of nucleic acid bases.
[0024] In another aspect, provided herein is a kit configured to perform the methods described herein.
[0025] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure.Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE
[0026] All publications, patents, and patent applications mentioned in this specification, including U.S. Provisional Patent Application No. 63 / 689,137, filed on August 30, 2024, are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:
[0028] FIG. 1 schematically illustrates an example of producing tagmented DNA segments, as described herein.
[0029] FIG. 2 schematically illustrates an example of amplification and sequencing of tagmented DNA segments, as described herein.
[0030] FIG. 3 schematically illustrates an example of primer design for piggybac insertions, as described herein.
[0031] FIG. 4 schematically illustrates an example of characterizing transformants, as described herein.
[0032] FIG. 5 schematically illustrates an example of alignment to confirm insertion location, as described herein.DETAILED DESCRIPTION
[0033] Traditional methods for identifying vector insertional or mutational modifications include in situ hybridization, whole genome shotgun sequencing, circularization followed by inverse PCR, microarray hybrid capture and chromosome walking. Difficult and complicated approaches for determining an integration site or investigating a distinct gene region also exist, such as Targeted Locus Amplification (TLA), but the protocols can be cumbersome to perform andperform poorly.
[0034] With prior methods, it also can be difficult to determine the copy number of transgene vector insertions, or more global genome modifications induced by transgene integration or selective DNA modification, due to read-length limitations of past sequencing technologies. These limitations can be somewhat alleviated by employing long-read sequencing techniques such as nanopore sequencing (PromethlON, Oxford Nanopore Technologies).
[0035] The location of the integration of a transgene provides relevant information such as the prediction of potential phenotypic complications and rearrangements of the transgene or the host genome at the insertion site. This could serve as a filtering process, preventing the use of any transgenic organism with an unintentional activation of oncogenes. However, prior PCR-based techniques used to identify the insertion site (IS) of the integration event do not adequately address this need. For example, the presence of multiple integration events or the use of large transgenes may reduce the reliability of this prior detection method.
[0036] Droplet digital PCR (ddPCR) is sometimes employed to quantify the number of vector inserts to wild-type (no-insert) alleles per cell or within a transgenic organism. Flow cytometry of cells from transgenic animals shows consistent phenotypes to the results of ddPCR, however, the results are often inconclusive.
[0037] In contrast, the methods described herein can work with a lack of knowledge of the sequences involved and do not suffer from the limitations of the methods described above. The method involves using the random ends created by tagmetation to amplify any locus. One might use the approach to verify that an insert is retained in progeny from a cross from individuals that have already been characterized or to check that multiple different insertions are retained following a cross involving many transgenic constructs. One can use the method to check for new integration events following another round of insertional mutagenesis.
[0038] The methods described herein are a simple and rapid protocol that amplifies long fragments from a specific location in the genome, or from withing a foreign sequence such as transgene vector integration, out of a pool of transposon-tagged DNA molecules. DNA transposons move from one genomic location to another by a cut-and-paste mechanism. They are powerful forces of genetic change and have played a significant role in the evolution of many genomes. As genetic tools, DNA transposons can be used to introduce a piece of foreign DNA into a genome. Tagmentation, or in vitro transposition to fragment DNA and add adapters, results in a pool of long DNA fragments, which are modified with transposon sequences attached at both ends of the molecules.
[0039] The sequences introduced by a vector, as well as endogenous sequences in the genome, can be targeted as priming sites for PCR out of this pool. With a little prior knowledge about thesequence of the inserted elements, or the expected location of an induced insertion or mutation in a genome, you can amplify long fragments of DNA using a region-specific primer paired with a universal primer that primes off the transposon sequence. This will amplify all fragments from locations in the genome that contain a priming site for the locus focus primer. Since the vector sequence could be located anywhere along the fragments created by random tagmentation, a smear of amplified fragments will result. The resulting PCR product can then be processed for long-read sequencing to identify the locus of focus.
[0040] Tagmentation of genomic DNA leads to fragments with adapter sequences attached at both ends and potentially the region of interest (ROI) somewhere along the fragment sequence. During tagmentation, when genomic DNA is mixed with a transposome complex (a hyperactive [MuA] transposase enzyme combined with adapter oligonucleotides), it will simultaneously fragment the target DNA in a random manner and append adapter sequences at both ends of the molecules. These adapter sequences, along with locus-specific primerss, can be used to amplify any locus out of the resulting pool of sheared fragments using PCR.
[0041] Turning to FIG. 1, provided herein is a method for characterizing a genomic locus.
[0042] The method can comprise providing a quantity of genomic DNA 100 having a locus of interest 102 therein. The method can include contacting the genomic DNA with a transposase 104. The transposase includes one or more polypeptides 106 associated with DNA fragments that comprise transposon DNA 108 and adapters 110 that enable sequencing of DNA on a sequencing platform (e.g., a nanopore sequencing platform).
[0043] Continuing with FIG. 1, the transposase fragments 112 the genomic DNA and appends transposon DNA 114 on a terminus of the fragments, thereby creating tagmented DNA fragments 116. For clarity, FIG. 1 shows only tagmentation of a single copy of genomic material, however a sample will typically have many copies of the genome, e.g., resulting in many tagmented fragments 116 that are fragmented at different genomic positions. Only a portion of the tagmented DNA fragments will include the locus of interest 102. Fragments not having the locus of interest are not shown in FIG. 1.
[0044] Sequences introduced into the genome by vector integration, or endogenous sequences in the genome, can be used as priming sites to target these integrations with long-range PCR. A locus-specific primer, paired with a universal primer that primes off the MuA transposon sequence attached at the ends of all tagmented DNA fragments, will amplify all locations in the genome that contain a priming site for the locus focusing primer. Since the vector sequence could be located anywhere along the fragments created by random tagmentation, a smear of amplified fragments will result. The resulting PCR product can then be processed for long-read sequencing to find the locus of focus.
[0045] Continuing to FIG. 2, the method continues by performing a PCR reaction utilizing the tagmented DNA fragments 200 (only one of which is shown here for simplicity). The PCR reaction uses a first primer 202 that anneals to a portion of the transposon DNA (e.g., to the adapter portion of the transposon DNA) and a second primer 204 that anneals to a portion of the locus of interest 206. In the melting of the double strands and first round of PCR reaction 208, one of the strands will not 210 be properly primed because the tagmentation process leaves a gap 212 in one of the strands. On the other strand, the first-round synthesis fills in the gap such that subsequent PCR rounds 212 results in exponential amplification of the desired product. The resultant library 214 contains (many copies of each of) a diversity of DNA molecules having a portion of the locus of interest on a first end, followed by genomic DNA flanking the locus of interest, with transposase DNA and adapters for sequencing on the other end. These molecules can be used for DNA sequencing 216.
[0046] In some embodiments, the locus of interest is a transgenic segment.
[0047] In some embodiments, the method further comprises, prior to sequencing the product of the PCR reaction, enriching the product of the PCR reaction for sequences having at least a minimum number of base pairs.
[0048] In some embodiments, the method further comprises aligning the sequences of the product of the PCR reaction to characterize the genomic DNA adjacent to the locus of interest.
[0049] In some embodiments, the transposase is a hyper-active transposase.
[0050] In some embodiments, the transposase has at least about 75% amino acids in common with a MuA or TN5 transposase.
[0051] In some embodiments, the transposase is MuA transposase.
[0052] In some embodiments, the transposon DNA is double stranded.
[0053] In some embodiments, the transposon DNA comprises an adapter.
[0054] In some embodiments, the tagmented DNA fragments are double stranded.
[0055] In some embodiments, the tagmented DNA fragments have a single strand nick in proximity to a location at which the transposon DNA is appended to the genomic DNA.
[0056] In some embodiments, the first primer anneals to the adapter.
[0057] In some embodiments, the second primer anneals to a portion of the locus of interest that is in proximity to a terminus of the locus of interest.
[0058] In some embodiments, the sequencing is performed using a method that results in sequence contigs having at least about 200 base pairs.
[0059] In some embodiments, the amplicons are passed through a nanopore, thereby generating a plurality of signals corresponding to an identity of a series of nucleic acid bases.
[0060] In another aspect, provided herein is a kit configured to perform the methods describedherein.Use Cases
[0061] With the introduction of CRISPR, researchers obtained an inexpensive and effective tool for targeted mutagenesis. Despite some limitations, CRISPR has been widely adopted in research settings and has made inroads into medical applications. Successful genome editing relies on the ability to confidently identify induced mutations after repair through nonhomologous end-joining (NHEJ) or homology directed repair (HDR). Insertions or deletions (indels) are often identified by sequencing the targeted loci and comparing the sequenced reads to a reference sequence. Deep sequencing has the advantage of both capturing the nature of the indel, readily identifying frameshift mutations or disrupted regulatory elements, and characterizing the heterogeneity of the introduced mutations in a population. This is of particular importance when the aim is allelespecific editing, or the experiment can result in mosaicism. These aforementioned use cases related to CRISPR technology can be achieved using the methods described herein.
[0062] Short tandem repeat (STR) analysis is a common molecular biology method used to compare allele repeats at specific loci in DNA between two or more samples. The methods described herein can be employed to target these repeat regions and look for length polymorphisms.
[0063] The methods described herein can be used with transposable element genotyping. I.e., identifying and genotyping all the retrotransposon sequences in a genome. The use of molecular markers has become an essential part of molecular genetics through their application in numerous fields, which includes identification of genes associated with targeted traits, operation of backcrossing programs, modem plant breeding, genetic characterization, and marker-assisted selection. Transposable elements are a core component of all eukaryotic genomes, making them suitable as molecular markers. One could design a locus focus primer with degenerate bases that targets all transposable elements of a given subtype. You would discover the locations of the elements as well as be able to identify SNPs associated with specific elements that distinguish individuals from each other.Additional considerations
[0064] The present disclosure recognized that the ability to amplify long fragments off the sheared end of transposed DNA, after tagmentation with the MuA transposase enzyme, led to the discovery of our ability to optimize amplification off specific loci in a DNA sequence using just a “single” PCR primer targeting a region of interest (whether it be a vector inserted in the genome of a transgenic organism, the location of an induced mutation or other genomic anomalies). This allows for the rapid amplification of any locus in the genome with only a minimum of sequence information.
[0065] MuA transposase enzyme simultaneously catalyzes fragmentation of double-stranded target DNA and tagging of the fragment ends with transposon DNA sequences, which can then be targeted as a universal priming site for locus focus. An optimized selective primer and high- quality DNA are also required.
[0066] Many insertions are not simple, single copy, “between the ITRs” cassette insertions. You will see concatemeric vector insertions where the cassettes are lined up in a row, or back-to- back, as well as any combination of the two, with full-length or deletion riddled pieces of the vector sequences themselves.
[0067] Nanopore sequencing, for longer reads, are useful for this approach to work efficiently.
[0068] Amplification off the transposon end coupled with a selective primer for long-read sequencing is a feature of the methods described herein. This can be done in a similar fashion for short-read Illumina sequencing.Example 1 - Primer design for piggybac vector
[0069] Provided herein is an example of primer design and expected output for a locus focus project on piggybac vectors. We designed a primer targeting the WPRE element in a piggybac vector, which has been integrated into 10 different human IPSC cell lines (FIG. 3). The long reads that are produced should emanate from the locus-specific primer, proceed through some vector sequence, and then into the insert region of the host genome.
[0070] A primer targeting the end of a WPRE cassette 301 in the vector used for insertion, can amplify from that location out to the edge of the expected insert 302 and then into the insertion location in the genome 303.
[0071] The protocol steps followed were: Tagment genomic DNA to shear and attach adapters. Perform PCR using MuA and the WPRE locus-specific primer. Purify large fragments and use a commercial library preparation kit to attach sequencing adapters for long-read sequencing. Sequencing output can be aligned to the genome of interest and viewed with the Integrated Genome Viewer (IGV, Broad Institute).
[0072] The present method also enables the targeted sequencing of any locus of interest or even to verify the vector sequence inserted at a specific location. The technique can be used to target reads to a given region of the genome, or to tile across that region with multiple primers to investigate SNPs, deletions, or translocations, etc. Output would look much the same as in FIG. 3, except the reads would emanate directly from the priming site located in the endogenous genome.Example 2 - Characterizing transformants
[0073] FIG. 4 shows the alignments from locus focus amplification targeting piggybac insertions in 10 different human cell lines, all potentially carrying many piggybac integrationevents. The reads in individuals that have an insertion at this location show reads that amplify from within the transgene, cross 1 ,2kb of vector sequence, before heading out into the surrounding DNA of the integration site, which happens to be on chromosome 4 in the RNF150 gene region. Cell line, pig8 does not appear to have this insertion, while all other lines do.
[0074] The reads show the location of integration. Alignments from locus focus data using primers targeting piggybac vector sequence insertions in human cell lines shown in IGV display. Reads in individuals with an insertion show reads heading out from the vector and into the surrounding DNA of the integration site. Cell line, pig8 does not appear to have this insertion. Example 3 - Whole genome alignment
[0075] We have performed long-read whole genome sequencing (WGS) on many of the same samples to verify integrations discovered using locus focus. FIG. 5 shows data displayed in IGV using a primer targeting the Kanamycin resistance (KanR) cassette of the vector used in Agrobacterium-infected Arabidopsis plants. The locus focus reads are displayed in the top panel and show a complex integration site with reads from the primer emanating out from both sides, indicating a back-to-back architecture of at least two KanR priming sites. Panel A shows the reads from the Locus Focus PCR aligned to the mouse reference. Panel B shows whole genome long-reads from the same sample as in A, indicating the same insertion site as locus focus with reads that terminate at the junction site. This is what you expect to see, these reads continue into the sequences of the integration site. Panel C shows the WGS reads from the wild-type plant that does not possess the insertion with all reads spanning the Agrobacterium integration site.
[0076] Locus Focus data is recapitulated with whole genome long-read sequencing. Locus Focus data displayed in IGV from a primer targeting the Kanamycin resistance (KanR) cassette of the vector used in transgenic Arabidopsis plants. Locus Focus read alignments are displayed in the top panel and show a complex integration site with reads from the primer emanating out from both sides. Panel A shows the read alignments from the Locus Focus PCR aligned to the mouse reference. Panel B shows whole genome sequencing long-read alignments from the same sample as A. Panel C shows the WGS alignments from the wild-type plant that does not possess the insertion.
[0077] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” may apply to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 may be equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.
[0078] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “lessthan,” or “less than or equal to” may apply to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 may be equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.
[0079] The term “at least one of A and B” and "at least one of A or B" may be understood to mean only A, only B, or both A and B. The term "A and / or B" may be understood to mean only A, only B, or both A and B.
[0080] The term “about” as used herein, generally refers to a quantity that is within twenty percent (20%) of the stated quantity.
[0081] Unless indicated otherwise, all percentages when used in the context of concentration are expressed as mole fractions (i.e., mol%).
[0082] As used herein, a "transposable element" (TE, transposon, or jumping gene) is a nucleic acid sequence in DNA that can change its position within a genome, sometimes creating or reversing mutations and altering the cell's genetic identity and genome size. Transposition often results in duplication of the same genetic material.
[0083] As used herein, a "DNA transposase" is an enyzme that move discrete segments of DNA called transposons from one location in the genome (often called the donor site) to a new site.
[0084] As used herein, "tagmentation", is a method in which a hyperactive transposase is used to simultaneously fragment target DNA and append universal adapter sequences, is used to prepare high-throughput sequencing libraries. Tagmentation effectively replaced a series of processing steps in traditional workflows with one single reaction. It is the simplicity, coupled with the high efficiency of tagmentation, that has made it a favored means of sequencing library construction and fueled a diverse range of adaptations to assay a variety of molecular properties.
[0085] As used herein, a "vector", as related to molecular biology, is a DNA molecule (often plasmid or virus) that is used as a vehicle to carry a particular DNA segment into a host cell as part of a cloning or recombinant DNA technique.
[0086] As used herein, a "transgene" is an experimentally constructed piece of DNA that has integrated into the genome of a recipient organism. Once integrated into germ cells, subsequent generations inherit the transgene, referred to as stable transgene transmission.
[0087] CRISPR (short for “clustered regularly interspaced short palindromic repeats”) is a technology that research scientists use to selectively modify the DNA of living organisms. CRISPR was adapted for use in the laboratory from naturally occurring genome editing systems found in bacteria.
[0088] Cas9 (or "CRISPR-associated protein 9") is an enzyme that uses CRISPR sequences as a guide to recognize and open up specific strands of DNA that are complementary to the CRISPR sequence. Cas9 enzymes together with CRISPR sequences form the basis of a technology knownas CRISPR-Cas9 that can be used to edit genes within the organisms. This editing process has a wide variety of applications including basic biological research, development of biotech products, and treatment of diseases.
[0089] The polymerase chain reaction (PCR) is a method widely used to make millions to billions of copies of a specific DNA sample rapidly, allowing scientists to amplify a very small sample of DNA (or a part of it) sufficiently to enable detailed study.
[0090] Genotyping is the process of determining differences in the genetic make-up (genotype) of an individual by examining the individual's DNA sequence using biological assays and comparing it to another individual's sequence or a reference sequence.
[0091] Polymorphism, as related to genomics, refers to the presence of two or more variant forms of a specific DNA sequence that can occur among different individuals or populations. The most common type of polymorphism involves variation at a single nucleotide (also called a single-nucleotide polymorphism, or SNP). Other polymorphisms can be much larger, involving longer stretches of DNA.
[0092] A single-nucleotide polymorphism (SNP) is a germline substitution of a single nucleotide at a specific position in the genome. SNPs can help explain differences in susceptibility to a wide range of traits or diseases across a population.
[0093] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A method for characterizing a genomic locus, the method comprising: a. providing a quantity of genomic DNA having a locus of interest therein; b. contacting the genomic DNA with a transposase, which transposase fragments the genomic DNA and appends transposon DNA on a terminus of the fragments, thereby creating tagmented DNA fragments; c. performing a PCR reaction utilizing (i) the tagmented DNA fragments; (ii) a first primer that anneals to a portion of the transposon DNA; and (iii) a second primer that anneals to a portion of the locus of interest; and d. sequencing a product of the PCR reaction.
2. The method of Claim 1, wherein the locus of interest is a transgenic segment.
3. The method of Claim 1, further comprising, prior to sequencing the product of the PCR reaction, enriching the product of the PCR reaction for sequences having at least a minimum number of base pairs.
4. The method of Claim 1, further comprising aligning the sequences of the product of the PCR reaction to characterize the genomic DNA adjacent to the locus of interest.
5. The method of Claim 1, wherein the transposase is a hyper-active transposase.
6. The method of Claim 1, wherein the transposase has at least about 75% amino acids in common with a MuA or TN5 transposase.
7. The method of Claim 1, wherein the transposase is MuA transposase.
8. The method of Claim 1, wherein the transposon DNA is double stranded.
9. The method of Claim 1, wherein the transposon DNA comprises an adapter.
10. The method of Claim 1, wherein the tagmented DNA fragments are double stranded.
11. The method of Claim 1, wherein the tagmented DNA fragments have a single strand nick in proximity to a location at which the transposon DNA is appended to the genomic DNA.
12. The method of Claim 1, wherein the first primer anneals to the adapter.
13. The method of Claim 1, wherein the second primer anneals to a portion of the locus of interest that is in proximity to a terminus of the locus of interest.
14. The method of Claim 1, wherein the sequencing is performed using a method that results in sequence contigs having at least about 200 base pairs.
15. The method of Claim 1, wherein the amplicons are passed through a nanopore, thereby generating a plurality of signals corresponding to an identity of a series of nucleic acid bases.
16. A kit configured to perform the method of any of the preceding claims.
Citation Information
Patent Citations
Transposon end compositions and methods for modifying nucleic acids
US20100120098A1
Methods and transposon nucleic acids for generating a DNA library
US20130017978A1
In vitro method for providing templates for DNA sequencing
US6593113B1