Capture and analysis of target genomic regions

The use of programmable endonucleases and crosslinked oligos addresses the challenges of capturing and analyzing long genomic regions, enabling efficient detection and sequencing of target genomic sequences up to 50 kb.

JP7854386B2Active Publication Date: 2026-05-01LGC GENOMICS LLC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
LGC GENOMICS LLC
Filing Date
2020-05-13
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing methods for capturing and analyzing target genomic regions, particularly long stretches of DNA sequences, are difficult to optimize, not suitable for multiplexing, and challenging for sequencing amplification products, especially for sequences longer than 1-50 kb.

Method used

A method involving the use of programmable endonucleases like Cas9 to cleave target genomic regions at specific sites, followed by hybridization with crosslinked oligos for capture and subsequent amplification and sequencing, allowing for the analysis of long genomic regions up to 50 kb.

Benefits of technology

Enables efficient capture and analysis of long genomic regions, facilitating detection and sequencing of target genomic sequences, particularly through techniques like nanopore sequencing and PCR, with improved specificity and scalability for multiple targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007854386000002
    Figure 0007854386000002
  • Figure 0007854386000003
    Figure 0007854386000003
  • Figure 0007854386000004
    Figure 0007854386000004
Patent Text Reader

Abstract

The present disclosure relates to materials and methods for capturing target genomic regions, including cleaving the target genomic region using an endonuclease at a specific recognition site comprising a sequence of at least about 10 to about 30 nucleotides, and capturing the target genomic region by hybridizing the target genomic region to a bridging oligonucleotide. The disclosure also relates to analyzing the captured target genomic region. The endonuclease used for cleavage can be one or more programmable endonucleases that specifically bind to and direct cleavage at the recognition site. The captured target genomic region is preferably amplified via polymerase chain reaction or rolling circle amplification and can be detected or sequenced. Furthermore, the present invention relates to kits for carrying out the methods of the present invention, comprising one or more endonucleases designed to cleave one or more target genomic regions and bridging oligonucleotides designed to capture one or more target genomic regions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference to related applications This application claims the priority of U.S. Provisional Application No. 62 / 846,988, filed on May 13, 2019, and is incorporated by reference in its entirety including all drawings and content containing amino acid or nucleic acid sequences.

[0002] The sequence listing of this application was created on May 11, 2020, and is labeled "Seq-List.txt" with a size of 3 KB. The entire content of the sequence listing is incorporated herein by reference in its entirety.

Background Art

[0003] Capture and analysis of sequencing etc. of target genomic regions are extremely important in various biotechnology fields ranging from agronomy, taxonomy to medicine. For example, sequencing of target regions in the genome is important in precision and personalized medicine where understanding genetic diversity in specific regions of the genome can be linked to phenotypic information regarding the health of a subject and as a result can direct treatment options. Also, sequencing of target regions of the genome is important in agronomic applications for improving desirable traits in plants such as seeds and food production.

[0004] Several methods exist for analyzing target genomic DNA or RNA sequences. For example, pre-sequencing targeted amplification is common, and multiple methods are available for isolating and / or amplifying target regions of the genome. Common themes include the use of polymerase chain reaction (PCR), ligation chain reaction, or probe hybridization to isolate and amplify target genomic regions. Current methods for isothermal detection of genetic material are performed using, for example, loop-mediated isothermal amplification (LAMP) or strand displacement amplification. However, these techniques are difficult to optimize for different targets and are not suitable for analyzing long targets. Furthermore, these isothermal techniques are difficult or impossible to multiplex for numerous targets and are not suitable for easily sequencing amplification products. For example, there is particular interest in identifying, and especially detecting and sequencing, longer stretches of DNA sequences, such as at least about 1–50 kilobases (kb). However, such detection and sequencing are not easily achieved by conventional methods. Therefore, there is a need for improved methods for identifying, for example, detecting and sequencing, target genomic regions, and especially for detecting and sequencing long stretches of target genomic regions. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Dahl, F., Stenberg, J., Fredriksson, S., Welch, K., Zhang, M., Nilsson, M., et al. (2007). Multigene amplification and massively parallel sequencing for cancer mutation discovery. Proc. Natl. Acad. Sci. USA 104, 9387-92. doi:10.1073 / pnas.0702165104. [Non-Patent Document 2] Fredriksson , S. , Baner , J. , Dahl , F. , Chu , A. , Ji , H. , Welch , K. , et al. (2007). Multiplex amplification of all coding sequences within 10 cancer genes by Gene-Collector. Nucleic Acids Res. 35, e47. doi:10.1093 / nar / gkm078.

Outline Building1

Outdoor Tools3

Outdoor Tools 4

Direct Environment 5

Outdoor Configuration6

Direct Environment 7

Outdoor Tools 8

[0006] Certain embodiments of the present invention provide materials and methods for capturing a target genomic region and optionally further analyzing it, for example, by detection and / or sequencing. The materials and methods disclosed herein are particularly suitable for analyzing a target genomic region comprising at least about 1 kb, preferably at least about 10 kb, more preferably at least about 30 kb, or most preferably at least about 50 kb.

[0007] According to the method disclosed herein, a target genomic region is isolated based on cleavage at two specific recognition sites, each recognition site containing at least about 10 to about 30 nucleotides. The two recognition sites are adjacent to the target genomic region. The cleaved target genomic region is then captured using an oligonucleotide referred to herein as “crosslinked oligo”. The crosslinked oligo has 3' and 5' terminal sequences that hybridize to the 3' and 5' terminal sequences of the cleaved target genomic region, respectively. The captured target genomic region is then detected and / or sequenced, for example, by amplification and sequencing.

[0008] Therefore, a particular embodiment of the present invention is a method for capturing a target genomic region from genetic material, a. Using one or more endonucleases containing a sequence of at least approximately 10 to 30 nucleotides and having a first recognition site and a second recognition site adjacent to the target genomic region, the target genomic region is cleaved from the genetic material. b. The cleaved genetic material is transformed into a single-stranded form, c. The single-stranded target genome region is captured by hybridizing it to a cross-linked oligo containing 3' and 5' sequences that hybridize to the 3' and 5' ends of the single-stranded target genome region, respectively. This includes the following.

[0009] The captured target genomic region is further analyzed, for example, by detection or sequencing. Such analyses are performed. d. A step of generating a single-stranded circular target genome region hybridized to a cross-linked oligo by ligating the free end of the single-stranded target genome region hybridized to the cross-linked oligo, e. Optionally, a step to degrade the non-cyclic genetic material, f. Optionally, a step of amplifying a target genome region by nucleic acid amplification to generate multiple copies of the target genome region, g. A step to analyze the amplified target genomic region. Includes.

[0010] In a preferred embodiment, cleavage of the target genomic region at the recognition site is performed using first and second programmable endonucleases, such as Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-related protein 9 endonuclease (Cas9 endonuclease), for example, a first Cas9 endonuclease containing a first guide RNA (gRNA) having a sequence complementary to the first recognition site, and a second Cas9 endonuclease containing a second gRNA having a sequence complementary to the second recognition site.

[0011] In further embodiments, the target genomic region hybridized to a crosslinked oligo is amplified via an amplification reaction, such as an isothermal amplification reaction, preferably via a polymerase chain reaction or a rolling circle amplification (RCA) reaction. When only the crosslinked oligo is used as a primer for amplification, a single-stranded amplification product is generated via an RCA containing a linked copy of the target genomic region in single-stranded form. In addition to the crosslinked oligo, one or more primers are also used in the RCA reaction to generate a double-stranded amplification product containing a linked copy of the target genomic region in double-stranded form. The target molecule can also be amplified using a conventional PCR reaction with appropriate primers that bind to the target region or a common region introduced by the crosslinked oligo.

[0012] The amplified target genomic region is detected using techniques known in the art, for example, using labeled probes complementary to the sequence within the target genomic region. The amplified target genomic region is also sequenced using techniques known in the art, such as nanopore sequencing (Oxford Nanopore Technologies®), reversible dye-terminator sequencing (Illumina®), and single-molecule real-time sequencing (SMRT) (PacBio®).

[0013] The materials and methods disclosed herein can be modified, for example, to capture and optionally analyze multiple target genomic regions in a multiple reaction. In such embodiments, multiple pairs of target recognition sites are designed to cleave multiple target recognition sites, and these multiple target genomic regions are captured via multiple cross-linking oligos, each specifically designed to capture the target genomic regions. In certain embodiments, the multiple cross-linking oligos can be immobilized on a solid substrate such as a chip. The thus captured multiple target genomic regions can be detected or sequenced using techniques known in the art.

[0014] A further embodiment of the present invention also provides a kit for implementing the method of the present invention. The kit of the present invention includes one or more endonucleases designed to cleave one or more target genomic regions, and specific crosslinking oligos designed to capture one or more target genomic regions.

[0015] In certain embodiments, the kit of the present invention includes one or more guide molecules in the form of DNA or RNA designed to cleave one or more target genomic regions, and specific crosslinking oligos designed to capture one or more target genomic regions. Such a kit can also include one or more Cas9 endonucleases.

[0016] The kit can further include polymerase, ligase, primers and other reagents for amplifying one or more captured target genomic regions. Additionally, the kit can provide instructions for implementing the method of the present invention.

[0017] This patent or patent application documents include at least one drawing created in color. Copies of this patent or patent application publication with color drawings will be provided by the Patent and Trademark Office upon request and payment of the necessary fees.

Brief Description of the Drawings

[0018] [Figure 1] Schematic diagrams of two examples of methods for capturing a target genomic region and analyzing, i.e., detecting or sequencing, the target genomic region. [Figure 2] An example of sequencing a target genomic region captured using single molecule real-time (SMRT) sequencing (PacBio (trademark)). [Figure 3] An example of detecting multiple target genomic regions using the sequences of crosslinking oligos designed to capture multiple target genomic regions. [Figure 4]An example of amplification of a cyclized target molecule by PCR. An example of a primer containing an adapter necessary for sequencing using reversible dye-terminator sequencing (Illumina®) is shown. A) Process overview. B) Structure of the cross-linking oligo required to capture the sequence and provide a landing site for the primer. C) Primer structure enabling simultaneous amplification of the molecule, incorporation of a unique sample identifier, and incorporation of structures necessary for dye-terminator sequencing. [Figure 5] An example of designing recognition sites and crosslinking oligos based on two gRNAs that bind to the same strand of a double-stranded genome. [Figure 6] An example of designing recognition sites and crosslinking oligos based on two gRNAs that bind to opposite strands of a double-stranded genome. [Figure 7] Another example of designing recognition sites and crosslinking oligos based on two gRNAs that bind to opposite strands of a double-stranded genome. [Figure 8] Another example of designing recognition sites and crosslinking oligos based on two gRNAs that bind to the same strand of a double-stranded genome. [Figure 9] Examples of various approaches to form a cyclic single-chain structure by ligating the ends of a target molecule. [Figure 10] An example of an RCA reaction arising from a cross-linked oligo having a non-complementary 5' end. [Figure 11] Examples of forming cyclic single-chain molecules and cross-linked oligos using a single programmable endonuclease and primers extending from the target. [Modes for carrying out the invention]

[0019] As used herein, the singular forms “one” and “it” are intended to include the plural form unless the context explicitly indicates otherwise. Furthermore, to the extent that “including,” “having,” “equipped with,” or variations thereof are used in any part of the detailed specification and / or claims, such terms are intended to encompass the term “comprising.” Transitional terms / phrases (and any grammatical variations thereof) including “including,” “consisting of,” and “essentially consisting of” are used without distinction.

[0020] The phrase "essentially derived from" indicates that the described embodiments include embodiments that involve specific materials or processes, and embodiments that do not substantially affect the basic and novel features of the described embodiments.

[0021] The term "approximately" means an acceptable margin of error for a particular value that can be determined by those skilled in the art, which depends in part on how the value is measured or determined, i.e., the limits of the measurement method. With respect to the length of a polynucleotide in which the term "approximately" is used, the polynucleotide contains a stated number of bases or base pairs with a variation of 0–10% around the value (X ± 10%).

[0022] In this disclosure, ranges are described concisely to avoid setting a length and describing each and all values ​​within that range. Any suitable value within a range is selected, where appropriate, as the upper limit, lower limit, or end of the range. For example, the range 0.1 to 1.0 represents the end values ​​of 0.1 and 1.0, the intermediate values ​​of 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, and 0.9, and all intermediate ranges contained within 0.1 to 1.0, such as 0.2 to 0.5, 0.2 to 0.8, and 0.7 to 1.0. Values ​​with at least two significant figures are assumed within a range; for example, the range 5 to 10 represents all values ​​between 5.0 and 10.00, and all values ​​between 5.00 and 10.00, including the end values. Where ranges such as the size of a polynucleotide are used herein, combinations of ranges and subcombinations (e.g., sub-ranges within the disclosed range), and specific embodiments therein, are expressly included.

[0023] As used herein, the term “organism” includes viruses, bacteria, fungi, plants, and animals. Further examples of organisms are known to those skilled in the art, and such embodiments are within the scope of the present invention. The assays described herein are useful for analyzing any genetic material obtained from any organism.

[0024] "Genome" or "genetic material" and other grammatical variations used herein refer to genetic material of any living organism. Genetic material is the genetic material found in organelles, such as viral genomic DNA or RNA, nuclear genetic material such as genomic DNA, or mitochondrial DNA or chloroplast DNA. It also refers to a mixture of natural or artificial genetic material or genetic material derived from some organism.

[0025] As used herein, the term “long target genome region” refers to a target genome region having at least about 1 kb, preferably at least about 10 kb, more preferably at least about 30 kb, or most preferably at least about 50 kb.

[0026] The materials and methods disclosed herein for identifying target genomic regions, particularly long target genomic regions, address the challenges associated with conventional methods for capturing and detecting target genomic regions, especially long target genomic regions.

[0027] In certain embodiments, the present invention provides a method for capturing a target genomic region from genetic material. This method is a. A step of cleaving a target genomic region from genetic material using one or more endonucleases having a first recognition site and a second recognition site, which contain a sequence of at least approximately 10 to approximately 30 nucleotides and are adjacent to the target genomic region. b. A process to denature the cleaved genetic material into a single-stranded form, c. A step of capturing a single-stranded target genome region by hybridizing the target genome region to a cross-linked oligo containing 3' and 5' terminal sequences that hybridize to the 3' and 5' ends, respectively, of the single-stranded target genome region. Includes.

[0028] The captured target genomic region is further analyzed, for example, by detection or sequencing. Such analyses are performed. d. A step of generating a single-stranded circular target genome region hybridized to a cross-linked oligo by ligating the free end of the single-stranded target genome region hybridized to the cross-linked oligo, e. Optionally, a step to degrade the non-cyclic genetic material, f. Optionally, a step of amplifying a target genome region by nucleic acid amplification to generate multiple copies of the target genome region, g. A step to analyze the amplified target genomic region. Includes.

[0029] As used herein, the “target genomic region” refers to the region in the genome of an organism. Such a region is adjacent to a first recognition site and a second recognition site. Each of the first and second recognition sites contains a sequence of at least about 10 to about 30 nucleotides.

[0030] The first and second recognition sites are selected based on the target genomic region. Typically, unique sequences adjacent to the target genomic region are selected as the first and second recognition sites. The uniqueness of these sequences ensures that they are unlikely to occur elsewhere in the genome, thus minimizing nonspecific cleavage and avoiding capture of regions outside the target genomic region. The minimum length of the first and second recognition sites is preferably about 10 to about 30 nucleotides, which ensures that these sequences occur rarely in the genome, thus minimizing nonspecific cleavage and capture of regions outside the target genomic region. The sequences of the first and second recognition sites may be identical or different. Those skilled in the art can determine appropriate sequences for the first and second recognition sites based on the sequence of the target genomic region and available genomic sequences for a particular organism, for example, from a genomic sequence database.

[0031] Cleavage of genetic material at the recognition site is carried out using one or more endonucleases that cleave phosphodiester bonds (cleavage) in the genetic material at a specific recognition site. The endonucleases are restriction endonucleases. Specific restriction endonucleases that cleave recognition sites containing at least about 10 to about 30 nucleotides, e.g., up to about 80 nucleotides, include meganucleases such as homing endonucleases from the LAGLIDADG, GIY-YIG, HNH, His-Cys box, and PD-(D / E)XK series. Further examples of restriction endonucleases that cleave recognition sites containing at least about 10 to about 30 nucleotides, e.g., up to about 80 nucleotides, are known in the art, and such embodiments are within the scope of the present invention.

[0032] In a preferred embodiment, cleavage at the recognition site is performed using one or more programmable endonucleases. The term programmable endonucleases is used to describe a different class of enzymes that are targeted to cleave specific regions of DNA or RNA molecules. Thus, a programmable endonuclease is an endonuclease designed or programmed to cleave a nucleotide sequence. For example, a programmable endonuclease comprises a target recognition portion and an endonuclease portion, where a common endonuclease portion combines with any target recognition portion to cleave the nucleotide sequence.

[0033] In one embodiment, programmable endonucleases are targeted by guide RNA (gRNA), guide DNA (gDNA), or by a structure formed between a guide molecule and a target (Non-Patent Literature 7). For example, Cas9 is a programmable endonuclease because it cleaves double-stranded genetic material by performing a double-strand break at a specific location of the recognition site (Non-Patent Literature 4). A gRNA with a specific sequence complementary to the recognition site directs the Cas9 endonuclease to the recognition site. It has also been shown that several types of algorithmic proteins can be targeted using gDNA or gRNA molecules (Non-Patent Literature 2, Non-Patent Literature 6). Further examples of programmable endonucleases include, in particular, Cpf1, C2c1, C2c2, C2c, RNA- or DNA-guided algorithmic proteins, and structure-guided endonucleases.

[0034] Novel proteins can also be designed to cleave DNA based on the recognition of specific DNA structures formed between gDNA and a target sequence, such as a 3' end mismatch (Non-Patent Literature 8). Typically, in the method of the present invention, two Cas9 endonucleases are used to cleave a target genomic region from genetic material. The genetic material is, i.e., a first Cas9 endonuclease that cleaves the DNA at a first recognition site based on a first gRNA having a sequence complementary to the first recognition site, and a second Cas9 endonuclease that cleaves the DNA at a second recognition site based on a second gRNA having a sequence complementary to the second recognition site. The complex containing the gRNA and Cas9 endonuclease components is called a ribonucleoprotein (RNP) complex. In some examples, only one RNP is required to cleave one strand of double-stranded DNA, and the selection point on the other strand of double-stranded DNA is the result of a primer extension reaction (Figure 11).

[0035] The use of programmable endonucleases, such as Cas9, offers a significant improvement over other methods known in the art for generating, capturing, or manipulating cyclic molecules from target regions (Non-Patent Document 1, Non-Patent Document 2). Firstly, the use of conventional restriction enzymes limits the target genomic regions that can be analyzed because there are no combinations of restriction enzymes that cleave near the target region without cleaving within it. This becomes particularly problematic as the number of sites increases. Furthermore, even when such combinations of restriction enzymes exist, they may change each time a new target is envisioned, making it extremely difficult to scale this process. Conversely, the use of programmable endonucleases allows for targeting of virtually any region in a scalable process because the design and synthesis of unique guide oligos are well known in the art. Secondly, using PCR instead of endonucleases to extract the site to be cyclized, as done in Non-Patent Document 2, limits the size of the target molecule because PCR tends to work best on fragments smaller than 1 Kb. Furthermore, PCR requires complex optimization, and performing it on multiple target regions in parallel is extremely difficult, too costly, and ultimately impossible.

[0036] The sequences of the first and second recognition sites are expressed in this disclosure in the 5' to 3' direction. Therefore, the sequence toward the 5' end of the gRNA is complementary to the sequence of the recognition site. In addition to the sequence complementary to the recognition site, the Cas9-gRNA RNP complex also recognizes a short conserved sequence motif of about 2–5, typically 3 nucleotides, located on the non-complementary strand of target DNA, adjacent to the 3' end of the gRNA sequence complementary to the recognition site. This is called a protospacer adjacent motif (PAM). PAMs are important for gRNA binding to the recognition site as well as Cas9 cleavage of the target genetic material, but a Cas9 enzyme design that does not require this PAM site is envisioned.

[0037] One example of a PAM is the sequence NGG, but it is known that different Cas9 enzymes require different PAM sites, and the enzymes have been modified to accommodate variable PAM sites (Non-Patent Literature 3). Accordingly, in certain embodiments of gRNAs useful in the materials and methods disclosed herein, the recognition site is designed such that the sequence of the recognition site on the genetic material is immediately followed by the sequence 5'-NGG-3' on the 3' side of the non-complementary strand.

[0038] When gRNA binds to the recognition site, Cas9 endonuclease performs a double-strand break of three nucleotides toward the 5' end of the NGG sequence on the non-complementary strand in the double-stranded genetic material. That is, starting from the 5' end and moving toward the 3' end of the recognition site, Cas9 endonuclease performs a double-strand break between the third and fourth nucleotides.

[0039] Once the genetic material is cleaved using an appropriate endonuclease, the resulting mixture of target genomic regions and the cleaved fragments of the genetic material are denatured and converted to a single-stranded form, for example, by denaturing conditions. Typically, denaturing conditions for cleaved genetic material involves exposure to an appropriate temperature at an appropriate pH in the presence of an appropriate compound such as salt, dimethyl sulfoxide, or sodium hydroxide. DNA can also be denatured using chemical treatment with NaOH or high salt concentrations. In preferred embodiments, the cleaved genetic material is denatured by exposure to a temperature between 75°C and 115°C, preferably between 80°C and 110°C, more preferably between 85°C and 105°C, even more preferably between 90°C and 100°C, and most preferably around 95°C. Those skilled in the art can determine appropriate denaturing conditions, and such embodiments are within the scope of the present invention.

[0040] To capture a target genomic region isolated from the genetic material, a specially designed cross-linked oligonucleotide is brought into contact with the cleaved and denatured genetic material. Typically, the cross-linked oligonucleotide is a single-stranded oligonucleotide. The cross-linked oligonucleotide contains 3' and 5' terminal sequences that hybridize to the 3' and 5' terminal sequences of the single-stranded target genomic region, respectively. The cross-linked oligonucleotide may also be a double-stranded oligonucleotide having 3' and 5' terminal overhangs that hybridize to the 3' and 5' terminal sequences of the single-stranded target genomic region, used to capture both strands of the target DNA molecule. The cross-linked oligonucleotide can be protected from nuclease degradation by different methods known in the art, such as by adding one or more phosphorothioate bonds to the 3' and / or 5' ends, and adding inverted dT and inverted ddT at its 3' and 5' ends, so that it remains present during the reaction after treatment with a given nuclease.

[0041] Cross-linked oligonucleotides serve several purposes. They circulate the correct target genomic region generated by on-target cuts with endonucleases. Since off-target cuts are expected in the reaction with endonucleases, cross-linked oligonucleotides provide further specificity by allowing only the correct molecules to be circulated. Additionally, the cross-linked oligonucleotides themselves act as primers for subsequent amplification.

[0042] Furthermore, cross-linked oligos can be designed to provide additional functionality, such as preparing molecules for sequencing. For example, cross-linked oligos can be immobilized on solid surfaces for detection on chips or biotinylated for recovery with streptavidin-conjugated beads. The 5' and 3' ends of the cross-linked oligo also have additional “tail” sequences that are non-complementary to the target region, which can generate common sequences used to link further functionality to the molecule, such as providing a site for primer binding for PCR, adding biotin molecules, or other modifications known in the art (Figure 10).

[0043] Between the two terminal sequences that hybridize with the target genomic region, the cross-linking oligo may further include sequences such as restriction sites specific to rare-cutter restriction endonucleases, primer-binding sequences, or targets of programmable endonucleases. Using the rare-cutter restriction sites or programmable endonuclease targets, individual target genomic regions can be cleaved from linked copies of the target genomic region generated after nucleic acid amplification. Examples of rare-cutter restriction endonucleases include, but are not limited to, those described in their entirety, particularly in PCT Publication WO2009 / 079488, which is incorporated herein by reference in Table 1.

[0044] As used herein, "rare cutter restriction endonuclease" is an endonuclease whose restriction site occurs rarely in the genetic material. For example, in the human genome, rare cutter restriction endonucleases are endonucleases whose restriction sites occur at an average interval of 50 to 100 kb, preferably every 100 to 200 kb, more preferably every 200 to 400 kb, and even more preferably every 400 to 600 kb. Examples of rare cutter restriction endonucleases and their restriction sites in the human genome are shown in Table 1 below.

[0045] [Table 1]

[0046] Further rare-cutter endonucleases are described, for example, in Restriction Endonucleases (Nucleic Acids and Molecular Biology), Springer; 1st ed (2004) by Pingoud (Editor). Many rare-cutter endonucleases, such as the roaming class endonucleases from New England BioLabs (Beverly, MA), are also commercially available. Yet another example of rare-cutter endonucleases is known in the art, and such embodiments are within the scope of this invention.

[0047] The primer-binding sequences in the cross-linked oligo facilitate the use of additional primers in the RCA reaction to amplify the target genomic region and generate a linked copy of the target genomic region for double-strand formation via superbranching of the original molecule.

[0048] The first and second recognition sites can be located on the same strand of genetic material or on opposite strands of genetic material.

[0049] Figure 5 shows an example of designing recognition sites and crosslinking oligos based on two gRNAs that bind to the same strand of a double-stranded genome. As shown in Figure 5, the lower strand of double-stranded genomic DNA is selected to design the recognition site. The first gRNA binds to the first recognition site, and the Cas9 endonuclease cleaves three nucleotides of the genomic DNA downstream of the 5' end of the first recognition site. The second gRNA binds to the second recognition site, and the Cas9 endonuclease cleaves three nucleotides of the genomic DNA downstream of the 5' end of the second recognition site. In this way, the target genomic region is cleaved from the genomic DNA in a double-stranded form. This target genomic region has all but the first three nucleotides from the first recognition site at the end facing the first recognition site, and the first three nucleotides from the second recognition site at the end facing the second recognition site. This double-stranded target genomic region is converted to a single-stranded form, and one or more crosslinking oligos are designed to capture either or both strands.

[0050] To capture the lower chain as shown in Figure 5, the crosslinked oligo is designed to have, toward the 3' end, a sequence complementary to the first three nucleotides of the second recognition site, and additional nucleotides complementary to the sequence located on the lower chain beyond the 5' end of the second recognition site and adjacent to it. This crosslinked oligo has, toward the 5' end, a sequence complementary to the first recognition site excluding the first three nucleotides, and optionally, additional nucleotides complementary to the sequence located on the lower chain beyond the 3' end of the first recognition site and adjacent to it.

[0051] To capture the upper chain as shown in Figure 5, the crosslinked oligo is designed to have, toward the 3' end, the sequence of the first recognition site excluding the first three nucleotides, and optionally, a sequence located on the lower chain beyond and adjacent to the 3' end of the first recognition site. This crosslinked oligo has, toward the 5' end, the sequence of the first three nucleotides of the second recognition site, and additional nucleotides located on the lower chain beyond and adjacent to the 5' end of the second recognition site.

[0052] Figure 6 shows an example of designing recognition sites and crosslinking oligos based on two gRNAs that bind to the opposite strand of a double-stranded genome. As shown in Figure 6, the lower strand of the double-stranded genomic DNA is selected to design the first recognition site, and the upper strand of the double-stranded genomic DNA is selected to design the second recognition site. The first gRNA binds to the first recognition site on the lower strand, and the Cas9 endonuclease cleaves three nucleotides of the genomic DNA downstream of the 5' end of the first recognition site. Similarly, the second gRNA binds to the second recognition site on the upper strand, and the Cas9 endonuclease cleaves three nucleotides of the genomic DNA downstream of the 5' end of the second recognition site. In this way, the target genomic region is cleaved from the genomic DNA in a double-stranded form. This target genomic region has the sequence of the first recognition site, excluding the first three nucleotides, at the end facing the first recognition site, and the sequence of the second recognition site, excluding the first three nucleotides, at the end facing the second recognition site. This double-stranded target genomic region is converted to a single-stranded form, and one or more cross-linking oligos are designed to capture either or both strands.

[0053] To capture the lower chain from Figure 6, the crosslinked oligo is designed to have, toward the 3' end, the sequence of the second recognition site excluding the first three nucleotides, and optionally, a sequence located on the upper chain beyond and adjacent to the 3' end of the second recognition site. This crosslinked oligo has, toward the 5' end, a sequence complementary to the first recognition site excluding the first three nucleotides, and optionally, a sequence complementary to the sequence located on the lower chain beyond and adjacent to the 3' end of the first recognition site.

[0054] To capture the upper chain from Figure 6, the crosslinked oligo is designed to have, toward the 3' end, the sequence of the first recognition site excluding the first three nucleotides, and optionally, a sequence located on the lower chain beyond the 3' end of the first recognition site and adjacent to it. This crosslinked oligo has, toward the 5' end, a sequence complementary to the second recognition site excluding the first three nucleotides, and optionally, a sequence complementary to the sequence located on the upper chain beyond the 3' end of the second recognition site and adjacent to it.

[0055] Figure 7 shows an example of designing recognition sites and crosslinking oligos based on two gRNAs that bind to the opposite strand of a double-stranded genome. As shown in Figure 7, the upper strand of the double-stranded genomic DNA is selected to design the first recognition site, and the lower strand of the double-stranded genomic DNA is selected to design the second recognition site. The first gRNA binds to the first recognition site on the upper strand, and the Cas9 endonuclease cleaves three nucleotides of the genomic DNA downstream of the 5' end of the first recognition site. Similarly, the second gRNA binds to the second recognition site on the lower strand, and the Cas9 endonuclease cleaves three nucleotides of the genomic DNA downstream of the 5' end of the second recognition site. In this way, the target genomic region is cleaved from the genomic DNA in a double-stranded form. This target genomic region has the first three nucleotides of the first recognition site at the end facing the first recognition site, and the first three nucleotides of the second recognition site at the end facing the second recognition site. This double-stranded target genome region can be converted to a single-stranded form, and one or more cross-linking oligos can be designed to capture one or both strands.

[0056] To capture the lower chain as shown in Figure 7, the crosslinked oligo is designed to have a sequence complementary to the first three nucleotides of the second recognition site toward the 3' end, and additional nucleotides complementary to the sequence located on the lower chain beyond and adjacent to the 5' end of the second recognition site. This crosslinked oligo has, toward the 5' end, the sequence of the first three nucleotides of the first recognition site, and further nucleotides located on the upper chain beyond and adjacent to the 5' end of the first recognition site.

[0057] To capture the upper chain as shown in Figure 7, the crosslinked oligo is designed to have a sequence complementary to the first three nucleotides of the first recognition site toward the 3' end, and an additional nucleotide complementary to the sequence located on the upper chain beyond the 5' end of the first recognition site and adjacent to it. This crosslinked oligo has a sequence of the first three nucleotides of the second recognition site toward the 5' end, and an additional nucleotide located on the lower chain beyond the 5' end of the second recognition site and adjacent to it.

[0058] Figure 8 shows another example of designing recognition sites and crosslinking oligos based on two gRNAs that bind to the same strand of a double-stranded genome. As shown in Figure 8, the upper strand of double-stranded genomic DNA is selected to design the first and second recognition sites. The first gRNA binds to the first recognition site on the upper strand, and the Cas9 endonuclease cleaves three nucleotides of the genomic DNA downstream of the 5' end of the first recognition site. Similarly, the second gRNA binds to the second recognition site on the upper strand, and the Cas9 endonuclease cleaves three nucleotides of the genomic DNA downstream of the 5' end of the second recognition site. In this way, the target genomic region is cleaved from the genomic DNA in a double-stranded form. This target genomic region has the first three nucleotides at the 5' end of the first recognition site at the end facing the first recognition site, and all but the first three nucleotides at the 5' end of the second recognition site at the end facing the second recognition site. This double-stranded target genomic region is converted to a single-stranded form, and one or more cross-linking oligos are designed to capture either or both strands.

[0059] To capture the lower chain from Figure 8, the crosslinked oligo is designed to have, toward the 3' end, the sequence of the second recognition site excluding the first three nucleotides, and optionally, a sequence located on the upper chain beyond and adjacent to the 3' end of the second recognition site. This crosslinked oligo has, toward the 5' end, the sequence of the first three nucleotides of the first recognition site, and further nucleotides located on the upper chain beyond and adjacent to the 5' end of the first recognition site.

[0060] To capture the upper chain as shown in Figure 8, the crosslinked oligo is designed to have a sequence complementary to the first three nucleotides of the first recognition site toward the 3' end, and additional nucleotides complementary to the sequence located adjacent to and beyond the 5' end of the first recognition site on the upper chain. This crosslinked oligo has a sequence complementary to the second recognition site toward the 5' end, except for the first three nucleotides, and optionally a sequence complementary to the sequence located adjacent to and beyond the 3' end of the second recognition site on the upper chain.

[0061] As shown in Figures 5-8 and described in the previous paragraph, the sequences of the cross-linked oligos do not need to be perfectly complementary or identical to the corresponding sequences on the target genomic region, as long as the cross-linked oligos can hybridize with the corresponding sequences. Therefore, some degree of mismatch is acceptable, and such modifications are within the scope of the present invention.

[0062] The term "hybridize" indicates that two sequences are sufficiently complementary to each other, enabling hybridization between them. Sequences that hybridize to each other are perfectly complementary but have some degree of mismatch. Thus, the 5' and 3' terminal sequences of a crosslinking oligo may have a slight mismatch with the corresponding 5' and 3' terminal sequences of the target genomic region, insofar as the crosslinking oligo hybridizes with and captures the target genomic region. Depending on the stringency of hybridization, a mismatch of about 5% to about 20% between two complementary sequences enables hybridization between the two sequences. Typically, high-stringency conditions have higher temperatures and lower salt concentrations, while low-stringency conditions have lower temperatures and higher salt concentrations. High-stringency conditions are preferred for hybridization, and therefore, it is preferable that the 3' and 5' terminal sequences of the crosslinking oligo are perfectly complementary to the 3' and 5' terminal sequences of the target genomic region, respectively.

[0063] The captured target genomic region is further analyzed, for example, by detection or sequencing. Such analyses are performed. d. A step of generating a single-stranded circular target genome region hybridized to a cross-linked oligo by ligating the free end of the single-stranded target genome region hybridized to the cross-linked oligo, e. Optionally, a step to degrade the non-cyclic genetic material, f. Optionally, a step of amplifying a target genome region by nucleic acid amplification to generate multiple copies of the target genome region, g. A step to analyze the amplified target genomic region. Includes.

[0064] In some embodiments, ligating the free end of a single-stranded target genome region hybridized to a crosslinked oligo to generate a single-stranded circular target genome region hybridized to the crosslinked oligo is performed using a ligase (Figure 9A). If further sequences are present in the crosslinked oligo beyond the complementary sequence at the end of the target genome region, ligating the free end of the single-stranded target genome region also involves nucleic acid synthesis to fill the gap between the two ends of the target genome region using a suitable DNA polymerase enzyme (Figure 9B). The gap between the ends of the target can also be filled by hybridizing a single-stranded oligo to the crosslinked oligo in the gap and ligating the resulting molecules together (Figure 9C). Suitable ligase enzymes for ligating the free end of a single-stranded target genome are known in the art and include T4 DNA ligase, ampligase, T7 DNA ligase, and Taq DNA ligase.

[0065] In certain embodiments, a properly circularized target sequence is converted into a double-stranded vector. This can be done by supplementing the reaction with DNA polymerases and DNA ligases that lack strand substitution ability, thereby synthesizing a second strand complementary to the circularized target (Figure 4).

[0066] In certain embodiments, once the target genomic region is properly circularized, the remaining genetic material, such as uncleaved genomic DNA and off-target genomic fragments, is removed, for example, degraded. This can be done by a combination of exonucleases that degrade nucleic acids from the exposed 3' or 5' ends, regardless of whether the DNA or RNA is single-stranded or double-stranded. Thus, treatment with appropriate exonucleases leaves only the circulating target molecule, significantly reducing the difficulty of subsequent steps.

[0067] In certain embodiments, a appropriately circularized target sequence, separated from the remainder of the genetic material via treatment with an exonuclease and hybridized to a crosslinked oligo, is amplified via nucleic acid amplification. In certain preferred embodiments, amplification can be performed via PCR using the sequence in the crosslinked oligo to design primers that can amplify the circular molecules and convert them back to linear molecules. This is particularly suitable for short stretches of DNA (<10Kb) where PCR is most efficient. Crosslinked oligos for multiple targets may include common sites that are non-complementary to the targets, used to drive the PCR reaction with a pair of common primers for all fragments (Figure 4). In other preferred embodiments, such amplification is isothermal amplification, e.g., rolling circle amplification (RCA). Isothermal amplification facilitates the amplification of long stretches of DNA, for example, which can produce multiple copies of a target genomic region, each copy containing at least about 10–50kb.

[0068] In conventional RCA reactions, a padlock synthetic probe is used to hybridize the target nucleic acid sequence and generate a circle (Non-Patent Literature 5). This circle is then amplified by an RCA reaction. In the present invention, a cyclic target sequence captured using a crosslinked oligo is amplified via an RCA reaction. In such an RCA reaction, the crosslinked oligo acts as a primer for the RCA reaction, which linearly amplifies the target sequence and generates long single-stranded tandem repeats of the target sequence. The amplified single-stranded target sequence is then used for subsequent analysis. The amplified target sequence can also be converted to a double-stranded form, for example, by capturing both target strands and re-annealing their RCA products.

[0069] In some embodiments, exponential amplification of the target cyclization molecule is obtained through a superbranched RCA using one or more additional primers (beyond the crosslinking oligo itself) to generate double-stranded tandem repeats of the target sequence. The primers beyond the crosslinking oligo are designed based on the sequence of the crosslinking oligo, based on sequences present in the target genomic region, or consist of random base combinations to amplify multiple positions of the circle.

[0070] Further approaches for converting single-stranded target genomic regions into double-stranded forms are known in the art, and such embodiments are within the scope of the present invention.

[0071] Since target genomic regions of different sizes can be generated and amplified by RCA, the limitations of conventional methods that rely on PCR to amplify nucleic acids and have not been effective in amplifying long target genomic regions are avoided. A key advantage provided by the methods disclosed herein is the detection and analysis of long target genomic regions, for example, genomic regions containing at least about 1 kb, preferably at least about 15 kb, more preferably at least about 30 kb, or most preferably at least about 50 kb.

[0072] In certain embodiments, the products of the RCA reaction are analyzed to detect or sequence a target genomic region. For example, amplification products can be detected based on reaction turbidity, fluorescence detection, or labeled molecular beacons.

[0073] The term "label" refers to a molecule that can be detected by spectroscopic, photochemical, biochemical, immunochemical, chemical, or other physical means. For example, useful labels include fluorescent dyes (fluorophores), fluorescent quenchers, luminescent agents, high electron density reagents, biotin, digoxigenin, 32This includes phosphorus (P) and other isotopes, or other molecules that become detectable, for example, by being incorporated into oligonucleotides. This term also includes combinations of labeling agents, such as combinations of phosphors, each of which provides a unique detectable signature, for example, at a specific wavelength or combination of wavelengths.

[0074] Examples of phosphors include Alexa dyes (e.g., Alexa350, Alexa430, Alexa488, etc.), AMCA, BODIPY630 / 650, BODIPY650 / 665, BODIPY-FL, BODIPY-R6G, BODIPY-TMR, BODIPY-TRX, Cascade Blue, Cy2, Cy5, Cy5.5, Cy7, Cy7.5, Dylight dyes (Dylight405, Dylight488, Dylight549, Dylight550, Dylight649, Dylight680, Dylight Examples include, but are not limited to, ght750, Dylight800), 6-FAM, fluorescein, FITC, HEX, 6-JOE, Oregon Green 488, Oregon Green 500, Oregon Green 514, Pacific Blue, REG, Rhodamine Green, Rhodamine Red, ROX, R-phycoerythrin (R-PE), Starbright Blue dyes (e.g., Starbright Blue 520, Starbright Blue 700), TAMRA, TET, tetramethylrhodamine, Texas Red, and TRITC.

[0075] A particular embodiment of the present invention detects five or more different targets in parallel. For example, different crosslinked oligos can be immobilized on a substrate, such as a chip, where the coordinates of each crosslinked oligo, i.e., the coordinates of each target genomic region, are known. Using such a substrate, tens, hundreds, or even thousands of target sequences can be detected and analyzed in multiple reactions.

[0076] In further embodiments, the products of PCR or RCA reactions are sequenced. Various sequencing methods can be used to sequence the PCR and RCA products, including using a portable Nanopore Minion® or benchtop machine, Nanopore Promethion®, PacBio Sequel®, or Illumina HiSeq®. The sequencing step can also be used for multiple detection and / or polymorphism detection of several targets.

[0077] In certain embodiments, short, appropriately circularized target sequences are prepared for sequencing after PCR amplification (Figure 4). PCR primers are either specific to the target region or, preferably, universal to all targets by designing them to amplify a common region introduced by a cross-linking oligo (Figures 4B-4C). Once amplified, adapter molecules are added to the molecule to fit it with the sequencer using them, according to the recommendations of the specific manufacturer. Alternatively, the adapters required during PCR can be incorporated using primers with a “tail” that is not complementary to the cross-linking oligo, thereby making the method more efficient (Figures 4A and 4C). These primers also contain a short, unique sequence (4-16 nucleotides) added during PCR to link the sequencing data to each sample, which is commonly known as an index or barcode (Figure 4C).

[0078] In certain embodiments, ligated copies can be cleaved to generate individual copies of the target genomic region, for example, by using a rare cutter restriction enzyme to cut the ligated copies at restriction sites introduced via cross-linking oligos. Such cleavage generates multiple copies of the target genomic region, each having sticky ends.

[0079] Sticky ends can be used to ligate target genomic regions to adapter sequences. For example, an adapter containing overhangs complementary to restriction sites specific to rare cutter restriction enzymes can be mixed with a copy of the target genomic region to generate double-stranded DNA containing the target genomic region adjacent to the adapter.

[0080] As used herein, the term “adapter” refers to a known nucleotide sequence of 4 to 100 nucleotides, preferably 10 to 20 nucleotides, and more preferably about 15 nucleotides, depending on the sequencing technique used. Once incorporated into the ends of an amplified copy of a target genomic region, the adapter sequence facilitates sequencing of the target genomic region, for example, by providing a binding site for a primer. In one embodiment, the target genomic region adjacent to the adapter is sequenced using end-to-end sequencing.

[0081] As used herein, the term “paired-end sequencing” refers to a sequencing technique in which both ends of a fragment are sequenced using specific primer binding sites located at each end of a double-stranded polynucleotide. Paired-end sequencing generates high-quality sequencing data that is aligned using a computer software program to generate the sequences of polynucleotides adjacent to two primer binding sites. Sequencing from only one end of a molecule degrades the quality of sequencing when long sequencing reads are performed; sequencing from both ends of a double-stranded molecule enables high-quality data from both ends of the double-stranded molecule.

[0082] In pair-end sequencing, the double-stranded polynucleotides generated at the adapter incorporation end are sequenced using specific primers that bind to the two ends of the double-stranded target genomic region adjacent to the adapter. An overview and principle of pair-end sequencing are described in Illumina Sequencing Technology, Illumina, Publication No. 770-2007-002, which is incorporated herein by reference in its entirety.

[0083] Examples of end-to-end sequencing technologies include, but are not limited to, Illumina MiSeq®, Illumina MiSeqDx®, and Illumina MiSeqFGx®. Further examples of end-to-end sequencing technologies used in the assays of the present invention are known in the art, and such embodiments are within the scope of the present invention.

[0084] In certain embodiments, the sticky ends of a cleaved copy of the target genomic region can be used to ligate the target genomic region to a hairpin adapter. For example, a hairpin adapter containing an overhang complementary to a restriction site specific to a rare cutter restriction enzyme can be mixed with a copy of the target genomic region to generate double-stranded DNA containing the target genomic region adjacent to the hairpin adapter.

[0085] As used herein, the term “hairpin adapter” refers to a polynucleotide comprising a double-stranded stem and a single-stranded hairpin loop. The single-stranded hairpin loop region of the hairpin adapter provides a primer binding site for sequencing. Thus, when the hairpin adapter hybridizes with both sticky ends of a target genome sequence, a double-stranded DNA template is generated containing the target genome region in a double-stranded region capped at both ends by hairpin loops. Such a template is used to sequence the target genome region via single-molecule real-time (SMRT) sequencing (PacBio®).

[0086] The description and principles of SMRT sequencing are described in Pacific Biosciences (2018), Publication No.: BR108-100318, and the entire content is incorporated herein by reference.

[0087] In further embodiments, nanopore technology is used to sequence the target genomic region. In certain such embodiments, for example, a copy of the target genomic region is processed to sequence it, as described in the Nanopore Technology Brochure, Oxford Nanopore Technologies (2019) and the Nanopore Product Brochure, Oxford Nanopore Technologies (2018). The entire contents of both of these brochures are incorporated herein by reference.

[0088] In certain embodiments, multiple target genomic regions are captured and optionally further analyzed, such as by detection or sequencing. In such embodiments, a pair of gRNAs is designed for each target genomic region. To design multiple gRNAs for multiple target genomic regions, the gRNA sequences are selected so that the Cas9 endonuclease for one target genomic region does not disrupt other target genomic regions. Based on the known genomic sequence of the organism, a person skilled in the art can determine an appropriate combination of recognition sites for specifically cleaving and isolating multiple target genomic regions. gRNAs can also be designed and synthesized using degenerate or wobble bases to enable hybridization to multiple or uncertain locations. For example, this can be done to add multiple degenerate bases to randomly hybridize to multiple locations in the genome to compensate for known polymorphisms at gRNA target sites, or to design gRNAs that hybridize to unknown genes encoding known amino acid sequences based on the degeneracy of the genetic code. In such cases, if necessary, crosslinking oligos can also be adapted to contain degenerate bases.

[0089] Therefore, certain embodiments of the present invention provide a method for capturing multiple target genomic regions from genetic material. This method is a. A step of cleaving multiple target genomic regions from genetic material using multiple pairs of endonucleases, wherein each pair of endonucleases has a first recognition site and a second recognition site, each recognition site contains a sequence of at least approximately 17 to approximately 24 nucleotides, and each pair of recognition sites is adjacent to the target genomic regions from the multiple target genomic regions. b. A process of denaturing the cleavage genetic material into a single-stranded form, c. A step of capturing multiple single-stranded target genome regions by hybridizing the target genome region to multiple cross-linking oligos, wherein each cross-linking oligo includes 3' and 5' terminal sequences that hybridize from the multiple single-stranded target genome regions to the 3' and 5' terminals of the target genome region, respectively. Includes.

[0090] The embodiments described above, such as designing a recognition site, designing an endonuclease used to cleave genetic material, and designing a cross-linking oligo for capturing a cleaved target genomic region, are also applicable to the present method for capturing multiple target genomic regions.

[0091] In one embodiment for capturing multiple target genomic regions, a substrate is provided having multiple cross-linked oligos bound to the substrate at specific known locations. The genetic material is cleaved with multiple pairs of endonucleases and converted into a single-stranded form, and the multiple single-stranded target genomic regions are brought into contact with the substrate under appropriate conditions to enable hybridization between the substrate-bound cross-linked oligos and the multiple target genomic regions. Once the target genomic regions have hybridized with the corresponding cross-linked oligos and are circularized, the target genomic regions located at specific locations are further analyzed. In some embodiments, the captured target genomic regions are amplified using RCA and further detected, for example, using laser excitation and emission detection methods.

[0092] In certain embodiments, multiple target genomic regions are further analyzed, for example, by detection or sequencing. Such analyses are performed. d. A step of ligating the free ends of single-stranded target genome regions hybridized to multiple cross-linked oligos to generate multiple single-stranded circular genetic materials, each containing a target genome region from multiple target genome regions hybridized to the corresponding cross-linked oligos, e. Optionally, a step to remove acyclic genetic material from the substrate, f. Optionally, a step of amplifying multiple target genomic regions by nucleic acid amplification to generate multiple copies of each target genomic region, g. A process of analyzing multiple amplified target genomic regions. Includes.

[0093] The above embodiments, which involve analyzing a target genomic region, such as ligating the free end of a single-stranded target genomic region, detecting a target genomic region, or sequencing a target genomic region, are also applicable to the present method for analyzing multiple target genomic regions.

[0094] In certain embodiments, the non-cyclic genetic material can be removed by degrading it with an exonuclease or by washing it off the substrate.

[0095] Further embodiments of the present invention provide a kit for performing the assay of the present invention. The kit of the present invention includes a specific endonuclease necessary for performing the assay of the present invention, a specific guide molecule designed to target one or more target genomic regions, a computer software program designed to process sequencing data obtained from the assay, and optionally, instructions for performing the assay. In one embodiment, the kit of the present invention is: a) One or more guide molecules, b) One or more cross-linking oligos designed to capture one or more target genomic regions Includes.

[0096] The kit may further include one or more programmable endonucleases.

[0097] Furthermore, the kit may include primers for PCR or RCA reactions of circularized target genomic regions, DNA ligases, polymerases and other reagents for PCR or RCA, restriction endonucleases, such as rare cutters for cleaving linked copies of target genomic regions, and sequencing reagents.

[0098] In certain embodiments, the kit of the present invention can be customized for one or more specific target genomic regions. For example, if a user provides sequences of one or more target genomic regions, a kit is generated to perform the assay of the present invention for one or more target sequences.

[0099] All patents, patent applications, provisional applications, and publications referenced or cited herein, including figures and tables, are incorporated herein by reference in their entirety, provided they do not conflict with the express teachings herein.

[0100] The following are examples illustrating the procedure for carrying out the present invention. These examples are not limiting. Unless otherwise specified, all percentages are by weight and all solvent mixing ratios are by volume.

[0101] Example 1 - Design of recognition site and capture of target gene region This example describes a typical procedure for capturing and analyzing a target genomic region according to the materials and methods disclosed herein. For example, the following sequences are identified for analysis in this example. 5'ACCCACTGTTGAGAGTCAGTGGCAAGAGAAGTCTCGTCTTTTACGTCCCGTATCAAAATGAGTGTACAATACAT[A / G]TCTATGTGCGAGTGAGA|GCA CGGGTCTGGGCCTGGAGAGTGGACCACCTACAATTGGCCATTTCGGTTTGCGGAAGCTGTCAAGTCAACGCGAGTCCTAG[G / A]ATCTCATAGTCTTCGCATTAACCCGTATTAAGTGGACTCGCCTACAGTTTGTCTTATGCTAGCAACCCAGGCA[T / C]AGTCTGTACGCGCGATCGTCCACTGG|TAG TGG GGTCCTTCCTGGAGTGT[C / G]GTATGGGCCATCCGGTCTACCTTACTAACTTAGGCTTTAAGCGCATTTCTATGTGCGTGAGGTGTCGCATTCATACTTAGTCTGGTCCTAAGTCTGTCACC 3' (SEQ ID NO: 1).

[0102] The sequence of Sequence ID No. 1 provides a non-complementary strand sequence; that is, the gRNA binds to the opposite strand. Therefore, the recognition site is located on the opposite strand, and the regions corresponding to the recognition site are "TCTATGTGCGAGTGAGA|GCA" from positions 76 to 95 of Sequence ID No. 1 and "GCGCGATCGTCCACTGG|TAG" from positions 260 to 279 of Sequence ID No. 1, with the NGG PAM motif underlined and the Cas9 endonuclease cleavage site indicated by a vertical bar. The first gRNA is designed to bind to a first recognition site having a sequence inversely complementary to the sequence of TCTATGTGCGAGTGAGAGCA (Sequence ID No. 2), and the second gRNA is designed to bind to a second recognition site having a sequence inversely complementary to the sequence of GCGCGATCGTCCACTGGTAG (Sequence ID No. 3). The two recognition sites are adjacent to genomic sequences having specific single nucleotide polymorphisms (SNPs) indicated by nucleotides in solid parentheses. This sequence is cleaved via Cas9, resulting in a 184 bp sequence. Sequencing this fragment provides information about the SNPs present in this region.

[0103] Cas9-mediated cleavage of genetic material containing the sequence of Sequence ID No. 1 generates the following fragment, which is the target genomic region. 5' GCACGGGTCTGGGCCTGGAGAGTGGACCACCTACAATTGGCCATTTCGGTTTGCGGAAGCTGTCAAGTCAACGCGAGTCCTAG[G / A]ATCTCATAGTCTTCGCATTAACCCGTATTAAGTGGACTCGCCTACAGTTTGTCTTATGCTAGCAACCCAG GCA[T / C]AGTCTGTACGCGCGATCGTCCACTGG 3' (Sequence ID 4).

[0104] To generate the cross-linking oligo for capturing the target genomic region of sequence number 4, select the following sequence from the end of the target genomic region, which is underlined in the reproduced sequence above. 5'GCACGGGTCTGGGCCTGGAGAGTGGACCAC(Sequence 5) 5'GCATAGTCTGTACGCGCGATCGTCCACTGG (Sequence ID 6).

[0105] Sequences complementary to these sequences are generated at the ends of the cross-linking oligo to capture the target genomic region of sequence number 4. 5'GTGGTCCACTCTCCAGGCCCAGACCCGTGC3' (Sequence ID 7) 5'CCAGTGGACGATCGCGCGTACAGACTATGC3'(Sequence No. 8)

[0106] Sequences 7 and 8 are concatenated to generate a cross-linked oligo with the following sequence. 5'GTGGTCCACTCTCCAGGCCCAGACCCGTGCCCAGTGGACGATCGCGCGTACAGACTATGC3' (Sequence No. 9).

[0107] The examples and embodiments described herein are illustrative, and various modifications and changes taking them into account are suggested to those skilled in the art, and these are also included in the spirit and scope of this application and the scope of the appended claims. Furthermore, any element or limitation of the invention or embodiment disclosed herein may be combined with any and / or all other elements or limitations disclosed herein (separately or in any combination) or any other invention or embodiment thereof, and all such combinations are considered to be within the scope of the present invention without limitation. This is intended to be the case.

[0108] All patents, patent applications, and other published reference materials cited herein are incorporated herein by reference in their entirety. [Sequence Listing Free Text]

[0109] A brief explanation of array keys Sequence ID 1: Exemplary areas to be studied according to the method of the present invention. Sequence ID 2: Recognition site of the first example. Sequence ID 3: Exemplary second recognition site. Sequence ID 4: Exemplary target genome region. Sequence ID 5: The sequence of the 5' end of the target genomic region of Sequence ID 4. Sequence ID 6: Sequence of the 3' end of the target genome region of Sequence ID 4. Sequence ID 7 is the 5' end sequence of a cross-linking oligo designed to capture the target genomic region of Sequence ID 4. Sequence ID 8 is the 3' terminal sequence of a cross-linking oligo designed to capture the target genomic region of Sequence ID 4. Sequence ID 9 is a cross-linking oligo designed to capture the target genomic region of Sequence ID 4.

Claims

1. A method for capturing a target genomic region from genetic material, a) A step of cleaving the target genomic region from the genetic material using one or more endonucleases having a sequence of 17 (-10%) to 24 (+10%) nucleotides and a first recognition site and a second recognition site adjacent to the target genomic region, b) A step of denaturing the cleaved genetic material into a single-stranded form, c) A step of capturing the single-stranded target genome region by hybridizing the target genome region to a cross-linked oligo containing 3' and 5' terminal sequences that hybridize to the 3' and 5' ends of the single-stranded target genome region, respectively. Methods that include...

2. The method according to claim 1, wherein the first recognition site and the second recognition site each comprise a sequence of 17 (-10%) to 24 (+10%) nucleotides, and the first recognition site and the second recognition site are adjacent to the target genome region.

3. The method according to claim 1, wherein the one or more endonucleases that cleave the genetic material at a specific recognition site are restriction endonucleases, meganucleases, or programmable endonucleases.

4. The method according to claim 3, wherein the one or more programmable endonucleases are selected from Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-related protein 9 endonuclease (Cas9 endonuclease), Cpf1, C2c1, C2c2, C2c, RNA or DNA-induced Argonaut proteins, and structure-induced endonucleases.

5. The method according to claim 4, wherein the programmable endonuclease comprises a first programmable nuclease that cleaves DNA at the first recognition site based on a first guide molecule having a sequence complementary to the first recognition site, and a second programmable endonuclease that cleaves DNA at the second recognition site based on a second guide molecule having a sequence complementary to the second recognition site.

6. The method according to claim 1, wherein the cross-linked oligo is a single-stranded oligonucleotide having 3' and 5' terminal sequences that hybridize to the 3' and 5' terminal sequences of the target genome region in the single-stranded form, respectively.

7. The method according to claim 1, wherein the cross-linked oligo is a double-stranded oligonucleotide having 3' and 5' terminal overhangs that hybridize to the 3' and 5' terminal sequences of the single-stranded target genome region, respectively.

8. The method according to claim 6, wherein the cross-linked oligo further comprises a restriction site specific to a rare cutter-restricted endonuclease, a cleavage site for a programmable endonuclease, and / or a primer-binding sequence.

9. The method according to claim 6, wherein the crosslinked oligo is immobilized on a solid substrate or biotinylated.

10. d) Ligating the free end of the single-stranded target genome region hybridized to the cross-linked oligo to generate a single-stranded circular target genome region hybridized to the cross-linked oligo, e) Decompose the non-cyclic genetic material, f) Amplify the target genome region by nucleic acid amplification to generate multiple copies of the target genome region, g) Analyzing the amplified target genomic region. The method according to claim 1, further comprising analyzing the target genomic region including the following.

11. The method according to claim 10, comprising amplifying the target genomic region by a rolling circle amplification (RCA) reaction to generate a plurality of linked copies of the target genomic region.

12. The method according to claim 10, wherein the analysis comprises sequencing the target genomic region.

13. The method according to claim 12, wherein the sequencing method includes nanopore sequencing, reversible dye terminator sequencing, or single-molecule real-time (SMRT) sequencing.

14. A method for capturing multiple target genomic regions from genetic material, a) Multiple pairs of endonucleases are used to cleave multiple target genomic regions from the genetic material, each pair of endonucleases having a first recognition site and a second recognition site, each recognition site containing a sequence of 17 (-10%) to 24 (+10%) nucleotides, and each pair of recognition sites is adjacent to the target genomic regions from the multiple target genomic regions. b) The cleaved genetic material is modified into a single-stranded form, c) A method for capturing multiple target genomic regions in single-stranded form by hybridizing them to multiple cross-linking oligos, wherein each cross-linking oligo contains 3' and 5' terminal sequences that hybridize from the multiple single-stranded target genomic regions to the 3' and 5' terminals of the target genomic regions, respectively.

15. d) Ligating the free ends of the single-stranded target genome regions hybridized to the plurality of cross-linked oligos to generate a plurality of single-stranded circular genetic materials, each containing a target genome region from the plurality of target genome regions hybridized to the corresponding cross-linked oligos, e) Remove acyclic genetic material, f) The plurality of target genomic regions are amplified by nucleic acid amplification to generate multiple copies of each target genomic region, g) Analyze the amplified target genome region. The method according to claim 14, further comprising analyzing the plurality of target genomic regions, including the above.

16. The method according to claim 14, wherein the plurality of crosslinked oligos are immobilized on a solid substrate.

17. The method according to claim 16, wherein the plurality of target genomic regions are amplified by nucleic acid amplification to generate a plurality of copies of each target genomic region.

18. The method according to claim 14, comprising amplifying the target genomic region by polymerase chain reaction (PCR) to generate a plurality of copies of the target genomic region.

19. The method according to claim 18, wherein the PCR primer can be specific to the target region or can be made common to all targets by being designed to bind to a common region introduced by the cross-linked oligo.

Citation Information

Patent Citations

  • Uses of cocoa butter in cooking

    JP2008515448A

  • Single molecule arrays for genetic and chemical analysis

    JP2011234723A

  • Methods for identifying and counting nucleic acid sequence, expression, copy, or methylation changes in dna using a combination of nucleases, ligases, polymerases, and sequencing reactions

    JP2017516487A