Selective capture of target DNA sequences
By employing capture and blocking nucleic acid molecules, the method selectively targets desired genomic regions, enhancing sequencing accuracy and specificity by preventing the capture of non-specific sequences, thus addressing the inefficiencies of current hybridization capture baits.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- PERSONAL GENOME DIAGNOSTICS INC
- Filing Date
- 2021-05-10
- Publication Date
- 2026-04-13
AI Technical Summary
Current hybridization capture baits fail to selectively capture unique genomic regions, leading to the sequencing of thousands of non-specific regions, resulting in wasted sequence capacity and potential misalignment issues, particularly in intronic regions.
The use of capture nucleic acid molecules and blocking nucleic acid molecules, where the blocking nucleic acid molecules are designed to preferentially bind to non-target regions, preventing their capture, while the capture nucleic acid molecules specifically target the desired sequence.
This approach enhances the sequencing accuracy and specificity of highly repetitive and related genomic regions by minimizing the capture of unwanted sequences, thereby improving the coverage and reducing false positives.
Smart Images

Figure 0007844349000001 
Figure 0007844349000002 
Figure 0007844349000003
Abstract
Description
Technical Field
[0001] Background of the Invention Cross - Reference to Related Applications This application claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. Application No. 63 / 022,968, filed May 11, 2020, the entire content of which is incorporated herein by reference. Incorporation of Sequence Listing
[0002] The content of the accompanying sequence listing is incorporated herein by reference. The accompanying sequence listing text file is named PGDX3140 - 1WO_SL.txt, was created on May 5, 2021, and is 8 kb. This file can be accessed using Microsoft Word® on a computer using the Windows® operating system. Field of the Invention
[0003] The present invention generally relates to the sequencing of highly repetitive regions of genomes, and more specifically, to the use of capture nucleic acid molecules and blocking nucleic acid molecules for sequencing regions of genomes having repetitive DNA sequences.
Background Art
[0004] Background Information Current hybridization capture baits work effectively with unique sequences, enabling the capture and deep sequencing of genomic regions thought to be involved in cancer and other diseases. When the captured sequence is not unique and is very similar to other sequences, the capture of thousands of regions from across the genome is unavoidable and, undesirably, sequenced anyway. Currently, this results in many sequence reads that are not properly mapped or aligned and must be discarded. This means that sequence capacity that could be used elsewhere is wasted. Even worse, some of these sequences may be misaligned, potentially causing false positives in the region of interest. This is particularly problematic in intronic regions, which are often needed to identify translocations. [Overview of the project] [Means for solving the problem]
[0005] This invention is based on the original discovery that the accuracy of sequencing highly repetitive and / or related regions of the genome can be increased by using capture nucleic acid molecules and blocking nucleic acid molecules.
[0006] In one embodiment, the present invention provides a method for sequencing a target sequence, comprising the steps of: hybridizing a sample nucleic acid with a capture nucleic acid molecule and a blocking nucleic acid molecule, wherein the sample nucleic acid has a target sequence and a non-target sequence; isolating the capture nucleic acid molecule hybridized to the target sequence; and sequencing the isolated target sequence.
[0007] In one embodiment, the non-target sequence has repeating regions and / or associated regions of nucleic acid. In another embodiment, the blocking nucleic acid molecule and the capture nucleic acid molecule have at least about 60 to at least about 120 nucleic acids. In a further embodiment, the capture nucleic acid is labeled with a detectable label, including but not limited to a radioactive phosphate, biotin, fluorophores, enzymes, or a combination thereof. In a further embodiment, the blocking nucleic acid molecule is present in about 10 times excess of the capture nucleic acid molecule. In one embodiment, the blocking nucleic acid molecule has at least about 70% sequence identity with respect to the capture nucleic acid molecule. In another embodiment, the blocking nucleic acid molecule has at least about 4 nucleic acid molecules that are different from the capture nucleic acid molecule. In a further embodiment, the target sequence has at least about 60 to at least about 120 nucleic acids. In a further embodiment, sequencing is performed by next-generation sequencing.
[0008] In another embodiment, the present invention provides a method for improving the sequencing specificity and / or accuracy of a target sequence by the steps of: hybridizing a sample nucleic acid with a capture nucleic acid molecule and a blocking nucleic acid molecule, wherein the sample nucleic acid has a target sequence and a non-target sequence; isolating the capture nucleic acid molecule hybridized with the target sequence; and sequencing the target sequence.
[0009] In one embodiment, the non-target sequence has repeating and / or related regions of nucleic acid. In another embodiment, the blocking nucleic acid molecule and the capture nucleic acid molecule have at least about 60 to at least about 120 nucleic acids. In a further embodiment, the capture nucleic acid is labeled with a detectable label, including but not limited to a radioactive phosphate, biotin, fluorophores, enzymes, or a combination thereof. In a further embodiment, the blocking nucleic acid molecule is present in about 10 times excess of the capture nucleic acid molecule. In one embodiment, the blocking nucleic acid molecule has at least about 70% sequence identity with respect to the capture nucleic acid molecule. In another embodiment, the blocking nucleic acid molecule has at least about 4 nucleic acid molecules that are different from the capture nucleic acid molecule. In a further embodiment, the target sequence has at least about 60 to at least about 120 nucleic acids. In a further embodiment, sequencing is performed by next-generation sequencing. [Modes for carrying out the invention]
[0010] Detailed description of the invention This invention is based on the original discovery that the accuracy of sequencing highly repetitive and / or related regions of the genome can be increased by using capture nucleic acid molecules and blocking nucleic acid molecules.
[0011] Before describing the compositions and methods of the present invention, it should be understood that the present invention is not limited to the specific compositions, methods, and experimental conditions described, as such compositions, methods, and conditions may vary. Furthermore, since the scope of the present invention is limited only to the appended claims, it should be understood that the terms used herein are for describing specific embodiments and are not intended to limit them.
[0012] As used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly indicates otherwise. Thus, for example, a reference to “method” includes one or more methods and / or steps of the type described herein, which would be obvious to those skilled in the art by reading this disclosure, etc.
[0013] All publications, patents, and patent applications referenced herein are incorporated by reference to the same extent as any individual publication, patent, or patent application is specifically and individually indicated as being incorporated by reference.
[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art in which this invention pertains. Any methods and materials similar or equivalent to those described herein may be used in carrying out or testing the present invention, but modifications and variations will be understood to be included in the spirit and scope of this disclosure. Preferred methods and materials are described below.
[0015] This invention provides a method for adding a blocking bait, designed to preferentially bind to unwanted regions and prevent their capture, to hybridization capture. Without interfering homologous repeat sequences, the desired region can then be selectively captured for sequencing. In addition to a moderate sequence difference between the targeting bait and the blocking bait, the blocking bait contains neither biotin nor other tags necessary for capture. The blocking bait interferes with the capture of unwanted sequences but has minimal impact on the capture of the desired region.
[0016] Many regions of genomic DNA are very similar to other regions of the genome, and therefore it is very difficult to capture them without also capturing similar, undesirable regions. This results in extra sequencing of unintended regions, reducing coverage of the desired region. To minimize the capture of unintended regions, blocking baits are designed to prevent the capture of similar but undesirable fragments. This allows for more targeted sequencing of the desired region. Blocking baits differ from capture baits in that they have a moderately different sequence that preferentially binds to undesirable DNA, and do not contain biotin or other modifications, and therefore remain after the capture bait is selected.
[0017] In one embodiment, the present invention provides a method for sequencing a target sequence, comprising the steps of: hybridizing a sample nucleic acid with a capture nucleic acid molecule and a blocking nucleic acid molecule, wherein the sample nucleic acid has a target sequence and a non-target sequence; isolating the capture nucleic acid molecule hybridized with the target sequence; and sequencing the isolated target sequence. In one embodiment, the non-target sequence has a repeating region and / or related region of the nucleic acid. In another embodiment, the blocking nucleic acid molecule and the capture nucleic acid molecule have at least about 60 to at least about 120 nucleic acids. In a further embodiment, the capture nucleic acid is labeled with a label, including but not limited to a radioactive phosphate, biotin, fluorophores, enzymes, or a combination thereof. In a further embodiment, the blocking nucleic acid molecule is present in about 10 times excess of the capture nucleic acid molecule. In one embodiment, the blocking nucleic acid molecule has at least about 70% sequence identity with respect to the capture nucleic acid molecule. In another embodiment, the blocking nucleic acid molecule has at least about 4 nucleic acid molecules distinct from the capture nucleic acid molecule. In a further embodiment, the target sequence has at least about 60 to at least about 120 nucleic acids. In a further embodiment, sequencing is performed by next-generation sequencing.
[0018] As used herein, the terms “nucleic acid” or “nucleic acid sequence” refer to oligonucleotides, nucleotides, polynucleotides, or fragments thereof, or to genomically or synthetically derived DNA or RNA, which may be single-stranded or double-stranded, or represent a sense strand or antisense strand, or to peptide nucleic acids (PNA), or to any DNA-like or RNA-like material of natural or synthetic origin. Examples of the terms “nucleic acid” or “nucleic acid sequence” include oligonucleotides, nucleotides, polynucleotides, or fragments thereof, which may be single-stranded or double-stranded, or represent a sense strand or antisense strand, or to genomically or synthetically derived DNA or RNA (e.g., mRNA, rRNA, tRNA, iRNA), or to any DNA-like or RNA-like material of natural or synthetic origin (e.g., iRNA, ribonucleoprotein (e.g., double-stranded iRNA, e.g., iRNP)). The term encompasses nucleic acids, i.e., oligonucleotides, including known analogues of natural nucleotides. This term also encompasses nucleic acid-like structures with a synthetic skeleton; see, for example, Mata (1997) Toxicol. Appl. Pharmacol. 144:189-197; Strauss-Soukup (1997) Biochemistry 36:8692-8698; Samstag (1996) Antisense Nucleic Acid Drug Dev 6:153-156.
[0019] As used herein, the terms “target region” or “target sequence” refer to the nucleic acid sequence that is the target of sequencing. The target sequence may originate from any source, such as a DNA library, genomic DNA, or other DNA source.
[0020] In one embodiment, the target sequence is at least about 60–120 nucleic acid lengths. In a particular embodiment, the target sequence is at least about 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, or 120 nucleic acid lengths.
[0021] As used herein, the term “non-target sequence” refers to any nucleic acid sequence that is not a target sequence. Non-target sequences may have repetitive sequences and / or related sequences. In repetitive DNA, stretches of DNA repeats occur in the genome either in tandem along the genome or scattered throughout. These sequences do not encode proteins. One class called highly repetitive DNA consists of short sequences, e.g., 5 to 100 nucleotides, that are repeated thousands of times in a single stretch, and includes satellite DNA. Related sequences are sequences that have high sequence identity with the target sequence but are not identical.
[0022] As used herein, "sample nucleic acid" refers to nucleic acid containing both target and non-target sequences.
[0023] The term "hybridization" means the process by which nucleic acid strands bind to complementary strands through base pairing. The hybridization reaction can be sensitive and selective so that a particular sequence of interest can be identified even in a sample where it is present at low concentration. Suitable stringent conditions can be defined, for example, by the concentration of salt or formamide in the prehybridization solution and the hybridization solution, or by the hybridization temperature, and are well known in the art. In particular, the stringency can be increased by lowering the salt concentration, increasing the formamide concentration, or raising the hybridization temperature. In alternative embodiments, the nucleic acids of the invention are defined by their ability to hybridize under various stringency conditions (e.g., high, medium, and low) as defined herein.
[0024] For example, hybridization under high stringency conditions can occur in about 50% formamide at about 37°C to 42°C. Hybridization can occur at about 30°C to 35°C in about 35% to 25% formamide under reduced stringency conditions. In particular, hybridization can occur under high stringency conditions at 42°C in 50% formamide, 5×SSPE, 0.3% SDS and 200 mg / ml of sheared and denatured salmon sperm DNA. Hybridization can occur, as described above, but under reduced stringency conditions at a reduced temperature of 35°C in 35% formamide. The temperature range corresponding to a particular level of stringency can be further narrowed by calculating the ratio of purines to pyrimidines of the nucleic acid of interest and adjusting the temperature accordingly. Modifications of the above ranges and conditions are well known in the art.
[0025] The terms "capture nucleic acid molecule" or "capture bait" are used interchangeably herein and refer to a nucleic acid molecule designed to hybridize to a target sequence. The capture nucleic acid molecule is designed to hybridize specifically to the target sequence and isolate it.
[0026] In one embodiment, the captured nucleic acid molecule has at least about 60 to 120 nucleic acid lengths. In a particular embodiment, the captured nucleic acid molecule has at least about 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, or 120 nucleic acid lengths.
[0027] As used herein, the terms “sequence identity” or “sequence homology” may be used interchangeably and refer to the exact nucleotide-to-nucleotide correspondence between two polynucleotide sequences. Typically, techniques for determining sequence identity involve determining the nucleotide sequence of a polynucleotide and / or the amino acid sequence encoded thereby, and comparing these sequences to a second nucleotide or amino acid sequence. Two or more sequences can be compared by determining their “identity percentage,” also referred to as the “homology percentage.” The identity percentage relative to a reference sequence, which may be a longer intramolecular sequence, may be calculated by dividing the number of perfect matches between two optimally aligned sequences by the length of the reference sequence and multiplying by 100. The identity percentage may also be determined by comparing sequence information using an advanced BLAST computer program, including version 2.2.9, available from the National Institutes of Health, for example. The BLAST program is based on the alignment method of Karlin and Altschul, Proc. Natl. Acad. Sci. USA 87:2264-2268 (1990) and is discussed in Altschul, et al., J. Mol. Biol. 215:403-410 (1990); Karlin and Altschul, Proc. Natl. Acad. Sci. USA 90:5873-5877 (1993); and Altschul et al., Nucleic Acids Res. 25:3389-3402 (1997). Simply put, the BLAST program defines identity as the number of identically aligned symbols (i.e., nucleotides or amino acids) divided by the total number of symbols in the shorter of the two sequences. This program can be used to determine the identity percentage over the entire length of the sequences being compared. Default parameters are provided to optimize the search for short query sequences in programs such as blastp.This program also allows the use of a SEG filter that masks off segments of the query sequence determined by the SEG program in Wootton and Federhen, Computers and Chemistry 17: 149-163 (1993). The desired degree of sequence identity ranges from approximately 80% to 100% and integer values in between. The percentage of identity between the disclosed sequence and the sequence of the present invention may be at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, or at least 99.9%. Generally, a perfect match indicates 100% identity over the length of the reference sequence. In some cases, the reference to the percentage of sequence identity refers to sequence identity measured using BLAST (Basic Local Alignment Search Tool). In other cases, ClustalW can be used for multiple sequence alignment. Other programs for comparing sequences and / or evaluating sequence identity include the Needleman-Wunsch algorithm and the Smith-Waterman algorithm (see, for example, the EMBOSS Water aligner). Optimal alignment can be evaluated using any appropriate parameters of the selected algorithm, including default parameters.
[0028] As used herein, the terms “sequence identity percentage (%)” or “identity percentage (%)” are defined as the percentage of nucleotides in a candidate sequence that are identical to the nucleotides of a reference sequence, after aligning the sequences to reach the maximum sequence identity percentage and introducing gaps as necessary, without considering any conservative substitutions as part of the sequence identity, and including “homology.” Optimal alignment of sequences for comparison can be performed manually, or by local homology algorithms as described in Smith and Waterman, 1981, Ads App. Math. 2, 482; Neddleman and Wunsch, 1970, J. Mol. Biol. 48, 443; similarity search methods as described in Pearson and Lipman, 1988, Proc. Natl. Acad. Sci. USA 85, 2444; or by computer programs using these algorithms (GAP, BESTFIT, FASTA, BLAST P, BLAST N, and TFASTA from the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Drive, Madison, Wis.).
[0029] In an additional embodiment, the captured nucleic acid molecule has at least 70% sequence identity with respect to the complement of the target sequence. In a specific embodiment, the captured nucleic acid has at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher, or complete (100%) sequence identity with respect to the complement of the target sequence.
[0030] In a further embodiment, the captured nucleic acid molecules are labeled or tagged. Examples of nucleic acid labels include radioactive phosphates, biotin, fluorophores, and enzymes. In certain embodiments, labeling is used to isolate captured nucleic acid molecules that hybridize to a target sequence.
[0031] Methods for labeling nucleic acids are well known in the field. Examples of nucleic acid labeling include γ- 32 5' end labeling of DNA with P rATP; α- 32 Labeling by PCR with P dNTP, Biotin-dNTP, and Fl-dNTP; α- 32 DNA 3' labeling with P dNTPs, Biotin-dNTPs, and Fl-dNTPs with single nucleotide labeling at the Fl terminator nucleotide; α- 32 Random priming by PCR with P dNTP, Biotin-dNTP, and Fl-dNTP; and α- 32 Nick translation by PCR using P dNTPs, Biotin-dNTPs, and Fl-dNTPs is one example.
[0032] Nucleic acid molecules can be isolated using labeling or tagging. For example, nucleic acid molecules labeled or tagged with biotin can be isolated using streptavidin and / or avidin. Biotin binds to streptavidin and avidin with extremely high affinity, rapid on-rate, and high specificity, and these interactions are used to isolate the biotinylated molecule of interest.
[0033] The terms "blocking nucleic acid molecule" or "blocking bait" are used interchangeably and refer to nucleic acid molecules designed to be similar to, but not identical to, the captured nucleic acid molecule. Blocking nucleic acid molecules are designed to hybridize to non-target sequences.
[0034] In one embodiment, the blocking nucleic acid molecule is at least about 60 to 120 nucleic acid lengths. In a particular embodiment, the captured nucleic acid molecule is at least about 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, or 120 nucleic acid lengths.
[0035] In an additional embodiment, the blocking nucleic acid molecule has at least about 70% sequence identity with respect to the captured nucleic acid molecule. In a specific embodiment, the captured nucleic acid has at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher, or complete (100%) sequence identity with respect to the captured nucleic acid molecule.
[0036] In a further embodiment, the blocking nucleic acid has at least about four nucleic acids distinct from the captured nucleic acid molecule. In a particular embodiment, the blocking nucleic acid molecule has at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleic acids distinct from the captured nucleic acid molecule.
[0037] In one embodiment, blocking nucleic acid molecules are present in an excess of at least about 10 times the amount of captured nucleic acid molecules. In a particular embodiment, blocking nucleic acid molecules are present in an excess of at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 times the amount of captured nucleic acid molecules.
[0038] In one embodiment, the isolated target sequence is sequenced. Sequencing can be carried out by any method known in the art. Exemplary sequencing methods include, for example, next-generation sequencing (NGS). Exemplary NGS methodologies include the Roche 454 sequencer, Life Technologies SOLiD system, Life Technologies Ion Torrent, BGI / MGI system, Genapsys system, and Illumina systems such as the Illumina Genome Analyzer II, Illumina MiSeq, Illumina HiSeq, Illumina NextSeq, and Illumina NovaSeq instruments. Sequence determination is, for example, at least 2x coverage, at least 10x coverage, at least 20x coverage, at least 30x coverage, at least 40x coverage, at least 50x coverage, at least 60x coverage, at least 70x coverage, at least 80x coverage, at least 90x coverage, at least 100x coverage, at least 200x coverage, at least 300x coverage, at least 400x coverage, at least 500x coverage, at least 600x coverage, at least 700x coverage, at least This can be done with deep coverage for each nucleotide, including 800x coverage, at least 900x coverage, at least 1000x coverage, at least 2000x coverage, at least 3000x coverage, at least 4000x coverage, at least 5000x coverage, at least 6000x coverage, at least 7000x coverage, at least 8000x coverage, at least 9000x coverage, at least 10000x coverage, at least 15000x coverage, at least 20000x coverage, and any number or range in between.
[0039] In another embodiment, the present invention provides a method for improving the sequencing specificity and / or accuracy of a target sequence by the steps of: hybridizing a sample nucleic acid with a capture nucleic acid molecule and a blocking nucleic acid molecule, wherein the sample nucleic acid has a target sequence and a non-target sequence; separating the capture nucleic acid molecule hybridized with the target sequence; and sequencing the target sequence. In one embodiment, the non-target sequence has repeating regions and / or related regions of the nucleic acid. In another embodiment, the blocking nucleic acid molecule and the capture nucleic acid molecule have at least about 60 to at least about 120 nucleic acids. In a further embodiment, the capture nucleic acid is labeled with a radioactive phosphate, biotin, fluorophores, enzymes, or a combination thereof. In a further embodiment, the blocking nucleic acid molecule is present in about 10 times excess of the capture nucleic acid molecule. In one embodiment, the blocking nucleic acid molecule has at least about 70% sequence identity with respect to the capture nucleic acid molecule. In another embodiment, the blocking nucleic acid molecule has at least about 4 nucleic acid molecules distinct from the capture nucleic acid molecule. In a further embodiment, the target sequence has at least about 60 to at least about 120 nucleic acids. In a further embodiment, sequencing is performed by next-generation sequencing.
[0040] The following embodiments are provided to further illustrate embodiments of the present invention, but are not intended to limit the scope of the invention. These are typical examples of what may be used, but other procedures, methodologies, or techniques known to those skilled in the art may be used instead. [Examples]
[0041] Example 1 Identification of blocking bait in ROS intron 31 Currently, generating high coverage of repeating regions is difficult due to interference from related sequences. This method does not improve the selective capture of identical regions longer than the bait length (often 60-120 nt), but it benefits many short, highly related sequences that cause problems due to their sheer number rather than their identity.
[0042] To identify potential blocking bait sequences, we examined bait within a capture panel containing some of the most highly repeated regions. Primers for position chr6:117654887~117655006 (ROS intron 31) were analyzed using BLAT, and hundreds of highly similar sequences were observed across the genome. The 18 closest matches were aligned to each other using ClustalW. Thirteen of the desired ROS1 intron sequences differed from the consensus of 19 sequences. These mismatches were modified to create a consensus sequence optimized for binding to other genomic regions, reducing the likelihood of binding to ROS1. When the resulting sequences were analyzed against the genome using BLAT, no perfect matches were found. All genomic sequences had at least four mismatches. Furthermore, the ROS1 sequence did not appear in the list of the top 200 sequences. Therefore, this sequence (Sequence ID 1 GAACCAAAGACAAAAACCACATGATTATCTCAATAGATGCAGAAAAGGCCTTTGATAAAATTCAACATCCCTTCATGTTAAAAACTCTCAATAAACTAGTTATTGATGGAACATATCTCA) can function as a blocking bait more than 10 times in excess to prevent unwanted sequences from being captured.
[0043] The sequence used to generate the consensus sequence chr6:117654887~117655006
[0044] Table 1 [Table 1] Example 2 Identification of blocking bait in ROS intron 31
[0045] A second bait from ROS1 intron 31 (chr6:117654108~117654227) was examined using a similar method. Twelve modifications were made to a 120nt bait, and again, a consensus sequence that did not perfectly match the genome was obtained. The top 200 positions that matched the consensus sequence did not contain any ROS1 sequences. Therefore, this consensus sequence (SEQ ID NO: 2 AGAGCAAACAAATTCAAAAGCTAGCAGAAGACAAGAAATAACTAAGATCAGAGCAGAATTGAAGGAGATAGAGACACAAAAAACCCTCCAAAAAAAATCAACGAATCCAGGAGCTGTTTT) can also be used as a blocking bait.
[0046] The sequence used to generate the consensus sequence chr6:117654108~117654227
[0047] Table 2 [Table 2]
[0048] Applying the same method to the entire region of a targeted ROI where the relevant homologous sequence is problematic can result in better coverage and fewer false positives for those regions, thus increasing the value of the sequencing assay. Generally, these untagged baits can be added at high molar concentrations, and their molar concentrations can be adjusted individually or as a group to optimize performance. Individual optimization can be performed empirically or using bait sequence characteristics such as the number of mismatches, expected Tm, GC content, homolog frequency in the genome, and / or other means. Example 3 Identification of blocking bait in ROS intron 31
[0049] The selected baits in the repeating region of ROS1 intron 31 are shown below. Proposed sequences for modified blocking baits are listed below each actual bait, modified with a lowercase letter or dash. These can be used to test the improvements obtained by using blocking baits.
[0050] Table 3 [Table 3]
[0051] Although the present invention has been described with reference to the above embodiments, it will be understood that modifications and variations are included within the spirit and scope of the invention. Accordingly, the present invention is limited only by the following claims. The present invention provides, for example, the following items: (Item 1) A method for determining the sequence of a target sequence, a) A step of hybridizing a sample nucleic acid with a capture nucleic acid molecule and a blocking nucleic acid molecule, wherein the sample nucleic acid includes a target sequence and a non-target sequence; b) the step of isolating the captured nucleic acid molecule hybridized to the target sequence; and c) The process of sequencing the isolated target sequence. Methods that include... (Item 2) The method according to item 1, wherein the non-target sequence includes repeating regions and / or related regions of nucleic acid. (Item 3) The method according to item 1, wherein the blocking nucleic acid molecule and the captured nucleic acid molecule contain at least about 60 to at least 120 nucleic acids. (Item 4) The method according to item 1, wherein the captured nucleic acid is labeled. (Item 5) The method according to item 4, wherein the captured nucleic acid is labeled with a label selected from the group consisting of a radioactive phosphate, biotin, fluorophores, enzymes, or combinations thereof. (Item 6) The method according to item 1, wherein the blocking nucleic acid molecule is present in a 10-fold excess compared to the captured nucleic acid molecule. (Item 7) The method according to item 1, wherein the blocking nucleic acid molecule has at least about 70% sequence identity with respect to the captured nucleic acid molecule. (Item 8) The method according to item 1, wherein the blocking nucleic acid molecule has at least about four nucleic acid molecules that are different from the captured nucleic acid molecule. (Item 9) The method according to item 1, wherein the target sequence comprises at least about 60 to at least about 120 nucleic acids. (Item 10) The method according to item 1, wherein the sequence determination step includes next-generation sequence determination. (Item 11) A method for improving the sequencing specificity and / or accuracy of a target sequence, a) A step of hybridizing a sample nucleic acid with a capture nucleic acid molecule and a blocking nucleic acid molecule, wherein the sample nucleic acid includes a target sequence and a non-target sequence; b) the step of isolating the captured nucleic acid molecule hybridized to the target sequence; and c) A step of sequencing the target sequence. Methods that include... (Item 12) The method according to item 11, wherein the non-target sequence includes a repeating region and / or related region of nucleic acid. (Item 13) The method according to item 11, wherein the blocking nucleic acid molecule and the captured nucleic acid molecule contain at least about 60 to at least 120 nucleic acids. (Item 14) The method according to item 11, wherein the captured nucleic acid is labeled. (Item 15) The captured nucleic acid is a radioactive phosphate, biotin, fluorophore, enzyme or similar The method described in item 14, which is marked with a marker selected from a group consisting of combinations. (Item 16) The method according to item 11, wherein the blocking nucleic acid molecule is present in a 10-fold excess compared to the captured nucleic acid molecule. (Item 17) The method according to item 11, wherein the blocking nucleic acid molecule has at least about 70% sequence identity with respect to the captured nucleic acid molecule. (Item 18) The method according to item 11, wherein the blocking nucleic acid molecule has at least about four nucleic acid molecules different from the captured nucleic acid molecule. (Item 19) The method according to item 11, wherein the target sequence comprises at least about 60 to at least 120 nucleic acids. (Item 20) The method according to item 11, wherein the sequence determination step includes next-generation sequence determination.
Claims
1. A method for determining the sequence of a target sequence, a) A step of hybridizing a sample nucleic acid with a capture nucleic acid molecule and a blocking nucleic acid molecule, wherein the sample nucleic acid comprises the target sequence and the non-target sequence, the blocking nucleic acid molecule is identified using a BLAST-like alignment tool (BLAT), the blocking nucleic acid molecule has at least 70% sequence identity with respect to the capture nucleic acid molecule, and has at least four bases different from the capture nucleic acid molecule; b) the step of isolating the captured nucleic acid molecule hybridized to the target sequence; and c) The step of sequencing the isolated target sequence. Methods that include...
2. The method according to claim 1, wherein the non-target sequence includes a repetitive sequence and / or related sequence of nucleic acid.
3. The method according to claim 1, wherein the blocking nucleic acid molecule and the captured nucleic acid molecule contain 60 to 120 nucleic acids.
4. The method according to claim 1, wherein the captured nucleic acid is labeled.
5. The method according to claim 4, wherein the captured nucleic acid is labeled with a label selected from the group consisting of a radioactive phosphate, biotin, a fluorophore, an enzyme, or a combination thereof.
6. The method according to claim 1, wherein the blocking nucleic acid molecule is present in an excess of 10 times the amount of the captured nucleic acid molecule.
7. The method according to claim 1, wherein the target sequence comprises 60 to 120 nucleic acids.
8. The method according to claim 1, wherein the step of determining the sequence includes determining the next generation sequence.
9. A method for improving the sequencing specificity and / or accuracy of a target sequence, a) A step of hybridizing a sample nucleic acid with a capture nucleic acid molecule and a blocking nucleic acid molecule, wherein the sample nucleic acid comprises the target sequence and the non-target sequence, the blocking nucleic acid molecule is identified using a BLAST-like alignment tool (BLAT), the blocking nucleic acid molecule has at least 70% sequence identity with respect to the capture nucleic acid molecule, and has at least four bases different from the capture nucleic acid molecule; b) the step of isolating the captured nucleic acid molecule hybridized to the target sequence; and c) Step of sequencing the target sequence. Methods that include...
10. The method according to claim 9, wherein the non-target sequence includes a nucleic acid repetitive sequence and / or related sequence.
11. The method according to claim 9, wherein the blocking nucleic acid molecule and the captured nucleic acid molecule contain 60 to 120 nucleic acids.
12. The method according to claim 9, wherein the captured nucleic acid is labeled.
13. The method according to claim 12, wherein the captured nucleic acid is labeled with a label selected from the group consisting of a radioactive phosphate, biotin, fluorophores, enzymes, or combinations thereof.
14. The method according to claim 9, wherein the blocking nucleic acid molecule is present in an excess of 10 times the amount of the captured nucleic acid molecule.
15. The method according to claim 9, wherein the target sequence comprises 60 to 120 nucleic acids.
16. The method according to claim 9, wherein the sequence determination step includes next-generation sequence determination.
Citation Information
Patent Citations
Nucleic acid sequencing
JP2007530026A
Reduction of off-target capture in sequencing techniques
JP2018529371A
Multiplex capture of nucleic acids
US20110171644A1