Method for Targeted Genome Analysis

The method of creating a tagged DNA library with random nucleic acid tags and multifunctional adapters, combined with capture probes, addresses the challenge of unreliable genetic diagnostic tests, enabling precise detection of gene sequences and copy numbers for personalized medicine.

JP7710016B2Active Publication Date: 2025-07-17RESOLUTION BIOSCIENCE INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023184894
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2013-03-15
Filing Date
2023-10-27
Publication Date
2025-07-17
Estimated Expiration
2033-12-10

AI Technical Summary

Technical Problem

Current genetic diagnostic tests are not robust enough to reliably determine the genetic status of relevant genes, hindering the realization of personalized medicine by integrating patient-specific genetic information with treatment options.

Method used

A method involving the creation of a tagged DNA library through fragmentation, end-repair, and ligation with random nucleic acid tags and multifunctional adapter modules, followed by hybridization with capture probes and enzymatic processing to isolate and amplify specific genomic regions for targeted genetic analysis.

Benefits of technology

Enables sensitive and specific detection of target gene sequences and determination of gene copy numbers, facilitating personalized medicine by providing accurate genetic information for treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710016000101
    Figure 0007710016000101
  • Figure 0007710016000102
    Figure 0007710016000102
  • Figure 0007710016000103
    Figure 0007710016000103
Patent Text Reader

Abstract

To provide methods for targeted genomic analysis.SOLUTION: The invention provides a method for genetic analysis in an individual to reveal both the genetic sequence and chromosomal copy number of a targeted specific genomic locus in a single assay. The invention further provides a method for detecting a target gene sequence and a gene expression profile with high sensitivity and specificity. In particular, the method of the invention comprises the steps of: (a) treating a fragmented genomic DNA with an end-repairing enzyme to prepare fragmented and end-repaired genomic DNA; and (b) ligating a random nucleic acid tag sequence and optionally a sample code sequence and / or a PCR primer sequence to the fragmented and end-repaired genomic DNA to prepare a tagged genomic library.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Citation of Related Applications This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Application No. 61 / 794,049, filed on Mar. 15, 2013, and U.S. Provisional Application No. 61 / 735,417, filed on Dec. 10, 2012. These applications are hereby incorporated by reference in their entirety.

[0002] Description of the Sequence Listing The sequence listing associated with this application is provided in text format instead of a paper copy, and is hereby incorporated by reference into this specification. The name of the text file containing the sequence listing is CLFK_001_02WO_ST25.txt. This text file is 188 KB, was created on Dec. 10, 2013, and was electronically submitted via EFS-Web.

[0003] Background Technical Field The present invention generally relates to methods for genetic analysis in an individual that reveal both the genetic sequence and chromosomal copy number of a targeted specific genomic locus in a single assay. In particular, the present invention relates to methods that provide for sensitive and specific detection of a target gene sequence or gene transcript, and methods for revealing both variant sequences and overall gene copy number in a single assay.

Background Art

[0004] Description of Related Applications Both the complete human genome sequences of individual human subjects and partial genomic resequencing studies have revealed the underlying theme that every human is thought to carry a genome that is not complete. In particular, it has been found that normal healthy human subjects have hundreds of genetic lesions, if not thousands, within their genome sequences. Many of these lesions are known or predicted to abrogate the function of the genes in which they reside. Normal diploid humans carry two functional copies of most genes on average, but it is implied that in many cases there is only one (or zero) functional gene copy present in any given human. Similarly, cases where genes are present in excess due to gene duplication / amplification events are also encountered with significant frequency.

[0005] One important feature in biological networks is functional redundancy. Normal healthy individuals can tolerate the average burden of genetic lesions because they carry two copies of all genes on average, such that the loss of one copy has a low impact. Furthermore, gene sets often perform similar functions such that minor perturbations in specific gene functions are generally compensated for within the larger network of functional elements. Functional compensation in biological systems is a general theme, but there are many cases where specific gene losses can trigger acute disruptive events. For example, cancer is thought to be the result of a genetic disease where the combined effects of multiple individual lesions result in uncontrolled cell growth. Similarly, prescription drugs are often specific chemicals that are transported, metabolized, and / or eliminated by very specific genes. Perturbations in such genes generally have low impact under normal circumstances but can manifest as adverse events (e.g., side effects) in chemotherapy.

[0006] The central goal of "personalized medicine," which is increasingly being referred to as "precision medicine," is to integrate patient-specific genetic information with treatment options that are compatible with the individual's genetic profile. However, the vast potential of personalized medicine has not yet been realized. To realize this goal, there is a need for a clinically acceptable robust genetic diagnostic test that can reliably determine the genetic status of relevant genes. Summary of the Invention Means for Solving the Problems

[0007] Brief Summary Certain embodiments contemplated herein are methods for creating a tagged DNA library, the method comprising treating fragmented DNA with a terminal repair enzyme to create fragmented and end-repaired DNA, and ligating a random nucleic acid tag sequence and optionally a sample code sequence and / or a PCR primer sequence to the fragmented and end-repaired DNA to create the tagged DNA library.

[0008] In certain embodiments, the random nucleic acid tag sequence is from about 2 to about 100 nucleotides. In some embodiments, the invention provides that the random nucleic acid tag sequence is from about 2 to about 8 nucleotides.

[0009] In certain embodiments, the fragmented and end-repaired DNA contains blunt ends. In some embodiments, the blunt ends are further modified to contain a single base pair overhang.

[0010] In certain embodiments, the ligation step comprises ligating a multifunctional adapter module to the fragmented and end-repaired DNA to create the tagged DNA library, the multifunctional adapter molecule comprising i) a first region comprising a random nucleic acid tag sequence, and ii) a second region containing a sample code array, and iii) a third region containing a PCR primer array and.

[0011] In a further embodiment, the method further comprises hybridizing the tagged DNA library with at least one multifunctional capture probe module to form a complex, wherein the multifunctional capture probe module hybridizes to a specific target region in the DNA library.

[0012] In a further embodiment, the method further comprises isolating the tagged DNA library - multifunctional capture probe module complex.

[0013] In some embodiments, the method further comprises 3'-5' exonuclease enzyme processing of the isolated tagged DNA library - multifunctional capture probe module complex to remove the single-stranded 3' end. In some embodiments, the enzyme used in the 3'-5' exonuclease enzyme processing is T4 polymerase.

[0014] In certain embodiments, the method further comprises 5'-3' DNA polymerase extension of the isolated tagged DNA library - multifunctional capture probe module complex from the 3' end of the multifunctional capture probe, using the isolated tagged DNA library fragment as a template.

[0015] In certain embodiments, the method further comprises ligating the multifunctional capture probe and the isolated tagged DNA library fragment by the concerted action of 5' FLAP endonuclease, DNA polymerization and nick closure by DNA ligase.

[0016] In a further embodiment, the method further comprises performing PCR on the complex processed by the 3'-5' exonuclease enzyme, such that for generating a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is copied, and the hybrid nucleic acid molecule comprises the genomic target region capable of hybridizing with the multifunctional capture probe module and the complement of the multifunctional capture probe module tail sequence.

[0017] In various embodiments, methods for targeted genetic analysis are provided, the methods comprising a) hybridizing a tagged DNA library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes with a specific DNA target region in the DNA library; b) isolating the tagged DNA library - multifunctional capture probe module complex obtained from a); c) performing 3'-5' exonuclease enzyme processing on the isolated tagged DNA library - multifunctional capture probe module complex obtained from b) using an enzyme having 3'-5' exonuclease activity to remove the single-stranded 3' end; d) performing PCR on the complex processed by the enzyme obtained from c), such that for generating a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is copied, and the hybrid nucleic acid molecule comprises the target region capable of hybridizing with the multifunctional capture probe module and the complement of the multifunctional capture probe module tail sequence; e) performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from d) and comprising.

[0018] In various specific embodiments, a) hybridizing a tagged DNA library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes with a specific target region in the DNA library; b) isolating the tagged genomic library-multifunctional capture probe module complex obtained from a); c) using the isolated tagged DNA library fragment as a template to perform 5'-3' DNA polymerase extension of the multifunctional capture probe; d) performing PCR on the complex processed by the enzyme obtained from c), wherein a complement of the isolated target region is copied to generate a hybrid nucleic acid molecule, and the hybrid nucleic acid molecule comprises a complement of the DNA target region, a target-specific region of the multifunctional capture probe, and a multifunctional capture probe module tail sequence; and e) performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from d). A method for targeted genetic analysis is provided that includes these steps.

[0019] In certain embodiments, a method for targeted genetic analysis comprises: a) hybridizing a tagged DNA library to a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes to a specific target region in the DNA library; b) isolating the tagged DNA library-multifunctional capture probe module complex obtained from a); c) performing the generation of tagged DNA target molecules isolated by the hybrid multifunctional capture probe by the concerted action of 5' FLAP endonuclease, DNA polymerization and DNA ligation for nick closure; d) performing PCR on the complex processed by the enzyme obtained from c), wherein, to generate a hybrid nucleic acid molecule, the multifunctional capture probe molecule is ligated to the isolated tagged DNA target clone, and the hybrid nucleic acid molecule comprises a genomic target region capable of hybridizing to the multifunctional capture probe module and a complement of the multifunctional capture probe module; and e) performing a targeted genetic analysis on the hybrid nucleic acid molecule obtained from d).

[0020] In a particular embodiment, a method for determining the copy number of a specific target region, comprising: a) hybridizing a tagged DNA library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes with the specific target region in the DNA library; b) isolating the tagged DNA library-multifunctional capture probe module complex obtained from a); c) performing 3'-5' exonuclease enzymatic processing on the isolated tagged DNA library-multifunctional capture probe module complex obtained from b) using an enzyme having 3'-5' exonuclease activity to remove the single-stranded 3' end; d) performing a PCR reaction on the complex processed by the enzyme obtained from c), wherein, to generate a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is replicated, and the hybrid nucleic acid molecule comprises a target region capable of hybridizing with the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; e) performing PCR amplification of the hybrid nucleic acid in d); and f) quantifying the PCR reaction in d), wherein the quantification enables determination of the copy number of the specific target region.

[0021] In a particular embodiment, a method for determining the copy number of a specific target region, comprising: a) hybridizing a tagged DNA library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes with the specific target region in the DNA library; b) isolating the tagged DNA library - multifunctional capture probe module complex obtained from a); c) using the isolated tagged DNA library fragment as a template to perform 5'-3' DNA polymerase extension of the multifunctional capture probe; d) performing a PCR reaction on the complex processed by the enzyme obtained from c), wherein, to generate a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is replicated, and the hybrid nucleic acid molecule comprises a target region capable of hybridizing with the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; e) performing PCR amplification of the hybrid nucleic acid in d); and e) quantifying the PCR reaction in d), wherein the quantification enables determination of the copy number of the specific target region.

[0022] In a further embodiment, a method for determining the copy number of a specific target region, comprising: a) hybridizing a tagged DNA library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes with the specific target region in the DNA library; b) isolating the tagged DNA library - multifunctional capture probe module complex obtained from a); c) performing the production of a tagged DNA target molecule isolated by the hybrid multifunctional capture probe by the concerted action of 5' FLAP endonuclease, DNA polymerization, and nick closure by DNA ligase; d) performing a PCR reaction on the complex processed by the enzyme obtained from c), wherein, to produce a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is replicated, and the hybrid nucleic acid molecule comprises a target region capable of hybridizing with the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; e) performing PCR amplification of the hybrid nucleic acid in d); and e) quantifying the PCR reaction in d), wherein the quantification enables determination of the copy number of the specific target region.

[0023] In additional embodiments, a method for targeted genetic analysis is provided that includes: a) hybridizing a tagged DNA library to a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes to a specific target region in the DNA library; b) isolating the tagged DNA library-multifunctional capture probe hybrid module complex obtained from a); c) performing PCR on the complex obtained from b) to generate a hybrid nucleic acid molecule that replicates a region 3' relative to the sequence of the multifunctional capture probe, the hybrid nucleic acid molecule comprising the multifunctional capture probe hybrid module and a complement of a region of the tagged DNA library sequence located 3' relative to the multifunctional capture probe; and d) performing a targeted genetic analysis on the hybrid nucleic acid molecule obtained from c).

[0024] In certain embodiments, a method for targeted genetic analysis is provided that includes: a) hybridizing a tagged DNA library to a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes to a specific target region in the genomic library; b) isolating the tagged DNA library-multifunctional capture probe hybrid module complex obtained from a); c) using the isolated tagged DNA library fragment as a template to perform 5'-3' DNA polymerase extension of the multifunctional capture probe, the hybrid nucleic acid molecule comprising the multifunctional capture probe hybrid module and a complement of a region of the tagged DNA library sequence located 3' relative to the multifunctional capture probe; and d) performing a targeted genetic analysis on the hybrid nucleic acid molecule obtained from c).

[0025] In certain embodiments, a method for targeted genetic analysis is provided that includes: a) hybridizing a tagged DNA library to a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes to a specific target region in the DNA library; b) isolating the tagged DNA library - multifunctional capture probe hybrid module complex obtained from a); c) performing the generation of a tagged DNA target molecule isolated by the hybrid multifunctional capture probe by the concerted action of 5' FLAP endonuclease, DNA polymerization, and DNA ligation, wherein the hybrid nucleic acid molecule includes a complement of the multifunctional capture probe hybrid module and a region of the tagged DNA library sequence located 5' relative to the multifunctional capture probe; and d) performing a targeted genetic analysis on the hybrid nucleic acid molecule obtained from c).

[0026] In certain embodiments, a method for determining the copy number of a specific target region, comprising: a) hybridizing a tagged DNA library with a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes with the specific target region in the DNA library; b) isolating the tagged DNA library-multifunctional capture probe hybrid module complex obtained from a); c) performing PCR on the complex obtained from b) to produce a hybrid nucleic acid molecule, replicating a region that is 3' compared to the sequence of the multifunctional capture probe, wherein the hybrid nucleic acid molecule comprises the multifunctional capture probe hybrid module and a complement of a region of the tagged DNA library sequence located 3' compared to the multifunctional capture probe; d) performing PCR amplification of the hybrid nucleic acid in c); and e) quantifying the PCR reaction in d), wherein the quantification enables determination of the copy number of the specific target region.

[0027] In various embodiments, the targeted genetic analysis is sequence analysis.

[0028] In certain embodiments, the tagged DNA library is amplified by PCR to produce an amplified tagged DNA library.

[0029] In certain embodiments, the DNA is derived from a biological sample selected from the group consisting of blood, skin, hair, hair follicles, saliva, oral mucosa, vaginal mucosa, sweat, tears, epithelial tissue, urine, semen, seminal fluid, seminal plasma, prostatic fluid, bulbourethral fluid (Cowper's gland fluid), excreta, biopsy, ascites, cerebrospinal fluid, lymph fluid, and tissue extract samples or biopsy samples.

[0030] In a further embodiment, the tagged DNA library comprises tagged DNA sequences, and each tagged DNA sequence comprises: i) fragmented and end-repaired DNA; ii) a random nucleotide tag sequence; iii) a sample code sequence; and iv) a PCR primer sequence.

[0031] In additional embodiments, the hybrid-tagged DNA library comprises hybrid-tagged DNA sequences for use in targeted genetic analysis, and each hybrid-tagged DNA sequence comprises: i) fragmented and end-repaired DNA; ii) a random nucleotide tag sequence; iii) a sample code sequence; iv) a PCR primer sequence; and v) a multifunctional capture probe module tail sequence.

[0032] In a further embodiment, the multifunctional adapter module comprises: i) a first region comprising a random nucleotide tag sequence; ii) a second region comprising a sample code sequence; and iii) a third region comprising a PCR primer sequence.

[0033] In certain embodiments, the multifunctional capture probe module comprises: i) a first region capable of hybridizing to a partner oligonucleotide; ii) a second region capable of hybridizing to a specific target region; and iii) a third region comprising a tail sequence. In some embodiments, the first region of the capture probe module is bound to a partner oligonucleotide. In some embodiments, the partner oligonucleotide is chemically modified.

[0034] In one embodiment, the composition comprises a tagged DNA library, a multifunctional adapter module, and a multifunctional capture probe module.

[0035] In certain embodiments, the composition comprises a hybrid-tagged genomic library according to the method of the invention.

[0036] In certain embodiments, the composition comprises a reaction mixture for performing the methods contemplated herein.

[0037] In certain embodiments, a reaction mixture capable of generating a tagged DNA library comprises a) fragmented DNA and b) a DNA end repair enzyme for generating fragmented and end-repaired DNA.

[0038] In certain embodiments, the reaction mixture further comprises a multifunctional adapter module.

[0039] In additional embodiments, the reaction mixture further comprises a multifunctional capture probe module.

[0040] In some embodiments, the reaction mixture further comprises an enzyme having 3'-5' exonuclease activity and PCR amplification activity.

[0041] In one embodiment, the reaction mixture comprises a FLAP endonuclease, a DNA polymerase, and a DNA ligase.

[0042] In any of the foregoing embodiments, the DNA can be isolated genomic DNA or cDNA.

[0043] Provided are methods for generating a tagged genomic library, the methods comprising treating fragmented genomic DNA with an end repair enzyme to generate fragmented and end-repaired genomic DNA, and ligating a random nucleic acid tag sequence and optionally a sample code sequence and / or a PCR primer sequence to the fragmented and end-repaired genomic DNA to generate a tagged genomic library.

[0044] In certain embodiments, the random nucleic acid tag sequence is from about 2 to about 100 nucleotides.

[0045] In certain embodiments, the random nucleic acid tag sequence is from about 2 to about 8 nucleotides.

[0046] In additional embodiments, the fragmented and end-repaired genomic DNA contains blunt ends.

[0047] In further embodiments, the blunt ends are further modified to contain single base pair overhangs.

[0048] In some embodiments, the ligation step comprises ligating a multifunctional adapter module to the fragmented and end-repaired genomic DNA to create a tagged genomic library, and the multifunctional adapter molecule comprises a first region containing a random nucleic acid tag sequence, a second region containing a sample code sequence, and a third region containing a PCR primer sequence.

[0049] In certain embodiments, the methods contemplated herein include hybridizing the tagged genomic library to a multifunctional capture probe module to form a complex, and the multifunctional capture probe module hybridizes to a specific genomic target region in the genomic library.

[0050] In certain embodiments, the methods contemplated herein include isolating the tagged genomic library - multifunctional capture probe module complex.

[0051] In additional specific embodiments, the methods contemplated herein include processing the isolated tagged genomic library - multifunctional capture probe module complex with a 3'-5' exonuclease enzyme to remove the single-stranded 3' ends.

[0052] In further specific embodiments, the enzyme used in the 3'-5' exonuclease enzyme processing is T4 DNA polymerase.

[0053] In certain specific embodiments, the method contemplated herein is the step of performing PCR on a complex processed by a 3'-5' exonuclease enzyme obtained from the preceding claims, wherein, to create a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is copied, and this hybrid nucleic acid molecule comprises a genomic target region capable of hybridizing with the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence.

[0054] In various embodiments, (a) hybridizing a tagged genomic library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes to a specific genomic target region in the genomic library; (b) isolating the tagged genomic library-multifunctional capture probe module complex obtained from (a); (c) performing 3'-5' exonuclease enzyme processing on the isolated tagged genomic library-multifunctional capture probe module complex obtained from (b) using an enzyme having 3'-5' exonuclease activity to remove the single-stranded 3' end; (d) performing PCR on the complex processed by the enzyme obtained from (c), wherein, to create a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is copied, and this hybrid nucleic acid molecule comprises a genomic target region capable of hybridizing with the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; and (e) performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from (d). A method for targeted genetic analysis is provided.

[0055] In certain embodiments, steps a) to d) are repeated at least about two times, and the targeted genetic analysis in e) comprises an alignment of the hybrid nucleic acid molecule sequences obtained from the at least two d) steps.

[0056] In a further embodiment, at least two different multifunctional capture probe modules are used in at least two a) steps, and each of the at least two a) steps uses one type of multifunctional capture probe module.

[0057] In some embodiments, at least one multifunctional capture probe module hybridizes downstream of the genomic target region, and at least one multifunctional capture probe module hybridizes upstream of the genomic target region.

[0058] In various embodiments, methods are provided for determining the copy number of a specific genomic target region, the method comprising: (a) hybridizing a tagged genomic library to a multifunctional capture probe module complex, wherein the multifunctional capture probe module complex selectively hybridizes to a specific genomic target region in the genomic library; (b) isolating the tagged genomic library - multifunctional capture probe module complex obtained from a); (c) performing 3'-5' exonuclease enzyme processing on the isolated tagged genomic library - multifunctional capture probe module complex obtained from b) using an enzyme having 3'-5' exonuclease activity to remove the single-stranded 3' end; (d) performing a PCR reaction on the complex processed by the enzyme obtained from c), wherein, to generate a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is replicated, and the hybrid nucleic acid molecule comprises the genomic target region capable of hybridizing to the multifunctional capture probe module and the complement of the multifunctional capture probe module tail sequence; (e) performing PCR amplification of the hybrid nucleic acid molecule in d); (f)e) The step of quantifying the PCR reaction, wherein the quantification enables determination of the copy number of the specific genomic target region, and comprises.

[0059] In some embodiments, the methods contemplated herein include obtaining the sequence of the hybrid nucleic acid molecule obtained from step e).

[0060] In further embodiments, steps a)-e) are repeated at least about 2 times, and sequence alignment is performed using the hybrid nucleic acid molecule sequences obtained from said at least 2 e) steps.

[0061] In additional embodiments, at least two different multifunctional capture probe modules are used in said at least 2 a) steps, and each of said at least 2 a) steps uses one multifunctional capture probe module.

[0062] In certain embodiments, at least one multifunctional capture probe module hybridizes downstream of the genomic target region and at least one multifunctional capture probe module hybridizes upstream of the genomic target region.

[0063] In various embodiments, methods are provided for determining the copy number of a specific genomic target region, the method comprising (a) hybridizing a tagged genomic library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module complex selectively hybridizes with a specific genomic target region in the genomic library; and (b) isolating the tagged genomic library - multifunctional capture probe module complex obtained from a). (c) Using an enzyme having 3'-5' exonuclease activity, performing 3'-5' exonuclease enzymatic processing on the isolated tagged genomic library - multifunctional capture probe module complex obtained from b) to remove the single-stranded 3' end; (d) Performing a PCR reaction on the complex processed by the enzyme obtained from c), wherein, to create a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is replicated, and the hybrid nucleic acid molecule comprises the genomic target region capable of hybridizing with the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; (e) Performing PCR amplification of the hybrid nucleic acid molecule in d); comprising.

[0064] In certain embodiments, the methods contemplated herein include obtaining the sequence of the hybrid nucleic acid molecule obtained from step e).

[0065] In certain embodiments, steps a)-e) are repeated at least about two times, and sequence alignment is performed using the hybrid nucleic acid molecule sequences obtained from the at least two e) steps.

[0066] In some embodiments, at least two different multifunctional capture probe modules are used in the at least two a) steps, and each of the at least two a) steps uses one type of multifunctional capture probe module.

[0067] In additional embodiments, at least one type of multifunctional capture probe module hybridizes downstream of the genomic target region and at least one type of multifunctional capture probe module hybridizes upstream of the genomic target region.

[0068] In various embodiments, a method is provided for determining the copy number of a specific genomic target region, the method comprising: (a) hybridizing a tagged genomic library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module complex selectively hybridizes with a specific genomic target region in the genomic library; (b) isolating the tagged genomic library - multifunctional capture probe module complex obtained from (a); (c) performing 3'-5' exonuclease enzyme processing on the isolated tagged genomic library - multifunctional capture probe module complex obtained from (b) using an enzyme having 3'-5' exonuclease activity to remove single-stranded 3' ends; (d) performing a PCR reaction on the complex processed by the enzyme obtained from (c), wherein the tail portion of the multifunctional capture probe molecule is replicated to create a hybrid nucleic acid molecule, the hybrid nucleic acid molecule comprising the genomic target region capable of hybridizing with the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; (e) performing PCR amplification of the hybrid nucleic acid molecule in (d); (f) performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from (e). The method includes steps (a) to (f).

[0069] In certain embodiments, steps (a) to (e) are repeated at least about two times, and the targeted genetic analysis in (f) includes performing sequence alignment of the hybrid nucleic acid molecule sequences obtained from the at least two steps (e).

[0070] In certain embodiments, at least two different multifunctional capture probe modules are used in the at least two a) steps, and each of the at least two a) steps uses one type of multifunctional capture probe module.

[0071] In additional embodiments, at least one multifunctional capture probe module hybridizes downstream of the genomic target region and at least one multifunctional capture probe module hybridizes upstream of the genomic target region.

[0072] In various embodiments, methods for targeted genetic analysis are provided, the methods comprising: (a) hybridizing a tagged genomic library to a multifunctional capture probe hybridization module complex, wherein the multifunctional capture probe hybridization module selectively hybridizes to a specific genomic target region in the genomic library; (b) isolating the tagged genomic library - multifunctional capture probe hybridization module complex obtained from a); (c) performing 5' to 3' DNA polymerase extension of the multifunctional capture probe on the complex obtained from b) to generate a hybrid nucleic acid molecule, replicating a region of the captured tagged genomic target region at the 3' of the multifunctional capture probe, wherein the hybrid nucleic acid molecule comprises the multifunctional capture probe hybridization module and a complement of a region of the tagged genomic target region located in the 3' direction from the position where the multifunctional capture probe hybridization module hybridizes to the genomic target region; (d) performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from c). and

[0073] In a further embodiment, steps a) to c) are repeated at least about two times, and the targeted genetic analysis in d) comprises a sequence alignment of the hybrid nucleic acid molecule sequences obtained from the at least two steps of d).

[0074] In some embodiments, at least two different multifunctional capture probe modules are used in the at least two steps of a), and each of the at least two steps of a) uses one type of multifunctional capture probe module.

[0075] In certain embodiments, at least one multifunctional capture probe module hybridizes downstream of the genomic target region, and at least one multifunctional capture probe module hybridizes upstream of the genomic target region.

[0076] In various embodiments, a method for determining the copy number of a specific genomic target region is provided, the method comprising (a) hybridizing a tagged genomic library with a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes with a specific genomic target region in the genomic library; (b) isolating the tagged genomic library - multifunctional capture probe hybrid module complex obtained from a); (c) performing a 5' to 3' DNA polymerase extension of the multifunctional capture probe on the complex obtained from b) to generate a hybrid nucleic acid molecule, replicating a region of the captured tagged genomic target region at the 3' of the multifunctional capture probe, wherein the hybrid nucleic acid molecule comprises the multifunctional capture probe hybrid module and a complement of a region of the tagged genomic target region located in the 3' direction from the position where the multifunctional capture probe hybrid module hybridizes with the genomic target region; (d) Performing PCR amplification of the hybrid nucleic acid molecule in (c); (e) Quantifying the PCR reaction in (d), wherein the quantification enables determination of the copy number of the specific genomic target region; and comprising.

[0077] In certain embodiments, the methods contemplated herein include obtaining the sequence of the hybrid nucleic acid molecule obtained from step (d).

[0078] In certain embodiments, steps (a)-(d) are repeated at least about two times, and sequence alignment of the hybrid nucleic acid molecules obtained from the at least two steps (d) is performed.

[0079] In additional embodiments, at least two different multifunctional capture probe modules are used in the at least two steps (a), and each of the at least two steps (a) uses one type of multifunctional capture probe module.

[0080] In further embodiments, at least one multifunctional capture probe module hybridizes downstream of the genomic target region, and at least one multifunctional capture probe module hybridizes upstream of the genomic target region.

[0081] In some embodiments, the targeted genetic analysis is sequence analysis.

[0082] In certain embodiments, the tagged genomic library is amplified by PCR to generate an amplified tagged genomic library.

[0083] In certain related embodiments, the genomic DNA is derived from a biological sample selected from the group consisting of blood, skin, hair, hair follicles, saliva, oral mucosa, vaginal mucosa, sweat, tears, epithelial tissue, urine, semen, seminal fluid, seminal plasma, prostatic fluid, bulbourethral gland fluid (Cowper's gland fluid), excrement, biopsy, ascites, cerebrospinal fluid, lymph fluid, and tissue extract samples or biopsy samples.

[0084] In various embodiments, a tagged genomic library comprising tagged genomic sequences is provided, wherein each tagged genomic sequence comprises fragmented and end-repaired genomic DNA, a random nucleotide tag sequence, a sample code sequence, and a PCR primer sequence.

[0085] In various related embodiments, a tagged cDNA library comprising tagged cDNA sequences is provided, wherein each tagged cDNA sequence comprises fragmented and end-repaired cDNA, a random nucleotide tag sequence, a sample code sequence, and a PCR primer sequence.

[0086] In various specific embodiments, a hybrid-tagged genomic library comprising hybrid-tagged genomic sequences for use in targeted genetic analysis is provided, wherein each hybrid-tagged genomic sequence comprises fragmented and end-repaired genomic DNA, a random nucleotide tag sequence, a sample code sequence, a PCR primer sequence, a genomic target region, and a multifunctional capture probe module tail sequence.

[0087] In certain specific embodiments, a hybrid-tagged cDNA library comprising hybrid-tagged cDNA sequences for use in targeted genetic analysis is provided, wherein each hybrid-tagged cDNA sequence comprises fragmented and end-repaired cDNA, a random nucleotide tag sequence, a sample code sequence, a PCR primer sequence, a cDNA target region, and a multifunctional capture probe module tail sequence.

[0088] In various specific embodiments, a multifunctional adapter module is provided, which includes a first region containing a random nucleotide tag sequence, a second region containing a sample code sequence, and a third region containing a PCR primer sequence.

[0089] In various additional embodiments, a multifunctional capture probe module is provided, which includes a first region capable of hybridizing with a partner oligonucleotide, a second region capable of hybridizing with a specific genomic target region, and a third region containing a tail sequence.

[0090] In a specific embodiment, the first region is bound to the partner oligonucleotide.

[0091] In a specific embodiment, a multifunctional adapter probe hybrid module is provided, which includes a first region capable of hybridizing with a partner oligonucleotide and functioning as a PCR primer, and a second region capable of hybridizing with a specific genomic target region.

[0092] In a certain specific embodiment, the first region is bound to the partner oligonucleotide.

[0093] In some embodiments, the partner oligonucleotide is chemically modified.

[0094] In further embodiments, a composition is provided that includes a tagged genomic library, a multifunctional adapter module, and a multifunctional capture probe module.

[0095] In additional embodiments, a composition is provided that includes a hybrid-tagged genomic library or cDNA library according to any of the preceding embodiments.

[0096] In various embodiments, a reaction mixture for performing the method described in any one of the preceding embodiments is provided.

[0097] In certain embodiments, a reaction mixture capable of generating a tagged genomic library is provided, the reaction mixture comprising fragmented genomic DNA and a DNA end repair enzyme for generating fragmented and end-repaired genomic DNA.

[0098] In certain embodiments, a reaction mixture capable of generating a tagged genomic library is provided, the reaction mixture comprising fragmented cDNA and a DNA end repair enzyme for generating fragmented and end-repaired cDNA.

[0099] In certain embodiments, a reaction mixture comprising a multifunctional adapter module is provided.

[0100] In some embodiments, the reaction mixture comprises a multifunctional capture probe module.

[0101] In certain embodiments, the reaction mixture comprises an enzyme having 3'-5' exonuclease activity and PCR amplification activity.

[0102] In various embodiments, a method for DNA sequence analysis is provided, the method comprising obtaining one or more clones, each clone comprising a first DNA sequence and a second DNA sequence, wherein the first DNA sequence comprises a targeted genomic DNA sequence and the second DNA sequence comprises a capture probe sequence; performing a paired-end sequencing reaction on the one or more clones to obtain one or more sequencing reads; and ordering or clustering the sequencing reads of the one or more clones according to the probe sequence of the sequencing reads.

[0103] In certain embodiments, a method for DNA sequence analysis is provided, the method comprising obtaining one or more clones, each clone comprising a first DNA sequence and a second DNA sequence, wherein the first DNA sequence comprises a targeted genomic DNA sequence and the second DNA sequence comprises a capture probe sequence; performing a sequencing reaction on the one or more clones, wherein a single long sequencing read of greater than about 100 nucleotides is obtained and the read is sufficient to identify both the first DNA sequence and the second DNA sequence; and ordering or clustering the sequencing reads of the one or more clones according to the probe sequence of the sequencing reads.

[0104] In certain embodiments, the sequences of the one or more clones are compared to one or more human reference DNA sequences.

[0105] In additional embodiments, sequences that do not match the one or more human reference DNA sequences are identified.

[0106] In further embodiments, the non-matching sequences are used to create a de novo assembly from the non-matching sequence data.

[0107] In some embodiments, the de novo assembly is used to identify novel sequence rearrangements associated with the capture probes.

[0108] In various embodiments, a method for genomic copy number determination analysis is provided, the method comprising obtaining one or more clones, each clone comprising a first DNA sequence and a second DNA sequence, wherein the first DNA sequence comprises a random nucleotide tag sequence and a targeted genomic DNA sequence, and the second DNA sequence comprises a capture probe sequence; performing a paired-end sequencing reaction on the one or more clones to obtain one or more sequencing reads; and ordering or clustering the sequencing reads of the one or more clones according to the probe sequences of the sequencing reads.

[0109] In some embodiments, a method for genomic copy number determination analysis is provided, the method comprising obtaining one or more clones, each clone comprising a first DNA sequence and a second DNA sequence, wherein the first DNA sequence comprises a random nucleotide tag sequence and a targeted genomic DNA sequence, and the second DNA sequence comprises a capture probe sequence; performing a sequencing reaction on the one or more clones to obtain a single long sequencing read that is greater than about 100 nucleotides, the read being sufficient for identification of both the first DNA sequence and the second DNA sequence; and ordering or clustering the sequencing reads of the one or more clones according to the probe sequences of the sequencing reads.

[0110] In certain embodiments, the random nucleotide tag sequence is from about 2 to about 50 nucleotides in length.

[0111] In further embodiments, the methods contemplated herein determine the distribution of unique and redundant sequencing reads, count the number of times unique reads are encountered, and Analyzing any sequencing reads associated with a second read sequence by fitting a cloth to a statistical distribution, estimating the total number of unique reads, and normalizing the estimated total number of unique reads against the assumption that most human loci are generally diploid.

[0112] In additional embodiments, the estimated copy number of one or more targeted loci is determined.

[0113] In some embodiments, the one or more target loci that deviate from the expected copy number values are determined.

[0114] In further embodiments, the one or more targeted loci of a gene are grouped together in a collection of loci, and the copy number measurements obtained from the collection of targeted loci are averaged and normalized.

[0115] In additional embodiments, the estimated copy number of a gene is represented by the normalized average of all target loci representing this gene.

[0116] In certain embodiments, a method for creating a tagged RNA expression library is provided, the method including fragmenting a cDNA library, treating the fragmented cDNA library with a terminal repair enzyme to create fragmented and end-repaired cDNA, and ligating a multifunctional adapter molecule to the fragmented and end-repaired cDNA to create a tagged RNA expression library.

[0117] In certain embodiments, a method for creating a tagged RNA expression library is provided, the method comprising preparing a cDNA library from the total RNA of one or more cells, fragmenting the cDNA library, treating the fragmented cDNA with a terminal repair enzyme to create fragmented and terminal repaired cDNA, and ligating a multifunctional adapter molecule to the fragmented and terminal repaired cDNA to create a tagged RNA expression library.

[0118] In various embodiments, the cDNA library is an oligo dT primed cDNA library.

[0119] In certain embodiments, the cDNA library is primed with a random oligonucleotide comprising from about 6 to about 20 random nucleotides.

[0120] In certain embodiments, the cDNA library is primed with a random hexamer or a random octamer.

[0121] In additional embodiments, the cDNA library is fragmented to a size of from about 250 bp to about 750 bp.

[0122] In further embodiments, the cDNA library is fragmented to a size of about 500 bp.

[0123] In some embodiments, the multifunctional adapter module comprises a first region comprising a random nucleic acid tag sequence, optionally a second region comprising a sample code sequence, and optionally a third region comprising a PCR primer sequence.

[0124] In related embodiments, the multifunctional adapter module comprises a first region comprising a random nucleic acid tag sequence, a second region comprising a sample code sequence, and a third region comprising a PCR primer sequence.

[0125] In various embodiments, the methods contemplated herein include hybridizing a tagged cDNA library with a multifunctional capture probe module to form a complex, wherein the multifunctional capture probe module hybridizes to a specific target region in the cDNA library.

[0126] In some embodiments, the methods contemplated herein include isolating the tagged cDNA library - multifunctional capture probe module complex.

[0127] In certain embodiments, the methods contemplated herein include 3'-5' exonuclease enzyme processing of the isolated tagged cDNA library - multifunctional capture probe module complex to remove the single-stranded 3' end.

[0128] In some embodiments, the enzyme used in the 3'-5' exonuclease enzyme processing is T4 DNA polymerase.

[0129] In certain specific embodiments, the methods contemplated herein include performing PCR on the complex processed by the 3'-5' exonuclease enzyme, wherein the tail portion of the multifunctional capture probe molecule is copied to create a hybrid nucleic acid molecule, and the hybrid nucleic acid molecule includes the cDNA target region capable of hybridizing to the multifunctional capture probe module and the complement of the multifunctional capture probe module tail sequence.

[0130] In further embodiments, methods for targeted gene expression analysis are provided, the methods comprising (a) Step of hybridizing a tagged RNA expression library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes with a specific target region in the tagged RNA expression library, and (b) Step of isolating the tagged RNA expression library-multifunctional capture probe module complex obtained from (a); (c) Step of performing 3'-5' exonuclease enzyme processing on the isolated tagged RNA expression library-multifunctional capture probe module complex obtained from (b) using an enzyme having 3'-5' exonuclease activity to remove the single-stranded 3' end; (d) Step of performing PCR on the complex processed by the enzyme obtained from (c), wherein the tail portion of the multifunctional capture probe molecule is copied to produce a hybrid nucleic acid molecule, and the hybrid nucleic acid molecule includes the target region capable of hybridizing with the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; (e) Step of performing targeted gene expression analysis on the hybrid nucleic acid molecule obtained from (d) comprising.

[0131] In additional embodiments, a method for targeted gene expression analysis is provided, the method comprising: (a) hybridizing a tagged RNA expression library with a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes with a specific target region in the RNA expression library, and (b) isolating the tagged RNA expression library-multifunctional capture probe hybrid module complex obtained from (a); (c) To produce a hybrid nucleic acid molecule, performing 5' to 3' DNA polymerase extension of the multifunctional capture probe in the complex obtained from b) to replicate a region of the captured tagged target region at the 3' of the multifunctional capture probe, wherein the hybrid nucleic acid molecule comprises the multifunctional capture probe hybrid module and a complement of the tagged target region in the cDNA library located in the 3' direction of the position where the multifunctional capture probe hybrid module hybridizes with the target region; (d) Performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from c); comprising.

[0132] In various embodiments, a method for targeted gene expression analysis is provided, the method comprising: (a) hybridizing a tagged cDNA library with a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes with a specific target region in the cDNA library; (b) isolating the tagged cDNA library - multifunctional capture probe hybrid module complex obtained from a); (c) To produce a hybrid nucleic acid molecule, performing 5' to 3' DNA polymerase extension of the multifunctional capture probe in the complex obtained from b) to replicate a region of the captured tagged target region in the cDNA library at the 3' of the multifunctional capture probe, wherein the hybrid nucleic acid molecule comprises the multifunctional capture probe hybrid module and a complement of the tagged target region in the cDNA library located in the 3' direction of the position where the multifunctional capture probe hybrid module hybridizes with the target region; (d) Performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from c); comprising.

[0133] In certain embodiments, at least two different multifunctional capture probe modules are used in said at least two a) steps, and said at least two a) steps each use one kind of multifunctional capture probe module.

[0134] In certain embodiments, at least one multifunctional capture probe module hybridizes downstream of said target region and at least one multifunctional capture probe module hybridizes upstream of said target region.

[0135] In additional embodiments, a method for cDNA sequence analysis is provided, the method comprising: (a) obtaining one or more clones, each clone comprising a first cDNA sequence and a second cDNA sequence, said first cDNA sequence comprising a targeted genomic cDNA sequence and said second cDNA sequence comprising a capture probe sequence; (b) performing a paired-end sequencing reaction on said one or more clones to obtain one or more sequencing reads; (c) ordering or clustering said sequencing reads of said one or more clones according to said probe sequence of said sequencing reads. and.

[0136] In various embodiments, a method for cDNA sequence analysis is provided, the method comprising: (a) obtaining one or more clones, each clone comprising a first cDNA sequence and a second cDNA sequence, said first cDNA sequence comprising a targeted genomic DNA sequence and said second cDNA sequence comprising a capture probe sequence; (b) performing a sequencing reaction on said one or more clones, wherein a single long sequencing read of more than about 100 nucleotides is obtained and said read is sufficient to identify both said first cDNA sequence and said second cDNA sequence; (c) ordering or clustering the sequencing reads of the one or more clones according to the probe array of the sequencing reads comprising.

[0137] In certain embodiments, the methods contemplated herein determine the distribution of unique and overlapping sequencing reads, count the number of times a unique read is encountered, fit the frequency distribution of the unique reads to a statistical distribution, estimate the total number of unique reads, and convert the unique read count to transcript abundance using normalization to the total reads collected within each cDNA library sample, and analyzing any sequencing reads associated with the second read sequence.

[0138] In certain embodiments, methods for targeted genetic analysis are provided, the method comprising: (a) hybridizing a tagged DNA library to a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes to a specific target region in the DNA library; (b) isolating the tagged DNA library - multifunctional capture probe hybrid module complex obtained from (a); (c) performing concerted enzymatic processing of the tagged DNA library - multifunctional capture probe hybrid module complex obtained from (b) to produce a hybrid nucleic acid molecule, including 5' FLAP endonuclease activity, 5' to 3' DNA polymerase extension, and nick closure by DNA ligase, to ligate a complement of the multifunctional capture probe to the target region 5' to the multifunctional capture probe binding site, wherein the hybrid nucleic acid molecule comprises a complement of the multifunctional capture probe hybrid module and a region of the tagged target region located 5' to the position where the multifunctional capture probe hybrid module hybridizes to the genomic target region. (d) Performing a targeted genetic analysis on the hybrid nucleic acid molecule obtained from (c). It includes.

[0139] In various embodiments, steps a) to c) are repeated at least about twice, and the targeted genetic analysis in d) includes an alignment of the hybrid nucleic acid molecule sequences obtained from the at least two d) steps.

[0140] In a particular embodiment, at least two different multifunctional capture probe modules are used in the at least two a) steps, and each of the at least two a) steps uses one type of multifunctional capture probe module.

[0141] In a specific embodiment, at least one multifunctional capture probe module hybridizes downstream of the target region, and at least one multifunctional capture probe module hybridizes upstream of the target region.

[0142] In additional embodiments, a method for determining the copy number of a specific target region is provided, and this method includes (a) Hybridizing a tagged DNA library with a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes with a specific target region in the genomic library. (b) Isolating the tagged DNA library - multifunctional capture probe hybrid module complex obtained from (a). (c) To generate a hybrid nucleic acid molecule, performing concerted enzymatic processing of the tagged DNA library - multifunctional capture probe hybrid module complex obtained from b), including 5' FLAP endonuclease activity, 5' to 3' DNA polymerase extension, and nick closure by DNA ligase, to ligate a complement of the multifunctional capture probe to the target region at the 5' of the multifunctional capture probe binding site, wherein the hybrid nucleic acid molecule includes a complement of the multifunctional capture probe hybrid module and a region of the tagged target region located 5' of the position where the multifunctional capture probe hybrid module hybridizes to the target region; (d) Performing PCR amplification of the hybrid nucleic acid molecule in c); (e) Quantifying the PCR reaction in d), wherein the quantification enables determination of the copy number of the specific target region; comprising.

[0143] In various embodiments, the methods contemplated herein include obtaining the sequence of the hybrid nucleic acid molecule obtained from step d).

[0144] In certain embodiments, steps a) - d) are repeated at least about 2 times, and sequence alignment of the hybrid nucleic acid molecules obtained from at least 2 d) steps is performed.

[0145] In certain embodiments, at least 2 different multifunctional capture probe modules are used in at least 2 a) steps, and each of the at least 2 a) steps uses one type of multifunctional capture probe module.

[0146] In certain embodiments, at least one type of multifunctional capture probe module hybridizes downstream of the genomic target region, and at least one type of multifunctional capture probe module hybridizes upstream of the genomic target region.

[0147] In additional embodiments, the targeted genetic analysis is sequence analysis.

[0148] In further embodiments, the target region is a genomic target region and the DNA library is a genomic DNA library.

[0149] In some embodiments, the target region is a cDNA target region and the DNA library is a cDNA library. In certain embodiments, for example, the following are provided: (Item 1) A method for producing a tagged genomic library, comprising: (a) treating fragmented genomic DNA with a terminal repair enzyme to produce fragmented and terminal-repaired genomic DNA; (b) ligating a random nucleic acid tag sequence and optionally a sample code sequence and / or a PCR primer sequence to the fragmented and terminal-repaired genomic DNA to produce the tagged genomic library. A method comprising the steps of: (Item 2) The method according to any one of the preceding items, wherein the random nucleic acid tag sequence is about 2 to about 100 nucleotides. (Item 3) The method according to any one of the preceding items, wherein the random nucleic acid tag sequence is about 2 to about 6 nucleotides. (Item 4) The method according to any one of the preceding items, wherein the fragmented and terminal-repaired genomic DNA contains blunt ends. (Item 5) The method according to any one of the preceding items, wherein the blunt ends are further modified to contain a single base pair overhang. (Item 6) The ligation step includes ligating a multifunctional adapter module to the fragmented and end-repaired genomic DNA to produce the tagged genomic library, wherein the multifunctional adapter molecule (i) a first region containing a random nucleic acid tag sequence, and (ii) a second region containing a sample code sequence, and (iii) a third region containing a PCR primer sequence The method according to any one of the preceding items. (Item 7) The method according to any one of the preceding items, further comprising the step of hybridizing the tagged genomic library with a multifunctional capture probe module to form a complex, wherein the multifunctional capture probe module hybridizes with a specific genomic target region in the genomic library. (Item 8) The method according to any one of the preceding items, further comprising the step of isolating the tagged genomic library - multifunctional capture probe module complex. (Item 9) The method according to any one of the preceding items, further comprising the step of processing the isolated tagged genomic library - multifunctional capture probe module complex with a 3'-5' exonuclease enzyme to remove the single-stranded 3' end. (Item 10) The enzyme used in the 3'-5' exonuclease enzyme processing is T4 DNA polymerase. The method according to any one of the preceding items. (Item 11) The method according to any one of the preceding items, further comprising the step of performing PCR on the complex processed by the 3'-5' exonuclease enzyme obtained from the preceding item, wherein, to generate a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is copied, and the hybrid nucleic acid molecule comprises the genomic target region capable of hybridizing with the multifunctional capture probe module and the complement of the multifunctional capture probe module tail sequence. (Item 12) (a) Hybridizing a tagged genomic library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes with a specific genomic target region in the genomic library; (b) Isolating the tagged genomic library-multifunctional capture probe module complex obtained from (a); (c) Performing 3'-5' exonuclease enzyme processing on the isolated tagged genomic library-multifunctional capture probe module complex obtained from (b) using an enzyme having 3'-5' exonuclease activity to remove the single-stranded 3' end; (d) Performing PCR on the complex processed by the enzyme obtained from (c), wherein, to generate a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is copied, and the hybrid nucleic acid molecule comprises the genomic target region capable of hybridizing with the multifunctional capture probe module and the complement of the multifunctional capture probe module tail sequence; (e) Performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from (d); A method for targeted genetic analysis, comprising: (Item 13) The method according to item 12, wherein steps a) to d) are repeated at least about 2 times, and the targeted genetic analysis in e) includes an alignment of the hybrid nucleic acid molecule sequences obtained from the at least 2 steps of d). (Item 14) The method according to item 13, wherein at least two different multifunctional capture probe modules are used in at least two steps of a), and each of the at least two steps of a) uses one kind of multifunctional capture probe module. (Item 15) The method according to item 14, wherein at least one multifunctional capture probe module hybridizes downstream of the genomic target region, and at least one multifunctional capture probe module hybridizes upstream of the genomic target region. (Item 16) A method for determining the copy number of a specific genomic target region, comprising: (a) hybridizing a tagged genomic library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module complex selectively hybridizes with a specific genomic target region in the genomic library; (b) isolating the tagged genomic library - multifunctional capture probe module complex obtained from a); (c) performing 3'-5' exonuclease enzyme processing on the isolated tagged genomic library - multifunctional capture probe module complex obtained from b) using an enzyme having 3'-5' exonuclease activity to remove the single-stranded 3' end; (d) performing a PCR reaction on the complex processed by the enzyme obtained from c), wherein the tail portion of the multifunctional capture probe molecule is replicated to produce a hybrid nucleic acid molecule, and the hybrid nucleic acid molecule includes the genomic target region capable of hybridizing with the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence. Performing the PCR amplification of the hybrid nucleic acid molecule in (e)d); (f)Quantifying the PCR reaction in (e), wherein the quantification enables determination of the copy number of the specific genomic target region; A method comprising the above steps. (Item 17) The method according to item 16, further comprising obtaining the sequence of the hybrid nucleic acid molecule obtained from step (e). (Item 18) The method according to item 17, wherein steps (a) to (e) are repeated at least about twice, and sequence alignment is performed using the hybrid nucleic acid molecule sequences obtained from the at least two steps (e). (Item 19) The method according to item 18, wherein at least two different multifunctional capture probe modules are used in the at least two steps (a), and each of the at least two steps (a) uses one type of multifunctional capture probe module. (Item 20) The method according to item 19, wherein at least one multifunctional capture probe module hybridizes downstream of the genomic target region, and at least one multifunctional capture probe module hybridizes upstream of the genomic target region. (Item 21) A method for determining the copy number of a specific genomic target region, comprising: (a)Hybridizing a tagged genomic library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module complex selectively hybridizes with a specific genomic target region in the genomic library; (b)Isolating the tagged genomic library - multifunctional capture probe module complex obtained from (a); (c) Using an enzyme having 3'-5' exonuclease activity, performing 3'-5' exonuclease enzyme processing on the isolated tagged genomic library-multifunctional capture probe module complex obtained from b) to remove the single-stranded 3' end; (d) Performing a PCR reaction on the complex processed by the enzyme obtained from c), wherein, to create a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is replicated, and the hybrid nucleic acid molecule comprises the genomic target region capable of hybridizing to the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; (e) Performing PCR amplification of the hybrid nucleic acid molecule in d); A method comprising the above steps. (Item 22) The method according to Item 21, further comprising obtaining the sequence of the hybrid nucleic acid molecule obtained from step e). (Item 23) The method according to Item 22, wherein steps a) to e) are repeated at least about twice, and sequence alignment is performed using the hybrid nucleic acid molecule sequences obtained from the at least two e) steps. (Item 24) The method according to Item 23, wherein at least two different multifunctional capture probe modules are used in the at least two a) steps, and each of the at least two a) steps uses one kind of multifunctional capture probe module. (Item 25) The method according to Item 24, wherein at least one kind of multifunctional capture probe module hybridizes downstream of the genomic target region, and at least one kind of multifunctional capture probe module hybridizes upstream of the genomic target region. (Item 26) A method for determining the copy number of a specific genomic target region, (a) Step of hybridizing a tagged genomic library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module complex selectively hybridizes with a specific genomic target region in the genomic library; (b) Step of isolating the tagged genomic library - multifunctional capture probe module complex obtained from (a); (c) Step of performing 3'-5' exonuclease enzymatic processing on the isolated tagged genomic library - multifunctional capture probe module complex obtained from (b) using an enzyme having 3'-5' exonuclease activity to remove the single-stranded 3' end; (d) Step of performing a PCR reaction on the complex processed by the enzyme obtained from (c), wherein, to produce a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is replicated, and the hybrid nucleic acid molecule includes the genomic target region capable of hybridizing with the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; (e) Step of performing PCR amplification of the hybrid nucleic acid molecule in (d); (f) Step of performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from (e) A method comprising the above steps. (Item 27) The method according to Item 26, wherein steps a) to e) are repeated at least about twice, and the targeted genetic analysis in f) includes performing sequence alignment of the hybrid nucleic acid molecule sequences obtained from the at least two steps of e). (Item 28) The method according to Item 27, wherein at least two different multifunctional capture probe modules are used in the at least two steps of a), and each of the at least two steps of a) uses one type of multifunctional capture probe module. (Item 29) The method according to item 28, wherein at least one multifunctional capture probe module hybridizes downstream of the genomic target region and at least one multifunctional capture probe module hybridizes upstream of the genomic target region. (Item 30) (a) Hybridizing a tagged genomic library with a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes with a specific genomic target region in the genomic library; (b) Isolating the tagged genomic library - multifunctional capture probe hybrid module complex obtained from (a); (c) Performing a 5' to 3' DNA polymerase extension of the multifunctional capture probe on the complex obtained from (b) to produce a hybrid nucleic acid molecule, and replicating a region of the captured tagged genomic target region at the 3' of the multifunctional capture probe, wherein the hybrid nucleic acid molecule comprises the multifunctional capture probe hybrid module and a complement of a region of the tagged genomic target region located in the 3' direction from the position where the multifunctional capture probe hybrid module hybridizes with the genomic target region; (d) Performing a targeted genetic analysis on the hybrid nucleic acid molecule obtained from (c); A method for targeted genetic analysis, comprising: (Item 31) The method according to item 30, wherein steps (a) to (c) are repeated at least about twice, and the targeted genetic analysis in (d) comprises an alignment of the hybrid nucleic acid molecule sequences obtained from the at least two steps of (d). (Item 32) The method according to item 31, wherein at least two different multifunctional capture probe modules are used in the at least two steps of (a), and each of the at least two steps of (a) uses one kind of multifunctional capture probe module. (Item 33) The method according to item 32, wherein at least one multifunctional capture probe module hybridizes downstream of the genomic target region and at least one multifunctional capture probe module hybridizes upstream of the genomic target region. (Item 34) A method for determining the copy number of a specific genomic target region, comprising: (a) hybridizing a tagged genomic library with a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes with a specific genomic target region in the genomic library; (b) isolating the tagged genomic library - multifunctional capture probe hybrid module complex obtained from (a); (c) performing 5' to 3' DNA polymerase extension of the multifunctional capture probe on the complex obtained from (b) to produce a hybrid nucleic acid molecule, replicating a region of the captured tagged genomic target region at the 3' of the multifunctional capture probe, wherein the hybrid nucleic acid molecule comprises the multifunctional capture probe hybrid module and a complement of a region of the tagged genomic target region located in the 3' direction from the position where the multifunctional capture probe hybrid module hybridizes with the genomic target region; (d) performing PCR amplification of the hybrid nucleic acid molecule in (c); (e) quantifying the PCR reaction in (d), wherein the quantification enables determination of the copy number of the specific genomic target region A method comprising the above steps. (Item 35) The method according to item 34, further comprising obtaining the sequence of the hybrid nucleic acid molecule obtained from step (d). (Item 36) The method according to item 35, wherein steps a) to d) are repeated at least about twice, and sequence alignment of the hybrid nucleic acid molecules obtained from said at least two d) steps is performed. (Item 37) The method according to item 36, wherein at least two different multifunctional capture probe modules are used in said at least two a) steps, and each of said at least two a) steps uses one kind of multifunctional capture probe module. (Item 38) The method according to item 37, wherein at least one multifunctional capture probe module hybridizes with the downstream of said genomic target region, and at least one multifunctional capture probe module hybridizes with the upstream of said genomic target region. (Item 39) The method according to any one of the preceding items, wherein said targeted genetic analysis is sequence analysis. (Item 40) The method according to any one of the preceding items, wherein said tagged genomic library is amplified by PCR to produce an amplified tagged genomic library. (Item 41) The method according to any one of the preceding items, wherein said genomic DNA is derived from a biological sample selected from the group consisting of blood, skin, hair, hair follicle, saliva, oral mucosa, vaginal mucosa, sweat, tear, epithelial tissue, urine, semen, seminal fluid, seminal plasma, prostatic fluid, bulbourethral gland fluid (Cowper's gland fluid), excrement, biopsy, ascites, cerebrospinal fluid, lymph fluid and tissue extract sample or biopsy sample. (Item 42) A tagged genomic library containing tagged genomic sequences, wherein each tagged genomic sequence (a) fragmented and end-repaired genomic DNA, (b) a random nucleotide tag sequence, (c) a sample code sequence, (d) a PCR primer sequence and a genomic library. (Item 43) A hybrid-tagged genomic library containing hybrid-tagged genomic sequences for use in targeted genetic analysis, wherein each hybrid-tagged genomic sequence comprises (a) fragmented and end-repaired genomic DNA, and (b) a random nucleotide tag sequence, and (c) a sample code sequence, and (d) a PCR primer sequence, and (e) a genomic target region, and (f) a multifunctional capture probe module tail sequence and comprising a genomic library. (Item 44) (a) a first region comprising a random nucleotide tag sequence, and (b) a second region comprising a sample code sequence, and (c) a third region comprising a PCR primer sequence and comprising a multifunctional adapter module. (Item 45) (a) a first region capable of hybridizing with a partner oligonucleotide, and (b) a second region capable of hybridizing with a specific genomic target region, and (c) a third region comprising a tail sequence and comprising a multifunctional capture probe module. (Item 46) The multifunctional capture probe module according to any one of the preceding items, wherein the first region is bound to a partner oligonucleotide. (Item 47) (a) a first region capable of hybridizing with a partner oligonucleotide and capable of functioning as a PCR primer, and (b) a second region capable of hybridizing with a specific genomic target region and comprising a multifunctional adapter probe hybrid module. (Item 48) The multifunctional capture probe hybrid module according to any one of the preceding items, wherein the first region is bound to a partner oligonucleotide. (Item 49) The method according to any one of the preceding items, wherein the partner oligonucleotide is chemically modified. (Item 50) A composition comprising a tagged genomic library, a multifunctional adapter module, and a multifunctional capture probe module. (Item 51) A composition comprising a hybrid-tagged genomic library according to any one of the preceding items. (Item 52) A reaction mixture for carrying out the method according to any one of the preceding items. (Item 53) A reaction mixture capable of producing a tagged genomic library, comprising: (a) fragmented genomic DNA; and (b) a DNA end repair enzyme for producing fragmented and end-repaired genomic DNA. (Item 54) The reaction mixture according to any one of the preceding items, further comprising a multifunctional adapter module. (Item 55) The reaction mixture according to any one of the preceding items, further comprising a multifunctional capture probe module. (Item 56) The reaction mixture according to any one of the preceding items, further comprising an enzyme having 3'-5' exonuclease activity and PCR amplification activity. (Item 57) (a) obtaining one or more clones, each clone containing a first DNA sequence and a second DNA sequence, wherein the first DNA sequence contains a targeted genomic DNA sequence and the second DNA sequence contains a capture probe sequence; (b) Performing a paired-end sequencing reaction on the one or more clones to obtain one or more sequencing reads; (c) Ordering or clustering the sequencing reads of the one or more clones according to the probe sequences of the sequencing reads; A method for DNA sequence analysis, comprising: (Item 58) (a) Obtaining one or more clones, each clone containing a first DNA sequence and a second DNA sequence, wherein the first DNA sequence contains a targeted genomic DNA sequence and the second DNA sequence contains a capture probe sequence; (b) Performing a sequencing reaction on the one or more clones, obtaining a single long sequencing read exceeding about 100 nucleotides, the read being sufficient for the identification of both the first DNA sequence and the second DNA sequence; (c) Ordering or clustering the sequencing reads of the one or more clones according to the probe sequences of the sequencing reads; A method for DNA sequence analysis, comprising: (Item 59) The method according to Item 57 or Item 58, wherein the sequences of the one or more clones are compared with one or more human reference DNA sequences. (Item 60) The method according to Item 59, wherein sequences that do not match the one or more human reference DNA sequences are identified. (Item 61) The method according to Item 60, wherein de novo assembly is created from the non-match sequence data using the non-match sequences. (Item 62) The method according to Item 61, wherein the de novo assembly is used to identify novel sequence rearrangements related to the capture probes. (Item 63) (a) Obtaining one or more clones, each clone containing a first DNA sequence and a second DNA sequence, wherein the first DNA sequence comprises a random nucleotide tag sequence and a targeted genomic DNA sequence, and the second DNA sequence comprises a capture probe sequence; (b) Performing a paired-end sequencing reaction on the one or more clones to obtain one or more sequencing reads; (c) Ordering or clustering the sequencing reads of the one or more clones according to the probe sequence of the sequencing reads; A method for genomic copy number determination analysis, comprising the above steps. (Item 64) (a) Obtaining one or more clones, each clone containing a first DNA sequence and a second DNA sequence, wherein the first DNA sequence comprises a random nucleotide tag sequence and a targeted genomic DNA sequence, and the second DNA sequence comprises a capture probe sequence; (b) Performing a sequencing reaction on the one or more clones, obtaining a single long sequencing read exceeding about 100 nucleotides, and the read being sufficient for the identification of both the first DNA sequence and the second DNA sequence; (c) Ordering or clustering the sequencing reads of the one or more clones according to the probe sequence of the sequencing reads; A method for genomic copy number determination analysis, comprising the above steps. (Item 65) The method according to item 63 or item 64, wherein the random nucleotide tag sequence has a length of about 2 to about 50 nucleotides. (Item 66) (a) Determining the distribution of unique sequencing reads and overlapping sequencing reads; (b) Counting the number of times a unique read is encountered; (c) Fitting the frequency distribution of the unique reads to a statistical distribution; (d) Estimate the total number of unique reads, (e) Normalize the said total number of estimated unique reads against the assumption that most human loci are generally diploid The method according to item 63 or item 64, further comprising the step of analyzing any sequencing reads associated with the second read sequence. (Item 67) The method according to item 66, wherein the estimated copy number of one or more targeted loci is determined. (Item 68) The method according to item 67, wherein the one or more target loci deviating from the expected copy number value are determined. (Item 69) The method according to item 67, wherein the one or more targeted loci of the gene are grouped together in a collection of loci, and the copy number measurement values obtained from the collection of targeted loci are averaged and normalized. (Item 70) The method according to item 67, wherein the estimated copy number of the gene is represented by the said normalized average of all target loci representing this gene. (Item 71) A method for generating a tagged RNA expression library, comprising: (a) Fragmenting a cDNA library; (b) Treating the fragmented cDNA library with a terminal repair enzyme to produce a fragmented and terminal-repaired cDNA; (c) Ligating a multifunctional adapter molecule to the fragmented and terminal-repaired cDNA to produce a tagged RNA expression library and comprising. (Item 72) A method for generating a tagged RNA expression library, comprising: (a) Preparing a cDNA library from total RNA of one or more cells; (b) Fragmenting the cDNA library; (c) Treating the fragmented cDNA with a terminal repair enzyme to produce fragmented and end-repaired cDNA; (d) Ligating a multifunctional adapter molecule to the fragmented and end-repaired cDNA to produce a tagged RNA expression library A method comprising the steps of. (Item 73) The method according to item 71 or item 72, wherein the cDNA library is an oligo dT-primed cDNA library. (Item 74) The method according to item 71 or item 72, wherein the cDNA library is primed with a random oligonucleotide containing about 6 to about 20 random nucleotides. (Item 75) The method according to item 71 or item 72, wherein the cDNA library is primed with a random hexamer or a random octamer. (Item 76) The method according to item 71 or item 72, wherein the cDNA library is fragmented to a size of about 250 bp to about 750 bp. (Item 77) The method according to item 71 or item 72, wherein the cDNA library is fragmented to a size of about 500 bp. (Item 78) The multifunctional adapter module is (i) a first region containing a random nucleic acid tag sequence, and optionally, (ii) a second region containing a sample code sequence, and optionally, (iii) a third region containing a PCR primer sequence The method according to any one of items 71 to 77, comprising the steps of. (Item 79) The method according to any one of claims 71 to 78, wherein the multifunctional adapter module includes a first region containing a random nucleic acid tag sequence, a second region containing a sample code sequence, and a third region containing a PCR primer sequence. (Item 80) Hybridizing the tagged cDNA library with a multifunctional capture probe module to form a complex, the multifunctional capture probe module further comprising hybridizing to a specific target region in the cDNA library, the method according to any one of items 71 to 78. (Item 81) The method according to any one of items 71 to 78, further comprising isolating the tagged cDNA library - multifunctional capture probe module complex. (Item 82) The method according to any one of items 71 to 78, further comprising subjecting the isolated tagged cDNA library - multifunctional capture probe module complex to 3'-5' exonuclease enzyme processing to remove the single-stranded 3' end. (Item 83) The enzyme used in the 3'-5' exonuclease enzyme processing is T4 DNA polymerase, the method according to item 82. (Item 84) Performing PCR on the complex processed by the 3'-5' exonuclease enzyme, wherein in order to produce a hybrid nucleic acid molecule, the tail portion of the multifunctional capture probe molecule is copied, and the hybrid nucleic acid molecule can hybridize to the cDNA target region capable of hybridizing to the multifunctional capture probe module and the complement of the multifunctional capture probe module tail sequence, the method according to item 82 or item 83. (Item 85) (a) Step of hybridizing a tagged RNA expression library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes with a specific target region in the tagged RNA expression library, and (b) Step of isolating the tagged RNA expression library-multifunctional capture probe module complex obtained from (a); (c) Step of performing 3'-5' exonuclease enzyme processing on the isolated tagged RNA expression library-multifunctional capture probe module complex obtained from (b) using an enzyme having 3'-5' exonuclease activity to remove the single-stranded 3' end; (d) Step of performing PCR on the complex processed by the enzyme obtained from (c), wherein the tail portion of the multifunctional capture probe molecule is copied to create a hybrid nucleic acid molecule, and the hybrid nucleic acid molecule includes the target region capable of hybridizing with the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; (e) Step of performing targeted gene expression analysis on the hybrid nucleic acid molecule obtained from (d) A method for targeted gene expression analysis, comprising: (Item 86) (a) Step of hybridizing a tagged RNA expression library with a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes with a specific target region in the RNA expression library, and (b) Step of isolating the tagged RNA expression library-multifunctional capture probe hybrid module complex obtained from (a); (c) To produce a hybrid nucleic acid molecule, performing 5'-to-3' DNA polymerase extension of the multifunctional capture probe in the complex obtained from b) to replicate a region of the captured tagged target region that is 3' to the multifunctional capture probe, wherein the hybrid nucleic acid molecule comprises the multifunctional capture probe hybrid module and a complement of the tagged target region in the cDNA library that is located 3' in the direction of the position where the multifunctional capture probe hybrid module hybridizes to the target region; (d) Performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from c); A method for targeted gene expression analysis, comprising: (Item 87) (a) Hybridizing a tagged cDNA library with a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes with a specific target region in the cDNA library; (b) Isolating the tagged cDNA library - multifunctional capture probe hybrid module complex obtained from a); (c) To produce a hybrid nucleic acid molecule, performing 5'-to-3' DNA polymerase extension of the multifunctional capture probe in the complex obtained from b) to replicate a region of the captured tagged target region in the cDNA library that is 3' to the multifunctional capture probe, wherein the hybrid nucleic acid molecule comprises the multifunctional capture probe hybrid module and a complement of the tagged target region in the cDNA library that is located 3' in the direction of the position where the multifunctional capture probe hybrid module hybridizes to the target region; (d) Performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from c); A method for targeted gene expression analysis, comprising: (Item 88) At least two different multifunctional capture probe modules are used in said at least two a) steps, and each of said at least two a) steps uses one kind of multifunctional capture probe module, the method according to any one of items 85 to 87. (Item 89) The method according to item 88, wherein at least one multifunctional capture probe module hybridizes downstream of the target region and at least one multifunctional capture probe module hybridizes upstream of the target region. (Item 90) (a) Obtaining one or more clones, each clone containing a first cDNA sequence and a second cDNA sequence, wherein the first cDNA sequence contains a targeted genomic cDNA sequence and the second cDNA sequence contains a capture probe sequence; (b) Performing a paired-end sequencing reaction on said one or more clones to obtain one or more sequencing reads; (c) Ordering or clustering said sequencing reads of said one or more clones according to said probe sequence of said sequencing reads A method for cDNA sequence analysis, comprising: (Item 91) (a) Obtaining one or more clones, each clone containing a first cDNA sequence and a second cDNA sequence, wherein the first cDNA sequence contains a targeted genomic DNA sequence and the second cDNA sequence contains a capture probe sequence; (b) Performing a sequencing reaction on said one or more clones, wherein a single long sequencing read of more than about 100 nucleotides is obtained and said read is sufficient for the identification of both said first cDNA sequence and said second cDNA sequence; (c) Ordering or clustering the sequencing reads of said one or more clones according to said probe sequence of said sequencing reads A method for cDNA sequence analysis, comprising: (Item 92) (a) Determining the distribution of unique sequencing reads and overlapping sequencing reads; (b) Counting the number of times unique reads are encountered; (c) Fitting the frequency distribution of said unique reads to a statistical distribution; (d) Estimating the total number of unique reads; (e) Converting the unique read count to transcript abundance using normalization against the total reads collected within each cDNA library sample; The method according to item 90 or item 91, further comprising analyzing any sequencing reads associated with a second read sequence by: (Item 93) (a) Hybridizing a tagged DNA library to a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes to a specific target region in the DNA library; (b) Isolating the tagged DNA library - multifunctional capture probe hybrid module complex obtained from (a); (c) Performing concerted enzymatic processing of the tagged DNA library - multifunctional capture probe hybrid module complex obtained from (b) to ligate a complement of the multifunctional capture probe to the target region 5' to the multifunctional capture probe binding site, comprising 5' FLAP endonuclease activity, 5' to 3' DNA polymerase extension, and nick closure by DNA ligase, wherein the hybrid nucleic acid molecule comprises a complement of the multifunctional capture probe hybrid module and a region of the tagged target region located 5' to the position where the multifunctional capture probe hybrid module hybridizes to the genomic target region; (d) Performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from (c). A method for targeted genetic analysis, comprising: (Item 94) The method according to item 93, wherein steps a) to c) are repeated at least about twice, and the targeted genetic analysis in d) includes an alignment of the hybrid nucleic acid molecule sequences obtained from said at least two d) steps. (Item 95) The method according to item 94, wherein at least two different multifunctional capture probe modules are used in said at least two a) steps, and each of said at least two a) steps uses one type of multifunctional capture probe module. (Item 96) The method according to item 95, wherein at least one multifunctional capture probe module hybridizes downstream of the target region and at least one multifunctional capture probe module hybridizes upstream of the target region. (Item 97) A method for determining the copy number of a specific target region, comprising: (a) Hybridizing a tagged DNA library with a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes with a specific target region in the genomic library; (b) Isolating the tagged DNA library - multifunctional capture probe hybrid module complex obtained from a); (c) To produce a hybrid nucleic acid molecule, perform concerted enzymatic processing of the tagged DNA library - multifunctional capture probe hybrid module complex obtained from b), including 5' FLAP endonuclease activity, 5' to 3' DNA polymerase extension, and nick closure by DNA ligase, to ligate a complement of the multifunctional capture probe to the target region at the 5' of the multifunctional capture probe binding site, wherein the hybrid nucleic acid molecule comprises a complement of the multifunctional capture probe hybrid module and a region of the tagged target region located 5' to the position where the multifunctional capture probe hybrid module hybridizes to the target region. (d) Performing PCR amplification of the hybrid nucleic acid molecule in c). (e) Quantifying the PCR reaction in d), wherein the quantification enables determination of the copy number of the specific target region. A method comprising the above steps. (Item 98) The method according to item 97, further comprising obtaining the sequence of the hybrid nucleic acid molecule obtained from step d). (Item 99) The method according to item 98, wherein steps a) to d) are repeated at least about 2 times, and sequence alignment of the hybrid nucleic acid molecules obtained from at least 2 steps of d) is performed. (Item 100) The method according to item 99, wherein at least two different multifunctional capture probe modules are used in at least 2 steps of a), and each of the at least 2 steps of a) uses one type of multifunctional capture probe module. (Item 101) The method according to item 100, wherein at least one type of multifunctional capture probe module hybridizes downstream of the genomic target region, and at least one type of multifunctional capture probe module hybridizes upstream of the genomic target region. (Item 102) The method according to any one of items 93 to 101, wherein the targeted genetic analysis is sequence analysis. (Item 103) The method according to any one of items 93 to 102, wherein the target region is a genomic target region and the DNA library is a genomic DNA library. (Item 104) The method according to any one of items 93 to 103, wherein the target region is a cDNA target region and the DNA library is a cDNA library.

Brief Description of the Drawings

[0150]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46

Figure 47

Figure 48

Mode for Carrying Out the Invention

[0151] Detailed Description A. Overview The present invention is based, at least in part, on the discovery that the coordinated use of several important molecular modules can be utilized in the performance of targeted genetic analysis.

[0152] In the practice of the present invention, unless specifically indicated to the contrary, conventional methods of chemistry, biochemistry, organic chemistry, molecular biology, microbiology, recombinant DNA techniques, genetics, immunology, and cell biology can be utilized within the skill of those in the art, and many of these are described hereinafter for illustrative purposes. Such techniques are well described in the literature. For example, Sambrook, et al., Molecular Cloning: A Laboratory Manual (3rd Edition, 2001); Sambrook, et al., Molecular Cloning: A Laboratory Manual (2nd Edition, 1989); Maniatis et al., Molecular Cloning: A Laboratory Manual (1982); Ausubel et al., Current Protocols in Molecular Biology (John Wiley and Sons, updated July 2008); Short Protocols in Molecular Biology:A Compendium of Methods from Current Protocols in Molecular Biology, Greene Pub. Associates and Wiley-Interscience; Glover, DNA Cloning: A Practical Approach, vol. I & II (IRL Press, Oxford, 1985); Anand, Techniques for the Analysis of Complex Genomes, (Academic Press, New York, 1992); Transcription and Translation (B. Hames & S. Higgins, Eds., 1984); Perbal, A Practical Guide to Molecular Cloning (1984); as well as Harlow and Lane, Antibodies, (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1998).

[0153] All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety.

[0154] B. Definitions Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred embodiments of the compositions, methods, and materials are described herein. For the purposes of the present invention, the following terms are defined below.

[0155] As used herein, the articles “a,” “an,” and “the” are used to refer to one or more than one (i.e., at least one) of the grammatical objects of the article. By way of example, “an element” means one element or more than one element.

[0156] The use of alternatives (e.g., “or”) is to be understood to mean either one of the alternatives, both, or any combination thereof.

[0157] The term “and / or” is to be understood to mean either or both of the alternatives.

[0158] As used herein, the terms “about” or “approximately” refer to a content, level, value, number, frequency, percentage, dimension, size, amount, weight, or length that varies by about 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% relative to a reference content, level, value, number, frequency, percentage, dimension, size, amount, weight, or length. In one embodiment, the terms “about” or “approximately” refer to a range of content, level, value, number, frequency, percentage, dimension, size, amount, weight, or length that is ±15%, ±10%, ±9%, ±8%, ±7%, ±6%, ±5%, ±4%, ±3%, ±2%, or ±1% with respect to a reference content, level, value, number, frequency, percentage, dimension, size, amount, weight, or length.

[0159] Throughout this specification, unless the context requires otherwise, the word "comprise", "comprises" and "comprising" will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements. In particular embodiments, the terms "include", "has", "contains" and "comprise" are used synonymously. "Consisting of" means including and limited to whatever follows the phrase "consisting of". Thus, the phrase "consisting of" indicates that the listed elements are necessary or essential and that no other elements may be present.

[0160] "Consisting essentially of" means including any elements listed after this phrase, limited to other elements that do not interfere with or contribute to the activity or action specified in the disclosure of the listed elements. Thus, the phrase "consisting essentially of" indicates that the listed elements are necessary or essential, but that other elements may or may not be present, depending on whether or not they affect the activity or action of the listed elements.

[0161]

[0162] ​​Throughout this specification, references to "one embodiment", "an embodiment", "a particular embodiment", "related embodiments", "certain embodiments", "additional embodiments" or "further embodiments" or combinations thereof mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the invention. Thus, the appearances of the foregoing phrases in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0163] As used herein, the term "isolated" means a material that substantially or essentially does not contain the components that are normally associated with it in its native state. In certain embodiments, the terms "obtained" or "derived from" are used synonymously with isolated.

[0164] As used herein, the term "DNA" refers to deoxyribonucleic acid. In various embodiments, the term DNA refers to genomic DNA, recombinant DNA, synthetic DNA, or cDNA. In one embodiment, DNA refers to genomic DNA or cDNA. In certain embodiments, DNA contains a "target region". DNA libraries contemplated herein include genomic DNA libraries and cDNA libraries constructed from RNA, such as RNA expression libraries. In various embodiments, the DNA library contains one or more additional DNA sequences and / or tags.

[0165] "Target region" refers to a region of interest within a DNA sequence. In various embodiments, targeted genetic analysis is performed in the target region. In certain embodiments, the target region is sequenced or the copy number of the target region is determined.

[0166] C. Exemplary Embodiments The present invention, in part, contemplates a method for generating a tagged genomic library. In certain embodiments, the method comprises treating fragmented DNA, such as genomic DNA or cDNA, with a terminal repair enzyme to produce fragmented and end-repaired DNA, and subsequently ligating a random nucleic acid tag sequence to produce a tagged genomic library. In some embodiments, a sample code sequence and / or a PCR primer sequence are ligated to the fragmented and end-repaired DNA, if desired.

[0167] The present invention, in part, contemplates a method for generating a tagged DNA library. In certain embodiments, the method comprises treating fragmented DNA with a terminal repair enzyme to produce fragmented and end-repaired DNA, and subsequently ligating a random nucleic acid tag sequence to produce a tagged DNA library. In some embodiments, a sample code sequence and / or a PCR primer sequence are ligated to the fragmented and end-repaired DNA, if desired.

[0168] Illustrative methods for fragmenting DNA include, but are not limited to, shearing, sonication, enzymatic digestion including restriction digestion, and other methods. In certain embodiments, any method known in the art for fragmenting DNA can be used in conjunction with the present invention.

[0169] In some embodiments, fragmented DNA is processed by a terminal repair enzyme to create end-repaired DNA. In some embodiments, the terminal repair enzyme can result in, for example, blunt ends, 5'-overhangs, and 3'-overhangs. In some embodiments, the end-repaired DNA contains blunt ends. In some embodiments, the end-repaired DNA is processed to contain blunt ends. In some embodiments, the blunt ends of the end-repaired DNA are further modified to contain single base pair overhangs. In some embodiments, the end-repaired DNA containing blunt ends can be further processed to contain an adenine (A) / thymine (T) overhang. In some embodiments, the end-repaired DNA containing blunt ends can be further processed to contain an adenine (A) / thymine (T) overhang as a single base pair overhang. In some embodiments, the end-repaired DNA has a non-templated 3' overhang In some embodiments, the end-repaired DNA is processed to contain a 3'-overhang. In some embodiments, the end-repaired DNA is processed to contain a 3'-overhang by terminal transferase (TdT). In some embodiments, a G-tail can be added by TdT. In some embodiments, the end-repaired DNA is processed to contain overhanging ends using partial digestion with any known restriction enzyme (e.g., enzyme Sau3A or others).

[0170] In certain embodiments, the DNA fragments are tagged using one or more "random nucleotide tags" or "random nucleic acid tags". As used herein, the terms "random nucleotide tag" or "random nucleic acid tag" refer to polynucleotides of individual lengths, and the nucleotide sequences are made or selected randomly. In certain illustrative embodiments, the length of the random nucleic acid tag is from about 2 to about 100 nucleotides, from about 2 to about 75 nucleotides, from about 2 to about 50 nucleotides, from about 2 to about 25 nucleotides, from about 2 to about 20 nucleotides, from about 2 to about 15 nucleotides, from about 2 to about 10 nucleotides, from about 2 to about 8 nucleotides, or from about 2 to about 6 nucleotides. In certain embodiments, the length of the random nucleotide tag is from about 2 to about 6 nucleotides (see, e.g., FIG. 1). In one embodiment, the random nucleotide tag sequence is about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 nucleotides.

[0171] In certain embodiments, the random nucleotide tags of the invention can be added to fragmented DNA using methods known in the art. In some embodiments, "tagmentation" can be used. Tagmentation is a commercially available Nextera technology (Illumina and Epicenter, USA) that can be used to load the random nucleotide tags and / or multifunctional adapter modules of the invention onto a transposon protein complex. The loaded transposon complex can then be used in the generation of a tagged genomic library according to the described methods.

[0172] The DNA used in this method may be derived from any source known to those skilled in the art. The DNA can be collected from any source, synthesized from RNA as complementary DNA (cDNA), and processed into pure or substantially pure DNA for use in this method. In some embodiments, the size of the fragmented DNA ranges from about 2 to about 500 base pairs, from about 2 to about 400 base pairs, from about 2 to about 300 base pairs, from about 2 to about 250 base pairs, from about 2 to about 200 base pairs, from about 2 to about 100 base pairs, or from about 2 to about 50 base pairs.

[0173] The combination of the DNA fragment end sequence and the introduced "random nucleic acid tag(s)" constitutes a two-element combination hereinafter referred to as a "genomic tag" or "cDNA tag". In some embodiments, whether a "genomic tag" or "cDNA tag" is unique can be determined by the product of the diversity combination within the attached random nucleotide tag pool multiplied by the diversity of the DNA fragment end sequence pool.

[0174] The present invention also contemplates, in part, a multifunctional adapter module. As used herein, the term "multifunctional adapter module" refers to a polynucleotide comprising (i) a first region containing a random nucleotide tag sequence, (ii) a second region containing a sample code sequence, if desired, and (iii) a third region containing a PCR primer sequence, if desired. In certain embodiments, the multifunctional adapter module comprises a PCR primer sequence, a random nucleotide tag, and a sample code sequence. In certain specific embodiments, the multifunctional adapter module comprises a PCR primer sequence and a random nucleotide tag or a sample code sequence. In some embodiments, the second region containing the sample code is optional. In some embodiments, the multifunctional adapter module does not contain the second region, but instead contains only the first and third regions. The multifunctional adapter module of the present invention can include blunt or complementary ends suitable for the ligation methods used, including the other ends known to those skilled in the art for ligating the multifunctional adapter module to fragmented DNA, together with the ends disclosed elsewhere in this specification.

[0175] In various embodiments, the first region contains a random nucleotide tag sequence. In certain embodiments, the first region contains a random nucleotide tag sequence of about 2 to about 100 nucleotides, about 2 to about 75 nucleotides, about 2 to about 50 nucleotides, about 2 to about 25 nucleotides, about 2 to about 20 nucleotides, about 2 to about 15 nucleotides, about 2 to about 10 nucleotides, about 2 to about 8 nucleotides, or about 2 to about 6 nucleotides, or any intervening number of nucleotides.

[0176] In certain embodiments, the second region, when present as needed, contains a sample code sequence. As used herein, the term "sample code sequence" refers to a polynucleotide used for sample identification. In certain embodiments, the second region contains a sample code sequence of about 1 to about 100 nucleotides, about 2 to about 75 nucleotides, about 2 to about 50 nucleotides, about 2 to about 25 nucleotides, about 2 to about 20 nucleotides, about 2 to about 15 nucleotides, about 2 to about 10 nucleotides, about 2 to about 8 nucleotides, or about 2 to about 6 nucleotides, or any intervening number of nucleotides.

[0177] In certain embodiments, the third region, when present as needed, contains a PCR primer sequence. In certain embodiments, the third region contains a PCR primer sequence of about 5 to about 200 nucleotides, about 5 to about 150 nucleotides, about 10 to about 100 nucleotides, about 10 to about 75 nucleotides, about 10 to about 50 nucleotides, about 10 to about 40 nucleotides, about 20 to about 40 nucleotides, or about 20 to about 30 nucleotides, or any intervening number of nucleotides.

[0178] In certain embodiments, the ligation step involves ligating a multifunctional adapter module to the fragmented and end-repaired DNA. Using this ligation reaction, a tagged DNA library can be created that contains end-repaired DNA ligated to multifunctional adapter molecules and / or random nucleotide tags. In some embodiments, a single multifunctional adapter module is used. In some embodiments, two or more multifunctional adapter modules are used. In some embodiments, a single multifunctional adapter module of the same sequence is ligated to each end of the fragmented and end-repaired DNA.

[0179] The present invention also provides a multifunctional capture probe module. As used herein, the term "multifunctional capture probe module" refers to a polynucleotide comprising (i) a first region capable of hybridizing with a partner oligonucleotide, (ii) a second region capable of hybridizing with a specific target region, and optionally (iii) a third region comprising a tail sequence.

[0180] In one embodiment, the multifunctional capture probe module comprises a region capable of hybridizing with a partner oligonucleotide, a region capable of hybridizing with a DNA target sequence, and a tail sequence.

[0181] In one embodiment, the multifunctional capture probe module comprises a region capable of hybridizing with a partner oligonucleotide and a region capable of hybridizing with a genomic target sequence.

[0182] In certain embodiments, the multifunctional capture probe module optionally comprises a random nucleotide tag sequence.

[0183] In various embodiments, the first region comprises a region capable of hybridizing with a partner oligonucleotide. As used herein, the term "partner oligonucleotide" refers to an oligonucleotide complementary to the nucleotide sequence of the multifunctional capture probe module. In certain embodiments, the first region capable of hybridizing with a partner oligonucleotide has a sequence of about 20 to about 200 nucleotides, about 20 to about 150 nucleotides, about 30 to about 100 nucleotides, about 30 to about 75 nucleotides, about 20 to about 50 nucleotides, about 30 to about 45 nucleotides, or about 35 to about 45 nucleotides. In certain specific embodiments, the region is about 30 to about 50 nucleotides, about 30 to about 40 nucleotides, about 30 to about 35 nucleotides, or about 34 nucleotides, or any intervening number of nucleotides.

[0184] In certain embodiments, the second region, when present optionally, comprises a region capable of hybridizing to a specific DNA target region. As used herein, the term "DNA target region" refers to a genomic or cDNA region selected for analysis using the compositions and methods contemplated herein. In certain embodiments, the second region comprising a region capable of hybridizing to a specific target region is from about 20 to about 200 nucleotides, from about 30 to about 150 nucleotides, from about 50 to about 150 nucleotides, from about 30 to about 100 nucleotides, from about 50 to about 100 nucleotides, from about 50 to about 90 nucleotides, from about 50 to about 80 nucleotides, from about 50 to about 70 nucleotides or from about 50 to about 60 nucleotides in sequence. In certain specific embodiments, the second region is about 60 nucleotides or any intervening number of nucleotides.

[0185] In certain embodiments, the third region, when present optionally, comprises a tail sequence. As used herein, the term "tail sequence" refers to a polynucleotide at the 5'-end of a multifunctional capture probe module that can serve as a PCR primer binding site in certain embodiments. In certain embodiments, the third region comprises a tail sequence from about 5 to about 100 nucleotides, from about 10 to about 100 nucleotides, from about 5 to about 75 nucleotides, from about 5 to about 50 nucleotides, from about 5 to about 25 nucleotides or from about 5 to about 20 nucleotides. In certain specific embodiments, the third region is from about 10 to about 50 nucleotides, from about 15 to about 40 nucleotides, from about 20 to about 30 nucleotides or about 20 nucleotides or any intervening number of nucleotides.

[0186] In one embodiment, the multifunctional capture probe module includes a region capable of hybridizing with a partner oligonucleotide and a region capable of hybridizing with a genomic target sequence. In certain embodiments where the multifunctional capture probe module includes a region capable of hybridizing with a partner oligonucleotide and a region capable of hybridizing with a genomic target sequence, the partner oligo can also function as a tail sequence or a primer binding site.

[0187] In one embodiment, the multifunctional capture probe module includes a tail region and a region capable of hybridizing with a genomic target sequence.

[0188] In various embodiments, the multifunctional capture probe includes a specific member of a binding pair and enables the isolation and / or purification of one or more captured fragments of a tagged DNA library that hybridize with the multifunctional capture probe. In certain embodiments, the multifunctional capture probe is conjugated to biotin or another suitable hapten, such as dinitrophenol, digoxigenin.

[0189] The present invention further contemplates, in part, the step of hybridizing a tagged DNA library with a multifunctional capture probe module to form a complex. In some embodiments, the multifunctional capture probe module substantially hybridizes with a specific genomic target region in the DNA library.

[0190] Hybridization or hybridization conditions can include any reaction conditions under which two nucleotide sequences form a stable complex; for example, the conditions under which a tagged DNA library and a multifunctional capture probe module form a stable tagged DNA library - multifunctional capture probe module complex. Such reaction conditions are well known in the art, and one of ordinary skill in the art will recognize that such conditions can be appropriately modified within the scope of the present invention. Substantial hybridization can occur when the second region of the multifunctional capture probe complex exhibits 100%, 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 85%, 80%, 75% or 70% sequence identity, homology or complementarity to the region of the tagged DNA library.

[0191] In certain embodiments, the first region of the multifunctional capture probe module does not substantially hybridize to the region of the tagged DNA library to which the second region substantially hybridizes. In some embodiments, the third region of the multifunctional capture probe module does not substantially hybridize to the region of the tagged DNA library to which the second region of the multifunctional capture probe module substantially hybridizes. In some embodiments, the first and third regions of the multifunctional capture probe module do not substantially hybridize to the region of the tagged DNA library to which the second region of the multifunctional capture probe module substantially hybridizes.

[0192] In certain embodiments, the methods contemplated herein include isolating the tagged DNA library - multifunctional capture probe module complex. In certain embodiments, methods for isolating DNA complexes are well known to those of ordinary skill in the art, and any method considered appropriate by those of ordinary skill in the art may be used in conjunction with the methods of the present invention (Ausubel et al., Current Protocols in Molecular Biology, pages 2007 - 2012). In certain embodiments, the complex is isolated using biotin-streptavidin isolation techniques. In some embodiments, the partner oligonucleotide capable of hybridizing to the first region of the multifunctional capture probe module is modified to contain biotin at the 5' or 3' end that can interact with streptavidin linked to a column, bead, or other substrate used in the DNA complex isolation method.

[0193] In certain embodiments, the first region of the multifunctional capture probe module is bound to the partner oligonucleotide. In some embodiments, the multifunctional capture probe module is bound to the partner oligonucleotide prior to the formation of the tagged DNA library - multifunctional capture probe module complex. In some embodiments, the multifunctional capture probe module is bound to the partner oligonucleotide after the formation of the tagged DNA library - multifunctional capture probe module complex. In some embodiments, the multifunctional capture probe module is bound to the partner oligonucleotide simultaneously with the formation of the tagged DNA library - multifunctional capture probe module complex. In some embodiments, the partner oligonucleotide is chemically modified.

[0194] In certain embodiments, removal of the single-stranded 3' end from the isolated tagged DNA library - multifunctional capture probe module complex is contemplated. In certain embodiments, the method includes the step of processing the isolated tagged DNA library - multifunctional capture probe module complex with a 3'-5' exonuclease enzyme to remove the single-stranded 3' end.

[0195] In certain other embodiments, the method includes the step of performing 5'-3' DNA polymerase extension of the multifunctional capture probe using the isolated tagged DNA library fragment as a template.

[0196] In certain other embodiments, the method includes creating a tagged DNA target molecule isolated by a hybrid multifunctional capture probe by the concerted action of 5' FLAP endonuclease, DNA polymerization and nick closure by DNA ligase.

[0197] For 3'-5' exonuclease enzyme processing of the isolated tagged DNA library - multifunctional capture probe module complex, various enzymes can be used. Examples of suitable enzymes that exhibit 3'-5' exonuclease enzyme activity and can be used in certain embodiments include, but are not limited to, T4 or exonuclease I, III, V (Shevelev IV, Hubscher U., "The 3' 5' exonucleases", Nat Rev Mol Cell Biol. Vol. 3(5): 364-7 See also page 6 (2002)). In certain embodiments, the enzyme containing 3'-5' exonuclease activity is T4 polymerase. In certain embodiments, for example, an enzyme that exhibits 3'-5' exonuclease enzyme activity, including T4 or exonuclease I, III, V, and can perform primer-template extension can be used. Id. 3'5'

[0198] In some embodiments, the methods contemplated herein include performing PCR on the complex processed by the 3'-5' exonuclease enzyme described above and elsewhere herein. In certain embodiments, the tail portion of the multifunctional capture probe molecule is copied to create a hybrid nucleic acid molecule. In one embodiment, the hybrid nucleic acid molecule created includes a target region that can hybridize to the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence.

[0199] In various embodiments, methods for targeted genetic analysis are also contemplated. In certain embodiments, a method for targeted genetic analysis comprises: a) hybridizing a tagged DNA library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes to a specific target region in a genomic library; b) isolating the tagged DNA library - multifunctional capture probe module complex obtained from a); c) performing 3'-5' exonuclease enzymatic processing on the isolated tagged DNA library - multifunctional capture probe module complex obtained from b) using an enzyme having 3'-5' exonuclease activity to remove single-stranded 3' ends; d) performing PCR on the complex processed by the enzyme obtained from c), wherein the tail portion of the multifunctional capture probe molecule is copied to create a hybrid nucleic acid molecule that comprises a target region capable of hybridizing to the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; and e) performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from d).

[0200] In various embodiments, methods for determining the copy number of a specific target region are contemplated. In certain embodiments, a method for determining the copy number of a specific target region comprises: a) hybridizing a tagged DNA library to a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes to the specific target region in the DNA library; b) isolating the tagged DNA library-multifunctional capture probe module complex obtained from a); c) performing 3'-5' exonuclease enzymatic processing on the isolated tagged DNA library-multifunctional capture probe module complex obtained from b) using an enzyme having 3'-5' exonuclease activity to remove the single-stranded 3' end; d) performing a PCR reaction on the complex processed by the enzyme obtained from c), wherein the tail portion of the multifunctional capture probe molecule is replicated to create a hybrid nucleic acid molecule, and the hybrid nucleic acid molecule comprises a target region capable of hybridizing to the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; e) performing PCR amplification of the hybrid nucleic acid in d); and f) quantifying the PCR reaction in e), wherein the quantification enables determination of the copy number of the specific target region.

[0201] In various embodiments, methods for targeted genetic analysis are also contemplated. In certain embodiments, a method for targeted genetic analysis comprises: a) hybridizing a tagged DNA library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes to a specific target region in the genomic library; b) isolating the tagged DNA library - multifunctional capture probe module complex obtained from a); c) using the isolated tagged DNA library fragment as a template to perform 5'-3' DNA polymerase extension of the multifunctional capture probe; d) performing PCR on the complex processed by the enzyme obtained from c), wherein the tail portion of the multifunctional capture probe molecule is copied to create a hybrid nucleic acid molecule, and the hybrid nucleic acid molecule comprises a target region capable of hybridizing to the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; and e) performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from d).

[0202] In various embodiments, methods for determining the copy number of a specific target region are contemplated. In certain embodiments, a method for determining the copy number of a specific target region comprises: a) hybridizing a tagged DNA library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes with the specific target region in the DNA library; b) isolating the tagged DNA library - multifunctional capture probe module complex obtained from a); c) using the isolated tagged DNA library fragment as a template to perform 5'-3' DNA polymerase extension of the multifunctional capture probe; d) performing a PCR reaction on the complex processed by the enzyme obtained from c), wherein the tail portion of the multifunctional capture probe molecule is replicated to create a hybrid nucleic acid molecule, and the hybrid nucleic acid molecule comprises a target region capable of hybridizing with the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; e) performing PCR amplification of the hybrid nucleic acid in d); and f) quantifying the PCR reaction in e), wherein the quantification enables determination of the copy number of the specific target region.

[0203] In various embodiments, methods for targeted genetic analysis are also contemplated. In certain embodiments, a method for targeted genetic analysis comprises: a) hybridizing a tagged DNA library with a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes with a specific target region in the genomic library; b) isolating the tagged DNA library - multifunctional capture probe module complex obtained from a); (c) producing a tagged DNA target molecule isolated by the hybrid multifunctional capture probe by the concerted action of nick closure by 5'FLAP endonuclease, DNA polymerization and DNA ligase; d) performing PCR on the complex processed by the enzyme obtained from c), wherein the tail portion of the multifunctional capture probe molecule is copied to produce a hybrid nucleic acid molecule, and the hybrid nucleic acid molecule comprises a target region capable of hybridizing with the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; and e) performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from d).

[0204] In various embodiments, methods for determining the copy number of a specific target region are contemplated. In certain embodiments, a method for determining the copy number of a specific target region comprises: a) hybridizing a tagged DNA library to a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes to a specific target region in the DNA library; b) isolating the tagged DNA library - multifunctional capture probe module complex obtained from a); c) generating a tagged DNA target molecule isolated by a hybridized multifunctional capture probe by the concerted action of nick closure by 5' FLAP endonuclease, DNA polymerization, and DNA ligase; d) performing a PCR reaction on the complex processed by the enzyme obtained from c), wherein the tail portion of the multifunctional capture probe molecule is replicated to generate a hybrid nucleic acid molecule, and the hybrid nucleic acid molecule comprises a target region capable of hybridizing to the multifunctional capture probe module and a complement of the multifunctional capture probe module tail sequence; e) performing PCR amplification of the hybrid nucleic acid in d); and f) quantifying the PCR reaction in e), wherein the quantification enables determination of the copy number of the specific target region.

[0205] In certain embodiments, PCR can be performed using any standard PCR reaction conditions well known to those skilled in the art. In certain embodiments, the PCR reaction in e) uses two PCR primers. In one embodiment, the PCR reaction in e) uses a first PCR primer that hybridizes to the target region. In certain embodiments, the PCR reaction in e) uses a second PCR primer that hybridizes to the hybrid molecule at the target region / tail junction. In certain embodiments, the PCR reaction in e) uses a first PCR primer that hybridizes to the target region and a second PCR primer that hybridizes to the hybrid molecule at the target genomic region / tail junction. In certain embodiments, the second primer hybridizes to the target region / tail junction such that at least one or more nucleotides of the primer hybridize to the target region and at least one or more nucleotides of the primer hybridize to the tail sequence. In certain embodiments, the hybrid nucleic acid molecules obtained from step e) are sequenced and the sequences are horizontally aligned, i.e., aligned with each other but not with the reference sequence. In certain embodiments, steps a) to e) are repeated one or more times by one or more multifunctional capture probe module complexes. The multifunctional capture probe complexes may be the same or different and are designed to target any DNA strand of the target sequence. In some embodiments, when the multifunctional capture probe complexes are different, they hybridize near the same target region within a tagged DNA library. In one embodiment, one or more multifunctional capture probes hybridize within about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 bp or more of the target region in a tagged DNA library, including any intervening distances from the target region.

[0206] In some embodiments, the method can be performed using two multifunctional capture probe modules per target region, one of which hybridizes to the "Watson" strand (non-coding or template strand) upstream of the target region and one of which hybridizes to the "Crick" strand (coding or non-template strand) downstream of the target region.

[0207] In certain embodiments, the methods contemplated herein can be further performed multiple times by any number of multifunctional probe modules, e.g., using 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more multifunctional capture probe modules per target region, any number of which hybridize to the Watson or Crick strand in any combination. In some embodiments, the resulting sequences can be aligned with each other to identify any number of differences.

[0208] In certain embodiments, one or more multifunctional probe modules are used to interrogate multiple target regions in a single reaction, e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 10000, 50000, 100000, 500000 or more are interrogated.

[0209] Copy number can provide useful information regarding unique and duplicate reads and can assist in the search for variants of known reads. As used herein, the terms "read", "read sequence", or "sequencing read" are used synonymously and refer to a polynucleotide sequence obtained by sequencing a polynucleotide. In certain embodiments, DNA tags, e.g., random nucleotide tags, are used to determine the copy number of the nucleic acid sequences being analyzed.

[0210] In one embodiment, the multifunctional capture probe hybrid module comprises: (i) a first region capable of hybridizing with a partner oligonucleotide and functioning as a PCR primer; and (ii) a second region capable of hybridizing with a specific genomic target region.

[0211] In various embodiments, the first region of the multifunctional capture probe hybrid module comprises a PCR primer sequence. In certain embodiments, this first region comprises a PCR primer sequence of about 5 to about 200 nucleotides, about 5 to about 150 nucleotides, about 10 to about 100 nucleotides, about 10 to about 75 nucleotides, about 10 to about 50 nucleotides, about 10 to about 40 nucleotides, about 20 to about 40 nucleotides, or about 20 to about 30 nucleotides, including any intervening number of nucleotides.

[0212] In certain embodiments, the first region of the multifunctional capture probe hybrid module is bound to the partner oligonucleotide. In certain embodiments, the multifunctional capture hybrid probe module is bound to the partner oligonucleotide prior to the formation of the tagged DNA library - multifunctional capture probe hybrid module complex. In certain embodiments, the multifunctional capture probe hybrid module is bound to the partner oligonucleotide after the formation of the tagged DNA library - multifunctional capture probe hybrid module complex. In some embodiments, the multifunctional capture probe hybrid module is bound to the partner oligonucleotide simultaneously with the formation of the tagged DNA library - multifunctional capture hybrid probe module complex. In some embodiments, the partner oligonucleotide is chemically modified.

[0213] In various embodiments, the methods contemplated herein include copying a captured tagged DNA library sequence to produce a multifunctional capture probe hybrid module complex and a hybrid nucleic acid molecule that includes a sequence complementary to a region of the captured tagged DNA library sequence located 3' or 5' of the multifunctional capture probe sequence relative to the position where the hybrid module hybridizes to the genomic target, such that PCR can be performed on the tagged DNA library - multifunctional capture probe hybrid module complex. In certain embodiments, the copied target region is located anywhere from 1 to 5000 nt from the 3' or 5' end of the sequence at the position where the multifunctional capture probe hybrid module hybridizes to the genomic target. In certain embodiments, to produce the hybrid nucleic acid molecule, the complementary sequence of the region 3' relative to the position where the multifunctional capture probe hybrid module hybridizes is copied. The hybrid nucleic acid molecule produced includes the multifunctional capture probe hybrid module and a complement of a region of the captured tagged DNA library sequence located 3' or 5' from the position where the multifunctional capture probe hybrid module hybridizes to the target region.

[0214] In various embodiments, the methods contemplated herein include processing a tagged DNA library - multifunctional capture probe module complex to create a hybrid nucleic acid molecule (i.e., a tagged DNA target molecule isolated by a hybrid multifunctional capture probe). In certain embodiments, the hybrid nucleic acid molecule includes a multifunctional capture probe hybrid module and a complement of a region of the tagged DNA library sequence that is located 3' to the position where the multifunctional capture probe hybrid module hybridizes to the target region. In one non-limiting embodiment, the hybrid nucleic acid molecule is created by 3'-5' exonuclease enzyme processing to remove the single-stranded 3' end from the isolated tagged DNA library - multifunctional capture probe module complex and / or 5'-3' DNA polymerase extension of the multifunctional capture probe.

[0215] In other certain embodiments, the hybrid nucleic acid molecule includes a multifunctional capture probe hybrid module and a complement of a region of the tagged DNA library sequence that is located 5' to the position where the multifunctional capture probe hybrid module hybridizes to the target region. In one non-limiting embodiment, the hybrid nucleic acid molecule is created by the concerted action of 5' FLAP endonuclease, DNA polymerization, and nick closure by DNA ligase.

[0216] In various embodiments, methods for targeted genetic analysis are provided. In one embodiment, a method for targeted genetic analysis comprises: a) hybridizing a tagged DNA library to a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes to a specific target region in the DNA library; b) isolating the tagged DNA library-multifunctional capture probe hybrid module complex obtained from a); c) performing PCR on the complex obtained from b) to form a hybrid nucleic acid molecule; and d) performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from c). In certain embodiments, the hybrid nucleic acid molecule obtained from step c) is sequenced and the sequences are horizontally aligned, i.e., aligned with each other but not with a reference sequence. In certain embodiments, steps a)-c) are repeated one or more times with one or more multifunctional capture probe modules.

[0217] The multifunctional capture probe modules may be the same or different and are designed to hybridize to either strand of the genome. In some embodiments, when the multifunctional capture probe modules are different, they hybridize to any position from 1 to 5000 nt of the same target region in the tagged DNA library.

[0218] In certain embodiments, the method can be performed twice using two multifunctional capture probe modules, one of which hybridizes to the upstream of the genomic target region (i.e., at the 5' end; i.e., the forward multifunctional capture probe module or complex), and the other of which hybridizes to the downstream of the genomic target region on the opposite genomic strand (i.e., at the 3' end; i.e., the reverse multifunctional capture probe module or complex).

[0219] In one embodiment, one or more multifunctional capture probes hybridize within about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 bp or more of a target region in a tagged DNA library, including any intervening distances from the target region.

[0220] In some embodiments, the method can be further performed multiple times by any number of multifunctional probe modules, e.g., using 2, 3, 4, 5, 6, 7, 8, 9, 10 or more multifunctional capture probe modules per target region, any number of which hybridize to the Watson or Crick strand in any combination.

[0221] In certain embodiments, one or more multifunctional probe modules are used to interrogate multiple target regions in a single reaction, e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 10000, 50000, 100000, 500000 or more are interrogated.

[0222] In certain embodiments, to identify mutations, the sequences obtained by the method can be aligned to each other without aligning to a reference sequence. In certain embodiments, the obtained sequences can be aligned to a reference sequence, if desired.

[0223] In various embodiments, methods for determining the copy number of a specific target region are contemplated. In certain embodiments, a method for determining the copy number of a specific target region comprises: a) hybridizing a tagged DNA library to a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes to the specific target region in the DNA library; b) isolating the tagged DNA library - multifunctional capture probe hybrid module complex obtained from a); c) performing PCR on the complex obtained from b) to form a hybrid nucleic acid molecule; d) performing PCR amplification of the hybrid nucleic acid in c); and e) quantifying the PCR reaction in d), wherein the quantification enables determination of the copy number of the specific target region. In certain embodiments, the PCR can be performed using any standard PCR reaction conditions well known to those skilled in the art. In certain specific embodiments, the PCR reaction in d) uses two PCR primers. In certain embodiments, the PCR reaction in d) uses two PCR primers, each of which hybridizes to a region downstream of the position where the multifunctional capture probe hybrid module hybridizes to the tagged DNA library. In a further embodiment, the region to which the PCR primer hybridizes is located in the region amplified in step c). In various embodiments, the hybrid nucleic acid molecule obtained from step c) is sequenced and the sequences are horizontally aligned, i.e., aligned with each other but not with a reference sequence. In certain embodiments, steps a) - c) are repeated one or more times by one or more multifunctional capture probe modules. The multifunctional capture probe modules may be the same or different and are designed to hybridize to either strand of the genome.

[0224] In one embodiment, one or more multifunctional capture probes hybridize within about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 bp or more of a target region in a tagged DNA library, including any intervening distance from the target region.

[0225] In some embodiments, the method can be further performed multiple times by any number of multifunctional probe modules, e.g., using 2, 3, 4, 5, 6, 7, 8, 9, 10 or more multifunctional capture probe modules per target region, any number of which hybridize to the Watson or Crick strand in any combination.

[0226] In certain embodiments, one or more multifunctional probe modules are used to collate multiple target regions in a single reaction, e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 10000, 50000, 100000, 500000 or more are collated.

[0227] In certain illustrative embodiments, a tagged DNA library is amplified, e.g., by PCR, to create an amplified tagged DNA library.

[0228] All genomic target regions will have 5' and 3' ends. In certain embodiments, the methods described herein can be performed by two multifunctional capture probe complexes that result in amplification of the targeted genomic region from both the 5' and 3' directions. In one embodiment, one or more multifunctional capture probes hybridize within about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 bp or more of the target region in a tagged DNA library, including any intervening distances from the target region.

[0229] In some embodiments, the method can be further performed multiple times by any number of multifunctional probe modules, e.g., using 2, 3, 4, 5, 6, 7, 8, 9, 10 or more multifunctional capture probe modules per target region, where any number of them hybridize to the Watson or Crick strand in any combination.

[0230] In certain embodiments, one or more multifunctional probe modules are used to collate multiple target regions in a single reaction, e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 10000, 50000, 100000 or more are collated.

[0231] In certain embodiments, the targeted genetic analysis is a sequence analysis. In certain embodiments, the sequence analysis includes any analysis in which one sequence is distinguished from a second sequence. In various embodiments, the sequence analysis excludes any purely mental sequence analysis that is performed in the absence of a composition or method for sequencing. In certain embodiments, sequence analysis includes, but is not limited to, sequencing, single nucleotide polymorphism (SNP) analysis, gene copy number analysis, haplotype analysis, mutation analysis, methylation state analysis (determined, for example, by bisulfite conversion of unmethylated cytosine residues), targeted resequencing of DNA sequences obtained in chromatin immunoprecipitation experiments (CHIP-seq), paternity testing in the sequences of captured fetal DNA collected from maternal plasma DNA during pregnancy, microbial presence and population assessment in samples captured by microbial-specific capture probes, and fetal genetic sequence analysis (e.g., using fetal cells or extracellular fetal DNA in maternal samples).

[0232] Examples of copy number analysis include, but are not limited to, analysis that tests the copy number of mutations that occur in a particular gene or a given genomic DNA sample, and can further include the quantitative determination of the copy number of a given gene or the difference in sequences in a given sample.

[0233] Methods for array alignment analysis that can be performed without the need to align with a reference array are also contemplated herein, which are referred to herein as horizontal array analysis (e.g., illustrated in FIG. 20). Such analysis can be performed on any array produced by the methods contemplated herein or any other method. In certain embodiments, the array analysis includes performing sequence alignment on hybrid nucleic acid molecules obtained by the methods contemplated herein. In one embodiment, one or more multifunctional capture probes hybridize within about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 bp or more of a target region in a tagged DNA library, including any intervening distances from the target region.

[0234] In some embodiments, the method can be further performed multiple times by any number of multifunctional probe modules, e.g., using 2, 3, 4, 5, 6, 7, 8, 9, 10 or more multifunctional capture probe modules per target region, any number of which hybridize to the Watson or Crick strand in any combination.

[0235] In certain embodiments, one or more multifunctional probe modules are used to collate multiple target regions in a single reaction, e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 10000, 50000, 100000 or more are collated.

[0236] In certain embodiments, DNA can be isolated from any biological source. Exemplary sources of DNA include, but are not limited to, blood, skin, hair, hair follicles, saliva, oral mucosa, vaginal mucosa, sweat, tears, epithelial tissue, urine, semen, seminal fluid, seminal plasma, prostatic fluid, bulbourethral fluid (Cowper's gland fluid), excreta, biopsy, ascites, cerebrospinal fluid, lymph fluid or tissue extract samples or biopsy samples.

[0237] In one embodiment, a tagged DNA library for use by the methods contemplated herein is provided. In some embodiments, the tagged DNA library comprises tagged genomic sequences. In certain embodiments, each tagged DNA sequence comprises: i) fragmented and end-repaired DNA; ii) one or more random nucleotide tag sequences; iii) one or more sample code sequences; and iv) one or more PCR primer sequences.

[0238] In one embodiment, a hybrid-tagged DNA library is contemplated. In certain embodiments, the hybrid-tagged DNA library comprises hybrid-tagged DNA sequences. In certain embodiments, each hybrid-tagged DNA sequence comprises: i) fragmented and end-repaired DNA comprising a target region; ii) one or more random nucleotide tag sequences; iii) one or more sample code sequences; iv) one or more PCR primer sequences; and v) a multifunctional capture probe module tail sequence.

[0239] In various embodiments, kits and compositions of reagents for use in the methods contemplated herein are provided. In some embodiments, the composition comprises a tagged DNA library, a multifunctional adapter module, and a multifunctional capture probe module. In certain embodiments, the composition comprises a tagged genomic library. In certain embodiments, the composition comprises a hybrid-tagged genomic library.

[0240] In various embodiments, a reaction mixture for performing the methods contemplated herein is provided. In certain embodiments, the reaction mixture is a reaction mixture for performing any of the methods contemplated herein. In certain embodiments, the reaction mixture can generate a tagged DNA library. In some embodiments, the reaction mixture capable of generating a tagged DNA library comprises a) fragmented DNA and b) a DNA end repair enzyme for generating end-repaired fragmented DNA. In certain embodiments, the reaction mixture further comprises a multifunctional adapter module. In various embodiments, the reaction mixture further comprises a multifunctional capture probe module. In certain embodiments, the reaction mixture further comprises an enzyme having 3'-5' exonuclease activity and PCR amplification activity.

[0241] In various embodiments, methods for DNA sequence analysis are provided for the sequence of one or more clones contemplated herein. In one embodiment, the method comprises obtaining one or more or a plurality of tagged DNA library clones, each clone comprising a first DNA sequence and a second DNA sequence, the first DNA sequence comprising a targeted DNA sequence and the second DNA sequence comprising a capture probe sequence; performing a paired-end sequencing reaction on the one or more clones to obtain one or more sequencing reads; or performing a sequencing reaction on the one or more clones to obtain a single long sequencing read that exceeds about 100, 200, 300, 400, 500 or more nucleotides, the read being sufficient to identify both the first DNA sequence and the second DNA sequence; and ordering or clustering the sequencing reads of the one or more clones according to the probe sequences of the sequencing reads.

[0242] Array reads can be compared to one or more human reference DNA sequences. Array reads that do not match the reference sequence can be identified and used to create a de novo assembly from the non-matching sequence data. In certain embodiments, the de novo assembly is used to identify novel sequence rearrangements associated with the capture probes.

[0243] In various embodiments, a method for copy number determination analysis is provided that includes obtaining one or more or a plurality of clones, each clone including a first DNA sequence and a second DNA sequence, the first DNA sequence including a random nucleotide tag sequence and a targeted DNA sequence, and the second DNA sequence including a capture probe sequence. In related embodiments, a paired-end sequencing reaction is performed on the one or more clones to obtain one or more sequencing reads. In another embodiment, a sequencing reaction is performed on the one or more clones to obtain a single long sequencing read that is greater than about 100 nucleotides, the read being sufficient to identify both the first DNA sequence and the second DNA sequence. The sequencing reads of the one or more clones can be ordered or clustered according to the probe sequences of the sequencing reads.

[0244] In certain embodiments, a method for determining copy number is provided. In certain embodiments, the method includes obtaining one or more or a plurality of clones, each clone including a first DNA sequence and a second DNA sequence, the first DNA sequence including a random nucleotide tag sequence and a targeted DNA sequence, and the second DNA sequence including a capture probe sequence, and ordering or clustering the sequencing reads of the one or more clones according to the probe sequences of the sequencing reads. In certain embodiments, the random nucleotide tag is about 2 to about 50 nucleotides in length.

[0245] The method can further include determining unique sequencing reads and the distribution of overlapping sequencing reads to analyze any sequencing reads associated with a second read array; counting the number of times unique reads are encountered; fitting the frequency distribution of unique reads to a statistical distribution; estimating the total number of unique reads; and normalizing the estimated total number of unique reads against the assumption that humans are generally diploid.

[0246] In certain embodiments, the methods contemplated herein can be used to calculate the inferred copy number of one or more targeted loci, and if any, calculate the deviation of this calculation from the expected copy number value. In certain embodiments, one or more targeted loci of a gene are grouped together in a collection of loci, and the copy number measurements of the collection of targeted loci are averaged and normalized. In one embodiment, the inferred copy number of a gene can be represented by the normalized average of all targeted loci representing this gene.

[0247] In various embodiments, the compositions and methods contemplated herein are also applicable to the generation and analysis of RNA expression. Without wishing to be bound by any particular theory, either the methods and compositions used to generate the tagged gDNA library can be used to generate the tagged cDNA library, and it is contemplated that the target regions corresponding to the RNA sequences embodied in the cDNA for subsequent RNA expression analysis can be captured and processed, including without limitation sequence analysis.

[0248] In various embodiments, a method for generating a tagged RNA expression library includes first obtaining or preparing a cDNA library. Methods for cDNA library synthesis are known in the art and may be applicable to various embodiments. The cDNA library can be prepared from one or more same or different cell types, depending on the application. In one embodiment, the method includes fragmenting the cDNA library, treating the fragmented cDNA library with a terminal repair enzyme to produce fragmented and end-repaired cDNA, and ligating a multifunctional adapter molecule to the fragmented and end-repaired cDNA to produce a tagged RNA expression library.

[0249] In certain embodiments, a tagged RNA expression library (cDNA library) is prepared by obtaining or preparing a cDNA library from total RNA of one or more cells, fragmenting the cDNA library, treating the fragmented cDNA with a terminal repair enzyme to produce fragmented and end-repaired cDNA, and ligating a multifunctional adapter molecule to the fragmented and end-repaired cDNA to produce a tagged RNA expression library.

[0250] In certain embodiments, the cDNA library is an oligo dT-primed cDNA library.

[0251] In certain embodiments, the cDNA library is primed with a random oligonucleotide containing about 6 to about 20 random nucleotides. In certain preferred embodiments, the cDNA library is primed with a random hexamer or a random octamer.

[0252] To achieve the desired average library fragment size, the cDNA library can be sheared or fragmented using known methods. In one embodiment, the cDNA library is fragmented to an average size of about 250 bp to about 750 bp. In certain embodiments, the cDNA library is fragmented to an average size of about 500 bp.

[0253] In various embodiments, the RNA expression libraries contemplated herein can be captured, processed, amplified, sequenced, etc., using any of the methods contemplated herein for capturing, processing, and sequencing tagged genomic DNA libraries, with or without minor variations.

[0254] In one embodiment, a method for targeted gene expression analysis is provided, comprising hybridizing a tagged RNA expression library to a multifunctional capture probe module complex, wherein the multifunctional capture probe module selectively hybridizes to a specific target region in the tagged RNA expression library; isolating the tagged RNA expression library-multifunctional capture probe module complex; performing 3'-5' exonuclease enzyme processing and / or 5'-3' DNA polymerase extension on the isolated tagged RNA expression library-multifunctional capture probe module complex; performing PCR on the enzymatically processed complex, wherein the tail portion of the multifunctional capture probe molecule (e.g., the PCR primer binding site) is copied to create a hybrid nucleic acid molecule that comprises a complement of the target region, a specific multifunctional capture probe sequence, and a capture module tail sequence; and performing targeted gene expression analysis on the hybrid nucleic acid molecule.

[0255] In one embodiment, a method for targeted gene expression analysis comprises hybridizing a tagged RNA expression library to a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module selectively hybridizes to a specific target region in the RNA expression library; isolating the tagged RNA expression library-multifunctional capture probe hybrid module complex; and performing PCR on the complex to form a hybrid nucleic acid molecule.

[0256] In certain embodiments, at least two different multifunctional capture probe modules are used in at least two hybridization steps, each of which uses one multifunctional capture probe module. In certain embodiments, at least one multifunctional capture probe module hybridizes to the 5' of the target region and at least one multifunctional capture probe module hybridizes to the 3' of the target region.

[0257] In one embodiment, one or more multifunctional capture probes hybridize within about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 bp or more of the target region in a tagged RNA expression or cDNA library, including any intervening distance from the target region.

[0258] In some embodiments, the method can be further performed multiple times with any number of multifunctional probe modules, e.g., using 2, 3, 4, 5, 6, 7, 8, 9, 10 or more multifunctional capture probe modules per target region, any number of which hybridize to the Watson or Crick strand in any combination.

[0259] In certain embodiments, one or more multifunctional probe modules are used to interrogate multiple target regions in a single reaction, for example, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 10000, 50000, 100000 or more are interrogated.

[0260] In further embodiments, methods for cDNA sequence analysis are provided that enable one of ordinary skill in the art to perform gene expression analysis from a cDNA library. In certain embodiments, any of the sequencing methods contemplated herein can be adapted to the sequencing of a cDNA library with little or no deviation from its application to the sequencing of tagged genomic clones. As described above, the statistical distribution of tagged cDNA sequencing reads of target regions of cDNA in RNA expression analysis contemplated herein correlates with the level of gene expression of the target regions in the cells from which the cDNA library was prepared or obtained.

[0261] All publications, patent applications, and issued patents cited herein are hereby incorporated by reference as if each individual publication, patent application, or issued patent were specifically and individually indicated to be incorporated by reference.

[0262] For clarity of understanding, the foregoing invention has been described in detail by way of illustration and example, and it will be readily apparent to one of ordinary skill in the art that certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims in view of the teachings of the present invention. The following examples are presented by way of illustration and not by way of limitation. One of ordinary skill in the art will readily recognize various non-critical parameters that can be changed or modified to obtain essentially the same results.

Example

[0263] (Example 1) Preparation of Target Genomic Regions for Gene Analysis Summary In certain embodiments, the methods contemplated herein involve the organized utilization of several key molecular modules. Each module is described separately in the following sections. At the end of this section, the interrelationships of the modules are described.

[0264] Section 1: Tagging of Genomic DNA Fragments Genomic DNA obtained from an individual can be recovered, processed into pure DNA, and fragmented random nucleotide sequences of 1 nucleotide or more, in some embodiments, in the range of 2 - 100 nucleotides, or in the range of 2 - 6 nucleotides, are ligated to the random ends of the genomic DNA fragments (Figure 1). The combination of the introduced random nucleotide tag sequences with the genomic fragment end sequences constitutes a unique combination of two elements, which is hereinafter referred to as the first region of the multifunctional adapter module in the following of this specification. The uniqueness of the first region of the multifunctional adapter module is determined by the product of the diversity of the genomic fragment end sequences multiplied by the diversity within the pool of the first regions of the ligated multifunctional adapter modules.

[0265] Section 2: Addition of Sample-Specific Codes and Universal Amplification Sequences The multifunctional adapter molecule may further include a sample-specific code (referred to herein as the second region of the multifunctional adapter module) and a universal amplification sequence (referred to herein as the PCR primer sequence or the third region of the multifunctional adapter module). In addition to the introduced random nucleotides obtained from the first region of the multifunctional adapter module, each segment bound to fragmented genomic DNA may also include an additional set of nucleotides that are common to each sample but differ between samples, such that the DNA sequence of this region can be used to uniquely identify (i.e., barcode) a given sample sequence within a set of sequences in which multiple samples are combined together. Additionally, the bound nucleotide sequence may also contain a universal sequence that can be used to amplify (e.g., by PCR) the polynucleotide. The combination of the random nucleotide tag sequence, the sample code, and the elements of the universal amplification sequence most commonly constitutes an "adapter" (also referred to as a multifunctional adapter module) that is bound to fragmented genomic DNA by nucleotide ligation.

[0266] An illustrative example of a multifunctional adapter module ligated to fragmented gDNA is illustrated in FIG. 1, and an exemplary set of such sequences is shown in Table 1. In Table 1, the set of adapter sequences is clustered into sets of four adapter sequences. Within each column, all adapters sharing the same two base codes and all of the possible 16 random tags are represented. The possible 16 adapters are mixed prior to ligation to the fragment. Only the "ligation strand", which is the top strand of each adapter, is shown, as this is the strand that is covalently bound to the end-repaired DNA fragment. The partner strand, which is ultimately lost, is shown in FIG. 1 but not incorporated into Table 1.

[0267] [Table 1-1]

Table 1-2

[0268] The application of a single set of adapters (i.e., a set of a universal amplification array, a sample-specific code, and random tags; also referred to as a multifunctional adapter module) with a single amplification array to both ends of genomic fragments has several significant advantages, including the fact that it tags the same genomic fragment independently at its two ends. As described in some of the following sections, the two strands by any given fragment will ultimately be separated from each other and will behave as independent molecules in the present invention. Thus, the presence of two different tags at the two ends of the same fragment is an advantage rather than a drawback of the present invention. Additionally, the ligation event between adapters is also a major problem in the construction of next-generation libraries, where the initial goal is to create amplicons with heterogeneous ends. When using the method of the present invention, it is the latter half of the process that introduces this asymmetry, and thus, in the present invention, identical ends are also acceptable. An unexpected and surprising benefit of this method is that adapter dimers are not observed in the library construction method of the present invention. Without being bound by theory, the inventors assume that this is because rapidly formed, rare adapter dimer molecular species form tight hairpin structures in the denaturation and annealing steps required for PCR amplification, and it is further assumed that these hairpin structures are completely resistant to further amplification directed by primers. The possibility of creating an adapter dimer-free library is an important technical feature in extremely low-input applications such as single-cell to few-cell genomic analysis, circulating DNA analysis (for applications such as fetal diagnosis, monitoring of tissue transplant rejection, or cancer screening), or single-cell transcriptome analysis. This method itself provides significant utility in such applications. An even more remarkable feature of the amplicon with a single primer is that it is possible to turn on the amplification with a 25-nucleotide PCR primer and turn off the amplification with a longer 58-nucleotide primer. This will be described in more detail in Section 6-5 below, and the importance to the present invention will also be emphasized.

[0269] Overview By an adapter strategy that uses a single universal amplification array at both ends of the target fragment, problems regarding adapter dimers are resolved. This is clearly demonstrated, for example, in Example 3: Construction of a genomic library with a single adapter.

[0270] Section 3: Quantification by library A further aspect of this method for genomic analysis strategies is that the "coverage depth" is known, i.e., the average genomic copy number present in the library is known or can be determined. The coverage depth is measured using the purified ligation reaction product before bulk amplification of the library required for subsequent steps. For illustrative purposes, if 50 genome equivalents of DNA are introduced into the library scheme of an embodiment of the present invention and ligation of the adapter to both ends of the fragment is done with 100% efficiency, each adapter end acts independently of the other adapter end, and thus, since 2 ends × 50 genomes = 100 coverage, the coverage depth is 100. In the universal PCR primers envisioned herein, the simple fact that adapter dimers are not amplified and fragments adapter-treated at both ends are amplified means that quantification by library becomes simply a matter of measuring the complexity of the library by quantitative PCR (qPCR) using universal primers and calibrating the results against a reference substance for which the coverage depth is known. In this specification, the terms "genomic copy" and "coverage depth" mean the same thing and can be used interchangeably. This method feeds a coverage depth of 4 to 1000, preferably 20 to 100 times, into sample processing according to the present invention, which is the next phase.

[0271] Section 4: Amplification by library In certain embodiments, a portion of the adapter-ligated genomic fragment library corresponding to a coverage depth of 20 to 100-fold will be amplified using standard PCR methods with a single universal primer sequence that drives the amplification. In certain embodiments, it is advantageous at this stage to convert a few picograms of material in the initial library into several micrograms of amplified material, which implies a 10,000-fold amplification.

[0272] Section 5: Hybridization of Target Library Fragments with Capture Probes Advances in oligonucleotide synthesis chemistry have created new opportunities for sophisticated genomic capture strategies. In particular, long oligonucleotides (100 - 200 nucleotides in length), which now have reasonable synthesis costs per base, relatively high yields, and excellent base accuracy, are commercially available from various vendors. This ability allows the inventors to create multifunctional capture probes (Figure 2). Exemplary elements of a multifunctional capture probe include the following. Region 1 is a 34-nucleotide region common to all probes and includes a region that hybridizes with a modified complementary oligonucleotide (also referred to as a partner oligonucleotide). This modified oligonucleotide further includes a biotin-TEG modification, which is biotin that can tightly bind to streptavidin protein at the 5' end and has a long hydrophilic spacer arm that alleviates steric hindrance to biotin binding. At the 3' end, the oligonucleotide terminates with a dideoxycytosine residue that renders this partner oligonucleotide inert to primer extension. This element of the probe design allows for adapter treatment of a non-limiting number of probes with biotin capture functionality without directly modifying the probe.

[0273] Region 2 contains a custom 60-nucleotide region that is target-specific and interacts with the gDNA fragment molecule. This region, which is designed via a computer-based method, is unique in the genome, promotes consideration of the presence of common SNPs that can impair binding efficiency, and secondary structure.

[0274] Region 3 contains a 20-nucleotide segment that is used as a PCR primer binding site in subsequent fragment amplification. This feature is described in more detail in the following paragraphs. The multiplicity of probes can be used to capture the genomic region of interest (multiplexing of probes). At least two types of probes can be employed to fully query a typical coding exon 100-150 bp in length. As an example, it is indicated that this would enable capturing 10 typical exon genes using 20 probes and screening a 100-gene panel using a total of 2000 probes. Hybridization of genomic library fragments to the probes can be performed by reannealing following heat denaturation. In one embodiment, the steps include the following.

[0275] 1. Combining the genomic library fragments with a pooled probe sequence (in this case, "probe sequence" refers to the combination of an equimolar amount of each individual probe with a highly modified partner oligonucleotide) at a specific target-to-probe ratio ranging from 1 target per 1 probe to 1 target per 1,000,000 probes. In one embodiment, the optimal ratio is approximately 1 fragment per 10,000 probes.

[0276] 2. Heating the combined fragment + probe in a solution containing 1 M NaCl, 10 mM Tris, pH 8.0, 1 mM EDTA, and 0.1% Tween 20 (a non-ionic detergent) to 95°C for >30 seconds to denature all double-stranded DNA structures.

[0277] 3. Step of controlling the combined probes and fragments stepwise, for example, cooling by decreasing the temperature by 1 °C every 2 minutes until <60 °C. This slow cooling will result in the formation of double-strands between the target genomic fragments and the probe array.

[0278] 4. Step of binding the probe:fragment complex to paramagnetic beads coated with carboxyl and modified with streptavidin, and "pulling out" these beads using a strong magnet.

[0279] 5. Step of washing the bound complex with a solution containing 25 (v / v)% formamide, 10 mM Tris, pH 8.0, 0.1 mM EDTA, and 0.05% Tween 20. In certain embodiments, the washing step is performed at least twice.

[0280] 6. Step of resuspending the washed beads in a solution suitable for subsequent enzyme processing steps.

[0281] Capture reaction Embodiments of the capture reaction are presented in Example 3 (Construction of genomic libraries with single adapters), developed in Example 5 (Verification of the PLP1 qPCR assay), and further described, invoking the qPCR assay.

[0282] Section 6: Enzyme processing of the hybridized probe:target complex As currently practiced in the art, hybridization-based array capture methods generally result in suboptimal enrichment of target sequences. From the literature and commercially available publications, it can be estimated that, at most, only about 5% - 10% of the reads map to their intended target sequences. The remaining reads often map in the vicinity of the intended target, and commercial vendors have redefined "hits" as reads that fall within about 1000 bases of the intended locus. The reason for this "spreading" effect is not fully understood and is likely the result of normal sequence hybridization events (see, for example, Figure 3).

[0283] The enzymatic processing of the complex envisioned herein focuses the captured sequences more sharply on the exact target region. In this step, a DNA polymerase that also possesses 3'-5' exonuclease activity is employed. An exemplary example of such an enzyme is T4 DNA polymerase. This enzyme will "chew back" the dangling tail sequences from the double-stranded regions formed between the probe and the target sequences. T4 DNA polymerase will then copy the tail segments on the probe. See, for example, Figure 4. The benefits provided by this step include, but are not limited to, the following.

[0284] 1. By employing this type of enzymatic processing, only the fragments that hybridize directly to the probe and form double-stranded structures are advanced. The final sequencing library is a chimeric (hybrid) set of molecules obtained from both the fragments and the probe.

[0285] 2. The probes are strand-specific, and thus the captured target has a unique orientation with respect to the probe (illustrated in Figure 5). This means that only one of the two strands made from a single fragment interacts with the probe, and processing focuses the read to the 5’ region of the probe array. In this regard, the complementary strand of the fragment becomes a completely independent molecular species. By placing the directed probe on one side of the target region (e.g., an exon), the technique enables highly specific focusing of sequencing reads on the target region (Figure 6).

[0286] 3. Target molecules that hybridize normally to the target fragment (but do not cross-hybridize to the probe; Figure 3) do not acquire the essential probe sequence and are thus lost in subsequent amplification steps.

[0287] 4. The actual “tail” sequence of the probe is copied to the target fragment as part of the amplification sequence. All commercially available sequencing platforms capable of execution (e.g., a sequencing platform based on reversible terminator chemistry manufactured by Illumina) require a sequencing library in which the target fragments have asymmetric ends, which are often referred to as “forward” adapter sequences and “reverse” adapter sequences, or in the jargon of sequencing laboratories, often as “P1” and “P2”. In certain embodiments, up to this point, the fragment library envisioned herein has a single molecular species at the ends and is designated as “P1”. The enzymatic processing step accomplishes two things. First, the enzymatic processing step “erases” one of these P1 ends (by 3’-5’ exonuclease activity). Second, the enzymatic processing step “adds” a base of a P2 end that is heterologous to P1 (via copying of the probe tail sequence by DNA polymerase).

[0288] 5. Target molecules enzymatically modified at the regular P1-P2 termini can be selectively enriched in the subsequent PCR amplification step following processing. This is achieved by the use of long PCR primers. In particular, long primers are necessary to add sufficient functionality required for next-generation sequencing and also impart selectivity to the amplification. Residual P1-P1 library fragments, which are "contaminants" obtained from the first round of amplification, are not amplified with long P1 primers. This is a significant advantage of the method. The initial P1-P1 library is effectively amplified with a single, 25-nucleotide PCR primer. Once the length of this primer is extended up to 57 nucleotides (with the addition of sequencing functionality), these same P1-P1 molecules are not amplified to any extent. Thus, the amplification of the initial library can be turned "on" with a 25-nucleotide primer and turned "off" with a 57-nucleotide primer.

[0289] Overview The non-amplification of the P1-insert-P1 library is demonstrated in Example 3 (Construction of a genomic library with a single adapter). The preferential amplification of the processed DNA fragments of P1-insert-P2 is shown in Example 3 (Construction of a genomic library with a single adapter). Example 3 further demonstrates a substantial improvement in target specificity associated with processing. Finally, Example 9 (Direct measurement of post-capture processing) demonstrates that the "sensitivity" of processing, which means the percentage of the initially processed complexes, is on the order of 10% of all the captured complexes.

[0290] Section 7: Amplification and Sequencing The core adapter array and primer array applied to the initial proof-of-concept experiment are shown in Table 2. The enzymatically processed complex obtained from Step 6 is added directly to a PCR amplification reaction containing a full-length forward PCR primer and a full-length reverse PCR primer. After amplification, the library can be purified, quantified, and loaded onto a high-throughput next-generation sequencer (in this embodiment, the library is configured for a reversible terminator-based platform by Illumina), and sequences of about several million fragments are determined. At this stage, single reads with a length of >36 nucleotides, preferably 72 or 100+ nucleotides, can be observed.

[0291]

Table 2-1

Table 2-2

[0292] Section 8: Data Analysis There are at least two major aspects to data analysis after sequencing. The first aspect is the identification of sequence variants (single nucleotide variants, microinsertions, and / or microdeletions relative to an established set of reference sequences). Although complex, these methods are well described in the art and those skilled in the art will understand such methods. The second aspect is the determination of copy number variations obtained from targeted sequencing data.

[0293] (Example 2) Determination of Copy Number Determination of copy number is used in various ways in the field of DNA sequencing. As a non-limiting example, massively parallel DNA sequencing technologies provide at least two opportunities to examine and analyze biological samples. One well-established aspect is the determination of DNA sequences that represent de novo sequences present in the sample (e.g., sequencing of newly isolated microorganisms) or re-sequencing of known regions for variants (e.g., searching for variants within a known gene). A second aspect of massively parallel sequencing is the potential for quantitative biology and counting the number of times a particular sequence is encountered. This would be a fundamental aspect of techniques such as "RNA-seq" and "CHIP-seq" that use counting to infer, respectively, gene expression or the association of a particular protein with genomic DNA. This example relates to the quantitative and counting-based aspects of DNA sequencing.

[0294] DNA fragments are very often counted as conformations of sequences sharing a high degree of similarity (i.e., the DNA fragments align with specific regions of known genomic sequences). Sequences within these clusters are often identical. It should be noted that DNA sequences with a) reads of different starting and ending DNA sequences, or b) high-quality sequence differences obtained from other reads in the set are often considered "unique reads". Thus, different starting sequence positions and sequence variations are a form of "tagging" used to differentiate unique events obtained from clones. In this example, in the course of library construction, random nucleotide tags (e.g., random 6-nucleotide sequences) are also introduced into genomic fragments. Tags are collectively constructed by combining 1) the random nucleotide tag sequence, 2) the starting point of the DNA sequencing read, and 3) the actual sequence of the read. This tag enables differentiation between convergent events where the same fragment is cloned twice (such fragments would have different random nucleotide tag sequences introduced in library construction) and fragments of the same origin replicated in library amplification (these "clones" would have the same random nucleotide segment and the same clone starting point). This type of tagging enables, among other things, further quantitative analysis for genomic DNA, and more generally, further quantitative analysis for DNA molecules (e.g., RNA-seq libraries).

[0295] By introducing random nucleotide tags (random N-mers combined with DNA clone ends) into a DNA sequencing library, in theory, each unique clone within the library can be identified by its unique tag sequence. By specifying "in theory", we recognize the confounding characteristics of normal experimental datasets that can occur, such as errors in sequencing, errors introduced during amplification by the library, and the introduction of contaminating clones from other libraries. All of these sources of confounding confound the theoretical considerations presented herein. In the context of sequence capture and targeted resequencing, tagging the library can enable quantitative analysis of the locus copy number within the captured library.

[0296] As a non-limiting example, consider a library constructed from an input equivalent to 100 diploid genomes created from a male subject. It is predicted that at each autosomal locus, there will be approximately 200 library clones, and at each X chromosome locus, there will be 100 clones. If the autosomal region is captured and sequenced 2000 times, all 200 tags will be encountered with a confidence interval exceeding 99% certainty. In the X chromosome region, in theory, 2000 reads will reveal a total of 100 tags. As an illustration, in this example, creating DNA tags within a DNA sequencing library supports the general concept that copy number differences can be preserved. This general framework can be applied to the methods described herein. Empirical evidence suggests that it may be necessary to adjust for differences in base cloning efficiency for each locus and for the sporadic introduction of artifact tags from experimental errors and the like. The practical implementation of this concept may vary in different contexts and may involve case-by-case sequence analysis methods, but the general principles summarized herein will underlie all such applications.

[0297] To date, the creation of tagged DNA libraries has been considered in the context of genomic DNA analysis, but it must be emphasized that this concept applies to all counting-based DNA sequencing applications. In certain embodiments, tagging can be applied to RNA-seq by cloning cDNA molecules made from mRNA samples by a method that creates tags. Such techniques can substantially increase the fidelity of sequence-based gene expression analysis. In certain embodiments, it is envisioned that tagging can increase the resolution of chromatin immunoprecipitation (CHIP (chromatin immunoprecipitation)-seq) experiments. In various embodiments, tagging will enhance the quantitative aspects of sequence counting used to determine the presence and abundance of microorganisms within microbiome compartments and environmental samples.

[0298] (Example 3) Construction of Genomic Libraries with a Single Adapter Purpose The goal of this example was to create a genomic DNA library from female hgDNA (≈200 bp) from Promega fragmented by acoustic treatment.

[0299] Summary The results clearly demonstrated the remarkable features of this method for adapter design. In particular, the ligation reaction with the adapter alone did not result in detectable adapter dimer molecular species. As with this method, the limit of input is invariably determined by the background level of adapter dimers, which was extremely important in the context of ultra-low input sequencing library preparation techniques. Attempts to maintain checks against adapter dimer contamination have applied highly specialized techniques. These include size exclusion methods such as column or gel purification, expensive custom oligonucleotide modifications designed to minimize adapter self-ligation events, and adapter sequence modifications that allow destruction of adapter dimers by restriction digestion after library construction.

[0300] Address the adapter dimer problem associated with a simple solution that evokes the basic principles of DNA structural principles, according to the simple, single adapter, single primer concept envisioned in this specification. This ultra-low input technology is useful for constructing genomic libraries for genomic analysis, and is also useful for transcriptome analysis of cloned double-stranded cDNA, for example, in RNA-seq applications for one or several special cells, and would also be useful for rescuing some intact fragments that may be present in highly modified, poorly conserved, formalin-fixed, paraffin-embedded (FFPE) nucleic acid samples.

[0301] Another essential feature of the adapter design of the present invention is the ability to turn "on" and "off" the PCR amplification of the target amplicon library by using different lengths of PCR primers. As clearly demonstrated, the primer length optimal for amplification by the library was a 25-nucleotide primer molecular species with an estimated Tm (under standard ionic strength conditions) of ≧55°C. Shorter, lower Tm primers presented low-efficiency amplification of amplicons and were considered favorable for amplicons with a small average insert size. There are many precedents where primers of this size class work well when paired with reverse primers of heterologous sequences.

[0302] In summary, these data demonstrated that the adapter and PCR amplification method of the present invention produces a fragment library free of adapter dimers with "tunable, on / off" amplification characteristics.

[0303] Method The primers received from IDT were hydrated to 100 μM in TEzero (10 mM Tris, pH 8.0, 0.1 mM EDTA).

[0304] Fragment repair: - 14 μl of water - 5 μl of hgDNA - 2.5 μl of 10x end repair buffer - 2.5 μl of 1 mM dNTP - A mixture of 1 μl of end repair enzyme and 0.5 μl of PreCR enzyme repair mix, mixed and added By combining them, the melted gDNA and 500 ng of gDNA were end-repaired.

[0305] The mixture was incubated at 20 °C for 30 minutes and at 70 °C for 10 minutes and held at 10 °C.

[0306] Adapter annealing: 68 μl of TEzero, 2 μl of 5 M NaCl, 20 μl of oligo 11, and 10 μl of oligo 12 were combined. It was heated to 95 °C for 10 seconds, heated at 65 °C for 5 minutes, and cooled to RT.

[0307]

Table 3

[0308] Ligation: Total volume of 20 μl: - 13 μl or 8 μl of water - 0 or 5 μl of end-repaired fragment = 100 ng. - 2 μl of 10x T4 ligase buffer - 3 μl of 50% PEG8000 - 1 μl of double-stranded 10 μM ACA2 adapter 23 - A mixture of 1 μl of T4 DNA ligase, mixed and added Among them, 1 = hgDNA without insert, 2 = 100 ng of end-repaired hgDNA were combined.

[0309] It was incubated at 23°C for 30 minutes and then at 65°C for 10 minutes. 80 μl of TEz and 120 μl of beads were added per reaction. It was mixed and incubated at RT for 10 minutes. It was washed twice with 200 μl aliquots of 70% EtOH:water (v / v) and resuspended in 50 μl of TEz.

[0310] PCR amplification: 10 μl aliquots of each ligation mix = 20 ng each of the library. It was planned to amplify for 18 cycles.

[0311]

Table 4

[0312] A 600 μl mix containing all components except primers and template was prepared. Six 80 μl aliquots were prepared. A ligation mix without insert (10 μl) was added to set 1 and a 10 μl hgDNA insert was added to set 2. 10 μM of the paired primers shown below were added to the ligation mix without insert and the ligation mix with hgDNA insert. It was mixed. Thermal cycling was carried out for 18 cycles at 94°C for 30 seconds, 60°C for 30 seconds, and 72°C for 60 seconds, and finally thermal cycling was carried out at 72°C for 2 minutes and held at 10°C.

Table 5

[0313] The PCR products were purified with 120 μl of beads. They were washed twice with 200 μl of 70% EtOH. The beads were dried and the DNA was eluted with 50 μl of TEz. 5 μl of each sample was analyzed on a 2% agarose gel.

[0314] Results The exact same gel image is shown in Figure 7 with four different color and contrast schemes. The samples loaded on the gel were 1. Ligation mix of only adapter without insert amplified by ACA2 20 2. Ligation mix of only adapter without insert amplified by ACA2 (normal 25-nucleotide PCR primer) 3. Ligation mix of only adapter without insert amplified by ACA2 FLFP (full-length forward primer) 4. Ligation mix of approximately 200 bp hgDNA insert + adapter amplified by ACA2 20, 20 ng 5. Ligation mix of approximately 200 bp hgDNA insert + adapter amplified by ACA2 (normal 25-nucleotide PCR primer), 20 ng 6. Ligation mix of approximately 200 bp hgDNA insert + adapter amplified by ACA2 FLFP (full-length forward primer), 20 ng as follows.

[0315] For ligation of only adapter → in the PCR products (lanes 1-3), it was clear that the material was not amplified. The short, 20-nucleotide ACA2 primer (lane 4) showed less efficient amplification compared to the "normal", 25-nucleotide ACA2 primer (lane 5). With the 58-nucleotide ACA2 FLFP primer (lane 6), only very faint traces of the material were visible by eye.

[0316] In a further embodiment, it may be useful to titrate with the amount of ACA2 primer and monitor the yield. Normal high-yield PCR primers carry 1 μM each of the forward and reverse primers per 2 μM total primer (per 100 μl of PCR reaction). Thus, by adding ACA2 up to 2 μM (since it is both the forward and reverse primers), the yield can be increased. Similarly, in certain embodiments, it may be useful to monitor the amplification characteristics of the library at primer annealing temperatures lower than 60°C.

[0317] (Example 4) Fragmentation of gDNA Purpose For the initial proof-of-concept experiment, sheared human gDNA obtained from males and females was required. In this example, human female and human male gDNA from Promega was utilized. Based on the amounts shown on the tubes, these were diluted to 1000 μl of DNA at 100 ng / μl and placed under Covaris conditions intended to create fragments in the range of 200 bp.

[0318] Overview There are at least two components in the laboratory research infrastructure. One is the ability to quantify DNA, and the other is the ability to visualize the size distribution of DNA on a gel. In this example, a Qubit 2.0 instrument manufactured by Life Technologies was employed to measure the DNA concentration. The recorded readings were found to be generally smaller than those of our previous experiments using Nanodrop. The Qubit readings were based on dsDNA-specific dye binding and fluorescence. One of the major advantages of Qubit is that it can be used to quantify DNA amplification reactions (e.g., PCR) without prior cleanup. In these experiments, it was found that Promega gDNA, which was thought to be 100 ng / μl, was measured at approximately 60 ng / μl with Qubit. Regarding the qualitative assessment of the gel and size distribution, electrophoresis was performed and it was recorded that the system worked effectively. In this example, it was found that the fragmented gDNA had an average size distribution centered around the desired approximately 200 bp.

[0319] Methods and Results After Covaris treatment, a Qubit instrument was used to measure the DNA concentration. The gDNA was diluted 10-fold and 2 μl was added to an assay solution with a final volume of 200 μl. The readings for both female and male samples were recorded at approximately 60 ng / mL, which means the starting solution was 60 ng / μl. This was below the initial expectation but was sufficiently within the appropriate range for a particular embodiment. Then, we loaded 2 μl (120 ng) and 5 μl (300 ng) of the material both before and after fragmentation onto a 2% agarose gel (Figure 8). The symbols in the top row represent M: male gDNA and F: female gDNA. The symbols in the bottom row represent U: unfragmented, and C: fragmented by Covaris. One important observation was that the average fragment size was an even distribution centered around the vicinity of 200 bp.

[0320] (Example 5) Verification of the PLP1 qPCR Assay Objective The proteolipid protein 1 (PLP1) gene on the X chromosome was examined for an initial proof-of-concept capture study. This gene is implicated in cancer and is located on the X chromosome, which was selected because it means having natural copy variations between males and females. The 187-nucleotide exon 2 region of NM000533.3, the Ref-Seq transcript of PLP1, was used as the target region. For the proof-of-principle study, the ability to monitor regions within and in the vicinity of PLP1 exon 2 by qPCR was required. This example presents the description of the design and validation of eight such assays.

[0321] Summary Eight qPCR assays (in this case, meaning simple primer pairs) were designed to monitor the capture of PLP1 exon 2. That five assays hit means that they are within the region targeted by the capture probe. That two are "near the target" means that one assay is located at a genomic coordinate 200 bp from the target region and one assay is located at a genomic coordinate 1000 bp from the target region on the reverse strand. These two assays were designed to quantify "spreading", a phenomenon in which regions near the target locus are pulled along as "hitchhikers" in the capture experiment. Finally, one assay was also designed for the chromosome 9 region, which is designed to monitor any non-homologous segment of human gDNA. This specification shows by way of example that PCR fragments matching amplicons of the predicted size are generated by all eight assays. The example appropriately shows that the specific activity per 1 ng of input gDNA is greater in females than in males by the PLP1 assay located on the X chromosome. These data verified the validity of the use of these assays in further experiments to monitor the capture of gDNA.

[0322] Methods, Results, and Discussion A 400-bp region centered around the vicinity of PLP1 exon 2 was excised into primer 3 for primers with an average length of 24 nucleotides and a Tm of 60°C to 65°C in order to create amplicons with a length of 80 - 100 bp. The search region was manipulated to obtain primer pairs (amplicons for qPCR) that "walk" from the 5' intron-exon boundary of exon 2, through the CDS, to the 3' exon-intron boundary. Near this, a proximity capture assay was also designed that is distal to exon 2, directed towards exon 3, and located approximately 200 nucleotides to approximately 1000 nucleotides from exon 2. These will be used to monitor "hitchhiker" genomic fragments that are captured in secondary hybridization events. Finally, on chromosome 9, one assay was also created to monitor the bulk genomic DNA level in the experiment. The primer sequences for these assays are shown below, and the details are in the appendix at the end of this example.

[0323]

Table 6-1

Table 6-2

[0324] To verify the performance of the primer pairs, PCR reactions containing male genomic DNA or female genomic DNA as templates were prepared. These were then amplified by real-time PCR on an Illumina Eco instrument or by conventional PCR. By qPCR, it was inferred that females show a slightly stronger PLP1 (X chromosome) signal than males. By conventional PCR, the inventors were able to check that the amplicons were of the correct size and unique. Both tests yielded data consistent with the interpretation that the performance of all eight assays was good.

[0325] Preparation of PCR reactions: For each female PCR reaction or each male PCR reaction, - 100 μl of water - 25 μl of 10× STD Taq buffer - 25 μl of 25 mM MgCl2 - 25 μl of sheared gDNA at 60 ng / μl (same concentration for both female and male as determined by Qubit) - 12.5 μl of DMSO - 12.5 μl of 10 mM dNTP - 6.25 μl of EvaGreen dye (Biotum) - 5 μl of ROX dye (InVitrogen) - 2.5 μl of Taq DNA polymerase, well - mixed and added A 250 - μl master mix containing the above was prepared on ice.

[0326]

Table 7

[0327] For the experiment, 24 μl of the mix was aliquoted into 2 sets of 8 strip tubes (female or male), and 6 μl of primer mix containing 10 μM of forward primer and reverse primer obtained from each assay was added. After mixing, three equal amounts of 5 μl each were aliquoted into the columns of a 48-well Eco PCR plate (three consecutive female samples on the upper side within the column, three consecutive male samples on the lower side within the column). The instrument was set to cycle over 40 cycles of 30 seconds at 95 °C, 30 seconds at 60 °C, and 30 seconds at 72 °C while monitoring SYBR and ROX. The JPG image of the amplification trace for Assay 6 is shown in Figure 9. The copy difference between female and male samples was clear. For female and male samples, all "Cq" values (the value when the fluorescence curve exceeds a certain automatically defined baseline) were accumulated, and then the difference between the average values of the three consecutive measurements was calculated. This is shown in Table 7 above (underline = male - female), where here, all values are positive except for the chromosome 9 assay. From the overall data, the performance of all eight assays was similar (Cq values of 22 - 24), and it was indicated that assays with the X chromosome generally had a stronger signal in females.

[0328] The conventional PCR reaction was cycled over 30 cycles of 30 seconds at 94 °C, 30 seconds at 60 °C, and 30 seconds at 72 °C, a pause of 2 minutes at 72 °C, and a hold at 10 °C. A total of 5 μl of the product was directly loaded onto a 2% agarose gel without purification, which is shown in Figure 10. The upper band of each doublet corresponded to the estimated mobility of the assay PCR product. The lower "fuzzy" material was likely the unused PCR primer.

[0329] From the results of real-time PCR and gel analysis following conventional PCR, it can be concluded that these eight assays are suitable for amplifying only their intended regions and monitoring fragment enrichment.

[0330] Appendix of Example 5: Details on Assay Design

[0331] PLP1 gene: Transcript ID: NM_000533.3; Exon 2: 187 nucleotides; The CDS2 CDS obtained from the UCSC browser is shown in bold uppercase letters with an underline, and the primer sequences are shaded. Flanking sequences are shown in lowercase letters.

Chem.

Chem.

Chem.

[0332] (Example 6) Capture of PLP1 Exon 2 Purpose In one embodiment, the DNA capture strategy by Clearfork Bioscience v1.0 involves the use of multifunctional probes targeted to specific genomic target regions. The goal was to validate a method using an Ultramer™ (Integrated DNA Technologies (IDT), Coralville, IA; Ultramer is a trademark given to specialized synthetic oligonucleotides ranging in length from 45 to 200 nucleotides) targeting PLP1 exon 2.

[0333] Overview In this example, a capture reaction was demonstrated. The Ultramer from IDT-DNA worked well for capture, the basic protocol was reasonable with respect to the stoichiometric ratio of the reagents throughout the capture step, and the molecular crowding agent PEG interfered with effective capture. After capture, subsequent efforts were made for enzymatic processing.

[0334] Brief description A multifunctional probe is schematically illustrated in Figure 2. The goal of this experimental data set was to examine all three features of these probes. Region 1 was the binding site for a 34-nucleotide, 5'-biotin-TEG modified and 3'-dideoxycytosine modified universal "pull-down" oligo. Two of these universal regions were designed to validate (validate / verify) equivalent (desirable) performance.

[0335] The sequences of these two universal oligos are shown in Table 8 below.

[0336]

Table 8

[0337] The following is a brief description of how these sequences were selected. The functional role of these oligos was to hybridize with the capture probe, thereby resulting in a stably bound biotin overhang that could be used to capture on magnetic beads modified with streptavidin.

[0338] Ten random sequences were generated by a random DNA sequence generation set to have a GC base composition of approximately 50%. The website used was www.faculty.ucr.edu / mmaduro / random.htm. Then, the ten sequences were subjected to BLAT against the hg19 of the human genome Screened against build. Only Array 3 showed significant alignment. Two sequences ending with "C" were selected because they could be blocked by ddC. Both sequences were analyzed by the IDT OligoAnalyzer. Array 1 has a GC of 47% and a melting temperature of 76 °C in 1 M NaCl. The GC content of Array 2 is 57% and the melting temperature in high salt concentration is 86 °C. Sequences 1 and 10, the selected sequences, are probe sequences complementary to the actual "universal" 5'-biotin TEG-ddC. The reverse complement strands of these were used as tails on the capture probes. Subsequently, these sequences were varied by adding the four bases, A, G, C, and T, and extended up to 34 bases in length. This length worked well for SBC and no firm reason to vary it was found. Second, some of the CGCG-type motifs were disrupted to reduce self-dimer formation.

[0339] Region 2 included the portion of the probe designed to contact the genomic sequence within the genomic library which was the sample. In this experiment, the target region was exon 2 of PLP1. The DNA sequence of PLP1 exon 2 is shown below. Exon 2 which is the CDS is highlighted in bold uppercase with underline. The capture probe sequences at equal intervals are shaded.

[0340]

Chemical formula

Chemical formula

[0341] Region 3 was complementary to a validated PCR primer called CAC3. The sequence of the CAC3 PCR primer is CACGGGAGTTGATCCTGGTTTTCAC (SEQ ID NO: 72).

[0342] The sequences of the Ultramers containing these probe regions are shown in Table 9.

[0343]

Table 9-1

Table 9-2

[0344] Examination of moles, micrograms, and molecules: The genomic library constructed in Example 3 (female hgDNA library from Promega) was used. The large-scale (800 μl) amplification of this library was initiated and carried out with 20 μl of ligation mix as the input. The final concentration of the purified library (400 μl) was 22 ng / μl. 1 microgram was used per experiment described herein. Further, based on all adapters being 50 bp and inserts being 150 - 200 bp, 75% of the library mass was assumed to be genomic DNA. Then, based on this assumption and the fact that the mass of one human genome is 3 pg, there were approximately 250,000 (750×10 -9 / 3×10 -12 =250,000) copies of any given genomic region present. Past experiments and literature suggested that a 10,000-fold molar excess of probe is a reasonable starting point. This implies 2,500,000,000 probe molecules. 2.5×10 9 molecules / mole 6.02×10 23 molecules / mole = 4.15×10 -15 moles = 4 attomoles of probe. Converting this to the volume of the stock solution, 4 nM (of each probe) 1 μl = 4 attomoles of probe. Finally, C1 beads coated with MyOne strep from Invitrogen bind approximately 1 picomole of 500 bp biotinylated dsDNA per 1 μl of beads. In this experiment, a total of 4 attomoles × 4 probes = 16 attomoles of probe were added. 1 μl of beads binds 1000 attomoles, and 1 μl is the practical amount for the beads to function, and 1 μl of beads has a binding capacity 60-fold in excess of the added probe. Thus, in this example, the following parameters: - Number of target molecules in the library per unit mass (250,000 copies of unique diploid loci per 1 μg of library); - Molar concentration of the probes required to engage the target locus with a 10,000-fold molar excess of probe (4 atto-moles of each probe, 16 atto-moles of total probe (4 types of probes), 1 μl of 4 nM probe solution); and - Amount of beads required to quantitatively capture all added probes (1 μl binds to 1000 atto-moles of dsDNA and / or unbound probe) were calculated.

[0345] Buffers and working solutions Solution 1: Binding probes: Probes for the universal binding partner and PLP1 were hydrated to 100 μM. In two separate tubes, 92 μl of TEz + 0.05% Tween-20 buffer was combined with 4 μl of universal oligo and 1 μl each of 4 cognate (to the universal oligo) probes. This produced 2 types out of the 1 μM stock solution of probes. 4 μl each of these were diluted into 1000 μl of TEz + Tween to yield a 4 nM probe working solution.

[0346] 4X concentration binding buffer = 4 M NaCl, 40 mM Tris, pH 8.0, 0.4 mM EDTA, and 0.4% Tween 20. 50 ml was prepared by combining 40 ml of 5 M NaCl, 2 ml of 1 M Tris, pH 8.0, 2 ml of 10% Tween20, 40 μl of 0.5 M EDTA, and 6 ml of water.

[0347] Wash buffer = 25% formamide, 10 mM Tris, pH 8.0, 0.1 mM EDTA, and 0.05% Tween 20. 50 ml was prepared by combining 37 ml of water, 12.5 ml of formamide, 500 μl of 1 M Tris, pH 8.0, 10 μl of 0.5 M EDTA, and 250 μl of 10% Tween 20.

[0348] Beads: 250 μl of 4-fold concentrated binding buffer was combined with 750 μl of water to prepare 1-fold concentrated binding buffer. 10 μl of beads was added to 90 μl of 1-fold concentrated binding buffer, pulled by a magnet, and the beads were washed twice with 100 μl of 1-fold concentrated binding buffer, and the washed beads were resuspended in 100 μl of 1-fold concentrated binding buffer. 10 microliters of the washed beads corresponds to 1 μl of beads when taken from the manufacturer's tube.

[0349] Method The following three parameters were examined. 1. Universal biotin oligo 1 compared with oligo 10; 2. Binding in 1-fold concentrated binding buffer compared with binding in 1-fold concentrated binding buffer with 7.5% PEG8000 (a molecular crowding agent that can enhance the annealing rate); 3. The enrichment factor of the PLP1 region after binding without processing and the enrichment factor of the PLP1 region after binding with enzymatic processing

[0350] To investigate these parameters, eight samples (2×2×2) were prepared. These samples contained 50 μl of 20 ng / μl genomic DNA, 25 μl of 4× binding buffer, 1 μl of binding probe, and 24 μl of water or 20 μl of 50% PEG8000 + 4 μl of water (four samples with PEG and four samples without PEG). According to the IDT DNA website for OligoAnalyzer, it was described that in high salt concentrations (e.g., 1 M NaCl), the Tm of the oligo shifts to a dramatically higher temperature. Therefore, the samples were melted at 95 °C and then the temperature was decreased to 60 °C in 1 °C decrements and 2-minute intervals (35 cycles by AutoX on an ABI2720 thermal cycler by the inventors, with each cycle decreasing by 1 °C and each cycle lasting for 2 minutes). After cooling the samples to room temperature (RT), 10 μl of washed beads were added per sample and incubated for 20 minutes. The beads were pulled out with a strong magnet, the solution was aspirated and discarded. The beads were washed four times with 200 μl of wash buffer and incubated at RT for 5 minutes each time the beads were resuspended. After the final wash, most of the remaining wash solution was sufficiently aspirated from the tube.

[0351] A set of four tubes was treated with T4 DNA polymerase. A cocktail was prepared by combining 10 μl of New England Biolab 10× Quick blunting buffer, 10 μl of 1 mM dNTP obtained from the same kit, 10 μl of water, and 1 μl of T4 DNA polymerase. 20 μl was added to a set of four tubes and the reaction was incubated at 20 °C for 15 minutes.

[0352] For PCR amplification after capture, non-T4 treated samples (captured only) were amplified with ACA2-25 (TGCAGGACCAGAGAATTCGAATACA; SEQ ID NO: 67) in a single primer reaction. T4 treated samples were amplified with ACA2FL primer and CAC3FL primer (AATGATACGGCGACCACCGAGATCTACACGTCATGCAGGACCAGAGAATTCGAATACA (SEQ ID NO: 69) and CAAGCAGAAGACGGCATACGAGATGTGACTGGCACGGGAGTTGATCCTGGTTTTCAC (SEQ ID NO: 74), respectively). The core reaction mix contained, per 400 μl of reaction, 1 20 μl of water, 40 μl of 10× STD Taq buffer (NEB), 40 μl of 25 mM MgCl2, 80 μl of 10 μM single primer, or 40 μl + 40 μl of F primer and R primer, 20 μl of DMSO, 20 μl of 10 mM dNTP, and 4 μl of Taq polymerase. 80 μl aliquots were added to beads (bound only) resuspended in 20 μl of TEz or 20 μl of T4 mix. The final volume was 100 μl. These samples were amplified by PCR for 30 cycles of 94 °C for 30 s, 60 °C for 30 s, and 72 °C for 60 s. Gel analysis (loading 5 μl of post-PCR material per lane) is shown in the results section. Readings by Qubit indicated that the concentration of each PCR reaction was approximately 20-25 ng / μl.

[0353] For post-amplification analysis, a 500 μl (final volume) master mix was prepared using a conventional PCR mix by combining 200 μl of water, 50 μl of 10x Taq buffer, 50 μl of 25 mM MgCl2, 25 μl of DMSO, 25 μl of 10 mM dNTP, 12.5 μl of EvaGreen (Biotum), and 5 μl of Taq polymerase (NEB). A 42 μl aliquot was dispensed into 8 tubes, and 12 μl of 10 μM F+R PLP1 primer mix (described in Example 5: Validation of the PLP1 qPCR assay for the assay) was added. 9 μl of the mix was dispensed into each of the 8 columns of the assay. A total of 6 samples were assayed with 1 μl of sample per well. These samples were Row 1: gDNA library starting material Row 2: Biotin oligo 1 capture material Row 3: Biotin oligo 1 + PEG capture material Row 4: Biotin oligo 10 capture material Row 5: Biotin oligo 10 + PEG capture material Row 6: TEz NTC control as follows.

[0354] The T4-treated samples were not assayed because gel analysis showed that only abnormal materials were processed by PCR amplification.

[0355] Results The capture-only libraries yielded a smear similar to the input genomic library as expected. The samples were, from left to right, (1) oligo 1, (2) oligo 1 + PEG, (3) oligo 10, and (4) oligo 10 + PEG. The T4-treated samples were contaminated with residual T4 polymerase (5 - 8). In certain embodiments, the T4 polymerase was heat inactivated.

[0356] The yields of the four capture-only libraries measured by Qubit are shown in Table 10 below.

[0357] [Table 10]

[0358] For qPCR, all eight validated PLP1 assays (Example 5) were used in columns and samples were used in rows. The sample array was Row 1: 1 μl of 25 ng / μl gDNA library Row 2: 1 μl of approximately 25 ng / μl C1 capture sample Row 3: 1 μl of approximately 25 ng / μl C1+P capture sample Row 4: 1 μl of approximately 25 ng / μl C10 capture sample Row 5: 1 μl of approximately 25 ng / μl C10+P capture sample Row 6: 1 μl of TEz (NTC) It was as follows.

[0359] In this configuration, with one sample per well, the data was intended to be a qualitative overview rather than a precise quantitative measurement. The data is shown in the table below. The table above shows the raw Cq values. The following table shows the Cq values converted to absolute values based on the assumption that all samples and assays fit the same 2-fold calibration curve. The following table shows the quotient of the Cq value of the captured sample divided by the Cq value of the gDNA library. This shows the meaning of the enrichment fold after capture.

[0360]

Table 11

[0361] Several conclusions were drawn from the data. (1) Capture worked. For C1, the average capture enrichment through hits assays 1 - 5 was 82,000 - fold. For C10, the average was 28,000 - fold. At the assay site, enrichments of about several hundred - to tens of thousands - fold were observed. This implies that the Ultramer worked and the basic probe design was effective. This meant that the basic stoichiometric ratio of gDNA to probe to beads was appropriate; (2) The two biotin designs worked similarly; (3) PEG did not enhance and inhibited the capture efficiency; and (4) In assay 6, significant "off - capture" of 200 bp from the target was observed. The deviation activity observed for regions 1000 bp apart was small.

[0362] In certain embodiments, it may be important to determine whether enzymatic processing of the captured complex contributes to the sensitivity (enrichment fold) and specificity (degree of "off - capture") in this scheme.

[0363] (Example 7) PLP1 qPCR Assay in SYBR Space Objective In some cases, it is useful for real - time conditions to exactly mimic non - real - time amplification conditions. In this example, this meant preparation on ice and a three - step relatively slow PCR reaction. Alternatively, some assays do not require replication of a set of amplification conditions and are intended to rigorously obtain quantitative measurements. For example, the PLP1 qPCR assay is preferably used only to measure the enrichment of the locus and not to generate fragments. In this type of situation, qPCR reactions prepared at room temperature and rapid cycling are advantageous. In this experiment, eight PLP1 assays with ABI 2× SYBR mix were examined. These are the same primer assays as those described in Example 5 (Verification of PLP1 qPCR Assay).

[0364] Summary These data suggested that at least six of the eight PLP1 qPCR assays could be used with SYBR Green qPCR mix and conditions.

[0365] Method The performance of the PLP1 assay on female gDNA libraries (Example 3: female hgDNA library from Promega) was measured. Per well of 10 μl, 5 μl of ABI 2× SYBR master mix, 0.2 μl of 10 μM F+R primer stock solution, 1 μl of gDNA library (20 ng / μl), and 3.8 μl of water were combined (a large volume master mix was made and aliquoted). Triplicate measurements for non-template controls and triplicate measurements for gDNA libraries were taken throughout each assay. Using standard two-step PCR (15 s at 95 °C, 45 s at 60 °C) with normalization by ROX passive reference dye, cycling was carried out for 40 cycles on an Illumina Eco real-time PCR.

[0366] Results The Cq values determined for each well are shown in Table 12 below. The NTC was extremely clean and the Cq of the gDNA was variable, which is highly likely due to pipetting. The general theme is that assays 1 and 7 had low performance, while the remaining assays worked reasonably well in the SYBR space. In Figure 11, the NTC trace (A) and +gDNA trace (B) were copied to present a qualitative portrait of the assay performance.

[0367] [Table 12]

[0368] (Example 8) Measurement of PLP1 exon 2 enrichment before and after complex enzymatic processing Objective In this example, the "specific activity" of PLP1 exon 2 DNA was measured in the capture complex before and after processing to directly examine the enzymatic processing of the complex in terms of yield. Ultramer supported excellent capture efficiency and the performance of the core capture protocol was also good.

[0369] Overview This experiment demonstrated that post-capture processing with T4-DNA polymerase dramatically improved the specificity of the capture reaction.

[0370] Background In Example 6 (capture of PLP1 exon 2), the success of the capture was described, but in the post-capture processing step without removing T4 polymerase before PCR, an artifact library was produced. In this example, the same basic experiment was repeated except that T4 was heat-inactivated at 95°C for 1 minute before PCR.

[0371] Methods, Results, Discussion In this experiment, four samples containing two universal biotin capture probes were prepared to evaluate the capture efficiency of the complex before and after enzyme processing. Each sample contained 50 μl of genomic DNA at 20 ng / μl, 20 μl of 4× binding buffer, 1 μl of binding probe, and 9 μl of water to make the final volume 80 μl. The samples were melted at 95°C for 1 minute and annealed by reducing the temperature to 60°C at a decrement of 1°C for 2 minutes (35 cycles by AutoX on an ABI2720 thermal cycler by the inventors), followed by cooling to RT. Then, a total of 10 μl of washed beads (corresponding to 1 μl of MyOne bead solution (C1 coated with streptavidin; Invitrogen)) per sample was added and incubated for 20 minutes. The beads were pulled out with a strong magnet, the solution was aspirated and discarded. The beads were washed four times with 200 μl of wash buffer each time and incubated at RT for 5 minutes each time the beads were resuspended. After the final wash, most of the remaining wash solution was aspirated well from the tube and the beads coated with the capture complex were left standing.

[0372] For the T4 processing of two samples, the inventors prepared a 50 μl enzyme processing mix containing 40 μl of water, 5 μl of 10× quick blunt buffer (New England Biolabs), 5 μl of 1 mM dNTP, and 0.5 μl of T4 DNA polymerase. Two aliquots of the complex were resuspended in 20 μl (each) of the T4 mix, incubated at 20°C for 15 minutes, at 95°C for 1 minute, and cooled to RT. A "non-treated" control was suspended in 20 μl of the same buffer (40 μl of water, 5 μl of 10× quick blunt buffer (New England Biolabs), 5 μl of 1 mM dNTP) lacking T4 polymerase.

[0373] To measure the specific activity, both the capture alone sample and the capture + processing sample were amplified by 30 cycles of PCR. The DNA was then quantified and the PLP1 assay signal was measured with the amplified DNA of the specific amount and the known amount. In this example, two amplification reactions were prepared. In capture alone, since these libraries are amplifiable with this single primer only, the amplification was carried out with ACA2-25 (TGCAGGACCAGAGAATTCGAATACA; SEQ ID NO: 67). In the enzyme-processed complex, the amplification was carried out with the ACA2FL primer and the CAC3FL primer (SEQ ID NO: 69 and SEQ ID NO: 74, respectively, AATGATACGGCGACCACCGAGATCTACACGTCATGCAGGACCAGAGAATTCGAATACA and CAAGCAGAAGACGGCATACGAGATGTGACTGGCACGGGAGTTGATCCTGGTTTTCAC). 100 μl of the PCR mi xture contained 10 μl of 10× STD Taq buffer (all reagents were obtained from NEB unless otherwise specified), 10 μl of 25 mM MgCl2, 20 μl of 10 μM single primer, or 10 μl + 10 μl of 10 μM dual primers, 20 μl of template (untreated control or T4 processing, beads and all templates), 5 μl of DMSO, 5 μl of 10 mM dNTP, and 1 μl of Taq DNA polymerase (all were prepared on ice before amplification). The samples were amplified by 30 cycles of PCR following a 3-step protocol of 95°C for 30 seconds, 60°C for 30 seconds, 72°C for 60 seconds, followed by 2 minutes at 72°C and paused at 10°C.

[0374] After amplification, the DNA yield was measured and the PCR-amplified material was examined by DNA gel electrophoresis. The yields measured by Qubit (Invitrogen) (DNA HS kit) are shown in Table 13 below. These data emphasize the basic feature that amplification with dual primers supports a greater overall yield than amplification with single primers.

[0375]

Table 13

[0376] Gel images (2% agarose, 100 ng of loaded material) are shown in Fig. 12. Processing had two notable effects. First, processing resulted in a predicted smear in addition to two faint bands at approximately 250 bp (upper arrow) and approximately 175 bp (lower arrow). The lower band corresponded to an unexpected cloning of the probe (115 bp adapter + 60 bp probe = 175 bp). Second, processing reduced the overall sample size distribution. This was notable because replacing the 50 bp single adapter with the 115 bp full-length adapter, which was predicted to result in an overall 65 bp upward shift of the processed material. Processing was interpreted to significantly reduce the average insert size of the library.

[0377] To measure the enrichment efficiency by qPCR, two attempts were made. In the first, more qualitative attempt, all eight PLP1 assays (described in detail in Example 5: Validation of the PLP1 qPCR Assay) were used on six samples: 1. 25 ng of starting gDNA library per assay 2. 0.25 ng of untreated C1 per assay 3. 0.25 ng of untreated C10 per assay 4. 0.25 ng of T4-treated C1 per assay 5. 0.25 ng of T4-treated C10 per assay 6. No-template control were measured.

[0378] The Cq values obtained from these single measurements are shown in Table 14 below. The performance of the gDNA and NTC controls was good (top and bottom; lightest shading) and was not further scored.

[0379]

Table 14

[0380] The signal of the T4-treated sample (dark shadow) was strong to the extent that quantitative analysis became less informative (Cq < 10). However, at the qualitative level, two trends were clear compared to the non-treated capture complex (medium shadow). One was that the hit signals obtained from Assays 1-5 increased dramatically (low Cq). The other was that when processed, the off-target signal 200 bp away from the target region obtained from Assay 6 was significantly attenuated. Although there were some fluctuations in the data, the central message was that processing greatly enhanced the specificity of the PLP1 exon 2 signal.

[0381] To capture the more quantitative aspects of this experiment, prior to qPCR, the untreated C10 capture amplicon was diluted 1000-fold and the processed C10 amplicon was diluted 15,000-fold. This was done to keep the Cq values within the measurable range. Next, the starting gDNA library was considered and these diluted samples within quadruplicate wells of the qPCR plate were examined through two in-target assays (Assays 2 and 5) and two off-target assays (Assays 6 and 7). The Cq values of the quadruplicate wells were averaged and these values are shown in Table 15 below. Again, the gDNA signal was weak, but since the goal of these experiments was to compare the PLP1 exon 2 signal in the unprocessed capture complex to the T4 polymerase-treated capture complex, the impact of the weak signal on data interpretation was not very significant. The Cq values were converted to absolute values using a “universal” calibration curve assuming 2-fold amplification per PCR cycle. The third section of the table shows the adjustment for the dilution factor. The ratio of the unprocessed and T4-treated complexes to gDNA in the fourth section is not very useful, but the quantitative ratio of the unprocessed complex compared to the T4-treated complex in the lower part of the table is useful. In Example 6, an 82,000-fold untreated capture enrichment for C1 and a 28,000-fold untreated capture enrichment for C10 were observed (as with all of these experiments, the denominator for gDNA was derived from extremely low-level signals, so the fold ranges were qualitatively affected by this), so it was reasonable to estimate that capture alone resulted in a 50,000-fold enrichment of the 300 bp PLP1 exon 2 region. Processing increased this enrichment by a further 50-fold (average of 83-fold and 24-fold obtained from Table 15), pushing the enrichment up to 2.5 million-fold to 10 million-fold (3 billion bases per genome / 300 bp target). Thus, at the level of qPCR measurements, capture + processing was considered to approach the best-case scenario for enrichment. It was noted that the off-target signal 200 bp away from the target monitored by Assay 6 was significantly enriched by capture alone (hitchhiker, cross-hybridization effect), but was significantly reduced by processing.

[0382]

Table 15

[0383] This experiment addressed the specificity (non-target qPCR signal) of capture + processing. The specific activity per ng of amplified DNA (obtained from PLP1 exon 2) was significantly enhanced by post-capture processing. This experiment did not address sensitivity, i.e., the percentage of capture complexes that are converted by the enzyme. In certain embodiments, it is also important to have a quantitative understanding of both the specificity and sensitivity of the method.

[0384] (Example 9) Direct measurement of post-capture processing

[0385] Objective In Example 8, it was determined that post-capture processing achieves the desired objective of substantially increasing the specificity of target capture. Another extremely important parameter to consider is sensitivity, i.e., the percentage of initially captured complexes that are recovered in the final sequencing library. In this example, the inventors demonstrated by direct measurement of sensitivity that enzymatic processing is effective for > 10% of the initially captured sequences.

[0386] Summary Data obtained from this experiment indicated that 10% of the on-target capture complexes are processed by T4 polymerase into post-capture sequencing library fragments.

[0387] Consideration As a reference, a schematic illustration of post-capture processing is shown in FIG. 4. In this figure, the sensitivity of the processing was measured in a 3-step procedure exemplified in the lower right of FIG. 13. First, a single PLP1 capture probe was used in an independent reaction to pull down / pull out the PLP1 exon 2-specific genomic DNA fragment from a female gDNA library (Example 3: female hgDNA library manufactured by Promega). Since there are 4 probes, 4 pull-downs were performed. As illustrated in FIG. 13(A), the amount of the captured material was measured using an adjacent PLP1 qPCR assay primer pair. As shown in FIG. 13(B), after the enzymatic processing of the complex, the amount of the processed complex was measured again via qPCR by using one PLP1-specific primer and one probe-specific primer. The ratio of the measured values at [B / A×100%] resulted in an estimated value of the processing efficiency. It is extremely important for the proper interpretation of the experimental results to extract the PCR products obtained from the real-time reaction and verify by gel analysis that amplicons of the predicted length were produced (FIG. 13(C)). This was possible because each PCR reaction had individual start and stop points. The processing efficiency was determined using pull-outs that provided interpretable data from A+B+C.

[0388] assay Each individual probe needed to be matched to the qPCR assay. Six combinations of probes were selected that were matched to the pre- and post-steps of the qPCR assay. These are shown below with the probe sequences in italics and the PLP1 exon 2-specific primers shaded. The primers shaded darkly are the primers paired with the CAC3 primers after processing. Also shown for each assay set is the predicted product size of the PCR amplicon.

[0389]

Chemical formula

Chemical formula

[0390] Method Probe: In these assays, the B10 universal oligo set of probes was selected (Experiment 4 on August 24, 2012: Capture of PLP1 exon 2). To generate individual capture probes, 1 μl of Universal Oligo 10 (100 μM) was combined with 1 μl of Ultramer probe at 100 μM and 98 μl of TEz + 0.05% Tween 20. This was further diluted into 4 μl to 996 μl of TEz + Tween to yield a 4 nM working solution.

[0391] Capture: For capture, 50 μl of 22 ng / μl gDNA library was combined with 20 μl of 4x binding buffer, 1 μl of probe, and 9 μl of water. There were six independent capture reactions (two reactions with probe 1, two reactions with probe 4, one reaction with probe 2, and one reaction with probe 3). These were heated to 95 °C for 1 minute and then cooled to 60 °C in -1 °C and 2-minute "cycles" of 35 as described above. After annealing, 10 μl of washed beads (= 1 μl of bead stock solution) was added and binding was incubated at RT for 20 minutes. The beads were then pulled to the side and washed four times with 200 μl aliquots of wash buffer for 5 minutes each. After the final wash, all remaining accessible fluid was aspirated from the beads.

[0392] Processing: The beads were resuspended in 10 μl of quick blunt solution (200 μl = 20 μl of 10× quick blunting buffer, 20 μl of 1 mM dNTP, and 160 μl of water). Each of six bead aliquots was split into two 5-μl aliquots. 5 μl of QB buffer without enzyme was added to one set of tubes (these are the capture-only aliquots). To the other 5-μl aliquots, 5 μl of QB buffer containing 0.025 μl of T4 polymerase (prepared by combining 100 μl of QB buffer with 0.5 μl of T4 polymerase and dispensing into 5-μl aliquots) was added. Both the capture-only tubes and the capture + processing tubes were incubated at 20 °C for 15 minutes, then at 98 °C for 1 minute, cooled to RT, and immediately placed on a magnet. Approximately 10 μl of supernatant was removed from six pairs of capture-only complexes and T4-processed complexes (for a total of 12 tubes). These supernatants were used directly in the qPCR described below.

[0393] qPCR: For these assays, a standard Taq reaction mix and 3-step thermal cycling were selected. Each consisted of - 14 μl of water - 4 μl of 10× STD Taq buffer - 4 μl of 25 mM MgCl2 - 4 μl blend of F primer and R primer at 10 μM each - 8 μl of template (supernatant obtained from above) - 2 μl of DMSO - 2 μl of 10 mM dNTP - 1 μl of EvaGreen - 0.8 μl of ROX - 0.4 μl of Taq polymerase Twelve 40-μl qPCR mixes were constructed, each containing

[0394] The reactants were distributed into four pairs and cycled over 40 cycles at 94 °C for 30 seconds, 55 °C for 30 seconds, and 72 °C for 60 seconds. After PCR, the reaction mixes were pooled from each of the four wells of the four pairs and 5 μl was analyzed on a 2% agarose gel.

[0395] Results To interpret the experimental data, the agarose gel shown in Figure 14 was examined. Under the cycling conditions used with the primers (etc.) employed, assay sets 3, 5, and 6 were observed to yield PCR products that corresponded to PLP1 (lower gel) after processing in light of the amplicon of the assay (upper gel) or the adapter amplicon. The better assay sets were - Probe 4 with assay 3 - Probe 2 with assay 5 - Probe 3 with assay 4 corresponding to.

[0396] The Cq values of qPCR are shown in Table 16 below. Assays 1 and 2 had poor gel analysis. Good assays are shown in assays 3, 5, and 4. To derive the processing % values, Cq was converted to an absolute value (in "Excel language", absolute value = power(10, log10(1 / 2)×Cq + 10)). The quotient of Cq by Cq after processing only by capture was then expressed as a percentage. This measurement assumed that the amplification efficiency of all amplicons was the same and conformed to an idealized calibration curve (presumably reasonably accurate). Then, assuming this to be appropriate, it was considered that approximately 10% of the captured material was processed.

[0397]

Table 16

[0398] (Example 10) Purpose of constructing extended code-processed male gDNA library and female gDNA library Construct a set of 16 coded male gDNA libraries and female gDNA libraries that are used to examine multiple capture parameters in a single MiSeq sequencing run.

[0399] Method Step 1: Prepared repaired gDNA. Step 2: Fabricated all 16 possible adapter codes. These codes have 4 base structures. The base positions at -4 and -3 (with respect to the insert) are random bases, and the base positions at -2 and -1 are sample codes. There are 4 "cluster" sample codes. These are - Cluster 1: AC, GA, CT, TG - Cluster 2: AA, GC, CG, TT - Cluster 3: AG, GT, CA, TC - Cluster 4: AT, GG, CC, TA respectively.

[0400] Clusters 2 - 4 were arrayed as 100 μM oligos in the plate. One set of the plates had ligation strands and one set of the plates had partner strands. The plate arrays were A1 - H1, A2 - H2, etc. To anneal the adapters within two sets of 96 - well PCR plates, a "annealing solution" containing 70 μl per well, 68 μl of TEz and 2 μl of 5 M NaCl was added to 20 μl of partner strand oligo and 10 μl of ligation strand oligo covered with tape, annealed at 95 °C for 10 seconds and 65 °C for 5 minutes, and cooled to RT. A set of 16 codes (random codes with the same sample code) was pooled into a set of 4 codes. Red = set AA, GC, CG, and TT. Purple = set AG, GT, CA, and TC. Blue = set AT, GG, CC, and TA (arranged in this order).

[0401] Step 3: It is easiest to create 16 ligation strands for female DNA and 16 ligation strands for male DNA, both of which accept the same set of 16 unique adapter types. This then allows, at a later stage, for maximum flexibility in determining which combination of samples to create. To do this, end-repaired gDNA obtained from the experiment was used. This will result in 32 ligation reactions per 20 μl per reaction being carried out as follows. Two gDNA cocktails, one female and one male, containing - 144 μl of water - 32 μl of 10x ligation buffer - 48 μl of 50% PEG8000 - 64 μl of gDNA were prepared.

[0402] The cocktails were mixed and aliquoted into 16 tubes with 18 μl each. 2 μl of adapter and 0.5 μl of HC T4 ligase were added and the resulting reactions were incubated at 22 °C for 60 minutes and then at 65 °C for 10 minutes and cooled to RT. 80 μl of TEz and then also 120 μl of Ampure beads were added to the reactions, mixed and incubated at RT for 10 minutes. The reactions were washed twice with 200 μl of 70% EtOH / water (v / v), air dried and resuspended in 100 μl of TEz.

[0403] Step 4: qPCR: - 175 μl of water - 50 μl of 10x STD Taq buffer - 50 μl of 25 mM MgCl2 - 100 μl of ACA2 primer (10 μM) - (50 μl of template: added later) - 25 μl of DMSO - 25 μl of 10 mM dNTP - 12.5 μl of Eva Green - 10 μl of ROX - 5 μl of Taq DNA polymerase A qPCR master mix containing the above was prepared.

[0404] 9 μl was dispensed into 48 wells of an Illumina Eco qPCR plate. Two serial dilutions of 10 pg / μl and 1 pg / μl of the library calibration reference material were prepared. The remaining wells of the plate were loaded with the libraries shown in the following table.

[0405]

Table 17

[0406] The second plate had the layout shown in Table 18 below.

[0407]

Table 18

[0408] Ligation efficiency was measured by the following cycling program: - 2 minutes at 72 °C - 40 cycles of 30 seconds at 94 °C, 30 seconds at 60 °C, and 60 seconds at 72 °C It was measured by.

[0409] Results Table 19 below shows the Cq values of the reference material and the samples (the average of two replicate measurements, except that (i) the experiment was repeated on Plate 2 and (ii) M1, M2, and M3 were measured in three sets of two replicates (the average of three measurements was taken)).

[0410]

Table 19

[0411] These were converted to arbitrary absolute values using the formula in Excel (shaded blue), molar mass = POWER(10, LOG10(1 / 2) * Cq + 8). Then, the absolute values were multiplied by 10 / 1583 (plate 1) or 10 / 1469 (plate 2) to standardize the values against a known standard (shaded red). The genome per 1 μl was calculated by multiplying by 7 / 8 (corresponding to the adapter mass) and then dividing by 3 pg per genome. Ligation efficiency was calculated (20 ng per ligation, 1 / 100 was measured = 200 pg was ligated), and the calculated efficiency indicated that approximately 5% conversion to the library was the approximate average value. This was the same for libraries made without embedding, suggesting that the reaction rate of the embedding reaction is rapid and can occur when the sample is heated to 94 °C in the first cycle.

[0412]

Table 20

[0413] The goal of this experiment was to create a ligation mix containing a gDNA library and to quantify the genomic equivalent per μl of the ligation mix so that the measured number of genomes could be amplified into microgram amounts of library material. Table 20 above shows the genomes per μl for each library produced. The goal of Table 21 shown below was to convert the samples shown (collected by random sampling) into libraries of 10 copies, 20 copies, 40 copies, 80 copies, etc. for the subsequent capture test. The table transposes the number of genomes per μl into the number of μl per PCR reaction to achieve the indicated coverage depth. The table assumes 200 μl of PCR and 40 μl of template input per sample. These experiments can be used as guidelines when creating and purifying an actual library.

[0414]

Table 21

[0415] (Example 11) Verification of 8 novel capture qPCR assays Objective Verify the performance of 8 new qPCR primer sets designed to pursue the capture efficiency of an extended probe collection.

[0416] Overview All 8 assays yielded amplicons of the size predicted when used to amplify human gDNA. Quantitative analysis of the X chromosome: 154376051 region (4 copies in females and 2 copies in males) showed a surprisingly tight correlation between the observed and predicted copy numbers.

[0417] Method Eight segments for assay design were selected representing sampling of 49 probe target regions. To design the assays, DNA segments within 200 bp from the 5' end of the probes were identified. Eight regions shown in Table 22 below were selected to be somewhat random selections of the target regions. The 200 bp segments were excised by the inventors into Primer3 for PCR primer picking that specifies amplicons of 50 - 100 bp with the primer Tm at 65°C (optimal Tm) and the primer length at 24 nucleotides (optimal length). Table 22 below shows the regions and unique genomic attributes, forward (F) primer sequences and reverse (R) primer sequences, the predicted amplicon lengths, and the actual amplicons in the context of the genomic sequences.

[0418] [Table 22]

[0419] The performance of each primer pair was explored by performing a 100 μl PCR reaction containing 200 ng (2 ng / μl) of female genomic DNA. The reaction mix contained, per 100 μl, 50 μl of water, 10 μl of 10× STD Taq buffer, 10 μl of 25 mM MgCl2, 10 μl of F+R primer blend with each primer present at 10 μM, 10 μl of 20 ng / μl gDNA, 5 μl of DMSO, 5 μl of 10 mM dNTP, and 1 μl of Taq polymerase. The reaction was prepared on ice. Amplification was carried out over 30 cycles of 94 °C for 30 seconds, 60 °C for 30 seconds, and 72 °C for 30 seconds, followed by an incubation at 72 °C for 2 minutes and held at 10 °C. 5 μl of the PCR product was examined on a 2% agarose gel.

[0420] The PCR product was purified on a Qiagen PCR purification column by combining 95 μl of the remaining PCR product with 500 μl of PB. The material was spun through the column at 6 KRPM for 30 seconds, washed with 750 μl of PE, and spun at 13.2 KRPM. The product was eluted from the column with 50 μl of EB and quantified by Qubit.

[0421] For qPCR analysis, the chrX-154376051 region (Assays 10 and 11) was examined in more detail. The purified PCR products were diluted to 100 fg / μl, 10 fg / μl, and 1 fg / μl. Genomic DNA was diluted to 10 ng / μl. 2 microliters of the reference material or gDNA was combined with 8 μl of PCR master mix per well of a 48-well Eco qPCR plate. The master mix contained, per 500 μl of the final reaction volume (corresponding to the addition of the template), 175 μl of water, 50 μl of 10× STD Taq buffer, 50 μl of 25 mM MgCl2, 50 μl of 10 μM F+R primer blend, 25 μl of DMSO, 25 μl of 10 mM dNTP, 12.5 μl of EvaGreen, 10 μl of ROX, and 5 μl of Taq polymerase. 32 μl of the mix was dispensed into 16 wells, and 8 μl of the template was added. These were then dispensed into the qPCR plate in 4 pairs. The plate layout is shown in Table 23 below.

[0422]

Table 23

[0423] Results and Discussion Gel analysis of the PCR products amplified from genomic DNA showed that all 8 PCR reactions yielded unique products of the predicted size (data not shown). The amplicons were clean enough (without residual bands and without residual primers) and were useful for creating a calibration curve for quantitative analysis. The amplicons were purified using Qiagen PCR spin columns and the products were eluted in 50 μl. The product yields were: Assay 9 - 18.4 ng / μl; Assay 10 - 26.1 ng / μl; Assay 11 - 13.9 ng / μl; Assay 12 - 26.6 ng / μl; Assay 13 - 7.9 ng / μl; Assay 14 - 19.2 ng / μl; Assay 15 - 23.1 ng / μl; and Assay 16 - 20.4 ng / μl.

[0424] Quantitative analysis was performed for assays 10 and 11 to account for potential segmental duplications on the X chromosome such that females had 4 copies and males had 2 copies.

[0425] The average Cq values are shown in Table 24 below. Using these, the calibration curves shown were created. The two reactions were essentially superimposable. Using these curves, the inventors calculated the absolute amounts in femtogram units for both the calibration curve wells and the genomic input wells. The data are shown below the calibration curve data in Table 24.

[0426] [Table 24]

[0427] One key point of this example was to emphasize the power of quantitative molecular biology. In this experiment, adding 2 μl of the reference material and sampling means that 1 fg / μl of the reference material actually exists as 2 fg in the qPCR reaction. This corresponds to 17,500 molecules by the 53 bp fragment of assay 10. 20 ng of genomic DNA was added to the reaction. This corresponds to DNA equivalent to 6667 genomes. Fragmenting the genomic DNA to an average size of 200 bp means that only 75% of the target region remains intact. Thus, the gDNA had approximately 5000 "qPCR - executable" genomic copies. Finally, in males, the predicted average number of overlapping regions of the X chromosome per genome was 1 copy, and in females, the predicted average was 2 copies. The predicted values compared to the observed values based on the number of molecules observed were as follows: male predicted value = 5000 copies; male observed value = 3500 copies; female predicted value = 10000 copies; and female observed value = 7000 copies.

[0428] [Table 25]

[0429] (Example 12) Additional post-capture processing strategies Objective Alternative methods for achieving post-capture processing (see Figure 15) were developed.

[0430] Overview The post-capture processing steps performed with the redesigned probe were thought to further enhance the already robust capture by a factor of 5 to 9. Overall, the tests were extremely successful.

[0431] Background In other embodiments of the assay design, it was envisioned to use an exonuclease step at the 3' end of the clone, before adding PCR priming sites and copying the probe tail sequence. In certain embodiments, it was further envisioned to shift from copying the probe to the clone. This reversal of polarity means that the inventors can use the 5' end of the probe as either a pull-down sequence or a reverse PCR primer sequence. The 3' end of the probe is left unmodified and the clone can then be copied using DNA polymerase. Conceptually, this approach has several advantages. First, there is a shift from a step that requires both exonuclease activity and polymerization to a simple polymerization step, so this step can be performed in concert with PCR. Further, this step can be performed at 72 °C with a thermostable polymerase enzyme, meaning that the potential secondary structure of the single-stranded clone is less of an issue. Finally, it was implied that the probe was shortened from 114 nucleotides to 95 nucleotides, resulting in cost savings.

[0432] Four well-behaved qPCR assays (Example 11: Validation of 8 new capture qPCR assays), assays 10, 14, 15, and 16, matched the "pointing" probes. It is important that the probes and the qPCR assays were in the vicinity of each other, but their DNA sequences did not overlap with each other (see Figure 16). The probe sequences and the corresponding assays are shown in Tables 26 and 27 below.

[0433]

Table 26

[0434]

Table 27

[0435] Method The gDNA library was re-made from samples F13 - F16 (Example 10) by combining 20 μl aliquots of ligation mix to a total of 80 μl and amplifying in a total of 800 μl. The beads were washed to 400 μl and the pool concentration was measured by Qubit to be 32 ng / μl.

[0436] The IDT oligos listed below were resuspended to 100 μM. Since Ultramer is obtained at 4 nanomolar, these were suspended in 40 μl of TEzero. Four 2 μl aliquots of the four test probes were combined with 8 μl of 100 μM universal tail sequence (obtained from the first 35 bases of full-length reverse primer 9) to yield 50 μM duplex tubes. This duplex was diluted from 10 μl to 990 μl of TEzero + Tween to 500 nM and then diluted again from 10 μl to 990 μl to 5 nM.

[0437] Forty microliters of combined gDNA were combined with 15 μl of 4x binding buffer and 5 μl of capture duplex. The reaction mix was annealed and captured on 2 μl of washed MyOne streptavidin-coated beads. The reaction was washed four times with wash buffer and the wash buffer was aspirated from the bead pellet. To measure the capture-only sample, one type of bead pellet was resuspended in 100 μl of PCR mix containing the single PCR primer ACA2. To measure the capture + processing sample, a different type of bead pellet was resuspended in 100 μl of PCR mix containing the full-length ACA2 forward primer (oligo 8) and the full-length CAC3 reverse primer (oligo 9). The latter sample was incubated at 72 °C for 2 minutes. Both samples were amplified over 25 cycles of 94 °C for 30 seconds, 60 °C for 30 seconds, and 72 °C for 60 seconds. Held at 72 °C for 2 minutes and cooled to RT, the PCR amplicons were purified on beads and resuspended in 50 μl of TEzero.

[0438] For qPCR, samples were assayed in assays 9 - 16 (assays 10, 14, 15, and 16 are the targets) using EvaGreen as the reporter dye, ROX as the reference dye, and three-step PCR over 40 cycles of 94 °C for 30 seconds, 60 °C for 30 seconds, and 72 °C for 60 seconds. The original gDNA library was present at a final concentration of 2 ng / μl. The captured samples and the capture + processing samples were present at a final concentration of 2 pg / μl (diluted in TEzero + 0.05% Tween20).

[0439] Results and Discussion The PCR yield by capture alone was 27.8 ng / μl, and the PCR yield by capture + processing was 40.4 ng / μl. These robust yields indicated that the amplification was complete. A 2% agarose gel image shows the starting input library, the captured library, and the capture + processing library (Figure 17). When processing worked, the average insert size of the library should have decreased, and it actually did. The fact that the lower end of the library somewhat formed a "band" indicates that some degree of priming off of the probe can occur. In this format, since the 3' end of our probe is exposed, it may be possible to eliminate the remaining, unbound probe with exonuclease I, a ssDNA-specific, 3'→5' exonuclease.

[0440] The important metrics in this experiment were the measurements of capture sensitivity and capture specificity by qPCR. The qPCR data are shown in Table 28 below.

[0441] **Table 28**

[0442] Regarding specificity, only the targeted regions (highlighted in light gray) showed significant enrichment. Furthermore, the processed library showed a significant increase in specific activity compared to capture alone for all target regions. These data indicated that this further probe design embodiment could be used for efficient post-capture processing.

[0443] (Example 13) Sequence Analysis of Post-Capture Processing Strategies Purpose The purpose of this experiment was to evaluate the enrichment and coverage of target regions within the sequencing library.

[0444] Overview The levels of enrichment and focusing of the target arrays were dramatically improved by coupling hybridization-based capture with enzymatic processing as compared to capture alone.

[0445] Background Previous experiments disclosed herein have demonstrated that post-capture processing increases the target content and the specific activity of the enriched libraries as measured by qPCR. In this experiment, next-generation DNA sequencing was used to compare the representation and distribution of target sequences within libraries made by capture alone or alternative processing methods.

[0446] Methods Two enriched library pools were constructed from an equal mix of male and female human genomic DNA using a set of 49 capture probes targeting sites within specific genes (KRAS, MYC, PLP1, CYP2D6 and AMY1) and overlapping regions on the X chromosome. The probe sequences are shown in Table 29 below.

[0447]

Table 29-1

Table 29-2

Table 29-3

Table 29-4

[0448] The first library pool was made as described for the "capture + processing" library in Example 12. The second pool was made as described for the "capture only" library in Example 12, with the following modifications: After capture PCR, a second round of PCR was performed to convert the single primer ACA2 amplified library into a dual primer heterogeneous end library suitable for Illumina sequencing. To do this, the library was diluted and amplified with the following primers: Primer 55 AATGATACGGCGACCACCGAGATCTACACGTCATGCAGGACCAGAG (SEQ ID NO: 199) and Primer 56 CAAGCAGAAGACGGCATACGAGATGTGACTGGCACGGGAGTTGAGAATTCGAATACA (SEQ ID NO: 200) The resulting mixture was re-amplified by

[0449] The 100 μl reaction mix contained 40 ng library, 10 μl 10× STD Taq buffer, 10 μl 25 mM MgCl2, 10 μl 55 primer and 10 μl 56 primer, both at 10 uM, 5 μl DMSO, 5 μl dNTPs, and 1 μl Taq DNA polymerase. Samples were amplified for 2 cycles of 94° C. for 30 sec, 50° C. for 30 sec, 52.5° C. for 30 sec, 55° C. for 30 sec, 57.5° C. for 30 sec, 60° C. for 30 sec, and 72° C.-1 min. Samples were then amplified for 8 cycles of 94° C. for 30 sec, 60° C. for 30 sec, and 72° C. for 60 sec, followed by 72° C. for 2 min. The PCR mix was bead purified and resuspended in 50 μl portions.

[0450] Results and Discussion Both pools were analyzed using an Illumina MiSeq Personal Sequencer. 50-nucleotide sequence reads obtained from each library pool were trimmed to remove the 4-base barcode sequences and mapped to the reference sequence of the human genome (hg19 variant) using the Bowtie sequence alignment program. Approximately 80% of the reads within both libraries aligned unambiguously to the reference sequence. Further characterization of the aligned reads revealed that hybridization-based capture, coupled with enzymatic processing, resulted in a 979,592-fold enrichment of the 4.9-kilobase target region compared to the input genomic DNA. This represented a 3-fold improvement in library content compared to the unprocessed "capture-only" approach. Generally, approximately 4 out of 5 sequences obtained by this alternative processing method mapped to genomic sites specifically targeted by the capture probes.

[0451] An overview of the alignment statistics for each library pool is shown in Table 30 below.

[0452]

Table 30

[0453] Reads obtained from each library pool were also visualized using the UCSC Genome Browser to assess the local sequence coverage and sequence distribution near the target sites. Magnified images of two segments of the X chromosome show that the processed library resulted in a much higher sequence coverage within the targeted sites compared to the "capture-only" library (Figure 18). Additionally, sequence mapping to the target region was more evenly distributed in the processed library compared to the unprocessed control. Overall, these data indicate that the alternative processing method dramatically improves the quantity and quality of the target sequences present within the enriched library.

[0454] (Example 14) Bioinformatics Summary Conventional next-generation sequencing (NGS) analysis is "vertical". The unique design of the molecules of the present invention envisioned herein enables a "horizontal" approach that revolutionizes clinical resequencing techniques.

[0455] As used herein in connection with sequence alignment, "vertical" refers to the approach illustrated by FIG. 19. The prior art approach to informatics analysis involves a first step of aligning short reads with a reference genome. After alignment, overlapping reads are analyzed for base changes that may indicate SNVs (single nucleotide variants). Herein, the approach is referred to by the moniker "vertical" because it often relies on alignments depicted as a vertical stacking of reads. A variety of programs allow for SNVs and indels (insertions / deletions), but the core approach is alignment-based.

[0456] In contrast, the paired-end read data obtained by the method envisioned herein will have DNA-tagged sequence information within read 1 and probe ID information within read 2. The first step in data analysis is to match the reads to the probes. Step 2 is to analyze the sequence information "horizontally" associated with each probe. See, for example, FIG. 20.

[0457] When the read depth is sufficient, the association of horizontal, probe-based arrays is alignment-independent. Instead, reads can be assembled into de novo contigs. All of the advantages of the method are that conventional, alignment-based methods struggle to detect and are most difficult in the situation of short sequence runs, insertions / deletions and multiple sequence changes, which are extremely robust against insertions / deletions and multiple sequence changes. Furthermore, by combining horizontal association with probing and tagging, more accurate hypothesis generation (i.e., determination of whether the observed sequence variant is likely to be a true sequence variant or a false sequence variant) becomes easier.

[0458] CNV and Structural Variation I In large-scale copy number variation (CNV: copy number variation) analysis, the method involves determining the number of unique reads associated with the captured sequence region. The majority of observed CNVs are "micro-CNVs" involving insertions and deletions of bases on the order of 2-100 bp in length. The vertical alignment method struggles with micro insertions / deletions (indels) because they require relaxation of the alignment strictness, which promotes a large number of false positive hypotheses. Horizontal methods and de novo contig assembly do not require such relaxation of alignment parameters and require the understanding of structural variations.

[0459] Consider the simple case of a small insertion within one allele of an exon illustrated in FIG. 21. In this example, horizontal alignment "forces" the association of reads with probes 1 and 2. The assembly will create two contigs, one with a wild-type exon structure and one with an insertion structure. Two principles emerge from this analysis: 1) overlapping reads obtained from adjacent probes serve as proof or disproof of hypotheses about indel-containing alleles of the captured exon, and 2) micro-CNV alleles outside the capture probes are easily detectable by the horizontal method.

[0460] CNV and Structural Variations II Verification of CNVs often involves vertical alignment methods. In these studies, alignments that are complete relative to a reference sequence are typically required. Such methods, which reject SNVs that cross reads and differ from the norm, are vulnerable to SNVs (such as common SNPs). The net result will be a habitual underestimation of the copy number. To proceed, the horizontal method made possible by the method of the present invention should be used.

[0461] Horizontal Hypothesis Testing for SNVs I Vertical, alignment-based methods for detecting SNVs are difficult to analyze. Alleles of homozygous variants involving a single base are straightforward to identify, but these changes are rare. More commonly, SNVs are heterozygous, and variants may occur at several consecutive positions or at positions with a proximate interval (error-prone repair tends to track positions where multiple bases are not consensus). When 49 reads carry an SNV and 47 reads carry the wild-type reference base (as a strictly hypothetical example), the heterozygous SNV hypothesis deviates from true high-coverage detection. When the read depth is shallow and the number of SNV reads compared to WT reads dissociates significantly from 50 / 50 (e.g., 8 are WT and 2 are variant, for a total of 10 reads), the determination becomes even more speculative. Hypotheses subjected to orthogonal verification invariably receive an arbitrary cutoff.

[0462] In certain embodiments that combine horizontal probe-based association with tagging, much greater accuracy in the SNV hypothesis is achieved. SNVs present on a single tag (tag = code + endpoints) are ignored, especially when the reads within the same tag are WT. See, for example, FIG. 22.

[0463] Horizontal Hypothesis Testing for SNVs II Even if the start sites of the reads are the same, the SNV hypotheses (A) that occur on two different tags, or the SNV hypotheses (B) that occur on different reads that horizontally associate with the same probe, or the SNV hypotheses (C) that result from the association of different probes in the same exon, are hypotheses that must necessarily be seriously considered. For example, refer to FIG. 23.

[0464] (Example 15) Molecular annotation Summary This example describes the interaction between "molecular annotation" (FIG. 24) for a sequencing library and the informatics used in subsequent steps to evaluate the resulting sequencing information. The reverse reads obtained from the probes are useful. Reverse read 2 that determines the region of the DNA probe sequence is significantly useful in all subsequent analytical considerations. For example, the usefulness can be found in the determination of mutants, and the results obtained therefrom are useful in the determination of copy number. These two aspects of data analysis are described below.

[0465] Probe sequence of read_2 A probe set is a unique, known sequence collection that can include one or two probes, or even tens of thousands of probes. This means that using read_2, any and all probes in an experiment can be identified. Of course, it is assumed that the length of read 2 is sufficient and the probes are designed such that the region scrutinized by read_2 constitutes a unique identifier. Table 31 describes a collection of 192 probes and the 10-nucleotide read_2 sequence used as a unique identifier for each probe. Note that two probes (CYP2C19_r5_F and CYP2C9_r5_F) of course share the same 10-nucleotide 5’ DNA sequence, and a 2-nucleotide code ("AG" or "CT") (shaded) that discriminates between them is added.

[0466]

Table 31-1

Table 31-2

Table 31-3

Table 31-4

Table 31-5

Table 31-6

Table 31-7

Table 31-8

Table 31-9

Table 31-10

Table 31-11

Table 31-12

Table 31-13

Table 31-14

Table 31-15

Table 31-16

Table 31-17

Table 31-18

Table 31-19

Table 31-20

Table 31-21

[0467] In paired-end sequencing experiments, reads_1 and reads_2 are derived from the same DNA clone. This implies that the genomic sequences of reads_1 (sub-figures (3) and (4) of Figure 24) exist because they have associated with a specific probe (sub-figure (5) of Figure 24). In summary, this data indicates that each DNA sequence present within a collection of next-generation sequences can associate with the probe sequence that targeted it. All DNA sequences that associate with a specific probe can be recovered.

[0468] This paradigm for next-generation re-sequencing analysis (targeted analysis or other forms of analysis) is to realign the reads to the reference genome again. The finding of targeting probe associations leads to a new workflow where reads are first sorted by the probe and then analyzed by alignment-based methods, de novo assembly methods, or both. As described in Example 14, PARSAR (probe-associated-read-scaffold-assembly) is one of the more complex and difficult problems in mutant discovery, where the most interesting mutants are those that deviate most significantly from the reference sequence, but these are precisely the sequences that are most refractory to conventional sequence-based alignment (Figure 25). Following probe association, the use of de novo local assembly allows such mutants to be easily identified.

[0469] Group the probe-based reads and use them along with other aspects of the molecular design to identify variants with a high level of first-pass reliability. As shown in Figure 26, the probes are generally designed to span the target region. The overlapping aspect of the reads enables potential variant sites to be queried by two-directionally independent reads. Additionally, this dual-probe design also ensures sequencing the adjacent probe-binding sites themselves. This is an important feature when the capture performance of the probes can be an issue. As an example, in this molecular design, single-nucleotide variants where the allele of a variant forms one of the bases in the capture probe sequence are identified and grasped in the subsequent informatics analysis.

[0470] A further aspect of the flow of information from molecular annotation to subsequent variant analysis involves a sequence "tag" (Figure 24's (1)+(3)) defined as a combination of labeling by three base sequences and the sequence start site of the sticky end. The sequence tag defines that each sequencing clone is unique. As illustrated in Figure 27, variants that occur within a collection of homologous clones and share the same sequence tag are likely to be false positives. In contrast, variants shared among sequences with different tags (even if occurring at low frequency) are likely to be true positive variants. This system of tagging sequences and using the tags to assign reliability predictions to variant determination has the prospect of substantially reducing the burden of subsequent variant verification (which can be costly and time-consuming). Molecular annotation is described in more detail in the literature of Example 16, which describes the sequencing platform by molecular techniques.

[0471] In summary, one of the prominent features of the technology platform envisioned in this specification is that all "annealing probe" events are copied into DNA clones that also carry additional molecular annotations. The sequences are sorted by the probe and sample labels into a collection belonging to the specific target region of the specific input sample. Then, a combination of alignment and de novo assembly can be used to detect variants. Finally, the redundancy of the appearance of candidate variants can be used to assign confidence in the determination of variants. In addition to variant analysis, a method for determining copy number is also provided. These two elements are tightly coupled, and the determination of copy number is specific because it depends on highly reliable sequencing reads. The overall scheme for determining copy number from sequence information is shown in FIG. 28.

[0472] (Example 16) Molecular technology sequencing platform Summary The genomic sequencing platform envisioned in this specification (1) addresses genomic samples obtained from multiple individuals in a single sequencing run; (2) detects single nucleotide variants (SNVs) as well as single nucleotide insertions and deletions (SNIDs) with high confidence; (3) detects large-scale and small-scale copy number variations (CNVs) in all queried gene environments; (4) detects micro-scale translocation, inversion, and insertion / deletion events in the queried gene environment; (5) develops a scalable technology system from ≧ exome-scale exploration (≧ 1-2% of the entire human genome sequence) to ≦ single gene-scale validation; (6) achieves high specificity (low false negative rate) and high sensitivity (low false positive rate) in genomic variant assays; (7) creates simple, portable, and scalable molecular and bioinformatics technologies in its execution; (8) provides a method that results in molecular methods that are easily adaptable to quality control measurements.

[0473] The overall scheme of genomic sequencing reads is shown in Figure 29. The description of each element is as follows.

[0474] (1) The "sequence label" is a set of (consecutive * ) nucleotides (i.e., a set of unique 3-mers) that, together with the read start position (3), is used to establish that each sequencing read is unique. In the pioneering literature, this combination of the label and the read start point was referred to as the "unique sequence tag". The sequence label is the first set of bases encountered, and in single-base synthesis (SBS) chemical reactions, since all four DNA bases must be equally presented at each read position, the constraints on the sequencing label are not only that it be unique, but also that the collection of bases used in the set of sequence labels must have all four bases present at all positions being sequenced. The use of unique sequence tags to determine local CNV is described in the bioinformatics section of this document.

[0475] (2) The "sample label" is a set of (consecutive * ) nucleotide codes that uniquely identify a specific sample within a set of multiplexed samples. Similar to the sequencing label, the collection of sample labels also contains all four bases to meet the requirements for sequence base determination by SBS. The sample code is intentionally placed adjacent to the genomic DNA fragment. The driving factor for this design is ligation bias, which means that there is a base preference for DNA ligation efficiency approximately 2 bases upstream and 1 - 2 bases downstream of the ligation junction. By placing the sample code at the ligation junction, all fragments within a specific sample are subject to the influence / bias of ligation.

[0476] * Without wishing to be bound by any particular theory, it is assumed that the sequence label and the sample label can be created as a comb-shaped nucleotide sequence.

[0477] (3) The "read start point" within a genomic fragment is one of two key elements that define a "unique sequence read". As discussed in section (1) above, the unique identification "tag" for each read includes the sequencing label and the read start point. As will be discussed in more detail below, the collection of unique [(1)+(3)] sequence tags is essential for determining large-scale CNVs. In this specification, "large-scale CNV" is defined as any CNV that involves the entire of at least one probe-binding region with some adjacent sequences added. A large-scale CNV can be as large as the acquisition or loss of an entire chromosome.

[0478] [(1)+(2)] sequence labels and sample labels are embedded within adapter sequences ligated to end-repaired genomic fragments at an early stage of the library construction process where the entire genomic library is created.

[0479] (4) Sequencing reads: The sequence information obtained from genomic fragments is, of course, the central focus of genomic assays. Each read is considered in the context of multiple, overlapping reads that are generated within the same assay.

[0480] (5) Probe level ("genomic indexing"): The overall genomic assay strategy is to combine multiple sequence labels with a composite "molecular annotation" that positions each sequencing read within a larger framework of genomic analysis. Within this operational paradigm, read 1 reveals the elements (1-4) of each annotated clone. Read 2 reveals the probe sequences that recovered each clone by hybridization-based capture and subsequent enzymatic processing. Since all reads are first clustered according to the probes that captured them, probe sequence information is central to the genomic strategy. Since each read is indexed against genomic probes prior to analysis, this clustering of probe-by-base information is referred to as "genomic indexing".

[0481] One interesting feature of the probe label is that the conformation of all probe sequences in the capture reactant is well-defined (it is known which probes are subjected to the capture reaction). This implies that read 2 need not necessarily target the entire 60-nucleotide probe sequence. Instead, read 2 need only be long enough to enable unambiguous identification of all probes in the specific reactant. As a non-limiting example, the probe set discussed in Example 15 consists of 192 probes that can be differentiated based on only 7 nucleotides of the 5' probe sequence (two probes with the same 7-nucleotide 5' end, tagged with a dinucleotide code, would be differentiable information-theoretically).

[0482] (6) Capture Label: The composition of the library is determined by the tight molecular interaction between the probe and the target sequence. The performance of each unique probe sequence can be monitored using a capture label, which can be as simple as a run of several (4 - 6) random bases. The diversity and statistical distribution of the capture labels detected in sequencing are a direct measure of probe performance. As an example, imagine a case where the sequences that associate with a particular probe are very few. One might attribute the lack of these sequences to poor probe performance and thus be tempted to initiate an iterative cycle of probe redesign. However, insufficient presentation of sequences can also be the result of sequences that do not ligate well to the adapter and / or do not amplify well with the particular PCR regimen used. The use of capture labels enables differentiation between these defective modes. If the performance of the probe is poor, the very few capture events that actually occur will be manifested as very few capture labels that appear multiple times. In contrast, if the presentation is poor due to reasons (ligation, PCR, end repair, etc.) prior to the actual capture reaction, a large configuration of uniquely presented capture labels will result. In certain embodiments, as the transition is made to the automated design of thousands of probes, the ability to quality control probe performance in an information - theoretic manner will become increasingly important.

[0483] (Example 17) Probe Selection and Implementation

[0484] Overview The selection of probe sequences and the methods of using them have not necessarily been developed in concert. This example describes the probe selection criteria in Section I and the laboratory methods that make them most effective in Section II. For example, refer to Figure 30.

[0485] Section I: Selection of Targeted Probes Most generally stated, the target enrichment probe is 60 nucleotides in length. The fact that the probes are generally directed means that they capture the sequence on one side (generally the 3' side) of their position. In addition to the 60-mer targeting core, a tail sequence is also added that adds additional functionality (e.g., a binding site for a complementary oligo that allows for PCR primer binding sites, pull-out by biotin, etc.). The 60-nucleotide targeting sequence is subject to the following constraints and criteria: (1) the probe should be placed at -100 to +50 nucleotides relative to the start point of the target sequence. In FIG. 30, the "start point" of the target sequence is the intron:exon junction; (2) the probe should be designed with redundancy such that the sequences obtained from the probe pairs overlap with reverse orientation, as illustrated; (3) the probe should be selected (if possible) to have a GC content of 33% or more (>20 G or C per 60-mer) and 67% or less (<40 G or C per 60-mer); (4) the probe should always be selected to avoid repeats if possible. This can be done with the help of REPEATMASKER and / or unique alignment criteria, both of which can be viewed on the UCSC Genome Browser; and (5) in the event that the position requirements, GC requirements, and uniqueness requirements cannot be met, the selection rules are selected by relaxing them in the following order (GC > position > uniqueness). In other words, the GC and positioning criteria are not strict, but the criterion of being unique is strict.

[0486] Section II: Laboratory Methods The inputs to target enrichment are the probes, gDNA library, and buffers described elsewhere in this specification. The first step in targeted enrichment is a melting step of the gDNA library, which begins in the form of double-stranded PCR fragments. This is achieved by denaturing the gDNA, preferably at a concentration of 100 ng / μl in a total volume of 10 μl, at 98 °C for 2 minutes, followed by rapid transfer to ice. The gDNA library is suspended in a low-salt buffer containing 10 mM Tris, pH 8.0 and 0.1 mM EDTA. The second step is to add 5 μl of enrichment binding buffer (4 M NaCl, 40 mM Tris, pH 8.0, 0.4 mM EDTA, and 0.4% Tween 20). These conditions are special, but the overall idea is that the salt concentration must be increased to 2 N osmolality to achieve the rapid reaction rate association of complementary DNA strands. 5 microliters of probe is also added ...

Claims

1. A method for targeted genetic analysis, the method comprising: (a) providing a tagged genomic library that hybridizes with a multifunctional capture probe hybrid module complex, wherein the multifunctional capture probe hybrid module complex comprises a multifunctional capture probe hybrid module that hybridizes with a partner oligonucleotide, wherein the multifunctional capture probe hybrid module is a polynucleotide, and the polynucleotide comprises: (i) a first region that hybridizes with the partner oligonucleotide and can function as a PCR primer, and (ii) a second region that can hybridize with a specific genomic target region, wherein the multifunctional capture probe hybrid module comprises the second region that selectively hybridizes with the specific genomic target region in the tagged genomic library, and (b) isolating the tagged genomic library - multifunctional capture probe hybrid module complex obtained from (a); (c) performing 5'-to-3' DNA polymerase extension of the multifunctional capture probe hybrid module on the complex obtained from (b) to generate a hybrid nucleic acid molecule, thereby replicating a region of the genomic target region located 3' to the multifunctional capture probe hybrid module, wherein the hybrid nucleic acid molecule comprises the multifunctional capture probe hybrid module and a complement of a region of the genomic target region located 3' in the direction from the position where the multifunctional capture probe hybrid module hybridizes with the genomic target region, and (d) performing targeted genetic analysis on the hybrid nucleic acid molecule obtained from (c). A method comprising the above steps.

2. The method according to claim 1, wherein the targeted genetic analysis is sequence analysis.

3. The method according to claim 1 or 2, wherein steps (a) to (c) are repeated at least twice, and the targeted genetic analysis in step (d) comprises sequence alignment of the hybrid nucleic acid molecule sequences obtained from the at least two steps (d).

4. At least two different multifunctional capture probe hybrid modules are used in said at least two (a) steps, and each of said at least two (a) steps uses one type of multifunctional capture probe hybrid module. The method according to claim 3.

5. At least one multifunctional capture probe hybrid module hybridizes downstream of said genomic target region, and at least one multifunctional capture probe hybrid module hybridizes upstream of said genomic target region. The method according to claim 4.

6. Said tagged genomic library is Comprising genomic DNA derived from a biological sample selected from the group consisting of blood, skin, hair, hair follicles, saliva, oral mucosa, vaginal mucosa, sweat, tears, epithelial tissue, urine, semen, sperm fluid, seminal plasma, prostatic fluid, bulbourethral gland fluid (Cowper's gland fluid), excreta, biopsy, ascites, cerebrospinal fluid, lymph fluid and tissue extract sample or biopsy sample. The method according to any one of claims 1 to 5.

7. In said step (b), said tagged genomic library - multifunctional capture probe hybrid module complex is isolated using a biotin - streptavidin isolation technique, and said partner oligonucleotide is modified to contain biotin at the 5' end or 3' end. The method according to any one of claims 1 to 6.

8. Said tagged genomic library is (1) Treating fragmented genomic DNA with a terminal repair enzyme to generate fragmented end - repaired genomic DNA, and (2) Ligating a random nucleic acid tag sequence to said fragmented end - repaired genomic DNA to generate said tagged genomic library and can be obtained thereby. The method according to any one of claims 1 to 7.

9. Said tagged genomic library is (1) Treating fragmented genomic DNA with a terminal repair enzyme to generate fragmented end - repaired genomic DNA, and (2) Ligating a multifunctional adapter module to said fragmented end - repaired genomic DNA to generate said tagged genomic library, wherein said multifunctional adapter module (i) a first region containing a random nucleic acid tag sequence, and (ii) a second region containing a sample code sequence, and (iii) a third region containing a PCR primer sequence comprising The method according to any one of claims 1 to 7, which can be obtained by

10. The method according to claim 9, wherein the sample code sequence is used to identify samples within a set of multiplexed samples.

11. The method according to claim 9 or 10, wherein the random nucleic acid tag sequence is used for copy number analysis.

12. Further comprising performing PCR amplification of the hybrid nucleic acid molecule using a nucleic acid comprising the nucleotide sequence of the first region of the multifunctional capture probe hybrid module as a reverse PCR primer, according to any one of claims 1 to 11 The method described in the item.

13. The method according to claim 12, wherein the amplification generates nucleic acids for sequencing.

14. One or more compositions, comprising (a) a multifunctional capture probe hybrid module, wherein the multifunctional capture probe hybrid module is a polynucleotide, and the polynucleotide is (i) a first region capable of hybridizing with a partner oligonucleotide, and (ii) a second region capable of hybridizing with a specific genomic target region, a multifunctional capture probe hybrid module (b) the partner oligonucleotide comprising A nucleic acid comprising the nucleotide sequence of the first region of the multifunctional capture probe hybrid module can amplify a hybrid nucleic acid molecule comprising the multifunctional capture probe hybrid module and a complement of a region of the genomic target region located 3' from the position where the multifunctional capture probe hybrid module hybridizes to the genomic target region. Composition.

15. The one or more compositions according to claim 14, wherein the hybrid nucleic acid molecule is generated by DNA polymerase extension from 5' to 3' of the multifunctional capture probe hybrid module to replicate a region of the genomic target region that is 3' of the multifunctional capture probe hybrid module.

16. The one or more compositions according to claim 14 or 15, wherein the nucleic acid comprising the nucleotide sequence of the first region of the multifunctional capture probe hybrid module can amplify the hybrid nucleic acid molecule to generate a nucleic acid for sequencing.

17. The one or more compositions, (a)a first composition comprising the multifunctional capture probe hybrid module, and (b)a second composition comprising the partner oligonucleotide The one or more compositions according to any one of claims 14 to 16, comprising.

18. The one or more compositions according to any one of claims 14 to 16, wherein the one or more compositions comprise a composition comprising a multifunctional capture probe hybrid module complex, the multifunctional capture probe hybrid module complex comprising the multifunctional capture probe hybrid module and the partner oligonucleotide, and the first region hybridizes with the partner oligonucleotide.

19. The partner oligonucleotide comprises a specific member of a binding pair configured to enable isolation and / or purification of a complex comprising the multifunctional capture probe hybrid module that hybridizes to a DNA fragment, and the specific member of the binding pair is a hapten. The one or more compositions according to any one of claims 14 to 18.

20. The partner oligonucleotide comprises a specific member of a binding pair configured to enable isolation and / or purification of a complex comprising the multifunctional capture probe hybrid module that hybridizes to a DNA fragment, and the specific member of the binding pair is biotin. The one or more compositions according to any one of claims 14 to 18.

21. The partner oligonucleotide comprises a specific member of a binding pair configured to enable isolation and / or purification of a complex comprising the multifunctional capture probe hybrid module that hybridizes to a DNA fragment, and the specific member of the binding pair is a dinitrophenol molecule or a digoxigenin molecule. The one or more compositions according to any one of claims 14 to 18.

22. One or more compositions according to any one of claims 14 to 21, wherein the first region of the multifunctional capture probe hybrid module comprises a tail sequence at the 5' end.

23. One or more compositions according to any one of claims 14 to 21, wherein the multifunctional capture probe hybrid module comprises a third region comprising a tail sequence at the 5' end.

24. A complex comprising: (a) a tagged DNA molecule comprising a DNA fragment; and (b) a hybrid nucleic acid molecule that hybridizes to the tagged DNA molecule, wherein the hybrid nucleic acid molecule comprises, in order from 5' to 3': (i) a tail sequence comprising a PCR primer binding site; (ii) a capture probe to which a portion of the DNA fragment hybridizes; and (iii) a complement of a region of the tagged DNA molecule located 3' to the capture probe. And a hybrid nucleic acid molecule; and (c) a partner oligonucleotide that hybridizes to the 5' end of the hybrid nucleic acid molecule. A complex.

25. A complex comprising: (a) a multifunctional capture probe hybrid module, wherein the multifunctional capture probe hybrid module is a polynucleotide, and the polynucleotide comprises, in order from 3' to 5': (i) a first region that hybridizes to a partner oligonucleotide; (ii) a second region that can hybridize to a specific genomic target region in a DNA fragment; and (iii) a third region comprising a tail sequence. And a multifunctional capture probe hybrid module; and (b) a tagged DNA molecule comprising the DNA fragment, wherein the tagged DNA molecule hybridizes to the multifunctional capture probe hybrid module, and the 3' end of the tagged DNA molecule comprises a complement of the third region of the multifunctional capture probe hybrid module. A tagged DNA molecule. A complex.

26. A complex comprising: (a) a tagged DNA molecule comprising a DNA fragment; and (b) a multifunctional capture probe hybrid module, wherein the multifunctional capture probe hybrid module is a polynucleotide, and the polynucleotide comprises: (i) a first region that hybridizes to a partner oligonucleotide; (ii) a second region capable of hybridizing with a specific target region in the DNA fragment and comprising a multifunctional capture probe hybrid module, wherein the tagged DNA molecule hybridizes with the multifunctional capture probe hybrid module A complex comprising

27. The complex according to claim 26, wherein the first region of the multifunctional capture probe hybrid module comprises a tail sequence.

28. The complex according to claim 26, wherein the multifunctional capture probe hybrid module comprises a third region comprising a tail sequence.

29. A nucleic acid comprising the nucleotide sequence of the first region of the multifunctional capture probe hybrid module is capable of amplifying a hybrid nucleic acid molecule, The hybrid nucleic acid molecule comprises the multifunctional capture probe hybrid module and a complement of the specific target region located in the 3' direction from the position where the multifunctional capture probe hybrid module hybridizes with the tagged DNA molecule. The complex according to any one of claims 26 to 28.

30. The complex according to claim 29, wherein the nucleic acid comprising the nucleotide sequence of the first region of the multifunctional capture probe hybrid module is capable of amplifying the hybrid nucleic acid molecule to generate a nucleic acid for sequencing.

31. The partner oligonucleotide comprises a specific member of a binding pair configured to enable isolation and / or purification of the complex, and the specific member of the binding pair is a hapten. The complex according to any one of claims 24 to 30.

32. The partner oligonucleotide comprises a specific member of a binding pair configured to enable isolation and / or purification of the complex, and the specific member of the binding pair is biotin. The complex according to any one of claims 24 to 30.

33. The partner oligonucleotide comprises a specific member of a binding pair configured to enable isolation and / or purification of the complex, and the specific member of the binding pair is a dinitrophenol molecule or a digoxigenin molecule. The complex according to any one of claims 24 to 30.

Citation Information

Patent Citations

  • Methods for detecting genetic variations in DNA samples

    WO2010129937A2