Compositions and methods for high-precision tumor assays
The composition of multiple primer reagents for tumor gene sequences addresses the challenge of detecting biomarkers in cancer therapies, enabling efficient and standardized assays for diagnosing and monitoring treatment responses in tumor samples.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-08-28
- Publication Date
- 2026-04-02
AI Technical Summary
Existing cancer therapies face challenges in effectively identifying responsive candidates and monitoring responses due to the complexity of tumor microenvironments and the need for higher-throughput, systematic, and standardized assays to detect multiple relevant biomarkers across various sample types.
A composition comprising multiple primer reagents targeting tumor gene sequences for rapid detection of low levels of actionable biomarkers, such as EGFR, ALK, BRAF, ROS1, HER2, MET, NTRK, and RET, in samples like FFPE tissue and plasma, using an integrated and automated workflow.
Enables efficient detection of key biomarkers, facilitating potential diagnoses, prognoses, candidate treatment regimens, and adverse event monitoring through multiplex assays and kits, supporting research and clinical applications with improved detection of gene fusions and mutations.
Smart Images

Figure 0007839727000001 
Figure 0007839727000002 
Figure 0007839727000003
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims priority and interest in U.S. Provisional Application No. 62 / 894,576, filed on 30 August 2019, which is incorporated herein by reference in its entirety.
[0002] Sequence List This application incorporates, by reference, the electronic sequence listing data submitted concurrently. The data in the electronic sequence listing was submitted as a text (.txt) file titled "LT01496_STX.txt" created on August 27, 2020, which has a file size of 550KB and is incorporated herein by reference in its entirety.
[0003] This invention relates to compositions of libraries of target nucleic acid sequences, methods for preparing them, and their uses. [Background technology]
[0004] Advances in cancer therapy are beginning to offer promising results across oncology. Targeted therapies, immune checkpoint inhibitors, cancer vaccines, and T-cell therapies are showing more sustainable outcomes in more responsive populations than conventional chemotherapy. However, effectively identifying responsive candidates and / or monitoring responses has proven challenging. There is an urgent need to better understand the tumor microenvironment, tumor evolution, and biomarkers of drug response. Higher-throughput, systematic, and standardized assay solutions are desired that can efficiently and effectively detect multiple relevant biomarkers across various sample types. [Overview of the project]
[0005] In one aspect of the present invention, a composition is provided for single-stream multiplexing of actionable tumor biomarkers in a sample. In some embodiments, the composition comprises multiple primer reagents directed to multiple target sequences for rapid and effective detection of low levels of targets in a sample. The provided composition targets tumor gene sequences selected from targets among DNA hotspot mutation genes, copy number variation (CNV) genes, intergene fusion genes, and intragene fusion genes. The provided composition maximizes the detection of key biomarkers, such as EGFR, ALK, BRAF, ROS1, HER2, MET, NTRK, and RET, from a variety of samples (e.g., FFPE tissue, plasma) in an integrated and automated workflow.
[0006] In some embodiments, multiple actionable target genes in a sample determine changes in tumor activity in the sample that indicate potential diagnoses, prognoses, candidate treatment regimens, and / or adverse events. In certain embodiments, the provided composition comprises multiple primer reagents selected from Table A. In some embodiments, a multiplex assay comprising the composition of the present invention is provided. In some embodiments, a test kit comprising the composition of the present invention is provided.
[0007] Another aspect of the present invention provides a method for determining actionable tumor biomarkers in a biological sample. Such a method comprises performing multiple amplification of multiple target sequences from a biological sample containing target sequences. Amplification involves contacting at least a portion of the sample containing the multiple target sequences of interest with multiple target-specific primers in the presence of polymerase under amplification conditions to generate multiple amplified target sequences. The method further comprises detecting the presence of each of the multiple target tumor sequences, and the detection of one or more actionable tumor biomarkers compared to a control sample determines changes in tumor activity in the sample that indicate potential diagnoses, prognoses, candidate treatment regimens, and / or adverse events. The method described herein utilizes the compositions of the present invention provided herein. In some embodiments, the target gene is selected from the group consisting of DNA hotspot mutation genes, copy number variation (CNV) genes, intergene fusion genes, and intragene fusion genes. In certain embodiments, the target gene is selected from the genes in Table 1. In certain embodiments, the target gene consists of the genes in Table 1.
[0008] Furthermore, the use of the provided compositions and kits containing the provided compositions for the analysis of nucleic acid library sequences is an additional embodiment of the present invention. In some embodiments, the analysis of the sequences of the resulting libraries enables the detection of low-frequency alleles, improved detection of gene fusions and novel fusions, and / or detection of gene mutations in a sample of interest and / or multiple samples of interest. In certain embodiments, manual, partially automated, and fully automated implementations of the use of the provided compositions and methods are contemplated. In certain embodiments, the use of the provided compositions is implemented in a fully integrated library preparation, templating, and sequencing system for the genetic analysis of samples. In certain embodiments, the use of the provided compositions and methods of the present invention benefits research and clinical applications, including first-line testing of tissue and / or plasma specimens, and continuous monitoring of specimens for the detection of biomarker recurrence and / or resistance.
[0009] All publications, patents, and patent applications referenced herein are incorporated by reference to the same extent that each individual publication, patent, or patent application is specifically and individually incorporated by reference.
[0010] An efficient method for generating targeted libraries containing actionable tumor biomarkers from complex samples is desirable for various nucleic acid analyses. The present invention provides, among other things, a method for preparing libraries of targeted nucleic acid sequences, enabling the rapid generation of highly multiplexed targeted libraries containing unique tag sequences, the resulting library compositions being useful for a variety of applications, including sequencing applications. The provided compositions are designed for the detection of mutations, copy number variations (CNVs), and gene fusions in tissue and plasma-derived samples. The provided compositions include targeted primer panels and reagents for use in high-throughput samples to bring about next-generation workflows for genetic analysis. In certain embodiments, use is implemented in a fully integrated sample-to-analysis system. Novel features of the present invention are described in detail in the appended claims, and a complete understanding of the features and advantages of the present invention will be obtained by referring to the following detailed description which describes exemplary embodiments in which the principles of the present invention are utilized. [Modes for carrying out the invention]
[0011] The section headings used herein are for structural purposes only and should not be construed as limiting the subject matter described herein. All documents and similar materials cited herein, including but not limited to patents, patent applications, articles, books, papers, and internet web pages, are expressly incorporated by reference in their entirety for any purpose. If the definitions of terms in the incorporated references differ from those provided in these instructions, the definitions provided in these instructions shall prevail. It will be understood that “approximately” is implied before temperatures, concentrations, times, etc., discussed herein, so that the deviations are very small and insignificant. In this application, the use of singular forms includes plural forms unless otherwise specifically stated. It should be noted that, when used herein, the singular forms “a,” “an,” and “the,” as well as any singular use of a word, include multiple referents unless explicitly and unquestionably limited to a single referent. Furthermore, the use of “comprise,” “comprises,” “comprising,” “contain,” “contains,” “containing,” “include,” “includes,” and “including” is not intended to be limiting. It should be understood that both general descriptions are merely illustrative and descriptive and do not limit the invention.
[0012] Unless otherwise defined, scientific and technical terms used in connection with the invention described herein shall have meanings generally understood by those skilled in the art. Furthermore, unless otherwise required by context, singular terms shall include plural terms and plural terms shall include singular terms. Generally, the terms and techniques used herein in connection with cell and tissue culture, molecular biology, and protein and oligo or polynucleotide chemistry and hybridization are well known and commonly used in the art. Unless otherwise indicated, the implementation of this subject matter may utilize conventional techniques and descriptions of organic chemistry, molecular biology (including recombinant techniques), cell biology, and biochemistry that are within the scope of the art. Such conventional techniques include, but are not limited to, the preparation of synthetic polynucleotides, polymerization techniques, chemical and physical analysis of polymer particles, preparation of nucleic acid libraries, nucleic acid sequencing and analysis, etc. Specific diagrams of suitable techniques can be used by referring to the examples provided herein. Other equivalent conventional procedures may also be used. Such conventional techniques and explanations can be found in standard laboratory manuals such as Genome Analysis: A Laboratory Manual Series (Vols. I-IV), PCR Primer: A Laboratory Manual, and Molecular Cloning: A Laboratory Manual (all from Cold Spring Harbor Laboratory Press), Hermanson, Bioconjugate Techniques, Second Edition (Academic Press, 2008), Merkus, Particle Size Measurements (Springer, 2009), and Rubinstein and Colby, Polymer Physics (Oxford University Press, 2003). When used in accordance with the embodiments provided herein, the following terms shall be understood to have the following meanings unless otherwise indicated.
[0013] As used herein, "amplify", "amplifying", or "amplification reaction", and derivatives thereof, generally refer to the action or process by which at least a portion of a nucleic acid molecule (referred to as a template nucleic acid molecule) is replicated or copied into at least one additional nucleic acid molecule. The additional nucleic acid molecule optionally contains a sequence that is substantially identical or substantially complementary to at least some portion of the template nucleic acid molecule. The template target nucleic acid molecule can be single-stranded or double-stranded. The additional resulting replicated nucleic acid molecules can independently be single-stranded or double-stranded. In some embodiments, amplification includes a template-dependent in vitro enzymatic catalytic reaction for the generation of at least one copy of at least some portion of the target nucleic acid molecule, or for the generation of at least one copy of a target nucleic acid sequence that is complementary to at least some portion of the target nucleic acid molecule. Amplification optionally includes linear or exponential replication of the nucleic acid molecule. In some embodiments, such amplification is performed using isothermal conditions, and in other embodiments, such amplification can include thermal cycling. In some embodiments, amplification is multiplex amplification that includes the simultaneous amplification of multiple target sequences in a single amplification reaction. At least some of the target sequences can be located on the same nucleic acid molecule or on different target nucleic acid molecules included in a single amplification reaction. In some embodiments, "amplify" includes the amplification of at least some portion of DNA-based nucleic acids and / or RNA-based nucleic acids, whether alone or in combination. The amplification reaction can include a single-stranded or double-stranded nucleic acid substrate and can further include any amplification process known to those skilled in the art. In some embodiments, the amplification reaction includes polymerase chain reaction (PCR). In some embodiments, the amplification reaction includes isothermal amplification.
[0014] As used herein, “amplification conditions” and its derivatives (e.g., amplification conditions) generally refer to conditions suitable for amplifying one or more nucleic acid sequences. Amplification can be linear or exponential. In some embodiments, amplification conditions include isothermal conditions, thermal cycling conditions, or a combination of isothermal and thermal cycling conditions. In some embodiments, conditions suitable for amplifying one or more target nucleic acid sequences include polymerase chain reaction (PCR) conditions. Typically, amplification conditions refer to a reaction mixture sufficient to amplify nucleic acids such as one or more target sequences, or to amplify target sequences linked to one or more adapters, e.g., adapter-linked target sequences. Generally, amplification conditions include a catalyst for amplification or nucleic acid synthesis, e.g., polymerase, primers having some complementarity to the nucleic acid to be amplified, and nucleotides such as deoxyribonucleoside triphosphates (dNTPs) that, when hybridized to the nucleic acid, promote primer elongation. Amplification conditions may require a denaturation step in which the primer hybridizes or anneales to the nucleic acid, extends the primer, and separates the extended primer from the nucleic acid sequence being amplified. Typically, but not always, amplification conditions may include thermal cycling. In some embodiments, amplification conditions include multiple cycles in which the steps of annealing, extension, and separation are repeated. Typically, amplification conditions include Mg ++ or Mn ++ It may contain cations such as (e.g., MgCl2) and optionally include various modifiers of ionic strength.
[0015] As used herein, the terms "target sequence", "target nucleic acid sequence", or "target sequence of interest" and derivatives generally refer to any single-stranded or double-stranded nucleic acid sequence that can be amplified or synthesized according to the present disclosure, including any nucleic acid sequence suspected or expected to be present in a sample. In some embodiments, the target sequence exists in double-stranded form and includes a specific nucleotide sequence, or at least a portion of its complement, that is amplified or synthesized prior to the addition of target-specific primers or additional adapters. The target sequence can include a nucleic acid to which a primer useful in an amplification or synthesis reaction can hybridize prior to elongation by a polymerase. In some embodiments, the term refers to a nucleic acid sequence whose nucleotide sequence identity, order, or position is determined by one or more of the methods of the present disclosure.
[0016] As used herein, the term "portion" and variations thereof, when used with respect to a given nucleic acid molecule, e.g., a primer or template nucleic acid molecule, include any number of contiguous nucleotides within the length of the nucleic acid molecule, including a portion or the full length of the nucleic acid molecule.
[0017] As used herein, “contact” and its derivatives, when used in relation to two or more components, refer to any process in which the approach, proximity, mixing, or incorporation of the referenced components is facilitated or achieved without necessarily requiring physical contact of such components, and includes mixing solutions containing any one or more of the referenced components with each other. The referenced components may be contacted in any particular order or combination, and the particular order of the enumeration of components is not limiting. For example, “contacting A with B and C” includes embodiments in which A is contacted first with B and then with C, as well as embodiments in which C is contacted with A and then with B, and embodiments in which a mixture of A and C is contacted with B, and so on. Furthermore, such contact does not necessarily require that the final result of the contact process be a mixture containing all of the referenced components, as long as all of the referenced components are present at the same time or are simultaneously contained in the same mixture or solution at some point during the contact process. For example, “contacting A with B and C” may include embodiments in which C is first contacted with A to form a first mixture, then the first mixture is contacted with B to form a second mixture, and subsequently C is removed from the second mixture, and A is also optionally removed, leaving only B. If one or more of the reference components to be contacted include multiple components (e.g., “contacting a target sequence with multiple target-specific primers and polymerases”), each of the multiple members can be seen as an individual component of the contact process, such that the contact may include contacting one or more of the multiple members with any of the other members of the multiple and / or any of the other reference components (e.g., some but not all of the multiple target-specific primers may be contacted with the target sequence, then the polymerase, and then the other members of the multiple target-specific primers) in any order or combination.
[0018] As used herein, the term “primer” and its derivatives generally refer to any polynucleotide that can hybridize to a target sequence of interest. In some embodiments, primers can also help initiate nucleic acid synthesis. Typically, a primer functions as a substrate from which a nucleotide can be polymerized by polymerase, but in some embodiments, a primer can provide a site into which another primer can hybridize to initiate the synthesis of a new chain complementary to the synthetic nucleic acid molecule. A primer may consist of any combination of nucleotides or their analogues that can be optionally linked to form a linear polymer of any preferred length. In some embodiments, the primer is a single-stranded oligonucleotide or polynucleotide. (For the purposes of this disclosure, the terms “polynucleotide” and “oligonucleotide” are used interchangeably herein and do not necessarily indicate a difference in length between the two). In some embodiments, the primer is double-stranded. In the case of a double-stranded primer, the primer is first treated to separate its strands before being used to prepare the extension product. Preferably, the primer is an oligodeoxyribonucleotide. The primer must be long enough to initiate the synthesis of the extension product. The length of the primer will depend on many factors, including temperature, the source of the primer, and the method used. In some embodiments, the primer acts as an initiation point for amplification or synthesis when exposed to amplification or synthesis conditions, and such amplification or synthesis may occur in a template-dependent manner, resulting in the formation of a primer extension product that is optionally complementary to at least a portion of the target sequence. Exemplary amplification or synthesis conditions may include contacting the primer with a polynucleotide template (e.g., a template containing the target sequence), nucleotides, and an inducer such as polymerase at a suitable temperature and pH to induce polymerization of nucleotides to the ends of a target-specific primer. In the case of double-stranded primers, the primer may optionally be treated to separate its strands before being used to prepare the primer extension product.In some embodiments, the primer is an oligodeoxyribonucleotide or oligoribonucleotide. In some embodiments, the primer may contain one or more nucleotide analogs. The exact length and / or composition, including the sequence, of a target-specific primer can affect many properties, including the melting temperature (Tm), GC content, secondary structure formation, repeating nucleotide motifs, the length of the predicted primer extension product, the degree of coverage across the nucleic acid molecule of interest, the number of primers present in a single amplification or synthesis reaction, and the presence of nucleotide analogs or modified nucleotides in the primer. In some embodiments, the primer can pair with a compatible primer in the amplification or synthesis reaction to form a primer pair consisting of a forward primer and a reverse primer. In some embodiments, the forward primer of the primer pair contains a sequence that is substantially complementary to at least a portion of the chain of the nucleic acid molecule, and the reverse primer of the primer pair contains a sequence that is substantially identical to at least a portion of the chain. In some embodiments, the forward and reverse primers can hybridize to the opposite chain of the nucleic acid double helix. Optionally, a forward primer initiates the synthesis of a first nucleic acid strand, and a reverse primer initiates the synthesis of a second nucleic acid strand, and the first and second strands can be substantially complementary to each other or hybridize to form a double-stranded nucleic acid molecule. In some embodiments, one end of the amplified or synthesized product is defined by the forward primer, and the other end of the amplified or synthesized product is defined by the reverse primer. In some embodiments, when amplification or synthesis of a long primer extension product is required, such as amplification of an exon, coding region, or gene, several primer pairs can be made that extend to the desired length to allow sufficient amplification of the region. In some embodiments, the primers can contain one or more cleavable groups. In some embodiments, the primer lengths are in the range of about 10 to about 60 nucleotides, about 12 to about 50 nucleotides, and about 15 to about 40 nucleotides.Typically, when exposed to amplification conditions in the presence of dNTPs and polymerase, primers can hybridize to the corresponding target sequence and undergo primer extension. In some cases, a specific nucleotide sequence or portion of a primer can be known at the start of the amplification reaction or determined by one or more of the methods disclosed herein. In some embodiments, the primers contain one or more cleavable groups at one or more positions within the primer.
[0019] As used herein, “target-specific primer” and its derivatives generally refer to single-stranded or double-stranded polynucleotides, typically oligonucleotides, that contain at least one sequence that is at least 50% complementary, typically at least 75% complementary, or at least 85% complementary, more typically at least 90% complementary, more typically at least 95% complementary, more typically at least 98%, or at least 99% complementary, or identical to at least a portion of a nucleic acid molecule containing a target sequence. In such cases, the target-specific primer and the target sequence are described as “corresponding” to each other. In some embodiments, a target-specific primer can hybridize to at least a portion of its corresponding target sequence (or its complement), and such hybridization can be optionally carried out under standard hybridization conditions or stringent hybridization conditions. In some embodiments, a target-specific primer cannot hybridize to the target sequence or its complement, but can hybridize to a portion of the nucleic acid chain containing the target sequence or its complement. In some embodiments, the target-specific primer includes at least one sequence that is at least 75% complementary, typically at least 85% complementary, more typically at least 90% complementary, more typically at least 95% complementary, more typically at least 98% complementary, or more typically at least 99% complementary to at least a portion of the target sequence itself. In other embodiments, the target-specific primer includes at least one sequence that is at least 75% complementary, typically at least 85% complementary, more typically at least 90% complementary, more typically at least 95% complementary, more typically at least 98% complementary, or more typically at least 99% complementary to at least a portion of nucleic acid molecules other than the target sequence. In some embodiments, the target-specific primer is substantially complementary to other target sequences present in the sample, and optionally, the target-specific primer is substantially complementary to other nucleic acid molecules present in the sample.In some embodiments, nucleic acid molecules present in a sample that do not contain or correspond to a target sequence (or complement of a target sequence) are referred to as “non-specific” sequences or “non-specific nucleic acids.” In some embodiments, target-specific primers are designed to contain a nucleotide sequence that is substantially complementary to at least a portion of its corresponding target sequence. In some embodiments, target-specific primers are at least 95% complementary, at least 99% complementary, or identical over their entire length to at least a portion of the nucleic acid molecule containing its corresponding target sequence. In some embodiments, target-specific primers may be at least 90%, at least 95%, at least 98%, or at least 99% complementary, or identical over their entire length to at least a portion of its corresponding target sequence. In some embodiments, forward target-specific primers and reverse target-specific primers define a pair of target-specific primers that can be used to amplify a target sequence via template-dependent primer extension. Typically, each primer in a target-specific primer pair contains at least one sequence that is substantially complementary to at least a portion of the nucleic acid molecule containing the corresponding target sequence, but less than 50% complementary to at least one other target sequence in the sample. In some embodiments, amplification can be carried out using multiple target-specific primer pairs in a single amplification reaction, each primer pair comprising a forward target-specific primer and a reverse target-specific primer, each containing at least one sequence that is substantially complementary to or substantially identical to the corresponding target sequence in the sample, and each primer pair has a different corresponding target sequence. In some embodiments, a target-specific primer may be substantially non-complementary at its 3' or 5' end to any other target-specific primer present in the amplification reaction. In some embodiments, a target-specific primer may include minimal cross-hybridization to other target-specific primers in the amplification reaction. In some embodiments, a target-specific primer may include minimal cross-hybridization to non-specific sequences in the amplification reaction mixture.In some embodiments, the target-specific primer contains minimal self-complementarity. In some embodiments, the target-specific primer may include one or more cleavable groups located at its 3' end. In some embodiments, the target-specific primer may include one or more cleavable groups located near or around the central nucleotide of the target-specific primer. In some embodiments, one or more target-specific primers may contain only non-cleavable nucleotides at their 5' end. In some embodiments, the target-specific primers may optionally contain minimal nucleotide sequence duplication at their 3' or 5' end compared to one or more different target-specific primers in the same amplification reaction. In some embodiments, one, two, three, four, five, six, seven, eight, nine, or ten or more target-specific primers in a single reaction mixture may include one or more of the embodiments described above. In some embodiments, substantially all of the multiple target-specific primers in a single reaction mixture may include one or more of the embodiments described above.
[0020] As used herein, the term “adapter” refers to a nucleic acid molecule that can be used for the manipulation of a polynucleotide of interest. In some embodiments, the adapter is used for the amplification of one or more target nucleic acids. In some embodiments, the adapter is used in a reaction for sequencing. In some embodiments, the adapter has one or more ends lacking a 5' phosphate residue. In some embodiments, the adapter includes, consists of, or is essentially derived from at least one priming site. Such a priming site containing an adapter may be called a “primer” adapter. In some embodiments, the adapter priming site may be useful in a PCR process. In some embodiments, the adapter includes a nucleic acid sequence that is substantially complementary to the 3' or 5' end of at least one target sequence in a sample, which is referred herein to as a gene-specific target sequence, target-specific sequence, or target-specific primer. In some embodiments, the adapter includes a nucleic acid sequence that is substantially incomplementary to the 3' or 5' end of any target sequence present in the sample. In some embodiments, the adapter includes a single-stranded or double-stranded linear oligonucleotide that is not substantially complementary to the target nucleic acid sequence. In some embodiments, the adapter includes a nucleic acid sequence that is substantially discomplementary to at least one, preferably some or all, nucleic acid molecules of the sample. In some embodiments, preferred adapter lengths are in the range of about 10–75 nucleotides, about 12–50 nucleotides, and about 15–40 nucleotides. Generally, the adapter may include any combination of nucleotides and / or nucleic acids. In some embodiments, the adapter includes one or more cleavable groups at one or more locations. In some embodiments, the adapter includes a sequence that is substantially identical or substantially complementary to at least a portion of a primer, e.g., a universal primer. In some embodiments, the adapter includes a tag sequence that assists in cataloging, identification, or sequencing.In some embodiments, the adapter acts as a substrate for amplification of a target sequence in the presence of polymerase and dNTPs, particularly under preferred temperatures and pH levels.
[0021] As used herein, “polymerase” and its derivatives generally refer to any enzyme capable of catalyzing the polymerization of nucleotides (including their analogues) into nucleic acid chains. Typically, though not necessarily, such nucleotide polymerization may occur in a template-dependent manner. Such polymerases may include, without limitation, native polymerases and any of their subunits and truncations, mutant polymerases, variant polymerases, recombinant, fusion, or otherwise manipulated polymerases, chemically modified polymerases, synthetic molecules or assemblies, and any analogues, derivatives, or fragments thereof that retain the ability to catalyze such polymerization. Optionally, a polymerase may be a mutant polymerase, comprising one or more mutations involving the substitution of one or more amino acids with other amino acids, the insertion or deletion of one or more amino acids from a polymerase, or the joining of two or more polymerase portions. Typically, a polymerase contains one or more active sites from which catalysis of nucleotide joining and / or nucleotide polymerization may occur. Some exemplary polymerases include, without limitation, DNA polymerases and RNA polymerases. As used herein, the term “polymerase” and its variations also refer to a fusion protein comprising at least two linked portions, the first portion comprising a peptide capable of catalyzing the polymerization of nucleotides into a nucleic acid chain, and linked to the second portion comprising a second polypeptide. In some embodiments, the second polypeptide may include a reporter enzyme or a processability-enhancing domain. Optionally, the polymerase may have 5' exonuclease activity or terminal transferase activity. In some embodiments, the polymerase may be optionally reactivated, for example, by heat, the use of chemicals, or by adding a new amount of polymerase to the reaction mixture. In some embodiments, the polymerase may include a hot-start polymerase and / or an aptamer-based polymerase that can be optionally reactivated.
[0022] As used herein, the terms “identity” and “identical” and their variations refer to the sequence similarity of two or more sequences (e.g., nucleotide or polypeptide sequences) when used in reference to two or more nucleic acid sequences. In the case of two or more homologous sequences, the identity or homology percentage of a sequence or a subsequence of a sequence indicates the percentage of all monomeric units (e.g., nucleotides or amino acids) that are the same (i.e., about 70% identity, preferably 75%, 80%, 85%, 90%, 95%, 98%, or 99% identity). The identity percentage may exceed a specified region when compared and aligned for the greatest match on a comparison window, or by using the BLAST or BLAST 2.0 sequence comparison algorithm with the default parameters described below, or by manual alignment and visual inspection. Sequences are said to be “substantially identical” if they have at least 85% identity at the amino acid or nucleotide level. Preferably, identity exists over a region of at least approximately 25, 50, or 100 residues in length, or over the entire length of at least one comparison sequence. Typical algorithms for determining the percentage of sequence identity and sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al, Nuc. Acids Res. 25:3389-3402 (1977). Other methods include algorithms such as those by Smith & Waterman, Adv. Appl. Math. 2:482 (1981), and Needleman & Wunsch, J. Mol. Biol. 48:443 (1970). Another indicator that two nucleic acid sequences are substantially identical is that the two molecules or their complements hybridize to each other under stringent hybridization conditions.
[0023] As used herein, the terms “complementary” and “complement,” and their variations thereof, refer to any two or more nucleic acid sequences (e.g., a portion or all of a template nucleic acid molecule, a target sequence, and / or primers) that can undergo cumulative base pairing at two or more individual corresponding positions in antiparallel orientation, such as a hybridized double helix. Such base pairing can proceed according to any set of established rules, e.g., according to the Watson-Crick base pairing rules, or according to some other base pairing paradigm. Optionally, there may be “complete” or “total” complementarity between a first nucleic acid sequence and a second nucleic acid sequence, where each nucleotide in the first nucleic acid sequence can undergo stabilizing base pairing interactions with the nucleotide at the corresponding antiparallel position on the second nucleic acid sequence. “Partial” complementarity describes nucleic acid sequences in which at least 20% but less than 100% of the residues in one nucleic acid sequence are complementary to residues in the other nucleic acid sequence. In some embodiments, at least 50% but less than 100% of the residues in one nucleic acid sequence are complementary to residues in another nucleic acid sequence. In some embodiments, at least 70%, 80%, 90%, 95%, or 98% but less than 100% of the residues in one nucleic acid sequence are complementary to residues in another nucleic acid sequence. If at least 85% of the residues in one nucleic acid sequence are complementary to residues in another nucleic acid sequence, the sequences are said to be "substantially complementary." In some embodiments, two complementary or substantially complementary sequences can hybridize to each other under standard or stringent hybridization conditions. "Discomplementary" describes nucleic acid sequences in which less than 20% of the residues in one nucleic acid sequence are complementary to residues in another nucleic acid sequence. If less than 15% of the residues in one nucleic acid sequence are complementary to residues in another nucleic acid sequence, the sequences are said to be "substantially discomplementary." In some embodiments, two non-complementary or substantially non-complementary sequences cannot hybridize to each other under standard or stringent hybridization conditions. “Mismatch” occurs when two opposing nucleotides are located at any position where they are not complementary.Complementary nucleotides include nucleotides that are efficiently incorporated by opposing DNA polymerases during DNA replication under physiological conditions. In typical embodiments, complementary nucleotides can form base pairs with each other between the nucleic acid bases of nucleotides and / or polynucleotides that are antiparallel to each other, such as AT / U and GC base pairs formed through specific Watson-Crick type hydrogen bonds, or base pairs formed through some other type of base pairing paradigm. The complementarity of other artificial base pairs can be based on other types of hydrogen bonding and / or hydrophobicity of the bases and / or shape complementarity between the bases.
[0024] As used herein, “amplified target sequence” and its derivatives generally refer to the amplification of a target sequence / a nucleic acid sequence produced by amplifying these using target-specific primers and the methods provided herein. The amplified target sequence may be either the same sense (positive strand produced in even-numbered amplifications from the second onward) or antisense (i.e., the negative strand produced in odd-numbered amplifications from the first onward) with respect to the target sequence. For the purposes of this disclosure, the amplified target sequence is typically less than 50% complementary to any portion of another amplified target sequence in the reaction.
[0025] As used herein, the terms “to ligate” and their derivatives generally refer to the act or process of covalently bonding two or more molecules together, for example, covalently bonding two or more nucleic acid molecules to one another. In some embodiments, ligation includes joining nicks between adjacent nucleotides of nucleic acids. In some embodiments, ligation includes forming a covalent bond between the ends of a first nucleic acid molecule and the ends of a second nucleic acid molecule. In some embodiments, for example, embodiments in which the nucleic acid molecules to be ligated contain conventional nucleotide residues, ligation may include forming a covalent bond between the 5' phosphate group of one nucleic acid and the 3' hydroxyl group of a second nucleic acid, thereby forming the ligated nucleic acid molecules. In some embodiments, any means can be used to join nicks or to bond the 5' phosphate to the 3' hydroxyl between adjacent nucleotides. In exemplary embodiments, enzymes such as ligases can be used.
[0026] As used herein, “ligase” and its derivatives generally refer to any agent capable of catalyzing the linking of two substrate molecules. In some embodiments, a ligase includes an enzyme capable of catalyzing the joining of a nick between adjacent nucleotides of a nucleic acid. In some embodiments, a ligase includes an enzyme capable of catalyzing the formation of a covalent bond between the 5' phosphate of one nucleic acid molecule and the 3' hydroxyl of another nucleic acid molecule, thereby forming a linked nucleic acid molecule. Preferred ligases may include, but are not limited to, T4 DNA ligase, T7 DNA ligase, Taq DNA ligase, and E. coli DNA ligase.
[0027] As defined herein, a “cleavable group” generally refers to any portion of a nucleic acid that, once incorporated, can be cleaved under appropriate conditions. For example, a cleavable group can be incorporated into a target-specific primer, amplified sequence, adapter, or nucleic acid molecule of a sample. In exemplary embodiments, a target-specific primer may contain a cleavable group that is incorporated into the amplification product and subsequently cleaved after amplification, thereby removing some or all of the target-specific primer from the amplification product. A cleavable group can be cleaved or otherwise removed from a target-specific primer, amplified sequence, adapter, or nucleic acid molecule of a sample by any acceptable means. For example, a cleavable group can be removed from a target-specific primer, amplified sequence, adapter, or nucleic acid molecule of a sample by enzymes, heat, photo-oxidation, or chemical treatment. In one embodiment, a cleavable group may include nucleic acid bases that do not exist in nature. For example, an oligodeoxyribonucleotide may include one or more RNA nucleic acid bases, such as uracil, which can be removed by uracil glycosylase. In some embodiments, the cleavable group may include one or more modified nucleic acid bases (such as 7-methylguanine, 8-oxo-guanine, xanthine, hypoxanthine, 5,6-dihydrouracil, or 5-methylcytosine) or one or more modified nucleosides (i.e., 7-methylguanosine, 8-oxo-deoxyguanosine, xanthosine, inosine, dihydrouridine, or 5-methylcytidine). The modified nucleic acid bases or nucleotides can be removed from the nucleic acid by enzymatic, chemical, or thermal means. In one embodiment, the cleavable group may include a moiety that can be removed from the primer after amplification (or synthesis) upon exposure to ultraviolet light (i.e., bromodeoxyuridine). In another embodiment, the cleavable group may include methylated cytosine. Typically, methylated cytosine can be cleaved from the primer, for example, after induction of amplification (or synthesis) upon treatment with sodium bisulfite. In some embodiments, the cleavable moiety may include a limiting site.For example, a primer or target sequence may contain nucleic acid sequences specific to one or more restriction enzymes, and following amplification (or synthesis), the primer or target sequence may be treated with one or more restriction enzymes to remove cleavable groups. Typically, a target-specific primer, amplified sequence, adapter, or nucleic acid molecule of a sample may contain one or more cleavable groups at one or more positions.
[0028] As used herein, “digestion,” “digestion step,” and its derivatives generally refer to any process by which a cleavable group is cleaved or otherwise removed from a target-specific primer, amplified sequence, adapter, or nucleic acid molecule of a sample. In some embodiments, the digestion step includes a chemical, thermal, photo-oxidative, or digestive process.
[0029] As used herein, the term “hybridization” is consistent with its use in the art and generally refers to the process by which two nucleic acid molecules undergo base-pairing interactions. Two nucleic acid molecules are said to be hybridized if any portion of one nucleic acid molecule is base-paired with any portion of the other nucleic acid molecule, and it is not necessarily required that the two nucleic acid molecules hybridize over their respective total lengths, and in some embodiments, at least one of the nucleic acid molecules may contain portions that do not hybridize with the other nucleic acid molecule. The phrase “hybridizing under stringent conditions” and its variations generally refer to conditions under which the hybridization of a target-specific primer to a target sequence occurs in the presence of a high hybridization temperature and low ionic intensity. As used herein, the phrase “standard hybridization conditions” and its variations generally refer to conditions under which the hybridization of a primer to an oligonucleotide (i.e., a target sequence) occurs in the presence of a low hybridization temperature and high ionic intensity. In one exemplary embodiment, standard hybridization conditions include an aqueous environment containing about 100 mM magnesium sulfate, about 500 mM tris sulfate at pH 8.9, and about 200 mM ammonium sulfate or its equivalent at about 50–55°C.
[0030] As used herein, the term “terminus” and its variations may include the terminal 30 nucleotides, terminal 20, and more typically terminal 15 nucleotides of a nucleic acid molecule, for example, when used in relation to a target sequence or an amplified target sequence. A linear nucleic acid molecule consisting of a series of linked nucleotides typically includes at least two ends. In some embodiments, one end of a nucleic acid molecule may include a 3' hydroxyl group or an equivalent, and may be referred to as the “3' end” and its derivatives. Optionally, the 3' end includes a 3' hydroxyl group that is not bonded to the 5' phosphate group of a mononucleotide pentose ring. Typically, the 3' end includes one or more 5' bonded nucleotides located adjacent to the nucleotide containing the unbonded 3' hydroxyl group, typically 30 nucleotides located adjacent to the 3' hydroxyl, typically terminal 20, and more typically terminal 15 nucleotides. Generally, one or more bonded nucleotides may be expressed as a percentage of the nucleotides present in the oligonucleotide, or they may be provided as several bonded nucleotides adjacent to the unbonded 3' hydroxyl. For example, the 3' end may contain less than 50% of the nucleotide length of the oligonucleotide. In some embodiments, the 3' end may contain any portion that does not contain any unbound 3' hydroxyl groups but can function as an attachment site for nucleotides by primer extension and / or nucleotide polymerization. In some embodiments, the term “3' end” may, for example, refer to a target-specific primer, contain terminal 10 nucleotides, terminal 5 nucleotides, terminal 4, 3, 2, or fewer nucleotides. In some embodiments, the term “3' end” may, for example, refer to a target-specific primer, contain nucleotides located 10 nucleotides or less from the 3' end. As used herein, “5' end” and its derivatives generally refer to the end of a nucleic acid molecule, for example, a target sequence or amplified target sequence containing a free 5' phosphate group or its equivalent.In some embodiments, the 5' end contains a 5' phosphate group that is not bound to the 3' hydroxyl of an adjacent mononucleotide pentose ring. Typically, the 5' end contains one or more bound nucleotides located adjacent to the 5' phosphate, typically 30 nucleotides located adjacent to the nucleotide containing the 5' phosphate group, typically 20 terminal nucleotides, and more typically 15 terminal nucleotides. Generally, one or more bound nucleotides can be expressed as a percentage of the nucleotides present in the oligonucleotide, or can be provided as several bound nucleotides adjacent to the 5' phosphate. For example, the 5' end may be less than 50% of the nucleotide length of the oligonucleotide. In another exemplary embodiment, the 5' end may contain about 15 nucleotides adjacent to the nucleotide containing the terminal 5' phosphate. In some embodiments, the 5' end may not contain any unbound 5' phosphate groups but may contain a 3' hydroxyl group or any portion that can function as an attachment site to the 3' end of another nucleic acid molecule. In some embodiments, the term “5' end” may, for example, refer to a target-specific primer, and the 5' end may contain terminal 10 nucleotides, terminal 5 nucleotides, terminal 4, 3, 2, or fewer nucleotides. In some embodiments, the term “5' end” may include nucleotides located 10 or less from the 5' end when referring to a target-specific primer. In some embodiments, the 5' end of a target-specific primer may include only non-cleavable nucleotides, e.g., nucleotides that do not contain one or more cleavable groups as disclosed herein, or cleavable nucleotides as readily determined by those skilled in the art. The “first end” and “second end” of a polynucleotide refer to the 5' end or 3' end of the polynucleotide. Either the first or second end of a polynucleotide may be the 5' end or the 3' end of the polynucleotide, and the terms “first” and “second” are not intended to indicate that the end is specifically the 5' end or the 3' end.
[0031] As used herein, “tag,” “barcode,” “unique tag,” or “tag sequence,” and its derivatives, generally refer to a unique short (6-14 nucleotide) nucleic acid sequence within an adapter or primer that can function as a “key” for distinguishing or separating multiple amplified target sequences in a sample. For the purposes of this disclosure, the barcode or unique tag sequence is incorporated into the nucleotide sequence of the adapter or primer. As used herein, “barcode sequence” refers to a nucleic acid immobilized sequence that is sufficient to enable the identification of a sample or source of the nucleic acid sequence of interest. The barcode sequence may, but does not necessarily, be a subdivision of the original nucleic acid sequence on which the identification is based. In some embodiments, the barcode is 5-20 nucleic acid lengths. In some embodiments, the barcode includes analogous nucleotides, e.g., L-DNA, LNA, PNA, etc. As used herein, “unique tag sequence” refers to a nucleic acid sequence having at least one random sequence and at least one immobilized sequence. A unique tag sequence, alone or in combination with a second unique tag sequence, is sufficient to enable the identification of a single target nucleic acid molecule in a sample. The unique tag sequence may, but does not necessarily, contain a sub-segment of the original target nucleic acid sequence. In some embodiments, the unique tag sequence is 2 to 50 nucleotides or base pairs long, or 2 to 25 nucleotides or base pairs long, or 2 to 10 nucleotides or base pairs long. The unique tag sequence may contain at least one random sequence interspersed with fixed sequences.
[0032] As used herein, “comparable highest and lowest melting temperatures” and its derivatives generally refer to the melting temperature (Tm) of each nucleic acid fragment of a single adapter or target-specific primer after digestion of cleavable groups. The hybridization temperatures of each nucleic acid fragment produced by the adapter or target-specific primer are compared to determine the highest and lowest temperatures required to prevent hybridization of the nucleic acid sequence from the target-specific primer or adapter or its fragment or portion to its respective target. Once the highest hybridization temperature is known, it is possible to manipulate the adapter or target-specific primer to achieve comparable highest and lowest melting temperatures for each nucleic acid fragment, for example, by shifting the position of one or more cleavable groups along the length of the primer, thereby optimizing the digestion and repair steps of library preparation.
[0033] As used herein, “add-only” and its derivatives generally refer to a series of steps in which reagents and components are added to a first or single reaction mixture. Typically, the series of steps excludes the removal of the reaction mixture from a first container to a second container in order to complete the series of steps. Generally, an add-only process excludes the handling of the reaction mixture outside the container containing the reaction mixture. Typically, add-only processes are suitable for automation and high throughput.
[0034] As used herein, “polymerization conditions” and its derivatives generally refer to conditions suitable for nucleotide polymerization. In typical embodiments, such nucleotide polymerization is catalyzed by polymerase. In some embodiments, polymerization conditions include conditions for primer extension in an optional, template-dependent manner, resulting in the generation of a synthesized nucleic acid sequence. In some embodiments, polymerization conditions include polymerase chain reaction (PCR). Typically, polymerization conditions are sufficient to synthesize nucleic acids and involve the use of a reaction mixture containing polymerase and nucleotides. Polymerization conditions can include conditions for annealing target-specific primers to a target sequence and for primer extension in a template-dependent manner in the presence of polymerase. In some embodiments, polymerization conditions can be carried out using thermal cycling. In addition, polymerization conditions can include multiple cycles in which the steps of annealing, extension, and separation of two nucleic acid chains are repeated. Typically, polymerization conditions include cations such as MgCl2. Generally, polymerization of one or more nucleotides to form a nucleic acid chain involves the nucleotides being linked to each other via phosphodiester bonds, although alternative bonding may be possible under the circumstances of certain nucleotide analogs.
[0035] As used herein, the term “nucleic acid” refers to natural nucleic acids, artificial nucleic acids, their analogues, or combinations thereof, including polynucleotides and oligonucleotides. As used herein, the terms “polynucleotide” and “oligonucleotide” are used interchangeably and are not limited to these, but mean single-stranded and double-stranded polymers of nucleotides, including 2'-deoxyribonucleotides (nucleic acids) and ribonucleotides (RNA) linked by internucleotide phosphodiester bonds, e.g., 3'-5' and 2'-5', reverse bonds, e.g., 3'-3' and 5'-5', branched structures, or analog nucleic acids. Polynucleotides are H + NH4 + , trialkylammonium, Mg 2+ kaNa +They have associated counterions such as . Oligonucleotides can consist of deoxyribonucleotides alone, ribonucleotides alone, or chimeric mixtures thereof. Oligonucleotides can consist of nucleic acid bases and sugar analogs. Polynucleotides typically range in size from a few monomer units, e.g., 5 to 40, to several thousand monomer nucleotide units, when they are more commonly referred to as oligonucleotides in the art, and for the purposes of this disclosure, however, both oligonucleotides and polynucleotides may be of any preferred length. Unless otherwise indicated, whenever an oligonucleotide sequence is expressed, it will be understood that the nucleotides are in 5' to 3' order from left to right, with "A" representing deoxyadenosine, "C" representing deoxycytidine, "G" representing deoxyguanosine, "T" representing thymidine, and "U" representing deoxyuridine. As discussed herein and known in the art, mononucleotides are reacted to form oligonucleotides via phosphodiesters or other preferred bonds, typically through the attachment of a 5' phosphate or equivalent group of one nucleotide to a 3' hydroxyl or equivalent group of its adjacent nucleotide, so oligonucleotides and polynucleotides are said to have a "5' end" and a "3' end".
[0036] As used herein, the term “polymerase chain reaction” (“PCR”) refers to the methods of KBMullis U.S. Patents 4,683,195 and 4,683,202, incorporated herein by reference, which describe a method for increasing the concentration of a desired polynucleotide segment in a mixture of genomic DNA without cloning or purification. This process for amplifying a desired polynucleotide consists of introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired polynucleotide, followed by thermal cycling in the exact order in the presence of DNA polymerase. The two primers are complementary to their respective strands of the double-stranded polynucleotide of interest. To perform amplification, the mixture is denatured and the primers are then annealed to their complementary sequences within the polynucleotide of the target molecule. Following annealing, the primers are extended with polymerase to form a new pair of complementary strands. The steps of denaturation, primer annealing, and polymerase extension can be repeated many times to obtain an amplified segment of the desired polynucleotide at a high concentration (i.e., denaturation, annealing, and extension constitute one “cycle,” and there can be many “cycles”). The length of the amplified segment (amplicon) of the desired polynucleotide is determined by the relative positions of the primers to each other, and therefore this length is a controllable parameter. By repeating the process, the method is called a “polymerase chain reaction” (hereinafter “PCR”). The desired amplified segment of the polynucleotide is said to be “PCR amplified” because it becomes the dominant nucleic acid sequence (in terms of concentration) in the mixture. When defined herein, target nucleic acid molecules in a sample containing multiple target nucleic acid molecules are amplified via PCR. In the modifications of the method discussed above, target nucleic acid molecules can be PCR amplified using multiple different primer pairs, in some cases one or more primer pairs per target nucleic acid molecule of interest, thereby forming a multiplex PCR reaction product. Using multiplex PCR, it is possible to amplify multiple target nucleic acid molecules from a sample simultaneously to form amplified target sequences.Several different methodologies are used (e.g., quantification by bioanalyzer or qPCR, hybridization with labeled probes, incorporation of biotinylated primers, followed by avidin enzyme-coupled detection, dCTP, or dATP). 32 It is also possible to detect the amplified target sequence by incorporating P-labeled deoxynucleotide triphosphates into the amplified target sequence. Any oligonucleotide sequence can be amplified with a suitable set of primers, thereby enabling the amplification of target nucleic acid molecules from genomic DNA, cDNA, formalin-fixed paraffin-embedded DNA, fine-needle biopsies, and various other sources. In particular, the amplified target sequences created by the multiplex PCR process disclosed herein are themselves efficient substrates for subsequent PCR amplification or various downstream assays or operations.
[0037] Where used herein, “multiplex amplification” refers to the selective and non-random amplification of two or more target sequences in a sample using at least one target-specific primer. In some embodiments, multiplex amplification is carried out so that some or all of the target sequences are amplified in a single reaction vessel. The “plexy” or “plex” of a given multiplex amplification generally refers to the number of different target-specific sequences amplified during that single multiplex amplification. In some embodiments, the plexy may be about 12 plexes, 24 plexes, 48 plexes, 96 plexes, 192 plexes, 384 plexes, 768 plexes, 1536 plexes, 3072 plexes, 6144 plexes or more.
[0038] composition To determine the state of tumors within a sample, we developed a single-stream multi-next-generation sequencing workflow for determining tumor biomarkers of actionable tumors within a sample. The tumor high-precision assay compositions and methods of the present invention provide specific and robust solutions for biomarker screening to understand the mechanisms involved in tumor immune responses. Accordingly, compositions for multi-library preparation are provided and can be used in combination with manual or automated next-generation sequencing technologies and workflow solutions (e.g., Ion Torrent® NGS workflows) to evaluate low-level biomarker targets of various sample types to assess tumor state.
[0039] Accordingly, compositions for single-stream multiplex determination of actionable tumor biomarkers in a sample are provided. In some embodiments, the composition comprises multiple sets of primer-pair reagents directed to multiple target sequences to detect low-level targets in the sample, and the target genes are selected from tumor response genes consisting of the following functions: DNA hotspot mutation genes, copy number variation (CNV) genes, intergene fusion genes, and intragene fusion genes. In some embodiments, the target genes are selected from oncogenes consisting of one or more functions from Table 1. In some embodiments, the target genes are selected from one or more actionable target genes in the sample that determine changes in tumor activity in the sample indicating potential diagnostic, prognosis, candidate treatment regimens, and / or adverse events. Overall, the diverse functions of the genes constituting the provided multiplex panel of the present invention provide a comprehensive concept that recommends an actionable approach to cancer therapy.
[0040] In certain embodiments, the target tumor sequence is directed to a sequence having cancer-related mutations. In some embodiments, the target sequence or amplified target sequence is directed to head and neck cancers (e.g., HNSCC, nasopharyngeal, salivary gland), brain cancers (e.g., glioblastoma, glioma, gliosarcoma, glioblastoma multiforme, neuroblastoma), breast cancers (e.g., TNBC, trastuzumab-resistant HER2+ breast cancer, ER+ / HER- breast cancer), gynecological cancers (e.g., uterine, ovarian, cervical, endometrial, Faropian cancer), colorectal cancer, gallbladder cancer, esophageal cancer, gastrointestinal cancer, stomach cancer, bladder cancer, prostate cancer, testicular cancer, urothelial carcinoma, liver cancer. The sequence is directed to have mutations associated with one or more solid tumor cancers selected from the group consisting of visceral cancers (e.g., hepatocyte, HCC), lung cancers (e.g., non-small cell lung, small cell lung), kidney (renal cell) cancer, pancreatic cancers (e.g., adenocarcinoma, glandular cancer), thyroid cancer, cholangiocarcinoma, pituitary tumors, Wilms' tumor, Kaposi's sarcoma, hairy cell carcinoma, osteosarcoma, thymic carcinoma, skin cancer, melanoma, cardiac cancer, oral and laryngeal cancers, neuroblastoma, mesothelioma, and other solid tumors (thymus, bone, soft tissue, oral SCC, myelofibrosis, synovial sarcoma). In one embodiment, the mutations may include substitutions, insertions, inversions, point mutations, deletions, mismatches, and translocations. In some embodiments, the target sequence or amplified target sequence is directed to a sequence having mutations associated with one or more blood / hematologic cancers selected from the group consisting of multiple myeloma, diffuse large B-cell lymphoma (DLBCL), lymphoma, Hodgkin lymphoma, non-Hodgkin lymphoma, follicular lymphoma, leukemia, acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), and myelodysplastic syndromes. In one embodiment, the cancer-associated mutational biomarker is located in at least one of the genes provided in Table 1.
[0041] In some embodiments, one or more mutant tumor sequences are located in at least one of the genes selected from Table 1. In some embodiments, one or more mutant sequences exhibit cancer activity.
[0042] In some embodiments, one or more variant sequences indicate the potential of a patient to respond to a therapeutic agent. In some embodiments, one or more variant tumor biomarker sequences indicate the potential of a patient to not respond to a therapeutic agent. In certain embodiments, the therapeutic agents in question may include, but are not limited to, kinase inhibitors, cell signaling inhibitors, checkpoint blockers, T-cell therapies, and therapeutic vaccines.
[0043] In some embodiments, the target sequence or variant target sequence is directed to cancer-related mutations. In some embodiments, the target sequence or variant target sequence is directed to head and neck cancers (e.g., HNSCC, nasopharyngeal, salivary gland), brain cancers (e.g., glioblastoma, glioma, gliosarcoma, glioblastoma multiforme, neuroblastoma), breast cancers (e.g., TNBC, trastuzumab-resistant HER2+ breast cancer, ER+ / HER- breast cancer), gynecological cancers (e.g., uterine, ovarian, cervical, endometrial, Faropian cancer), colorectal cancer, gallbladder cancer, esophageal cancer, gastrointestinal cancer, stomach cancer, bladder cancer, prostate cancer, testicular cancer, urothelial carcinoma The mutations are directed towards mutations associated with one or more solid tumor cancers selected from the group consisting of liver cancer (e.g., hepatocyte, HCC), lung cancer (e.g., non-small cell lung, small cell lung), kidney (renal cell) cancer, pancreatic cancer (e.g., adenocarcinoma, glandular cancer), thyroid cancer, bile duct cancer, pituitary tumor, Wilms' tumor, Kaposi's sarcoma, hairy cell carcinoma, osteosarcoma, thymic carcinoma, skin cancer, melanoma, cardiac cancer, oral and laryngeal cancer, neuroblastoma, mesothelioma, and other solid tumors (thymus, bone, soft tissue, oral SCC, myelofibrosis, synovial sarcoma). In one embodiment, the mutations may include substitutions, insertions, inversions, point mutations, deletions, mismatches, and translocations. In one embodiment, the mutations may include copy number variations. In one embodiment, the mutations may include germline mutations or somatic mutations. In some embodiments, the target sequence or amplified target sequence is directed to a sequence having one or more mutations associated with hematological / hematological cancers selected from the group consisting of multiple myeloma, diffuse large B-cell lymphoma (DLBCL), lymphoma, Hodgkin lymphoma, non-Hodgkin lymphoma, follicular lymphoma, leukemia, acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), and myelodysplastic syndromes.
[0044] In one embodiment, the cancer-associated mutation is located in at least one of the genes provided in Table 1. In some embodiments, the mutant target sequence is directed to one or more of the genes provided in Table 1. In some embodiments, the mutant target sequence includes any one or more amplicon sequences of the genes provided in Table 1. In some embodiments, the mutant target sequence consists of any one or more amplicon sequences of the genes provided in Table 1. In some embodiments, the mutant target sequence includes each of the amplicon sequences of the genes provided in Table 1.
[0045] In some embodiments, the composition comprises one or more of the tumor target-specific primer pairs provided in Table A. In some embodiments, the composition comprises all of the tumor target-specific primer pairs provided in Table A. In some embodiments, one or more of the tumor target-specific primer pairs provided in Table A can be used to amplify target sequences present in a sample, as disclosed by the method described herein.
[0046] In some embodiments, the tumor target-specific primers from Table A include 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 40, 60, 80, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, or more target-specific primer pairs. In some embodiments, the amplified target sequence may include one or more of the amplified target sequences produced using the target-specific primers provided in Table A. In some embodiments, at least one of the cancer-related target-specific primers is at least 90% identical to at least one nucleic acid sequence produced using a target-specific primer selected from SEQ ID NOs: 1-1563. In some embodiments, at least one of the tumor-related target-specific primers is complementary to at least one target sequence in the sample over its entire length. In some embodiments, at least one of the immune response-related target-specific primers includes an uncleavable nucleotide at its 3' end. In some embodiments, the uncleavable nucleotide at the 3' end includes a terminal 3' nucleotide. In one embodiment, the amplified target sequence is directed to one or more individual exons that have cancer-related mutations.
[0047] method The methods provided by the present invention include efficient procedures that enable the rapid preparation of highly multiplexed libraries suitable for downstream analysis. The methods optionally allow the incorporation of one or more unique tag sequences. Certain methods include streamlined, additive-only procedures that convey extremely rapid library generation.
[0048] Provided herein are methods for determining tumor activity in a sample. In some embodiments, the method comprises multiple amplification of multiple tumor sequences from a biological sample, the amplification comprising contacting at least a portion of the sample with a set of multiple primer-pair reagents directed to multiple target sequences and a polymerase under amplification conditions, thereby generating amplified target-expressed sequences. The method further comprises detecting the presence of mutations in one or more target sequences in the sample, and mutations in one or more tumor markers compared to a control determine a change in tumor activity in the sample. In some embodiments, the tumor sequences of the method are selected from tumor response genes consisting of the following functions: DNA hotspot mutation genes, copy number variation (CNV) genes, intergene fusion genes, and intragene fusion genes. In some embodiments, the target genes are selected from tumor genes consisting of one or more functions from Table 1. In some embodiments, the target genes are selected from one or more actionable target genes in the sample that determine a change in tumor activity in the sample indicating potential diagnostic, prognosis, candidate treatment regimens, and / or adverse events. Overall, the diverse functions of the genes constituting the multi-panel provided by this invention offer a comprehensive concept that promotes an actionable approach to cancer therapy.
[0049] In certain embodiments, the target tumor sequence of the method is directed to a sequence having cancer-related mutations. In some embodiments, the target sequence or amplified target sequence may be head and neck cancers (e.g., HNSCC, nasopharyngeal, salivary gland), brain cancers (e.g., glioblastoma, glioma, gliosarcoma, glioblastoma multiforme, neuroblastoma), breast cancers (e.g., TNBC, trastuzumab-resistant HER2+ breast cancer, ER+ / HER- breast cancer), gynecological cancers (e.g., uterine, ovarian, cervical, endometrial, Faropian cancer), colorectal cancer, gallbladder cancer, esophageal cancer, gastrointestinal cancer, stomach cancer, bladder cancer, prostate cancer, testicular cancer, urothelial carcinoma, liver cancer. The sequence is directed to have mutations associated with one or more solid tumor cancers selected from the group consisting of visceral cancers (e.g., hepatocyte, HCC), lung cancers (e.g., non-small cell lung, small cell lung), kidney (renal cell) cancer, pancreatic cancers (e.g., adenocarcinoma, glandular cancer), thyroid cancer, cholangiocarcinoma, pituitary tumors, Wilms' tumor, Kaposi's sarcoma, hairy cell carcinoma, osteosarcoma, thymic carcinoma, skin cancer, melanoma, cardiac cancer, oral and laryngeal cancers, neuroblastoma, mesothelioma, and other solid tumors (thymus, bone, soft tissue, oral SCC, myelofibrosis, synovial sarcoma). In one embodiment, the mutations may include substitutions, insertions, inversions, point mutations, deletions, mismatches, and translocations. In some embodiments, the target sequence or amplified target sequence is directed to a sequence having mutations associated with one or more hematological / hematological cancers selected from the group consisting of multiple myeloma, diffuse large B-cell lymphoma (DLBCL), lymphoma, Hodgkin lymphoma, non-Hodgkin lymphoma, follicular lymphoma, leukemia, acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), and myelodysplastic syndromes. In one embodiment, the cancer-associated mutational biomarker is located in at least one of the genes provided in Table 1.
[0050] In some embodiments, one or more mutant tumor sequences of this method are located in at least one of the genes selected from Table 1. In some embodiments, one or more mutant sequences exhibit cancer activity.
[0051] In some embodiments, one or more variant sequences of the method indicate the potential for patients to respond to the therapeutic agent. In some embodiments, one or more variant tumor biomarker sequences indicate the potential for patients to not respond to the therapeutic agent. In certain embodiments, the therapeutic agents in question may include, but are not limited to, kinase inhibitors, cell signaling inhibitors, checkpoint blockers, T-cell therapies, and therapeutic vaccines.
[0052] In some embodiments, the target sequence or variant target sequence of the method is directed to cancer-related mutations. In some embodiments, the target sequence or variant target sequence of the method is directed to head and neck cancers (e.g., HNSCC, nasopharyngeal cancer, salivary gland cancer), brain cancers (e.g., glioblastoma, glioma, gliosarcoma, glioblastoma multiforme, neuroblastoma), breast cancers (e.g., TNBC, trastuzumab-resistant HER2+ breast cancer, ER+ / HER- breast cancer), gynecological cancers (e.g., uterine cancer, ovarian cancer, cervical cancer, endometrial cancer, Fallopian cancer), colorectal cancer, gallbladder cancer, esophageal cancer, gastrointestinal cancer, stomach cancer, bladder cancer, prostate cancer, testicular cancer, urothelial cancer. The mutations are directed towards one or more solid tumor cancer-related mutations selected from the group consisting of cancer, liver cancer (e.g., hepatocyte, HCC), lung cancer (e.g., non-small cell lung, small cell lung), kidney (renal cell) cancer, pancreatic cancer (e.g., adenocarcinoma, glandular cancer), thyroid cancer, bile duct cancer, pituitary tumor, Wilms' tumor, Kaposi's sarcoma, hairy cell carcinoma, osteosarcoma, thymic carcinoma, skin cancer, melanoma, cardiac cancer, oral and laryngeal cancer, neuroblastoma, mesothelioma, and other solid tumors (thymus, bone, soft tissue, oral SCC, myelofibrosis, synovial sarcoma). In one embodiment, the mutations may include substitutions, insertions, inversions, point mutations, deletions, mismatches, and translocations. In one embodiment, the mutations may include copy number variations. In one embodiment, the mutations may include germline mutations or somatic mutations. In some embodiments, the target sequence or amplified target sequence is directed to a sequence having one or more mutations associated with hematological / hematological cancers selected from the group consisting of multiple myeloma, diffuse large B-cell lymphoma (DLBCL), lymphoma, Hodgkin lymphoma, non-Hodgkin lymphoma, follicular lymphoma, leukemia, acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), and myelodysplastic syndromes.
[0053] In one embodiment, the cancer-associated mutation is located in at least one of the genes provided in Table 1. In some embodiments, the mutant target sequence is directed to one or more of the genes provided in Table 1. In some embodiments, the mutant target sequence includes any one or more amplicon sequences of the genes provided in Table 1. In some embodiments, the mutant target sequence consists of any one or more amplicon sequences of the genes provided in Table 1. In some embodiments, the mutant target sequence includes each of the amplicon sequences of the genes provided in Table 1.
[0054] In some embodiments, the method includes the use of one or more of the tumor target-specific primer pairs provided in Table A. In some embodiments, the method includes the use of all of the tumor target-specific primer pairs provided in Table A. In some embodiments, the use of one or more of the tumor target-specific primer pairs provided in Table A can be used to amplify target sequences present in a sample, as disclosed by the method described herein.
[0055] In some embodiments, the method includes the use of tumor target-specific primers from Table A, which include 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 40, 60, 80, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, or more target-specific primer pairs. In some embodiments, the method, which includes the detection of amplified target sequences, may include one or more of the amplified target sequences produced using the target-specific primers provided in Table A. In some embodiments, the method includes the use of at least one cancer-related target-specific primer that is at least 90% identical to at least one nucleic acid sequence produced using a target-specific primer selected from SEQ ID NOs: 1 to 1563. In some embodiments, at least one of the tumor-related target-specific primers is complementary to at least one target sequence in the sample over its entire length. In some embodiments, at least one of the immune response-related target-specific primers includes an uncleavable nucleotide at its 3' end. In some embodiments, the uncleavable nucleotide at the 3' end includes a terminal 3' nucleotide. In one embodiment, the amplified target sequence is directed to one or more individual exons having cancer-related mutations.
[0056] In some embodiments, the method includes the detection and, optionally, the identification of clinically actionable markers. As defined herein, the term “clinically actionable marker” includes clinically actionable mutations and / or clinically actionable expression patterns that are known to or can be associated with the prognosis of cancer treatment by those skilled in the art. In one embodiment, the prognosis of cancer treatment includes the identification of mutations and / or expression patterns associated with the responsiveness or non-responsiveness of cancer to a drug, drug combination, or treatment plan. In one embodiment, the method includes the amplification of multiple target sequences from a population of nucleic acid molecules associated with or correlated with the onset, progression, or remission of cancer. In some embodiments, the provided method includes the selective amplification of one or more target sequences in a sample, as well as the detection and / or identification of cancer-related mutations. In some embodiments, the amplified target sequences include two or more nucleotide sequences of genes provided in Table 1. In some embodiments, the amplified target sequences may include any one or more amplified target sequences produced using target-specific primers provided in Table A. In one embodiment, the amplified target sequence includes amplicons of 10, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600 or more genes from Table 1.
[0057] One aspect of the present invention provides a method for preparing a library of target nucleic acid sequences. In some embodiments, the method includes contacting a nucleic acid sample with a plurality of adapters capable of amplifying one or more target nucleic acid sequences in the sample, under conditions that the target nucleic acid undergoes a first amplification; digesting the resulting first amplification product to reduce or eliminate the resulting primer dimers to prepare a partially digested target amplicon, thereby generating a gapped double-stranded amplicon. The method further includes repairing the partially digested target amplicon and then amplifying the repaired target amplicon using universal primers in a second amplification, thereby generating a library of target nucleic acid sequences. Each of the plurality of adapters used in the method herein includes a universal handle sequence, a target nucleic acid sequence, a cleavable portion, and optionally one or more tag sequences. Methods provide at least two and up to 100,000 target-specific adapter pairs, where the target nucleic acid sequence of each adapter includes at least one cleavable portion, and the universal handle sequence does not include a cleavable portion. In some embodiments, where any tag array includes at least one adapter, the severable portion is included in the adapter array adjacent to any end of the tag array.
[0058] One aspect of the present invention provides a method for preparing a tagged library of target nucleic acid sequences. In some embodiments, the method includes contacting a nucleic acid sample with a plurality of adapters capable of amplifying one or more target nucleic acid sequences in the sample under conditions that the target nucleic acid undergoes a first amplification; digesting the resulting first amplification product to reduce or eliminate the resulting primer dimers to prepare a partially digested target amplicon, thereby generating a gapped double-stranded amplicon. The method further includes repairing the partially digested target amplicon and then amplifying the repaired target amplicon using universal primers in a second amplification, thereby generating a library of target nucleic acid sequences. Each of the plurality of adapters used in the method herein includes a universal handle sequence, a target nucleic acid sequence, a cleavable portion, and one or more tag sequences. Methods provide at least two and up to 100,000 target-specific adapter pairs, where the target nucleic acid sequence of each adapter includes at least one cleavable portion, the universal handle sequence does not include a cleavable portion, and the cleavable portion is included adjacent to one end of the tag sequence.
[0059] In a particular embodiment, the comparable highest and lowest melting temperatures of each universal sequence are higher than the comparable highest and lowest melting temperatures of each target nucleic acid sequence and each tag sequence present in the adapter.
[0060] In some embodiments, each adapter comprises a unique tag sequence further described herein, each further comprising a cleavable base adjacent to any end of the tag sequence in each adapter. In some embodiments in which the unique tag sequence is used, each generated target-specific amplicon sequence comprises at least one different sequence and up to 10 7 It includes several different sequences. In a particular embodiment, each target-specific pair of multiple adapters includes up to 16,777,216 different adapter combinations, each containing a different tag sequence.
[0061] In some embodiments, the method involves contacting a plurality of gapped polynucleotide products simultaneously with a digestion reagent and a repair reagent. In some embodiments, the method involves sequentially contacting a plurality of gapped polynucleotide products with a digestion reagent, and then with a repair reagent.
[0062] Digestive reagents useful in the methods provided herein include any reagents capable of cleaving cleavable sites present in the adapter, and in some embodiments, but not limited to, one or a combination of uracil DNA glycosylase (UDG), aprine endonuclease (e.g., APE1), RecJf, formamidepyrimidine [fapy]-DNA glycosylase (fpg), Nth endonuclease III, endonuclease VIII, polynucleotide kinase (PNK), Taq DNA polymerase, DNA polymerase I, and / or human DNA polymerase beta.
[0063] Repair reagents useful in the methods provided herein include, but are not limited to, any reagents capable of repairing gapped amplicons, including, in some embodiments, one or a combination of Phusion DNA polymerase, Phusion U DNA polymerase, SuperFi DNA polymerase, Taq DNA polymerase, human DNA polymerase beta, T4 DNA polymerase and / or T7 DNA polymerase, SuperFiU DNA polymerase, E. coli DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, Taq DNA ligase, and / or 9°N DNA ligase.
[0064] Therefore, in certain embodiments, the digestion and repair reagents include one or a combination of uracil DNA glycosylase (UDG), aprine endonuclease (e.g., APE1), RecJf, formamidepyrimidine [fapy]-DNA glycosylase (fpg), Nth endonuclease III, endonuclease VIII, polynucleotide kinase (PNK), Taq DNA polymerase, DNA polymerase I and / or human DNA polymerase beta, as well as one or a combination of Phusion DNA polymerase, Phusion U DNA polymerase, SuperFi DNA polymerase, Taq DNA polymerase, human DNA polymerase beta, T4 DNA polymerase and / or T7 DNA polymerase, SuperFiU DNA polymerase, E. coli DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, Taq DNA ligase, and / or 9°N DNA ligase. In certain embodiments, the digestion and repair reagent comprises one or a combination of uracil DNA glycosylase (UDG), aprine endonuclease (e.g., APE1), Taq DNA polymerase, Phusion U DNA polymerase, SuperFiU DNA polymerase, and T7 DNA ligase. In certain embodiments, the digestion and repair reagent comprises one or a combination of uracil DNA glycosylase (UDG), formamidepyrimidine [fapy]-DNA glycosylase (fpg), Phusion U DNA polymerase, Taq DNA polymerase, SuperFiU DNA polymerase, T4 PNK, and T7 DNA ligase.
[0065] In some embodiments, the method includes a digestion and repair step performed in a single step. In other embodiments, the method includes digestion and repair steps performed at different temperatures and separated by time.
[0066] In some embodiments, the method of the present invention is performed, and one or more of the method steps are performed in manual mode. In certain embodiments, the method of the present invention is performed, and each of the method steps is performed manually. In some embodiments, the method of the present invention is performed, and one or more of the method steps are performed in automatic mode. In certain embodiments, the method of the present invention is performed, and each of the method steps is automated. In some embodiments, the method of the present invention is performed, and one or more of the method steps are performed in a combination of manual and automatic modes.
[0067] In some embodiments, the method of the present invention includes at least one purification step. For example, in a particular embodiment, the purification step is performed only after the second amplification of the repaired amplicon. In some embodiments, two purification steps are utilized, with the first purification step performed after digestion and repair, and the second purification step performed after the second amplification of the repaired amplicon.
[0068] In some embodiments, the purification step includes performing a solid-phase deposition reaction, a solid-phase immobilization reaction, or gel electrophoresis. In certain embodiments, the purification step includes separation performed using solid-phase reversible immobilization (SPRI) beads. In certain embodiments, the purification step includes separation performed using SPRI beads, where the SPRI beads include paramagnetic beads.
[0069] In some embodiments, the method includes contacting a nucleic acid sample with a plurality of adapters capable of amplifying one or more target nucleic acid sequences in the sample under conditions that the target nucleic acid undergoes a first amplification; digesting the resulting first amplification product to reduce or eliminate the resulting primer dimers to prepare a partially digested target amplicon, thereby generating a gapped double-stranded amplicon. The method further includes repairing the partially digested target amplicon, then purifying the repaired amplicon, then amplifying the repaired target amplicon using universal primers in a second amplification, thereby generating a library of target nucleic acid sequences; and then purifying the resulting library. Each of the plurality of adapters used in the method herein includes a universal handle sequence, a target nucleic acid sequence, a cleavable portion, and optionally one or more tag sequences. Methods are provided that offer at least two and up to 100,000 target-specific adapter pairs, wherein the target nucleic acid sequence of each adapter includes at least one cleavable portion, and the universal handle sequence does not include a cleavable portion. In some embodiments, where any tag array includes at least one adapter, the severable portion is included in the adapter array adjacent to any end of the tag array.
[0070] In some embodiments, the method includes contacting a nucleic acid sample with a plurality of adapters capable of amplifying one or more target nucleic acid sequences in the sample under conditions that the target nucleic acid undergoes a first amplification; digesting the resulting first amplification product to reduce or eliminate the resulting primer dimers to prepare a partially digested target amplicon, thereby generating a gapped double-stranded amplicon. The method further includes repairing the partially digested target amplicon, purifying the repaired amplicon, then amplifying the repaired target amplicon using universal primers in a second amplification, thereby generating a library of target nucleic acid sequences; and then purifying the resulting library. Each of the plurality of adapters used in the method herein includes a universal handle sequence, a target nucleic acid sequence, a cleavable portion, and one or more tag sequences. Methods are provided which include at least two and up to 100,000 target-specific adapter pairs, wherein the target nucleic acid sequence of each adapter includes at least one cleavable portion, the universal handle sequence does not include a cleavable portion, and the cleavable portion is included in any adjacent end of the tag sequence.
[0071] In some embodiments, the method includes contacting a nucleic acid sample with a plurality of adapters capable of amplifying one or more target nucleic acid sequences in the sample under conditions that the target nucleic acid undergoes a first amplification; digesting the resulting first amplification product to reduce or eliminate the resulting primer dimers to prepare a partially digested target amplicon, thereby generating a gapped double-stranded amplicon. The method further includes repairing the partially digested target amplicon, then purifying the repaired amplicon, then amplifying the repaired target amplicon using universal primers in a second amplification, thereby generating a library of target nucleic acid sequences; and then purifying the resulting library. Each of the plurality of adapters used in the method herein includes a universal handle sequence, a target nucleic acid sequence, a cleavable portion, and optionally one or more tag sequences. Methods are provided that offer at least two and up to 100,000 target-specific adapter pairs, wherein the target nucleic acid sequence of each adapter includes at least one cleavable portion, and the universal handle sequence does not include a cleavable portion. In some embodiments, where any tag array includes at least one adapter, the severable portion is included in the adapter array adjacent to any end of the tag array.In some embodiments, the digestion and repair reagents include one or a combination of uracil DNA glycosylase (UDG), aprine endonuclease (e.g., APE1), RecJf, formamidepyrimidine [fapy]-DNA glycosylase (fpg), Nth endonuclease III, endonuclease VIII, polynucleotide kinase (PNK), Taq DNA polymerase, DNA polymerase I and / or human DNA polymerase beta, as well as one or a combination of Phusion DNA polymerase, Phusion U DNA polymerase, SuperFi DNA polymerase, Taq DNA polymerase, human DNA polymerase beta, T4 DNA polymerase and / or T7 DNA polymerase, SuperFiU DNA polymerase, E. coli DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, Taq DNA ligase, and / or 9°N DNA ligase. In certain embodiments, the digestion and repair reagent comprises one or a combination of uracil DNA glycosylase (UDG), aprine endonuclease (e.g., APE1), Taq DNA polymerase, Phusion U DNA polymerase, SuperFiU DNA polymerase, and T7 DNA ligase. In certain embodiments, the digestion and repair reagent comprises one or a combination of uracil DNA glycosylase (UDG), formamidepyrimidine [fapy]-DNA glycosylase (fpg), Phusion U DNA polymerase, Taq DNA polymerase, SuperFiU DNA polymerase, T4 PNK, and T7 DNA ligase.
[0072] In some embodiments, the method includes contacting a nucleic acid sample with a plurality of adapters capable of amplifying one or more target nucleic acid sequences in the sample under conditions that the target nucleic acid undergoes a first amplification; digesting the resulting first amplification product to reduce or eliminate the resulting primer dimers to prepare a partially digested target amplicon, thereby generating a gapped double-stranded amplicon. The method further includes repairing the partially digested target amplicon, purifying the repaired amplicon, then amplifying the repaired target amplicon using universal primers in a second amplification, thereby generating a library of target nucleic acid sequences; and then purifying the resulting library. Each of the plurality of adapters used in the method herein includes a universal handle sequence, a target nucleic acid sequence, a cleavable portion, and one or more tag sequences. Methods are provided which include at least two and up to 100,000 target-specific adapter pairs, wherein the target nucleic acid sequence of each adapter includes at least one cleavable portion, the universal handle sequence does not include a cleavable portion, and the cleavable portion is included in any adjacent end of the tag sequence.In some embodiments, the digestion and repair reagents include one or a combination of uracil DNA glycosylase (UDG), aprine endonuclease (e.g., APE1), RecJf, formamidepyrimidine [fapy]-DNA glycosylase (fpg), Nth endonuclease III, endonuclease VIII, polynucleotide kinase (PNK), Taq DNA polymerase, DNA polymerase I and / or human DNA polymerase beta, as well as one or a combination of Phusion DNA polymerase, Phusion U DNA polymerase, SuperFi DNA polymerase, Taq DNA polymerase, human DNA polymerase beta, T4 DNA polymerase and / or T7 DNA polymerase, SuperFiU DNA polymerase, E. coli DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, Taq DNA ligase, and / or 9°N DNA ligase. In certain embodiments, the digestion and repair reagent comprises one or a combination of uracil DNA glycosylase (UDG), aprine endonuclease (e.g., APE1), Taq DNA polymerase, Phusion U DNA polymerase, SuperFiU DNA polymerase, and T7 DNA ligase. In certain embodiments, the digestion and repair reagent comprises one or a combination of uracil DNA glycosylase (UDG), formamidepyrimidine [fapy]-DNA glycosylase (fpg), Phusion U DNA polymerase, Taq DNA polymerase, SuperFiU DNA polymerase, T4 PNK, and T7 DNA ligase.
[0073] In certain embodiments, the method of the present invention is carried out in a single, additive workflow reaction, enabling the rapid generation of highly multiplexed targeted libraries. For example, in one embodiment, a method for preparing a library of target nucleic acid sequences includes contacting a nucleic acid sample with a plurality of adapters capable of amplifying one or more target nucleic acid sequences in the sample, under conditions that the target nucleic acid undergoes a first amplification; digesting the resulting first amplification product to reduce or eliminate the resulting primer dimers, thereby preparing a partially digested target amplicon, and thereby generating a gapped double-stranded amplicon. The method further includes repairing the partially digested target amplicon, and then amplifying the repaired target amplicon using universal primers in a second amplification, thereby generating a library of target nucleic acid sequences, and purifying the resulting library. In certain embodiments, the purification includes a single or repeated separation step that follows the generation of the library after the second amplification, and the other method steps are carried out in a single reaction vessel without the need to transfer any portion (aliquot) of the product generated in the steps to another reaction vessel. Each of the multiple adapters used in the methods herein comprises a universal handle sequence, a target nucleic acid sequence, a cleavable portion, and optionally one or more tag sequences. Methods provide at least two and up to 100,000 target-specific adapter pairs, where the target nucleic acid sequence of each adapter comprises at least one cleavable portion, and the universal handle sequence does not comprise a cleavable portion. In some embodiments in which any tag sequence is included in at least one adapter, the cleavable portion is included in an adapter sequence adjacent to any end of the tag sequence.
[0074] In another embodiment, a method is provided for preparing a tagged library of target nucleic acid sequences, comprising: contacting a nucleic acid sample with a plurality of adapters capable of amplifying one or more target nucleic acid sequences in the sample under conditions that the target nucleic acid undergoes a first amplification; digesting the resulting first amplification product to reduce or eliminate the resulting primer dimers to prepare a partially digested target amplicon, thereby generating a gapped double-stranded amplicon. The method further comprises repairing the partially digested target amplicon, then amplifying the repaired target amplicon using universal primers in a second amplification, thereby generating a library of target nucleic acid sequences, and purifying the resulting library. In certain embodiments, the purification comprises a single or repeated separation step, and the other method steps are optionally performed in a single reaction vessel without the need to transfer any portion of the products generated in the steps to another reaction vessel. Each of the plurality of adapters used in the method herein comprises a universal handle sequence, a target nucleic acid sequence, a cleavable portion, and one or more tag sequences. The method provides at least two and up to 100,000 target-specific adapter pairs, each adapter's target nucleic acid sequence containing at least one cleavable portion, and the universal handle sequence containing no cleavable portions, with any cleavable portions being included adjacent to either end of the tag sequence.
[0075] In one embodiment, a method for preparing a library of target nucleic acid sequences includes: contacting a nucleic acid sample with a plurality of adapters capable of amplifying one or more target nucleic acid sequences in the sample under conditions in which the target nucleic acid undergoes a first amplification; digesting the resulting first amplification product to reduce or eliminate the resulting primer dimers to prepare a partially digested target amplicon, thereby generating a gapped double-stranded amplicon; further comprising repairing the partially digested target amplicon, then amplifying the repaired target amplicon using universal primers in a second amplification, thereby generating a library of target nucleic acid sequences, and purifying the resulting library.
[0076] In certain embodiments, the digestion reagent includes one or a combination of any of the following: uracil DNA glycosylase (UDG), AP endonuclease (APE1), RecJf, formamidepyrimidine [fapy]-DNA glycosylase (fpg), Nth endonuclease III, endonuclease VIII, polynucleotide kinase, Taq DNA polymerase, DNA polymerase I, and / or human DNA polymerase beta. In certain embodiments, the digestion reagent comprises one or a combination of any of the following: uracil DNA glycosylase (UDG), AP endonuclease (APE1), RecJf, formamidepyrimidine [fapy]-DNA glycosylase (fpg), Nth endonuclease III, endonuclease VIII, polynucleotide kinase, Taq DNA polymerase, DNA polymerase I, and / or human DNA polymerase beta, and the digestion reagent lacks formamidepyrimidine [fapy]-DNA glycosylase (fpg).
[0077] In some embodiments, the digestion reagent comprises a single-stranded DNA exonuclease that degrades in the 5'-3' direction. In some embodiments, the cleavage reagent comprises a single-stranded DNA exonuclease that degrades debase sites. In some embodiments herein, the digestion reagent comprises RecJf exonuclease. In certain embodiments, the digestion reagent comprises APE1 and RecJf, and the cleavage reagent comprises an aprine / apyrimidine endonuclease. In certain embodiments, the digestion reagent comprises AP endonuclease (APE1).
[0078] In some embodiments, the repair reagent comprises at least one DNA polymerase, and the gap-filling reagent comprises one or a combination of any of the following: Phusion DNA polymerase, Phusion U DNA polymerase, SuperFi DNA polymerase, Taq DNA polymerase, human DNA polymerase beta, T4 DNA polymerase and / or T7 DNA polymerase and / or SuperFi U DNA polymerase. In some embodiments, the repair reagent further comprises a plurality of nucleotides.
[0079] In some embodiments, the repair reagent comprises an ATP-dependent or ATP-independent ligase, and the repair reagent comprises one or a combination of any of the following: E. coli DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, Taq DNA ligase, and 9°N DNA ligase.
[0080] In certain embodiments, the digestion and repair reagents include one or a combination of uracil DNA glycosylase (UDG), aprine endonuclease (e.g., APE1), RecJf, formamidepyrimidine [fapy]-DNA glycosylase (fpg), Nth endonuclease III, endonuclease VIII, polynucleotide kinase (PNK), Taq DNA polymerase, DNA polymerase I and / or human DNA polymerase beta, as well as one or a combination of Phusion DNA polymerase, Phusion U DNA polymerase, SuperFi DNA polymerase, Taq DNA polymerase, human DNA polymerase beta, T4 DNA polymerase and / or T7 DNA polymerase, SuperFiU DNA polymerase, E. coli DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, Taq DNA ligase, and / or 9°N DNA ligase. In certain embodiments, the digestion and repair reagents include one or a combination of uracil DNA glycosylase (UDG), aprine endonuclease (e.g., APE1), Taq DNA polymerase, Phusion U DNA polymerase, SuperFiU DNA polymerase, and T7 DNA ligase. In certain embodiments, purification includes a single or repeated separation step performed following the generation of a second amplified library, and the method steps are carried out in a single reaction vessel up to the first purification without the need to transfer any portion of the products generated in the steps to another reaction vessel. Each of the plurality of adapters used in the methods herein includes a universal handle sequence, a target nucleic acid sequence, a cleavable portion, and optionally one or more tag sequences. Methods are provided which include at least two and up to 100,000 target-specific adapter pairs, wherein the target nucleic acid sequence of each adapter includes at least one cleavable portion, and the universal handle sequence does not include a cleavable portion.In some embodiments, where any tag array includes at least one adapter, the severable portion is included in the adapter array adjacent to any end of the tag array.
[0081] Another embodiment provides a method for preparing a tagged library of target nucleic acid sequences, comprising: contacting a nucleic acid sample with a plurality of adapters capable of amplifying one or more target nucleic acid sequences in the sample under conditions that the target nucleic acid undergoes a first amplification; digesting the resulting first amplification product to reduce or eliminate the resulting primer dimers to prepare a partially digested target amplicon, thereby generating a gapped double-stranded amplicon. The method further comprises repairing the partially digested target amplicon, then amplifying the repaired target amplicon using universal primers in a second amplification, thereby generating a library of target nucleic acid sequences, and purifying the resulting library. In certain embodiments, the digestion and repair reagents include one or a combination of uracil DNA glycosylase (UDG), aprine endonuclease (e.g., APE1), RecJf, formamidepyrimidine [fapy]-DNA glycosylase (fpg), Nth endonuclease III, endonuclease VIII, polynucleotide kinase (PNK), Taq DNA polymerase, DNA polymerase I and / or human DNA polymerase beta, as well as one or a combination of Phusion DNA polymerase, Phusion U DNA polymerase, SuperFi DNA polymerase, Taq DNA polymerase, human DNA polymerase beta, T4 DNA polymerase and / or T7 DNA polymerase, SuperFiU DNA polymerase, E. coli DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, Taq DNA ligase, and / or 9°N DNA ligase. In certain embodiments, the digestion and repair reagents include one or a combination of uracil DNA glycosylase (UDG), aprine endonuclease (e.g., APE1), Taq DNA polymerase, Phusion U DNA polymerase, SuperFiU DNA polymerase, and T7 DNA ligase.In certain embodiments, purification includes a single or repeated separation step that follows the generation of a second amplified library, while other method steps are carried out in a single reaction vessel without the need to transfer any portion (aliquot) of the product generated in the step to another reaction vessel. Each of the multiple adapters used in the methods herein comprises a universal handle sequence, a target nucleic acid sequence, a cleavable portion, and one or more tag sequences. Methods are provided which offer at least two and up to 100,000 target-specific adapter pairs, where the target nucleic acid sequence of each adapter includes at least one cleavable portion, the universal handle sequence does not include a cleavable portion, and the cleavable portion is included adjacent to one end of the tag sequence.
[0082] In some embodiments, the adapter-dimer byproducts obtained from the first amplification step of the method are largely removed from the resulting library. In certain embodiments, the enriched population of amplified target nucleic acids contains reduced amounts of adapter-dimer byproducts. In certain embodiments, the adapter-dimer byproducts are eliminated.
[0083] In some embodiments, the library is prepared in less than 4 hours. In some embodiments, the library is prepared, concentrated, and sequenced in less than 3 hours. In some embodiments, the library is prepared, concentrated, and sequenced in 2 to 3 hours. In some embodiments, the library is prepared in about 2.5 hours. In some embodiments, the library is prepared in about 2.75 hours. In some embodiments, the library is prepared in about 3 hours.
[0084] composition Additional embodiments of the present invention include compositions comprising multiple nucleic acid adapters, and library compositions prepared according to the methods of the present invention. The provided compositions are useful in conjunction with the methods described herein, as well as for additional analyses and applications known in the art.
[0085] Therefore, a composition is provided comprising multiple nucleic acid adapters, each of which comprises a 5' universal handle sequence, optionally one or more tag sequences, and a 3' target nucleic acid sequence, each adapter comprising a cleavable portion, the target nucleic acid sequence of the adapter comprising at least one cleavable portion, and if a tag sequence is present, the cleavable portion comprising is adjacent to one of the ends of the tag sequence, and the universal handle sequence comprising no cleavable portions. At least two and up to 100,000 target-specific adapter pairs are provided in the composition. The composition enables the rapid generation of highly multiplexed targeted libraries.
[0086] In some embodiments, the provided composition comprises a plurality of nucleic acid adapters, each of which comprises a 5' universal handle sequence, one or more tag sequences, and a 3' target nucleic acid sequence, each adapter comprising a cleavable portion, the target nucleic acid sequence of the adapter comprising at least one cleavable portion, the cleavable portion comprising being adjacent to any end of the tag sequence, and the universal handle sequence not comprising a cleavable portion. At least two and up to 100,000 target-specific adapter pairs are contained in the provided composition. The provided composition enables the rapid generation of highly multiplexed tagged and targeted libraries.
[0087] The primer / adapter composition may be single-stranded or double-stranded. In some embodiments, the adapter composition includes a single-stranded adapter. In some embodiments, the adapter composition includes a double-stranded adapter. In some embodiments, the adapter composition includes a mixture of single-stranded and double-stranded adapters.
[0088] In some embodiments, the composition comprises multiple adapters capable of amplifying one or more target nucleic acid sequences, including multiple adapter pairs capable of amplifying at least two different target nucleic acid sequences, wherein the target-specific primer sequences are substantially non-complementary to other target-specific primer sequences in the composition. In some embodiments, the composition comprises at least 25, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1000, 1250, 1500, 1750, 2000, 2250, 2500, 2750, 3000, 3250, 3500, 3750, 4000, 4500, 5000, 5500, 6000, 7000, 8000, 9000, 10000, 11000, or 12000 or more target-specific adapter pairs. In some embodiments, the target-specific adapter pair comprises approximately 15 to 40 nucleotides in length, with at least one nucleotide replaced by a cleavable group. In some embodiments, the cleavable group is a uridine nucleotide. In some embodiments, the target-specific adapter pair is designed for amplification of exons, genes, exomes, or regions of the genome associated with clinical or pathological conditions, e.g., amplification of one or more sites containing one or more mutations (e.g., driver mutations) associated with cancer, e.g., lung cancer, colon cancer, breast cancer, etc., or amplification of mutations associated with genetic diseases, e.g., cystic fibrosis, muscular dystrophy, etc. In some embodiments, the target-specific adapter pair, when hybridized to a target sequence and amplified as provided herein, produces a library of adapter-linked amplified target sequences having a length of approximately 100 to 600 base pairs. In some embodiments, the adapter-linked amplified target sequences are not overexpressed in the library at a rate of more than 30% compared to the remainder of other adapter-linked amplified target sequences in the library. In some embodiments, the adapter-coupled amplification target sequence library is substantially homogeneous with respect to the GC content, amplification target sequence length, or melting temperature (Tm) of each target sequence.
[0089] In some embodiments, the target-specific primer sequences of the adapter pair in the composition of the present invention are target-specific sequences capable of amplifying a specific region of a nucleic acid molecule. In some embodiments, the target-specific adapter can amplify genomic DNA or cDNA. In some embodiments, the target-specific adapter can amplify mammalian nucleic acids, such as human DNA or RNA, mouse DNA or RNA, bovine DNA or RNA, canine DNA or RNA, horse DNA or RNA, or any other mammal of interest, but is not limited to these. In other embodiments, the target-specific adapter includes sequences directed to amplify plant nucleic acids of interest. In other embodiments, the target-specific adapter includes sequences directed to amplify infectious agents, such as bacterial and / or viral nucleic acids. In some embodiments, the amount of nucleic acid required for selective amplification is about 1 ng to 1 microgram. In some embodiments, the amount of nucleic acid required for selective amplification of one or more target sequences is about 1 ng, about 5 ng, or about 10 ng. In some embodiments, the amount of nucleic acid required for selective amplification of target sequences is about 10 ng to about 200 ng.
[0090] As described herein, each of the multiple adapters includes a 5' universal handle sequence. In some embodiments, the universal handle sequence includes one or a combination of any of the amplification primer binding sequence, the sequencing primer binding sequence, and / or the capture primer binding sequence. In some embodiments, the comparable highest and lowest melting temperatures of each adapter universal handle sequence are higher than the comparable highest and lowest melting temperatures of each target nucleic acid sequence and each tag sequence present in the same adapter. Preferably, the universal handle sequences of the provided adapter do not show significant complementarity and / or hybridization to any portion of the unique tag sequence and / or target nucleic acid sequence of interest. In some embodiments, the first universal handle sequence includes one or a combination of any of the amplification primer binding sequence, the sequencing primer binding sequence, and / or the capture primer binding sequence. In some embodiments, the second universal handle sequence includes one or a combination of any of the amplification primer binding sequence, the sequencing primer binding sequence, and / or the capture primer binding sequence. In certain embodiments, the first and second universal handle sequences correspond to forward and reverse universal handle sequences, and in certain embodiments, the same first and second universal handle sequences are included for each of a plurality of target-specific adapter pairs. Such forward and reverse universal handle sequences are targeted together with universal primers to carry out a second amplification of the repaired amplicon in the generation of a library according to the method of the present invention.In a particular embodiment, the first 5' universal handle sequence comprises two universal handle sequences (e.g., a combination of an amplification primer-binding sequence, a sequencing primer-binding sequence, and / or a capture primer-binding sequence), and the second 5' universal sequence comprises two universal handle sequences (e.g., a combination of an amplification primer-binding sequence, a sequencing primer-binding sequence, and / or a capture primer-binding sequence), and the 5' first and second universal handle sequences do not show significant hybridization with respect to any portion of the target nucleic acid sequence of interest.
[0091] The structure and properties of universal amplification primers or universal primers are well known to those skilled in the art and can be implemented for use in conjunction with methods and compositions provided to fit specific analytical platforms. The universal handle sequences of adapters provided herein are appropriately adapted to accommodate preferred universal primer sequences. For example, universal P1 and A primers having optional barcode sequences, as described herein, are described in the art and are used for sequencing on the Ion Torrent sequencing platform (Ion Xpress® adapter, Thermo Fisher Scientific). Similarly, additional and other universal adapter / primer sequences described and known in the art (for example, the Illumina universal adapter / primer sequence can be found, for example, at https: / / support.illumina.com / content / dam / illumina-support / documents / documentation / chemistry_documentation / experiment-design / illumina-adapter-sequences_1000000002694-01.pdf, and the PacBio universal adapter / primer sequence can be found, for example, at https: / / s3.amazonaws.com / files.pacb.com / pdf / Guide_Pacific_Biosciences_Template_Preparation_and_Sequencing.pdf) can be used in conjunction with the methods and compositions provided herein. Suitable universal primers of appropriate nucleotide sequences for use with the adapters of the present invention can be readily prepared using standard automated nucleic acid synthesis equipment and reagents in routine use in the art.One single type of universal primer or two different types of distinct universal primers (or even a mixture), for example, a pair of universal amplification primers suitable for amplification of the repaired amplicon in the second amplification, are included for use in the method of the present invention. The universal primers optionally include different tag (barcode) sequences, which do not hybridize to the adapter. The barcode sequences incorporated into the amplicon in the second universal amplification can be utilized, for example, for the effective identification of the sample source.
[0092] In some embodiments, the adapter further includes a unique tag sequence located between a 5' first universal handle sequence and a 3' target-specific sequence, and the unique tag sequence does not exhibit significant complementarity and / or hybridization to any part of the desired unique tag sequence and / or target nucleic acid sequence. In some embodiments, a plurality of primer-adapter pairs have combinations of 10 4 ~10 9 different tag sequences. Thus, in a particular embodiment, each generated target-specific adapter pair includes 10 4 ~10 9 different tag sequences. In some embodiments, the plurality of primer-adapters include each target-specific adapter that includes at least one different unique tag sequence and up to 10 5 different unique tag sequences. In some embodiments, the plurality of primer-adapters include each target-specific adapter that includes at least one different unique tag sequence and up to 10 5 different unique tag sequences. In a particular embodiment, each generated target-specific amplicon includes at least two and up to 10 different tag sequences, each having two different unique tag sequences. 9This includes a number of different adapter combinations. In some embodiments, the multiple primer adapters include each target-specific adapter containing 4096 different tag sequences. In a particular embodiment, each generated target-specific amplicon includes up to 16,777,216 different adapter combinations containing different tag sequences, each having two different unique tag sequences.
[0093] In some embodiments, each primer adapter in a plurality of adapters includes a unique tag sequence (e.g., contained in the tag adapter) which includes different random tag sequences alternating with a fixed tag sequence. In some embodiments, at least one unique tag sequence includes at least one random sequence and at least one fixed sequence, or includes random sequences adjacent to the fixed sequence on both sides, or includes fixed sequences adjacent to the random sequence on both sides. In some embodiments, the unique tag sequence includes a fixed sequence having a length of 2 to 2000 nucleotides or base pairs. In some embodiments, the unique tag sequence includes a random sequence having a length of 2 to 2000 nucleotides or base pairs.
[0094] In some embodiments, the unique tag array includes an array having at least one random array interspersed with fixed arrays. In some embodiments, each tag in a plurality of unique tags has structure (N) n (X) x (M) m (Y) yThe formula has the following characteristics, where "N" represents a random tag sequence generated from A, G, C, T, U, or I, and "n" representing the nucleotide length of the "N" random tag sequence is 2 to 10; "X" represents a fixed tag sequence, and "x" representing the nucleotide length of the "X" random tag sequence is 2 to 10; "M" represents a random tag sequence generated from A, G, C, T, U, or I, and the random tag sequence "M" is either different from or the same as the random tag sequence "N", and "m" representing the nucleotide length of the "M" random tag sequence is 2 to 10; "Y" represents a fixed tag sequence, and the fixed tag sequence "Y" is either the same as or different from the fixed tag sequence "X", and "y" representing the nucleotide length of the "Y" random tag sequence is 2 to 10. In some embodiments, the fixed tag sequence "X" is the same in multiple tags. In some embodiments, the fixed tag sequence "X" is different in multiple tags. In some embodiments, the fixed tag array "Y" is the same across multiple tags. In some embodiments, the fixed tag array "Y" is different across multiple tags. In some embodiments, the fixed tag array "(X)" across multiple adapters x " and "(Y) y " is an alignment anchor.
[0095] In some embodiments, random sequences within a unique tag sequence are represented by "N" and fixed sequences by "X". Thus, a unique tag sequence is represented as N1N2N3X1X2X3 or N1N2N3X1X2X3N4N5N6X4X5X6. Optionally, a unique tag sequence may have a random sequence in which some or all of the nucleotide positions are randomly selected from the group consisting of A, G, C, T, U, and I. For example, each nucleotide at each position in the random sequence may be independently selected from any one of A, G, C, T, U, or I, or selected from a subset of these six different types of nucleotides. Optionally, each nucleotide at each position in the random sequence may be independently selected from any one of A, G, C, or T. In some embodiments, the first fixed tag sequence "X1X2X3" is identical or different sequences in multiple tags. In some embodiments, the second fixed tag sequence "X4X5X6" is identical or different in multiple tags. In some embodiments, the first fixed tag array "X1X2X3" and the second fixed tag array "X4X5X6" within the multiple adapters are array alignment anchors.
[0096] In some embodiments, the unique tag sequence includes the sequence 5'-NNNACTNNNTGA-3', where "N" represents a position in a random sequence randomly generated from A, G, C, or T, and the number of distinct random tags that may result is 4 6 (or 4^6) is calculated to be approximately 4096, and the number of different combinations that can occur with two unique tags is 4 12 (or 4^12), which is approximately 16.78 million. In some embodiments, [ka] The underlined portion is the alignment anchor.
[0097] In some embodiments, the fixed sequences within the unique tag sequence are sequence alignment anchors that can be used to generate error-corrected sequencing data. In some embodiments, the fixed sequences within the unique tag sequence are sequence alignment anchors that can be used to generate a family of error-corrected sequencing reads.
[0098] The adapters provided herein include at least one cleavable portion. In some embodiments, the cleavable portion is located within the 3' target-specific sequence. In some embodiments, the cleavable portion is located at or near the junction between the 5' first universal handle sequence and the 3' target-specific sequence. In some embodiments, the cleavable portion is located at or near the junction between the 5' first universal handle sequence and the unique tag sequence, and at or near the junction between the unique tag sequence and the 3' target-specific sequence. The cleavable portion may be present in a modified nucleotide, nucleoside, or nucleic acid base. In some embodiments, the cleavable portion may include a nucleic acid base that is not naturally present in the target sequence of interest.
[0099] In some embodiments, at least one cleavable portion in the multiple adapters is a uracil base, uridine, or deoxyuridine nucleotide. In some embodiments, the cleavable portion is located within the junction between the 3' target-specific sequence and the 5' universal handle sequence and / or the unique tag sequence and / or 3' target-specific sequence, and at least one cleavable portion in the multiple adapters is cleavable with uracil DNA glycosylase (UDG). In some embodiments, the cleavable portion is cleaved, resulting in a sensitive debasing site, and at least one enzyme capable of reacting at the debasing site produces a gap containing an elongable 3' end. In certain embodiments, the resulting gap contains a 5'-deoxyribose phosphate group. In certain embodiments, the resulting gap contains an elongable 3' end and a 5'-linkable phosphate group.
[0100] In another embodiment, inosine can be incorporated into a DNA-based nucleic acid as a cleavable group. In one exemplary embodiment, EndoV can be used to cleave near an inosine residue. In another exemplary embodiment, the enzyme hAAG can be used to cleave an inosine residue from the nucleic acid to create an abasic site.
[0101] Where cleavable regions exist, the location of at least one cleavable region in the adapter does not significantly alter the melting temperature (Tm) of any given double-stranded adapter in the plurality of double-stranded adapters. The melting temperatures (Tm) of any two given double-stranded adapters from the plurality of double-stranded adapters are substantially the same, and the melting temperatures (Tm) of any two given double-stranded adapters do not differ from each other by more than 10°C. However, within each of the plurality of adapters, the melting temperatures of the sequence regions differ, for example, the comparable highest and lowest melting temperatures of the universal handle sequence are higher than the comparable highest and lowest melting temperatures of the unique tag sequence and / or target-specific sequence in any of the adapters. This localized difference in comparable highest and lowest melting temperatures can be adjusted to optimize the digestion and repair of the amplicon, and the ultimately improved efficacy of the method provided herein.
[0102] A composition comprising a nucleic acid library produced by the method of the present invention is further provided. Thus, a composition comprising a plurality of amplified target nucleic acid amplicons is provided, each of the plurality of amplicons comprising a 5' universal handle sequence, optionally a first unique tag sequence, an intermediate target nucleic acid sequence, optionally a second unique tag sequence, and a 3' universal handle sequence. The provided composition comprises at least two and up to 100,000 target-specific amplicons. The provided composition comprises a highly multiplexed targeted library. In some embodiments, the provided composition comprises a plurality of nucleic acid amplicons, each of the plurality of amplicons comprising a 5' universal handle sequence, a first unique tag sequence, an intermediate target nucleic acid sequence, a second unique tag sequence, and a 3' universal handle sequence. The provided composition comprises at least two and up to 100,000 target-specific tagged amplicons. The provided composition comprises a highly multiplexed tagged targeted library.
[0103] In some embodiments, the library composition comprises multiple target-specific amplicons containing multiples of at least two different target nucleic acid sequences. In some embodiments, the composition comprises at least 25, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1000, 1250, 1500, 1750, 2000, 2250, 2500, 2750, 3000, 3250, 3500, 3750, 4000, 4500, 5000, 5500, 6000, 7000, 8000, 9000, 10000, 11000, or 12000 or more target-specific amplicons. In some embodiments, the target-specific amplicon includes an amplicon containing one or more exons, genes, exomes, or regions of the genome associated with a clinical or pathological condition, such as one or more sites containing one or more mutations (e.g., driver mutations) associated with cancer, such as lung cancer, colon cancer, or breast cancer, or an amplicon containing mutations associated with a genetic disease, such as cystic fibrosis or muscular dystrophy. In some embodiments, the target-specific amplicon includes a library of adapter-linked amplicon target sequences having a length of approximately 100 to approximately 750 base pairs.
[0104] As described herein, each of the plurality of amplicons includes a 5' universal handle sequence. In some embodiments, the universal handle sequence includes one or a combination of any of the amplification primer binding sequence, the sequencing primer binding sequence, and / or the capture primer binding sequence. Preferably, the universal handle sequence of the provided adapter does not show significant complementarity and / or hybridization to any portion of the desired unique tag sequence and / or target nucleic acid sequence. In some embodiments, the first universal handle sequence includes one or a combination of any of the amplification primer binding sequence, the sequencing primer binding sequence, and / or the capture primer binding sequence. In some embodiments, the second universal handle sequence includes one or a combination of any of the amplification primer binding sequence, the sequencing primer binding sequence, and / or the capture primer binding sequence. In certain embodiments, the first and second universal handle sequences correspond to forward and reverse universal handle sequences, and in certain embodiments, the same first and second universal handle sequences are included for each of the plurality of target-specific amplicons. Such forward and reverse universal handle sequences are targeted together with universal primers to carry out a second amplification of a preliminary library composition in the generation of an amplification product obtained according to the method of the present invention. In certain embodiments, the first 5' universal handle sequence comprises two universal handle sequences (e.g., a combination of an amplification primer-binding sequence, a sequencing primer-binding sequence, and / or a capture primer-binding sequence), and the second 5' universal sequence comprises two universal handle sequences (e.g., a combination of an amplification primer-binding sequence, a sequencing primer-binding sequence, and / or a capture primer-binding sequence), and the 5' first and second universal handle sequences do not show significant hybridization to any portion of the target nucleic acid sequence of interest.
[0105] The structure and properties of universal amplification primers or universal primers are well known to those skilled in the art and can be implemented for use in conjunction with methods and compositions provided to fit specific analytical platforms. The universal handle sequences of adapters and amplicons provided herein are appropriately adapted to accommodate preferred universal primer sequences. For example, universal P1 and A primers having optional barcode sequences, as described herein, are described in the art and are used for sequencing on the Ion Torrent sequencing platform (Ion Xpress® adapter, Thermo Fisher Scientific). Similarly, additional and other universal adapter / primer sequences described and known in the art (for example, the Illumina universal adapter / primer sequences can be found, for example, at https: / / support.illumina.com / content / dam / illumina-support / documents / documentation / chemistry_documentation / experiment-design / illumina-adapter-sequences_1000000002694-01.pdf, and the PacBio universal adapter / primer sequences can be found, for example, at https: / / s3.amazonaws.com / files.pacb.com / pdf / Guide_Pacific_Biosciences_Template_Preparation_and_Sequencing.pdf) can be used in conjunction with the methods and compositions provided herein. Suitable universal primers of appropriate nucleotide sequences for use with the libraries of the present invention can be readily prepared using standard automated nucleic acid synthesis equipment and reagents in routine use in the art.One single type or two different types of universal primers (or even a mixture thereof), for example, a pair of universal amplification primers suitable for amplification of a preliminary library, may be used in the generation of the library of the present invention. The universal primers optionally include a tag (barcode) sequence, the tag (barcode) sequence does not hybridize to an adapter sequence or a target nucleic acid sequence. The barcode sequence incorporated into the amplicon in the second universal amplification can be used, for example, for effective identification of the sample source, thereby generating a barcoded library. Thus, the provided composition includes a highly multiplexed barcoded targeted library. The provided composition also includes a highly multiplexed barcoded tagged targeted library.
[0106] In some embodiments, the amplicon library includes a unique tag sequence located between a 5' first universal handle sequence and a 3' target-specific sequence, wherein the unique tag sequence does not exhibit significant complementarity and / or hybridization with any portion of the unique tag sequence and / or target nucleic acid sequence. In some embodiments, the amplicons include 10 4 ~10 9 It has a combination of 10 different tag arrays. Therefore, in a particular embodiment, each of the multiple amplicons in the library is 10 4 ~10 9 It includes several different tag arrays. In some embodiments, each of the multiple amplicons in the library has at least one different unique tag array and up to 10 5 It includes several different unique tag sequences. In a particular embodiment, each target-specific amplicon in the library contains at least two and up to 10 different tag sequences, each having two different unique tag sequences. 9This includes a number of different combinations. In some embodiments, each of the multiple amplicons in the library includes a tag sequence containing 4096 different tag sequences. In a particular embodiment, each target-specific amplicon in the library includes up to 16,777,216 different combinations of different tag sequences, each having two different unique tag sequences.
[0107] In some embodiments, individual amplicons in a library of multiple amplicons include a unique tag sequence (e.g., included in a tag adapter sequence) that contains different random tag sequences alternating with a fixed tag sequence. In some embodiments, at least one unique tag sequence includes at least one random sequence and at least one fixed sequence, or includes random sequences adjacent to a fixed sequence on both sides, or includes fixed sequences adjacent to random sequences on both sides. In some embodiments, the unique tag sequence includes a fixed sequence having a length of 2 to 2000 nucleotides or base pairs. In some embodiments, the unique tag sequence includes a random sequence having a length of 2 to 2000 nucleotides or base pairs.
[0108] In some embodiments, the unique tag array includes an array having at least one random array interspersed with fixed arrays. In some embodiments, each tag in a plurality of unique tags has structure (N) n (X) x (M) m (Y) yThe formula has the following characteristics, where "N" represents a random tag sequence generated from A, G, C, T, U, or I, and "n" representing the nucleotide length of the "N" random tag sequence is 2 to 10; "X" represents a fixed tag sequence, and "x" representing the nucleotide length of the "X" random tag sequence is 2 to 10; "M" represents a random tag sequence generated from A, G, C, T, U, or I, and the random tag sequence "M" is either different from or the same as the random tag sequence "N", and "m" representing the nucleotide length of the "M" random tag sequence is 2 to 10; "Y" represents a fixed tag sequence, and the fixed tag sequence "Y" is either the same as or different from the fixed tag sequence "X", and "y" representing the nucleotide length of the "Y" random tag sequence is 2 to 10. In some embodiments, the fixed tag sequence "X" is the same in multiple tags. In some embodiments, the fixed tag sequence "X" is different in multiple tags. In some embodiments, the fixed tag array "Y" is the same across multiple tags. In some embodiments, the fixed tag array "Y" is different across multiple tags. In some embodiments, the fixed tag array "(X)" within multiple amplicons x " and "(Y) y " is an alignment anchor.
[0109] In some embodiments, random sequences within a unique tag sequence are represented by "N" and fixed sequences by "X". Thus, a unique tag sequence is represented as N1N2N3X1X2X3 or N1N2N3X1X2X3N4N5N6X4X5X6. Optionally, a unique tag sequence may have a random sequence in which some or all of the nucleotide positions are randomly selected from the group consisting of A, G, C, T, U, and I. For example, each nucleotide at each position in the random sequence may be independently selected from any one of A, G, C, T, U, or I, or selected from a subset of these six different types of nucleotides. Optionally, each nucleotide at each position in the random sequence may be independently selected from any one of A, G, C, or T. In some embodiments, the first fixed tag sequence "X1X2X3" is identical or different sequences in multiple tags. In some embodiments, the second fixed tag sequence "X4X5X6" is identical or different in multiple tags. In some embodiments, the first fixed tag array "X1X2X3" and the second fixed tag array "X4X5X6" within a plurality of amplicons are array alignment anchors.
[0110] In some embodiments, the unique tag sequence includes the sequence 5'-NNNACTNNNTGA-3', where "N" represents a position in a random sequence randomly generated from A, G, C, or T, and the number of distinct random tags that may result is 4 6 (or 4^6) is calculated to be approximately 4096, and the number of different combinations that can occur with two unique tags is 4 12 (or 4^12), which is approximately 16.78 million. In some embodiments, [ka] The underlined portion is the alignment anchor.
[0111] In some embodiments, the fixed sequences within the unique tag sequence are sequence alignment anchors that can be used to generate error-corrected sequencing data. In some embodiments, the fixed sequences within the unique tag sequence are sequence alignment anchors that can be used to generate a family of error-corrected sequencing reads.
[0112] Kit, system Kits for use in the preparation of libraries of target nucleic acids using methods according to first or second embodiments of the present invention are further provided herein. Embodiments of the kit include a supply of at least one pair of target-specific adapters as defined herein, capable of producing a first amplification product, and optionally, a supply of at least one universal pair of amplification primers, capable of annealing to the universal handles of the adapters to initiate the synthesis of the amplification product, wherein the amplification product comprises a target sequence of interest ligated to the universal sequence. The adapters and / or primers may be supplied in a ready-to-use kit, or more preferably as concentrates requiring dilution before use, or further in a lyophilized or dry form requiring reconstitution before use. In certain embodiments, the kit further includes a supply of suitable diluents for dilution or reconstitution of the components. Optionally, the kit further includes a supply of reagents, buffers, enzymes, dNTPs, etc., for use when carrying out amplification, digestion, repair, and / or purification in the preparation of libraries provided herein. Non-limiting examples of such reagents are described in the Materials and Methods section of the accompanying illustrations. Further components optionally supplied in the kit include components suitable for the purification of libraries prepared using the methods provided. In some embodiments, a kit is provided for generating a target-specific library comprising a 5' universal handle sequence, a 3' target-specific sequence, and multiple target-specific adapters having cleavable groups, DNA polymerase, adapters, dATP, dCTP, dGTP, dTTP, and digestion reagents. In some embodiments, the kit further comprises one or more antibodies, repair reagents, a universal primer optionally containing a nucleic acid barcode, a purified solution, or a column.
[0113] Specific features of the adapters to be included in the kit are described elsewhere herein in relation to other aspects of the invention. The structures and properties of the universal amplification primers are well known to those skilled in the art and can be implemented for use with the methods and compositions provided to be adapted to specific analytical platforms (for example, as described herein, the universal P1 and A primers are described in the art and used for sequencing on the Ion Torrent sequencing platform). Similarly, additional and other universal adapter / primer sequences described and known in the art (e.g., Illumina universal adapter / primer sequences, PacBio universal adapter / primer sequences, etc.) can be used with the methods and compositions provided herein. Suitable primers of appropriate nucleotide sequences for use with the adapters included in the kit are readily prepared using standard automated nucleic acid synthesis equipment and reagents in routine use in the art. The kit includes a supply of one single type of universal primer or two different types of universal primers (or further mixtures), e.g., a pair of amplification primers suitable for amplification of a template modified with the adapter in a first amplification. The kit may include, in addition to at least one pair of adapters for the first amplification of the target sample according to the method of the present invention, at least two different amplification primers having different tag (barcode) sequences, the tag (barcode) sequences not hybridizing to the adapters. Using the kit, at least two different samples can be amplified, each sample being amplified separately according to the method of the present invention, and the second amplification includes using a single universal primer having a barcode, and then pooling the prepared sample library after library preparation. In some embodiments, the kit includes different pairs of universal primers for use in the second amplification step described herein. In this context, the “universal” primer pairs have substantially identical nucleotide sequences but may differ with respect to some other features or modifications.
[0114] Further provided are systems, for example, systems used to carry out the methods provided herein and / or systems comprising compositions provided herein. In some embodiments, the system facilitates methods carried out in an automated mode. In certain embodiments, the system facilitates a high-throughput mode. In certain embodiments, the system includes, for example, fluid handling elements, fluid-containing elements, heat sources and / or heat sinks for achieving and maintaining a desired reaction temperature, and / or robotic elements (e.g., multi-well plate handling elements) that can move components of the system from their respective locations as needed.
[0115] sample As defined herein, “sample” and its derivatives are used in their broadest sense and include any specimen, culture, and / or analogue suspected of containing the target nucleic acid. In some embodiments, a sample includes DNA, RNA, TNA, chimeric nucleic acids, hybrid nucleic acids, multiple forms of nucleic acids, or any combination of two or more of the aforementioned. In some embodiments, a sample useful in relation to the methods of the present invention includes any biological, clinical, surgical, agricultural, atmospheric, or water-based specimen containing one or more target nucleic acids of interest. In some embodiments, a sample includes nucleic acid molecules obtained from animals, such as human or mammalian sources. In other embodiments, a sample includes nucleic acid molecules obtained from non-mammalian sources, such as plants, bacteria, viruses, or fungi. In some embodiments, the source of nucleic acid molecules may be a storage or extinct specimen or species. In some embodiments, a sample includes isolated nucleic acid samples prepared from sources such as genomic DNA, RNA, or TNA, or prepared samples such as fresh-frozen or formalin-fixed paraffin-embedded (FFPE) nucleic acid specimens. It is also conceivable that the sample may be from a single individual, a collection of nucleic acid samples from genetically related members, multiple nucleic acid samples from genetically unrelated members, multiple nucleic acid samples (matching) from a single individual such as tumor and normal tissue samples, or genetic material from a single source containing two different forms of genetic material, such as maternal and fetal DNA obtained from a maternal subject, or from the presence of contaminating bacterial DNA in a sample containing plant or animal DNA. In some embodiments, the source of nucleic acid material includes nucleic acids obtained from a newborn (e.g., blood samples for newborn screening). In some embodiments, the method provided includes amplification of multiple target-specific sequences from a single nucleic acid sample. In some embodiments, the method provided includes target-specific amplification of two or more target sequences from two or more nucleic acid samples or species. In certain embodiments, the method provided includes amplification of highly multiplexed target nucleic acid sequences from a single sample. In certain embodiments, the method provided includes amplification of highly multiplexed target nucleic acid sequences from multiple samples, each from the same source organism.
[0116] In some embodiments, the sample comprises a mixture of target nucleic acids and non-target nucleic acids. In certain embodiments, the sample comprises a plurality of initial polynucleotides which may comprise a mixture of one or more target nucleic acids and one or more non-target nucleic acids. In some embodiments, the sample comprising the plurality of polynucleotides comprises a portion or aliquot of the original sample, and in some embodiments, the sample comprises the plurality of polynucleotides which constitute the entire original sample. In some embodiments, the sample comprises a plurality of initial polynucleotides isolated from the same source or the same subject at different time points.
[0117] In some embodiments, the nucleic acid sample includes cell-free nucleic acids from body fluids, nucleic acids from tissues, nucleic acids from biopsy tissues, nucleic acids from needle biopsies, nucleic acids from a single cell, or nucleic acids from two or more cells. In certain embodiments, a single reaction mixture contains 1 to 100 ng of multiple initial polynucleotides. In some embodiments, the multiple initial polynucleotides include formalin-fixed paraffin-embedded (FFPE) samples, genomic DNA, RNA, TNA, cell-free DNA or RNA or TNA, circulating tumor DNA or RNA or TNA, fresh-frozen samples, or a mixture of two or more of the above, and in some embodiments, the multiple initial polynucleotides include a nucleic acid reference standard. In some embodiments, the sample includes nucleic acid molecules obtained from biopsies, tumors, scrapes, swabs, blood, mucus, urine, plasma, semen, hair, laser-captured microanatomical specimens, surgical excisions, and other clinically or laboratory-obtained samples. In some embodiments, the sample is an epidemiological, agricultural, forensic, or pathogenic sample. In certain embodiments, the sample includes a reference. In some embodiments, the sample is normal tissue or a well-established tumor sample. In certain embodiments, the reference is a standard nucleic acid sequence (e.g., Hg19).
[0118] Targeted nucleic acid sequence analysis The methods and compositions of the present invention provided are particularly suitable for amplifying, optionally tagging, and preparing target sequences for subsequent analysis. Thus, in some embodiments, the methods provided herein include analyzing the resulting library preparations. For example, the methods include analyzing the polynucleotide sequence of the target nucleic acid and, where applicable, the analysis of any tag sequences added to the target nucleic acid. In some embodiments where multiple target nucleic acid regions are amplified, the methods provided include determining the polynucleotide sequences of the multiple target nucleic acids. The methods provided further optionally include using a second tag sequence, e.g., a barcode sequence, to identify the source of the target sequences (or provide other information about the sample source). In certain embodiments, the use of the prepared library composition is provided for the analysis of the sequences of the nucleic acid library.
[0119] In certain embodiments, the use of a prepared tagged library composition is provided for further analysis of the target nucleic acid library sequence. In some embodiments, sequencing involves determining the abundance of at least one target sequence in the sample. In some embodiments, determining low-frequency alleles in the sample is included in sequencing the nucleic acid library. In certain embodiments, determining the presence of mutant target nucleic acids in multiple polynucleotides is included in sequencing the nucleic acid library. In some embodiments, determining the presence of mutant target nucleic acids involves detecting the abundance level of at least one mutant target nucleic acid in multiple polynucleotides. For example, such determination includes detecting that at least one mutant target nucleic acid is present in the sample at a concentration of 0.05% to 1% of the original multiple polynucleotides, detecting that at least one mutant target nucleic acid is present in the sample at a concentration of about 1% to about 5% of the polynucleotides, and / or detecting at least 85% to 100% of the target nucleic acid in the sample. In some embodiments, determining the presence of mutant target nucleic acids includes detecting and identifying copy number variations and / or gene fusion sequences in the sample.
[0120] In some embodiments, nucleic acid sequencing of the amplified target sequence generated by teaching the present disclosure includes de novo sequencing or targeted rearrangement. In some embodiments, nucleic acid sequencing further includes comparing the nucleic acid sequencing result of the amplified target sequence with a reference nucleic acid sequence. In some embodiments, nucleic acid sequencing of the target library sequence further includes determining the presence or absence of mutations in the nucleic acid sequence. In some embodiments, nucleic acid sequencing includes the identification of genetic markers associated with a disease (e.g., cancer and / or genetic disease).
[0121] In some embodiments, the prepared library of target sequences of the disclosed method is used in various downstream analyses or assays, with or without further purification or manipulation. In some embodiments, the analysis includes sequencing by conventional sequencing reactions, high-throughput next-generation sequencing, targeted multiplex array sequence detection, or any combination of two or more of the aforementioned. In certain embodiments, the analysis is performed by high-throughput next-generation sequencing. In certain embodiments, the sequencing is performed in a bidirectional manner, thereby generating sequence reads in both the forward and reverse strands for any given amplicon.
[0122] In some embodiments, libraries prepared according to the methods provided herein are then further manipulated for additional analysis. For example, the prepared library sequences are used in downstream enrichment techniques known in the art, such as bridge amplification or emPCR, to generate template libraries used in next-generation sequencing. In some embodiments, target nucleic acid libraries are used in enrichment and sequencing applications. For example, sequencing of the provided target nucleic acid libraries is achieved using any suitable DNA sequencing platform. In some embodiments, library sequences of the disclosed methods or subsequently prepared template libraries are used for single nucleotide polymorphism (SNP) analysis, genotyping or epigenetic analysis, copy number variation analysis, gene expression analysis, and analysis of gene mutations, including, but not limited to, detection, prognosis and / or diagnosis, and analysis of rare or low-frequency allele mutations, and nucleic acid sequencing, including, but not limited to, de novo sequencing, targeted rearrangement, and synthetic assembly analysis. In one embodiment, the prepared library sequences are used to detect mutations at allele frequencies of less than 5%. In some embodiments, the methods disclosed herein are used to detect mutations in a population of nucleic acids at allele frequencies of 4%, 3%, less than 2%, or about 1%. In another embodiment, the library prepared as described herein is sequenced to detect and / or identify germline or somatic mutations from a population of nucleic acid molecules. In a particular embodiment, a sequencing adapter is ligated to the end of the prepared library to generate multiple libraries suitable for nucleic acid sequencing.
[0123] In some embodiments, methods for preparing target-specific amplicon libraries are provided for use in various downstream processes or assays, such as nucleic acid sequencing or clonal amplification. In some embodiments, the library is amplified using bridge amplification or emPCR to generate multiple clonal templates suitable for nucleic acid sequencing. For example, optionally, secondary and / or tertiary amplification processes are performed following target-specific amplification, including, but not limited to, library amplification steps and / or clonal amplification steps. "Clone amplification" refers to the generation of numerous copies of individual molecules. Various methods known in the art are used for clonal amplification. For example, emulsion PCR is one method that involves isolating individual DNA molecules with primer-coated beads in bubbles within an oil phase. Polymerase chain reaction (PCR) then coats each bead with a clonal copy of the isolated library molecule, and these beads are subsequently immobilized for sequencing. Emulsion PCR is used in methods published by Margulies et al. and Shendure and Porreca et al. (also known as "Polony sequencing," commercialized by Agencourt and recently acquired by Applied Biosystems). Margulies, et al. (2005) Nature 437:376-380, Shendure et al., Science 309(5741):1728-1732. Another method for clonal amplification is "bridge PCR," in which a fragment is amplified with primers attached to a solid surface. These methods, like other methods of clonal amplification, generate a number of physically isolated sites, each containing numerous copies derived from a single-molecule polynucleotide fragment. Thus, in some embodiments, one or more target-specific amplicons are amplified using bridge amplification or emPCR to generate multiple clonal templates suitable for nucleic acid sequencing, for example.
[0124] In some embodiments, at least one of the clonely amplified library sequences is attached to a support or particle. The support can be made of any preferred material and can have any preferred shape, for example, a plane, a spheroid, or fine particles. In some embodiments, the support is a scaffold polymer particle described in U.S. Patent Application Publication 2010 / 0304982, which is incorporated herein by reference in its entirety. In certain embodiments, the method comprises depositing at least a portion of the enriched population of library sequences onto a support (e.g., a sequencing support), the support comprising a series of sequencing reaction sites. In some embodiments, the enriched population of library sequences is attached to the sequencing reaction sites on the support, and the support comprises a series of 10 2 ~10 10 Includes a sequence determination reaction site.
[0125] Sequencing means determining information about the sequence of a nucleic acid and may include identifying or determining partial and complete sequence information of a nucleic acid. Sequence information may be determined with varying degrees of statistical certainty or reliability. In some embodiments, sequence analysis includes high-throughput, low-depth detection by methods such as qPCR, rtPCR, and / or array hybridization detection methodologies known in the art. In some embodiments, sequence analysis includes detailed sequence evaluation determination by methods such as Sanger sequencing or other high-throughput next-generation sequencing methods. Next-generation sequencing means sequencing using methods that determine a large number (typically thousands to billions) of nucleic acid sequences in an inherently large-scale parallel manner, for example, the large number of sequences are read out, for example, in parallel, or using an ultra-high-throughput continuous process that can itself be parallelized. Thus, in certain embodiments, the methods of the present invention include sequence analysis that includes large-scale parallel sequencing.Such methods, though not limited to these, include pyrosequencing (e.g., commercialized by 454 Life Sciences, Inc., Branford, Conn.), ligation sequencing (e.g., SOLiD® technology, commercialized by Life Technologies, Inc., Carlsbad, Calif.), synthetic sequencing using modified nucleotides (TruSeq® and HiSeg® by Illumina, Inc., San Diego, Calif., as well as MiSeq® and / or NovaSeq® technologies, HeliScope® by Helicos Biosciences Corporation, Cambridge, Mass., and PacBio Sequel® or RS systems by Pacific Biosciences of California, Inc., Menlo Park, Calif.), ion detection sequencing (e.g., Ion Torrent® technology, Life Technologies, Carlsbad, Calif.), and DNA nanoball sequencing (Complete This includes Genomics, Inc., Mountain View, Calif., nanopore-based sequencing techniques (e.g., developed by Oxford Nanopore Technologies, LTD, Oxford, UK), and similar highly parallelized sequencing methods.
[0126] For example, in certain embodiments, the libraries produced by the teachings of this disclosure are in sufficient yield to be used in a variety of downstream applications, including Ion Xpress® template kits using the Ion Torrent® PGM system (e.g., PCR-mediated addition of nucleic acid fragment libraries onto Ion Sphere® particles) (Life Technologies, Part No. 4467389) or the Ion Torrent Proton® PGM system. For example, instructions for preparing a template library from an amplicon library can be found in the Ion Xpress template kit user guide (Life Technologies, Part No. 4465884), which is incorporated herein by reference in its entirety. Instructions for loading the subsequent template library onto an Ion Torrent® chip for nucleic acid sequencing are found in the Ion sequencing user guide (Part No. 4467391), which is incorporated herein by reference in its entirety.
[0127] The starting point for a sequencing reaction may be provided by annealing a sequencing primer to the product of a solid-phase amplification reaction. In this regard, one or both of the adapters added during the formation of the template library may include a nucleotide sequence that enables the annealing of the sequencing primer to the amplification product derived from the whole genome or solid-phase amplification of the template library. Depending on the implementation of embodiments of the present invention, the tag sequence and / or target nucleic acid sequence may be determined in a single read from a single sequencing primer or in multiple reads from two different sequencing primers. In the case of two reads from two sequencing primers, the “tag read” and “target sequence read” are performed in either order, along with a preferred denaturation step to remove the annealed primer after the first sequencing read is completed.
[0128] In some embodiments, the sequencer is coupled to a server that applies parameters or software to determine the sequence of an amplified target nucleic acid molecule. In certain embodiments, the sequencer is coupled to a server that applies parameters or software to determine the presence of low-frequency mutant alleles present in the sample.
[0129] Example Example 1: Materials and Method Reverse transcription (RT) reactions (21 μL reaction) can be performed on samples in which RNA and DNA are to be analyzed, such as FFPE RNA and cfTNA. 1. Thaw the 5×URT buffer at room temperature for at least 5 minutes. (Note: Check for any white precipitate in the tube. Vortex to mix if necessary.) [Table 1] 2. Set up the RT reaction in a MicroAmp EnduraPlate 96-well plate by adding the following components. (5-15 ng of RNA or DNA / / 5-40 ng of cfTNA) [Table 2] 3. Mix the entire contents by vortexing or pipetting. Spin down easily. 4. Add 20 μl of Parol 40C oil to the top of each reaction mix. 5. Load the plate into a thermocycler (e.g., a SimpliAmp thermocycler) and run the following program. [Table 3]
[0130] Low-cycle tagging PCR (38 μL reaction volume ± 20 μL oil): Assemble the tagging PCR reaction in a 96-well PCR plate. FFPE DNA samples only 1. Add the following components to a MicroAmp EnduraPlate 96-well plate to assemble the reaction. a. Preparation of UDG mix: 1 µl + 5 µl 5 × URT buffer b. Add 6 μl of diluted UDG to 15 μl of FFPE DNA sample. c. Mix by vortexing. Briefly spin down to collect the reactants at the bottom of the well. d. Add 20 μL of Parol 40C oil to the top of each sample. e. The reaction is carried out as follows. [Table 4] 2. Preparation of the amplified master mix: [Table 5] 3. Add 17 μL of PCR master mix to 21 μL of UDG-Teat FFPE DNA sample. Set the pipette to a volume of 20 μL. Mix the reactants under the oil by pipetting up and down 20 times to ensure complete mixing without disturbing the oil phase. Briefly spin down the plate. FFPE RNA and cfTNA samples only 1. Add the components directly to the RT reactant from the RT step described above. [Table 6] 2. Set the pipette to a volume of 20 μL. Mix the reactants under the oil by pipetting up and down 20 times to completely mix the reactants without disturbing the oil phase. Briefly spin down the plate. 3. Perform 3 cycles of tagging PCR using SimpliAmp with the following cycling conditions. For FFPE DNA and RNA libraries: [Table 7] In the case of the cfTNA library, [Table 8]
[0131] Digestion and filling ligation (45.6 μL reaction volume ± 20 μL oil): 1. Add 7.6 μL of SUPA to each of the PCR reaction wells described above. Add SUPA directly to the sample below the oil layer. 2. Set the pipette to 25 μL. Mix the reaction below the oil layer by pipetting up and down 20 times. Briefly spin down the plate. 3. Load the plate into the thermocycler and run the following program. [Table 9]
[0132] Library amplification (approximately 51 μL reaction volume ± 20 μL oil) 1. Carefully transfer 30 μL of the reaction mixture after the digestion, filling, and ligation process described above to the AmpliSeq HD dual barcode. Mix thoroughly by pipetting up and down 20 times. Return all of the reaction mixture to the original well below the oil layer. 2. Set the pipette to 30 μL. Mix the entire reactant under the oil by pipetting up and down 20 times. Briefly spin down the plate. 3. Load the plate into the thermocycler and run the following program. [Table 10]
[0133] 2-Round AmpureXP Library Refinement The resulting restored sample is purified using two rounds of 36.8 µl of Ampure® beads (Beckman Coulter, Inc.) according to the manufacturer's instructions. In short: Transfer the 46 μL library reaction mixture from beneath the oil layer to a new, clean well on the PCR plate. Add 36.8 μl of Agencourt® AMPure® XP reagent to each sample, mix by pipetting, and incubate at room temperature for 5 minutes. Place the plate on the magnet until the solution in the wells becomes clear. Carefully remove the supernatant, and then remove any remaining supernatant. Add 150 μL of 80% ethanol to 10 mM pH 8 Tris-HCl. Be careful not to disturb the bead pellet. Switch the plate on the magnet three times at 5-second intervals. Remove the supernatant. Repeat the washing step once more. Use a pipette to remove any remaining buffer from the wells. Allow the wells to dry at room temperature for 5 minutes. Add 30 μL of low TE buffer to the well and resuspend the beads with a pipette. Incubate the solution at room temperature for 5 minutes, then place the plate on a magnet to make the solution clear. Transfer 30 μL of eluent to a clean well on the plate. Add 30 μL (1 × volume) of AmpureXP beads to the wells mentioned above. Pipette thoroughly to mix well. After the second purification, elute using 40 μL of low-TE buffer and repeat the steps described above. Transfer 40 μL of the library to a new, clean well.
[0134] Library normalization using individual equalizers First, warm all reagents in the Ion Library Equalizer® kit to room temperature. Vortex and centrifuge all reagents. Wash the Equalizer® beads (skip adding and washing the Equalizer® beads if done previously). 1. For each of the four reactions, add 12 μL of beads to a clean 1.5 mL tube and 24 μL of Equalizer® wash buffer per reaction. 2. Place the tube in the magnetic rack for 3 minutes, or until the solution is completely clear. 3. Carefully remove the supernatant without disturbing the pellets and discard it. 4. Remove from the magnet, add 24 μL of Equalizer® washing buffer after each reaction, and resuspend. Expand the library 5. Remove the plate containing the purified library from the magnet and add 10 μL of 5×DV-Amp Mix and 2 μL of Equalizer® primer (pink cap from the Equalizer kit). Total volume = 52 μL 6. Mix. 7. Gently add 20 μL of Parol 40C oil onto the sample. 8. Run the following program on the thermocycler. 98°C for 2 minutes For FFPE DNA / RNA, amplification is 9 cycles; for cfTNA, amplification is 6 cycles. 98C for 15 seconds 64°C for 1 minute Next Infinitely hold in 4C 9. After (optional) thermal cycling, the plate is centrifuged to collect the droplets. Add Equalizer(trademark) capture to the amplified library. 10. Add 10 μL of Equalizer Capture to each library amplified reactant below the oil layer. 11. Mix up and down 10 times. 12. Incubate at room temperature for 5 minutes. Wash with Equalizer® beads added. 13. Transfer the 60 μL of amplified library sample from beneath the oil layer to the well containing the washed beads. 14. Mix thoroughly. 15. Incubate at room temperature for 5 minutes. 16. Place the plate in a magnet, then incubate for 2 minutes or until the solution becomes clear. 17. Remove the supernatant. 18. Add 150 μL of Equalizer® washing buffer to each reactant. 19. With the plate still in the magnet, remove the supernatant and discard it. 20. Repeat the bead washing process to elute the equalized library. Elucidating the equalized library 21. Remove the plates from the magnet and add 100 μL of Equalizer® elution buffer to each pellet. 22. Mix by pipetting five times in a volume of 50 µl. 23. Elute the library by incubating it in a thermocycler at 32°C for 5 minutes. 24. Immediately remove the plate, place it in a magnet, and transfer the solution to a new well as soon as it becomes clear. 25. Perform qPCR and adjust the pool to 100 pM for template formation and sequencing.
[0135] Example 2 Composition and Method The first step of the provided method involves several amplifications, e.g., 3 to 6 cycles of amplification, and in a particular example, 3 cycles of amplification using forward and reverse adapters for each gene-specific target sequence. Each adapter contains a 5' universal sequence and a 3' gene-specific target sequence. In some embodiments, the adapter optionally includes a unique tag sequence located between the 5' universal and 3' gene-specific target sequences.
[0136] In specific embodiments where unique tag sequences are used, each gene-specific target adapter pair contains a large number of different unique tag sequences in each adapter. For example, each gene-specific target adapter contains up to 4096 TAGs. Thus, each target-specific adapter pair contains at least 4 and up to 16,777,216 possible combinations.
[0137] Each of the adapters provided contains uracil that can be cleaved in place of thymine at specific locations in the forward and reverse adapter sequences. The location of uracil (U) is consistent for all forward and reverse adapters with a unique tag sequence, and uracil (U) is present adjacent to the 5' and 3' ends of the unique tag sequence, if present, and U is present in each of the gene-specific target sequence regions, although the location of each gene-specific target sequence will inevitably vary. The uracil adjacent to each unique tag sequence (UT) and located in the gene-specific sequence region is designed, along with the sequence and the calculated Tm of such sequence, to facilitate fragment dissociation at a temperature lower than the melting temperature of the universal handle sequence, so that it remains hybridized at a selected temperature. Variation of U in adjacent sequences of the UT region is possible, but the design maintains a melting temperature lower than that of the universal handle sequence for each of the forward and reverse adapters. Exemplary adapter sequence structures include: [ka] Here, each N is a base selected from A, C, G, or T, and the constant section of the UT region is used as an anchor sequence to ensure correct identification of the variable (N) portion. The constant and variable regions of the UT can be significantly altered (e.g., alternative constant sequence, >3N per section) as long as the Tm of the UT region remains below that of the universal handle region. Importantly, cleavable uracil is not present in the respective forward (e.g., TCTGTACGGTGACAAGGCG) and reverse (e.g., TGACAAGGCGTAGTCACGG) universal handle sequences. In this embodiment, the universal sequence is designed to accommodate subsequent amplification and addition of the sequence on the ION Torrent platform, but those skilled in the art will understand that such a universal sequence can be adapted to use other universal sequences that can be applied by alternative sequencing platforms (e.g., ILLUMINA sequencing system, QIAGEN sequencing system, PACBIO sequencing system, BGI sequencing system, etc.).
[0138] The methods of using the provided compositions include library preparation using AmpliSeq HD technology, minor modifications thereof, and the use of reagents and kits available from Thermo Fisher Scientific. SuperFiU DNA includes modifications in the uracil-binding pocket (e.g., AA 36) and the family B polymerase catalytic domain (e.g., AA 762). SuperFiU is described in U.S. Provisional Patent Application No. 62 / 524,730, filed June 26, 2017, which is incorporated herein by reference. Polymerase enzymes may have limited ability to utilize uracil and / or any alternative cleavable residues (e.g., inosine, etc.) contained in the adapter sequence. In certain embodiments, it may be advantageous to use a mixture of polymerases to reduce enzyme-specific PCR errors.
[0139] The second step of the method involves partial digestion of the resulting amplicon and any unused uracil-containing adapter. For example, if uracil is incorporated as a cleavable site, digestion and repair involves enzymatic cleavage of uridine monophosphate from the resulting primer, primer dimer, and amplicon, and then repairing the gapped amplicon by thawing the DNA fragment and subsequently polymerase fill-in and ligation. This step reduces and potentially eliminates primer-dimer products that occur in multiplex PCR. In some examples, digestion and repair are performed in a single step. In certain examples, it may be desirable to temporally separate the digestion and repair steps. For example, a heat-unstable polymerase inhibitor may be used in conjunction with the method such that digestion occurs at a lower temperature (25–40°C) and repair is activated by increasing the temperature sufficiently to disrupt polymerase inhibitor interactions (e.g., polymerase-Ab), but not high enough to thaw the universal handle sequence.
[0140] Uracil can be removed using the uracil-DNA glycosylase (UDG) enzyme, leaving the abase site. This can be achieved by several enzymes or combinations of enzymes, including (but not limited to) APE1-aprine / apyrimidine endonuclease, FPG-formamidepyrimidine[fapy]-DNA glycosylase, Nth-endonuclease III, Endo VIII-endonuclease VIII, PNK-polynucleotide kinase, Taq-Thermus aquaticus DNA polymerase, DNA pol I-DNA polymerase I, and Pol beta-human DNA polymerase beta. In a particular implementation, the method uses human aprine / apyrimidine endonuclease, APE1. APE1 activity leaves the 3'-OH and 5'-deoxyribose-phosphate (5'-dRP). Removal of 5'-dRP can be achieved by many enzymes, including recJ, polymerase beta, Taq, DNA pol I, or any DNA polymerase with 5'-3' exonuclease activity. Removal of 5'-dRP by any of these enzymes creates a ligable 5'-phosphate terminus. In another implementation, UDG activity removes uracil, leaving an abasic site, which is removed by FPG, leaving 3' and 5'-phosphates. The 3'-phosphate is then removed by T4 PNK, leaving a polymerase-extendable 3'-OH group. The 5'-deoxyribose phosphate can then be removed by polymerase beta, fpg, Nth, Endo VIII, Taq, DNA pol I, or any other DNA polymerase with 5'-3' exonuclease activity. In a specific implementation, Taq DNA polymerase is utilized.
[0141] The repair fill-in process can be achieved by virtually any polymerase, in some cases by the amplification polymerase used for amplification in step 1, or by any polymerase added in step 2, including (but not limited to) Phusion DNA polymerase, Phusion U DNA polymerase, SuperFi DNA polymerase, SuperFi U DNA polymerase, TAQ, Pol beta, T4 DNA polymerase, and T7 DNA polymerase. Amplicon ligation repair can be carried out by many ligases, including (but not limited to) T4 DNA ligase, T7 DNA ligase, and Taq DNA ligase. In specific implementations of the method, Taq DNA polymerase is utilized and ligation repair is achieved by T7 DNA ligase.
[0142] The final step in library preparation involves amplifying the repaired amplicon by a standard PCR protocol using universal primers containing sequences complementary to the universal handle sequences on the 5' and 3' ends of the prepared amplicon. For example, the A-universal primer and the P1 universal primer, as well as parts of the Ion Express adapter kit (Thermo Fisher Scientific, Inc.), may optionally contain sample-specific barcodes. The final library amplification step can be performed by many polymerases, including, but is not limited to, Phusion DNA polymerase, Phusion U DNA polymerase, SuperFi DNA polymerase, SuperFi U DNA polymerase, Taq DNA polymerase, and Veraseq Ultra DNA polymerase.
[0143] Example 3 Assay Content and Method Along with primers directed to target sequences specific to the targets in Table 1, the adapter contains 4096 unique tag sequences for each gene-specific target sequence, resulting in an estimate of 16,777,216 different unique tag combinations for each pair of gene-specific target sequences.
[0144] The library was prepared according to the method described above. The prepared library was then prepared for template and sequencing and analyzed. Sequencing can be performed by various known methods, including, but not limited to, synthetic sequencing, concatenation sequencing, and / or hybridization sequencing. In the examples herein, sequencing was performed using the Ion Torrent platform (Thermo Fisher Scientific, Inc.), but the library can be prepared and adapted for analysis, e.g., sequencing, using any other platform, e.g., Illumina, Qiagen, PacBio, etc. The results can be analyzed using several metrics to evaluate performance, e.g., as follows: ○ Family count (ng of captured input DNA): The median family count is a measure of the number of families mapped to individual targets. In this case, each unique molecular tag is a family. Uniformity is a measure of the percentage of target bases covered by an average read depth of at least 0.2 times. This metric is used to ensure that the technique does not selectively under-amplify specific targets. ○ Positive / Negative: When control samples with known mutations are used and analyzed (e.g., Acrometrix Oncology Hotspot Control DNA, Thermo Fisher Scientific, Inc.), the number of true positives can be tracked. ■True Positives: The number of true positives indicates the number of mutations that exist and have been correctly identified. ■ False Positives (FP): (Hotspot and entire target) The number of false positives indicates the number of mutations that are known not to be present in the sample but are determined to be present. ■ False Negatives (FN) (when acrometric spikein is used): The number of false negatives indicates the number of mutations that were present but not identified. ○On / Off target is the percentage of mapped reads that are aligned / unaligned across the target region. This metric is used to verify that the technology primarily amplifies the panel's designed target. ○Low quality is tracked to ensure the data is worth analyzing. This metric is a general system metric and is not directly related to this technology. [Table 11]
[0145] Clinical evidence is defined as the number of instances in which a gene / variant combination appears on drug labels, guidelines, and / or clinical trials. Tables 2 and 3 show the top genes / variants and indications related to the provided assays, supported by clinical evidence. [Table 12]
[0146] The maximum of 29 gene and variant combinations covered by the provided assay are listed on the drug label and / or guidelines (NCCN and ESMO). [Table 13]
[0147] Results of Example 4 The primers were designed using the composition design approach provided herein, and the library amplification step utilized two primer pairs to enable bidirectional sequencing as described herein (with two universal sequences placed at each end of the amplicon, e.g., A-universal handle and P1-universal handle at each end), and targeted to tumor genes using those of the panel target genes as described above in Table 1. The prepared libraries were sequenced using the Ion Gene Studio template / and sequencing kit and instrumentation (Thermo Fisher Scientific, Inc.) and / or a new, fully integrated library preparation, template, and sequencing system. Implementation of the instant panel demonstrates that the technology can adequately detect the desired mutations, copy number variations, and fusions.
[0148] 4A. Fusion detection capability for various ALK and ROS isoforms from NSCLC FFPE samples The library was prepared and sequenced as described above. As expected, the detection of various fusion isoforms was demonstrated. [Table 14] 4B: Mutation detection in matching samples using GeneStudio S5 and a new sequencer [Table 15]
[0149] Library preparation, sequencing, and analysis were performed for mutation detection in matched samples as described above, using both manual preparation and sequencing in the ION GeneStudio S5, as well as an automated integrated library preparation, template, and sequencing system. High-precision assays demonstrated the detection of matched PIK3CA and KRAS mutations across matched tissue and plasma samples, both in the GeneStudio S5 with manual workflows and compared to the automated system.
[0150] 4C: Detection of DNA variants across various cancer indications Library preparation, sequencing, and analysis were performed for mutation detection across various sample types, as described above, using both manual preparation and sequencing in ION GeneStudio S5, as well as automated integrated library preparation, templating, and sequencing systems. High-precision assays demonstrated the detection of diverse driver mutations across various cancer indication sample types. [Table 16]
[0151] 4D: Detection using a cohort of matched FFPE and plasma samples Library preparation, sequencing, and analysis were performed to evaluate the assay's performance in detecting variants across a cohort of matched FFPE and plasma samples. The assay demonstrated the detection of various driver mutations across different cancer indication sample types. Using the assay with a set of matched FFPE and plasma samples, four out of eight samples had matched PIK3CA(1) and KRAS(3) mutations, while one out of eight detected a matched NO variant. [Table 17]
[0152] 4E: Detection of mutations in FFPE cancer samples with known variants To evaluate the performance of the ONCOMINE cfDNA assay in detecting variants in 16 FFPE samples (NSCLC, breast, and CRC), including known mutations previously identified using the assay, libraries were prepared, sequenced, and analyzed. Samples were tested in a single run (chip) using the assay in an integrated system. Eight samples had been previously characterized using the ONCOMINE cfDNA assay. In this cohort, 17 mutations were detected within EGFR, ERBB4, IDH1, KRAS, MET, PIK3CA, and TP53. In addition, three amplifications were detected in EGFR, ERBB2, and FGFR1, and finally, three fusions with the FGFR2 and RSPO3 driver genes were also detected. This assay was able to detect a variety of variants in this cohort of FFPE samples, including SNV mutations, CNV amplifications, and fusions.
[0153] 4F: Detection of SNVs and CNVs across multiple cancer types from FFPE This assay was used to detect various driver mutations across different cancer indications. All results were consistent with previous characterizations using different assays and systems. [Table 18]
[0154] 4G: Detection of fusions between ALK, ROS1, RET, NTRK1, NTRK2, and NTRK3 driver genes. To evaluate the assay's performance in detecting fusion variants, libraries were prepared, sequenced, and analyzed. The assay reproducibly detected 11 fusion isoforms representing six driver genes (ALK, BRAF FGFR3, NTRK1, NTRK3, RET, and ROS1), as well as 15 NTRK fusion isoforms representing three driver genes (NTRK1, NTRK2, and NTRK3) using targeted isoform detection. [Table 19] 4H: Detection of EGFR and KRAS variants in control materials [Table 20]
[0155] To evaluate the performance of the assay in detecting EGFR variants using Horizon EGFR gene-specific multiplex reference standards 5% and 1% FFPE controls, and in detecting RAS fusion variants using Horizon KRAS gene-specific multiplex reference standard 5% FFPE, libraries were prepared, sequenced, and analyzed. This assay was able to detect all EGFR variants at a 5% allele frequency using the Horizon FFPE control. At a 1% allele frequency, below the normal LOD, the assay detected 7 out of 8 cases in two replications. This assay was able to reproducibly detect 6 RAS mutations using the Horizon control.
[0156] 4I: Detection of KRAS, BRAF, KIT, and EGFR mutations using cfDNA controls To evaluate the performance of the assay in detecting KRAS, BRAF, KIT, and EGFR mutations using cfDNA controls: SeraCare Seraseq ctDNA reference material v2 AF 0.125% or Horizon Multiplex I cfDNA reference standard sets (1% and 0.1%), libraries were prepared, sequenced, and analyzed. This assay was able to detect mutations with allele frequencies down to 0.1% using the cfDNA controls. [Table 21]
[0157] 4J:Detection of copy number variations in MET and PTEN To evaluate the performance of the assay in detecting MET copy number increase and PTEN copy number decrease using control and cell lines, libraries were prepared, sequenced, and analyzed. The Structural Multiplex FFPE reference standard (Horizon) was used to detect MET, and the PTEN cell line (ATCC) was used to detect PTEN copy number variation. This assay successfully detected MET copy number increase and PTEN copy number decrease using control and cell lines, respectively. [Table 22]
[0158] 4K: Detection of NTRK1, FGFR3, and RET fusions in cell lines. To evaluate the performance of the assay in detecting NTRK1, FGFR3, and RET fusion cell lines, libraries were prepared, sequenced, and analyzed. KM12 cell line (ATCC), SW780 cell line (ATCC), and LC-2 / ad cell line (Sigma Aldrich) were used for nucleic acid preparation and evaluation. The assay was able to detect the TPM3-NTRK1 fusion isoform using both the assay's targeted isoform and unbalanced assay methods. The assay was able to detect the FGFR3-BAIAP2L1 fusion isoform using both the assay's targeted isoform and unbalanced assay methods. Interestingly, ALK unbalance was also detected in this cell line. Research is underway to further understand these results. This assay was able to detect the CCDC6-RET fusion isoform using both the assay's targeted isoform and unbalanced assay methods. [Table 23]
[0159] Detection of ALK and ROS1 fusions in 4L:FFPE samples To evaluate the performance of the assay in detecting ALK and ROS1 fusions in FFPE samples, library preparation, sequencing, and analysis were performed. This assay successfully detected ALK fusions using both targeted isoform and unbalanced assay methods, and ROS1 fusions using the targeted isoform method for FFPE samples. [Table 24]
[0160] Preferred embodiments of the present invention are shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided only as examples. Many variations, modifications, and substitutions will be conceivable to those skilled in the art without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be used in carrying out the present invention. The following claims define the scope of the present invention and are intended to encompass methods and structures within the scope of these claims, as well as their equivalents. [Table 25-1] [Table 25-2] [Table 25-3] [Table 25-4] [Table 25-5] [Table 25-6] [Table 25-7] [Table 25-8] Table 25-9 Table 25-10 Table 25-11 Table 25-12 Table 25-13 Table 25-14 Table 25-15 Table 25-16 Table 25-17 Table 25-18 Table 25-19 Table 25-20 Table 25-21 Table 25-22 Table 25-23 Table 25-24 Table 25-25 Table 25-26 Table 25-27 Table 25-28 Table 25-29 Table 25-30 Table 25-31 Table 25-32 Table 25-33 Table 25-34 Table 25-35 Table 25-36
[0161] Table 25-37 Another aspect of the present invention may be as follows: [1] A composition for single-stream multiplexing of actionable tumor biomarkers in a sample, wherein the composition comprises multiple sets of primer-pair reagents directed to multiple target sequences for detecting low-level targets in the sample, and the target gene is selected from the group consisting of the following functions: DNA hotspot mutation genes, copy number variation (CNV) genes, intergene fusion genes, and intragene fusion genes. [2] The composition according to [1], wherein one or more actionable target genes in the sample determine changes in tumor activity in the sample that indicate potential diagnoses, prognoses, candidate treatment regimens, and / or adverse events. [3] The composition according to [1], wherein the target gene is selected from the genes in Table 1. [4] The composition according to [1], wherein the target gene consists of the genes in Table 1. [5] The composition according to [1], wherein the plurality of target sequences include amplicon sequences detected by primers from Table A. [6] The composition according to [1], wherein each of the plurality of target sequences comprises an amplicon sequence detected by a primer from Table A. [7] The composition according to [1], wherein the plurality of primer reagents are selected from the primers in Table A. [8] The composition according to [1], wherein the plurality of primer reagents consist of each of the primers in Table A. [9] A multiple assay comprising the composition described in any of [1] to [8] above.
[10] A test kit comprising the composition described in any of [1] to [8] above.
[11] A method for determining the presence of one or more actionable tumor biomarkers in a biological sample, Multiple amplification of multiple target tumor sequences from a biological sample, wherein the amplification includes contacting at least a portion of the sample with a composition and polymerase described in any one of items [1] to [8] under amplification conditions to generate amplified target tumor sequences, A method comprising detecting each of the aforementioned plurality of target sequences, wherein the detection of one or more actionable tumor biomarkers compared to a control sample determines a change in tumor activity in the sample indicating a potential diagnosis, prognosis, candidate treatment regimen, and / or adverse event.
[12] The method according to
[11] , wherein the target gene is selected from the group consisting of DNA hotspot mutation genes, copy number variation (CNV) genes, intergenetic fusion genes, and intragenetic fusion genes.
[13] The method according to
[11] , wherein the target gene is selected from the genes in Table 1.
[14] The method according to
[11] , wherein the target gene consists of the genes in Table 1.
[15] The method according to
[11] , wherein the plurality of target sequences include amplicon sequences detected by primers from Table A.
[16] The method according to
[11] , wherein each of the plurality of target sequences includes an amplicon sequence detected by a primer from Table A.
[17] The method according to
[11] , wherein the plurality of primer reagents are selected from the primers in Table A.
[18] The method according to
[11] , wherein the plurality of primer reagents consist of each of the primers in Table A.
[19] The method according to
[11] , wherein the biological sample includes the tumor and / or surrounding tissue.
[20] The method according to
[11] , wherein the biological sample and the control sample are from the same individual.
Claims
1. A composition for single-stream multiplexing of actionable tumor biomarkers in a sample, wherein the composition comprises multiple sets of primer-pair reagents directed to multiple target sequences of the actionable tumor biomarker to detect low-level targets in the sample, the target gene containing the actionable tumor biomarker is selected from the group consisting of the following functions: DNA hotspot mutation genes, copy number variation (CNV) genes, intergene fusion genes, and intragene fusion genes, and the multiple sets of primer-pair reagents comprise two or more primers from SEQ ID NOs: 1 to SEQ ID NOs: 1563.
2. The composition according to claim 1, wherein one or more target genes containing actionable tumor biomarkers in the sample determine changes in tumor activity in the sample that indicate potential diagnosis, prognosis, candidate treatment regimens, and / or adverse events.
3. The composition according to claim 1, wherein the target gene is selected from the group consisting of AKT1, AKT2, AKT3, ALK, AR, ARAF, BRAF, CD274, CDK4, CDKN2A, CHEK2, CTNNB1, EGFR, ERBB2, ERBB3, ERBB4, ESR1, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MAP2K1, MAP2K2, MET, MTOR, NRAS, NRG1, NTRK1, NTRK2, NTRK3, NUTM1, PDGFRA, PIK3CA, PTEN, RAF1, RET, ROS1, RSPO2, RSPO3, SMO, and TP53.
4. The composition according to claim 1, wherein the target genes consist of the genes AKT1, AKT2, AKT3, ALK, AR, ARAF, BRAF, CD274, CDK4, CDKN2A, CHEK2, CTNNB1, EGFR, ERBB2, ERBB3, ERBB4, ESR1, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MAP2K1, MAP2K2, MET, MTOR, NRAS, NRG1, NTRK1, NTRK2, NTRK3, NUTM1, PDGFRA, PIK3CA, PTEN, RAF1, RET, ROS1, RSPO2, RSPO3, SMO, and TP53.
5. The composition according to claim 1, wherein the plurality of target sequences include amplicon sequences that are detected by any of the primers SEQ ID NOs: 1 to 1563.
6. The composition according to claim 1, wherein each of the plurality of target sequences includes an amplicon sequence detected by the primers SEQ ID NO: 1 to SEQ ID NO: 1563.
7. The composition according to claim 1, wherein a plurality of sets of primer-reagent pairs are selected from any of the primers of SEQ ID NO: 1 to SEQ ID NO: 1563.
8. The composition according to claim 1, wherein each of the plurality of primer-reagent sets comprises the primers of SEQ ID NO: 1 to SEQ ID NO: 1563.
9. A multiplex assay comprising the composition according to any one of claims 1 to 8.
10. A test kit comprising the composition according to any one of claims 1 to 8.
11. A method for determining the presence of one or more actionable tumor biomarkers in a sample in order to collect information relating to diagnosis, prognosis, candidate treatment regimens, and / or adverse events, Multiple amplification of multiple target sequences of the actionable tumor biomarker from the sample, wherein the amplification includes contacting at least a portion of the sample with the composition and polymerase described in any one of claims 1 to 8 under amplification conditions to generate amplified target sequences, A method comprising detecting each of the aforementioned plurality of target sequences, wherein the detection of one or more actionable tumor biomarkers compared to a control sample determines a change in tumor activity in the sample.
12. The method according to claim 11, wherein the target gene is selected from the group consisting of DNA hotspot mutation genes, copy number variation (CNV) genes, intergenetic fusion genes, and intragenetic fusion genes.
13. The method according to claim 11, wherein the target gene is selected from the group consisting of AKT1, AKT2, AKT3, ALK, AR, ARAF, BRAF, CD274, CDK4, CDKN2A, CHEK2, CTNNB1, EGFR, ERBB2, ERBB3, ERBB4, ESR1, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MAP2K1, MAP2K2, MET, MTOR, NRAS, NRG1, NTRK1, NTRK2, NTRK3, NUTM1, PDGFRA, PIK3CA, PTEN, RAF1, RET, ROS1, RSPO2, RSPO3, SMO, and TP53.
14. The method according to claim 11, wherein the target genes consist of the genes AKT1, AKT2, AKT3, ALK, AR, ARAF, BRAF, CD274, CDK4, CDKN2A, CHEK2, CTNNB1, EGFR, ERBB2, ERBB3, ERBB4, ESR1, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MAP2K1, MAP2K2, MET, MTOR, NRAS, NRG1, NTRK1, NTRK2, NTRK3, NUTM1, PDGFRA, PIK3CA, PTEN, RAF1, RET, ROS1, RSPO2, RSPO3, SMO, and TP53.
15. The method according to claim 11, wherein the plurality of target sequences include amplicon sequences that are detected by any of the primers SEQ ID NOs: 1 to 1563.
16. The method according to claim 11, wherein each of the plurality of target sequences includes an amplicon sequence detected by the primers SEQ ID NO: 1 to SEQ ID NO: 1563.
17. The method according to claim 11, wherein a plurality of sets of primer-pair reagents are selected from any of the primers of SEQ ID NO: 1 to SEQ ID NO: 1563.
18. The method according to claim 11, wherein the plurality of sets of primer-reagent pairs consist of each of the primers of SEQ ID NO: 1 to SEQ ID NO: 1563.
19. The method according to claim 11, wherein the sample includes the tumor and / or surrounding tissue.
20. The method according to claim 11, wherein the sample and the control sample are from the same individual.
Citation Information
Patent Citations
Cancer gene mutation and gene amplification detection
CN104630375B
Multiple PCR primers for detecting non-small cell lung cancer oncogene mutation based on high-throughput sequencing, kit and method
CN107723354A
Systems and Methods for Monitoring Lifelong Tumor Evolution Field of Invention
US20190010552A1
Methods and materials for assessing and treating cancer
WO2019067092A1