Multi-vector detection of variant sequences
The dPCR method with multiple probe sets and detectable labels creates multidimensional plots to enhance the detection of low-abundance nucleic acids, addressing sensitivity and specificity issues in detecting somatic variants, especially in cell-free DNA samples, by forming distinct clusters for accurate disease and tumor mutation analysis.
Patent Information
- Application Number
- JP2025513006
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-31
- Filing Date
- 2023-08-31
- Publication Date
- 2025-09-10
AI Technical Summary
Existing methods for detecting low-abundance variant nucleic acids, particularly somatic variants, are inadequate due to high input sample requirements, high costs, complex workflows, insufficient sensitivity, and inability to detect low-abundance mutant sequences in high background noise, especially in samples like circulating cell-free DNA and single-cell analysis.
A method using digital polymerase chain reaction (dPCR) with multiple sets of probes and detectable labels generates multidimensional plots to identify variant sequences, allowing for the detection of multiple variants by forming distinct clusters in a multi-vector representation, even in highly fragmented samples.
The method enhances sensitivity and specificity in detecting low-abundance nucleic acids, enabling the identification of disease states, minimal residual disease, and tumor mutations, particularly in cell-free DNA and formalin-fixed paraffin-embedded tissue samples, with reduced false positives and improved detection of low-abundance targets.
Smart Images

Figure 2025530035000001_ABST
Abstract
Description
[Technical Field]
[0001] Technical Field The present invention relates to the field of methods for the detection of low abundance target nucleic acids. [Background technology]
[0002] background Many diseases and conditions, especially cancer, are associated with variant DNA sequences. Many of these variant sequences are somatic variants, which occur in non-germline cells. Although many germline variants are associated with pathogenic conditions, they do not necessarily contribute to disease progression or symptomatology. Nevertheless, emerging evidence indicates that many diseases and conditions exhibit patterns of variant sequences, whether in individuals or across populations. Consequently, detecting variant sequences or patterns of variant sequences can provide insights, including information about disease progression / regression, treatment efficacy, minimal residual disease (MRD), tumor origin, specific defects in DNA repair, certain environmental exposures, and optimal treatment of conditions, and may be the basis for personalized medicine.
[0003] However, although detecting variant sequences has obvious utility, in many situations, especially in the context of somatic variants, these variants exist only in small amounts.Therefore, an exceptionally sensitive and specific method for mutation detection is needed, especially for low-input samples such as circulating cell-free DNA (cfDNA) and single-cell analysis.Traditional methods can suffer from various problems, including high input sample DNA requirements, high cost per sample, complex and laborious workflow, insufficient sensitivity and / or specificity, and the inability to detect low-abundance mutant sequences (so-called mutant allele fraction; MAF) in a high background of normal wild-type sequences.One conventional solution is digital PCR, in which PCR reactions are compartmentalized into many smaller individual reactions so that each reaction compartment contains zero to very few target sequence molecules.This potentially increases sensitivity. However, conventional dPCR often fails to detect the presence of low-abundance target nucleic acids in biological samples due to the relatively weak signals provided by low-abundance nucleic acids (e.g., variant sequences caused by cancer) and due to false positives from sources such as noise or DNA damage. The failure to detect low-abundance sequences is more pronounced when traditional dPCR is used to simultaneously detect multiple variant sequences in one assay. Summary of the Invention [Means for solving the problem]
[0004] Abstract The present invention provides methods and systems for rapid identification of variant sequences or patterns of variant sequences, despite their low abundance in a sample. According to the present invention, variant sequences are detected using digital polymerase chain reaction (dPCR), which generates identifiable fingerprints that enable variant detection using many, a few, only a few, or even a single detectable label to identify multiple variants. The method of the present invention utilizes multidimensional plots to provide a multi-vector representation of the variants in a sample. The plots are generated using dPCR using a first set of probes (each of which has a first detectable label and is specific for a different variant sequence); and additional sets of probes with second or third (or more) detectable labels. Amplicons are generated in compartments, from which the labels are detected. Any suitable detectable label can be used. For example, in some embodiments, the detectable labels are optical labels (e.g., from fluorescently labeled hydrolysis probes), and a multicolor plot can be generated from the detected optical signals to identify variant sequences based on deviations from the expected wild-type cluster plot. In this manner, the present invention is adaptable to more than two probe sets to elucidate one or more plots that can be compared to the expected wild-type cluster plot. While the preferred embodiment is presented in the form of a two-color plot for ease of demonstration, those skilled in the art will understand that additional multi-color plots can be used in accordance with the present invention.
[0005] In certain aspects, the present invention provides a method for detecting variant nucleic acids. According to the present invention, a sample suspected of containing a nucleic acid target is compartmentalized, and the nucleic acid in the compartment is amplified in the presence of a set of variant-specific probes, where each variant-specific probe contains a detectable label, and the entire set of probes contains a number of distinct detectable labels, which may be fewer than the number of variants to be detected. Variants are detected from the amplicons in the compartments. A plot of points representing the optical signals is generated, and the presence of variants in the sample is identified based on corresponding clusters of points in the plot.
[0006] A preferred method may include mapping the detected signals onto a space defined by the number of distinct detectable labels. A preferred method may include assigning a vector to each cluster of points, where each vector uniquely and specifically identifies one of the variants in the sample. In some embodiments, at least one of the probes detects all of the variants in the sample, and at least a second of the probes is specific to fewer than all of the variants and is present at a different concentration than the first probe. The second probe is used to distinguish between the variants by forming distinct clusters of points in the plot. Two or more of the different variant sequences may be present at positions on the target nucleic acid that are amplified by one primer pair in the amplifying step.
[0007] The presence of one or more variant sequences may indicate a disease state (e.g., cancer). In certain embodiments, the presence of one or more variant sequences indicates minimal residual disease. In one embodiment, the method of the present invention includes, prior to compartmentalizing the sample, obtaining an estimate of the relative abundance of the variants and designing the variant-specific probes based on the estimate. The presence or absence of one or more of the variant sequences indicates progression or regression of the disease state. The variants may include tumor mutations determined by sequencing tumor nucleic acids from tumor samples. In some embodiments, digital PCR further uses a probe specific to a wild-type sequence.
[0008] Some embodiments comprise: allocating the variants to a section, and for at least one section, providing at least one probe that detects all the variants in the section and at least one probe that distinguishes between the variants in the section.Section can be determined based on the information about genomic location or the relative abundance of the variants in the sample.For example, each section can be defined as a set of variants that can be amplified by one primer pair.
[0009] In certain embodiments, the detectable label is optical label.In some embodiments, the optical label for use in the present invention is selected from FAM, HEX, SUN, VIC, TAMRA, ATTO550, Cy5, ROX, ATTO700, Cy5.5, Yakima Yellow, ABY or JUN (for example, on fluorescent hydrolysis probe such as TAQMAN probe).Said sample can be cell-free DNA (cfDNA).
[0010] Before dividing into aliquots, the present invention may include identifying a first pair of first and second variants among the variants that form overlapping clusters on a 2D dPCR plot, and designing a probe set including: (i) detection probes for both variants of the pair, where both detection probes have a detectable optical label of a first color; and (ii) at least one discrimination probe having a discriminatory optical label of a second color. Amplification may be performed with the detection probes and the discrimination probes present at different concentrations. The detection probes may further include any number of additional probes and labels (e.g., a third probe specific to a third variant not of the pair, the third probe having a detectable label (e.g., the optical label of the first color)). The plot may include a first cluster, a second cluster, and a third cluster from the first variant, the second variant, and the third variant, respectively, and the vectors passing from the origin of the plot through the centers of gravity of the clusters are non-orthogonal. Preferably, the first variant, second variant, and third variant are all located at positions on the target nucleic acid such that they are amplified by one primer pair during the amplifying step.
[0011] In certain embodiments, the probes include a first probe specific for the first variant and having a first optical label; a second probe specific for the second variant and having the first optical label; a third probe specific for the third variant and having the first optical label; a WT probe specific for the wild-type sequence and having a second optical label; and a discrimination probe specific for the first variant and having a third optical label. A first two-color plot is generated from the first and second optical labels, and the presence of at least the first and second variants is identified from the first two-color plot. The method may further include generating a second two-color plot from the first and third optical labels, and distinguishing the first variant sequence from at least the second variant based on the deviation of the second two-color plot from the expected dPCR two-color plot.
[0012] In another aspect, the present application provides a method for detecting variant nucleic acids. An exemplary method includes compartmentalizing a sample containing a target nucleic acid into reaction compartments and performing digital PCR in the reaction compartments. The dPCR reaction can be performed using a first probe set (each specific to a different variant sequence and having a first optical label) and one or more probes specific to different sequences and having one or more different optical labels. Using a dPCR assay, optical signals are detected from the probes in each compartment. For example, the signals from the wild-type and variant sequence probes are used to generate a two-color plot. Subsequently, the presence of one or more variant sequences can be identified based on deviations of the plotted signals from the expected wild-type cluster plot. This deviation can be a recognized pattern in the multi-vector representation that is characteristic of a particular variant sequence, whether in a subject or historically associated with a particular disease or condition.
[0013] In certain methods, the first probe set comprises a probe that has the first optical label and has specificity for the first variant sequence, and a probe that has a third optical label and has specificity for the first variant sequence.This method can further comprise generating a two-color plot from the first optical label and the third optical label, and identifying the presence of the first variant sequence based on the shift in polarization.In this case, the third optical label can be used to help identify one or more of the variant sequences.
[0014] Optionally, in some embodiments, probe labels can be assigned to variants based on information about variant frequencies (e.g., population frequencies, disease frequencies, allele frequencies, or mutant allele fractions (MAFs). For example, variants can be assigned to segments or categories based on their frequencies or MAFs. This allows, for example, very rare (low frequency) variants to be assigned to specific labels. For example, two rare variants and a relatively abundant variant that are close to each other (e.g., polymorphic at the same base or within a few (e.g., 20) bases of each other), all three can be given the same optical label on their respective probes, and then Further distinguishing probes can be provided with different optical labels.In some of the above embodiments, the probe that comprises the third optical label is specific to the variant sequence that has a higher frequency / MAF than the different variant sequences of the first probe set.Such method can further comprise the use of a wild-type sequence-specific probe without label; a second probe set that is specific to each different variant sequence and has a third optical label; a third probe set that is specific to each different variant sequence and has a fourth optical label; and / or a fourth probe set that is specific to each different variant sequence and has a fifth optical label.
[0015] The method of the present invention can be used to detect two or more different variant sequences present at the same genomic location of target nucleic acid.In certain aspects, one or more variant sequences detected by the method and system of the present invention indicate a pathological condition (e.g., cancer).Alternatively or additionally, the presence of one or more variant sequences indicates minimal residual disease (MRD).
[0016] In some methods of the invention, the relative frequency or concentration (variant copy number per volume unit) of one or more variants occurring in a target nucleic acid is known for a particular disease or condition. These frequencies or MAFs or concentrations for the variants can be measured for a subject at a first time point, and the information can be used to design probes and labels for a subsequent assay (e.g., dPCR). The designed probes and labels can then be used to assay a sample from the subject for those variants at a second, later time point. Alternatively, or in addition, the variant sequences can be recurrent variant sequences of the pathological condition in a population and have a known historical frequency. In either situation, identifying the presence or absence of one or more of the variant sequences can indicate progression or regression of the pathological condition—for example, by detecting a change in MRD or frequency in response to treatment.
[0017] In a particular method of the present invention, performing digital PCR further comprises using a second probe set, each specific to a different variant sequence and having a third optical label, wherein the first and second probe sets are specific to distinct variant sequences.In this way, the multiplexing capability of the method of the present invention is expanded.In this situation, each probe set is specific to a different variant sequence segment that has a similar relative frequency during the pathological condition.The inventors have discovered that, although not essential, dividing the variant sequence into segments of variants with similar frequency helps provide more robust detection of the sequence by using one label for each segment.
[0018] In a preferred aspect, the labels used in the methods of the present invention include, but are not limited to, one or more labels selected from FAM, HEX, VIC, SUN, Yakima Yellow, CY5, CY5.5, ROX, ABY, TAMRA, and ATTO550.
[0019] The method of the present invention can be used when the sample containing target nucleic acid contains highly fragmented target nucleic acid.For example, the method of the present invention can be used when the sample contains cell-free DNA (cfDNA) and / or DNA from formalin-fixed paraffin-embedded (FFPE) tissue sample.Advantageously, the method of the present invention can be used with the probe that binds to the amplicon of less than 65bp in length during the digital PCR step.It is noted that this size is sufficiently smaller than the typical size for cfDNA, allowing the method to be used for these highly fragmented but critically important nucleic acids. [Brief explanation of the drawings]
[0020] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1-1] FIG. 1 illustrates an exemplary method of the present invention. [Figure 1-2] Same as above. [Figure 2] FIG. 2 summarizes the probes, variants, and labels of the exemplary method of FIG. [Figure 3] FIG. 3 provides a two-color plot from the dPCR reaction to identify variant sequences. [Figure 4] Figure 4 provides a two-color plot from the dPCR reaction to identify variant sequences. [Figure 5] Figure 5 provides a composite two-color plot from the dPCR reactions to identify variants. [Figure 6] FIG. 6 provides a schematic fingerprint or pattern for the dPCR results. [Figure 7] Figure 7 provides a composite two-color plot from the dPCR reactions to identify variants. [Figure 8] FIG. 8 provides a two-color plot from the dPCR reaction to identify variant sequences. [Figure 9] FIG. 9 provides a composite two-color plot from the dPCR reactions to identify variants. [Figure 10] FIG. 10 provides a two-color plot from the dPCR reaction to identify variant sequences. [Figure 11] FIG. 11 provides an exemplary workflow according to the present invention. [Figure 12] FIG. 12 provides exemplary probe concentrations for an exemplary design for identifying variants. DETAILED DESCRIPTION OF THE INVENTION
[0021] Detailed Description The present invention provides a method for detecting low-abundance nucleic acids, including structural variants and mutants. In a preferred embodiment, a sample is subjected to multiplex digital PCR (dPCR) with a probe set that uniquely identifies each variant in the sample. Multiplex dPCR according to the present invention reveals the presence and quantity of multiple variants in a sample, even when the number of variants is greater than the number of distinct optical labels used on the probes. In one illustrative example shown in Figure 7, two colors can separately identify the presence of three different variants. Continuing with Figure 7 as an introductory example, the three variants are specifically identified by two colors because each variant appears as a cluster on a 2D plot of fluorescence intensity for each imaged compartment. Briefly, amplicons containing the variants of interest are diluted and divided into compartments (e.g., droplets or microwells). Each compartment contains a fluorescently labeled hydrolysis probe specific to each variant and primers for amplifying the amplicon. The compartments are thermocycled, and the fluorescence intensities for the two colors are measured. Each compartment is depicted as a single point on a 2D plot (e.g., Figure 7). The points appear in well-defined clusters specific to the variant. The number of points in each cluster provides a measure of the amount of the variant in the sample. In one example, the number of points is multiplied by the dilution factor to obtain the amount (e.g., moles) of amplicons bearing the variant in the sample. Dividing this amount by the sample volume gives the molar concentration. Because the variant concentrations are given for many variants, the dPCR assay provides a measure of the relative abundance of the variants.
[0022] Note that each cluster on the 2D plot can be represented by a vector in 2D space, and the method of the present invention can be extended to higher dimensions. A probe set with three, four, or more distinct colors can be used in a single dPCR assay to quantify any larger number of variants (e.g., 7, or 10, or 15, or 25 or more). Fluorescence from each probe is measured from each section, and each section is represented as a point in a multidimensional space where the number of dimensions can be the number of probes of different colors (although some embodiments may use one or more channels as reference channels or for some other housekeeping purpose that is not used to form the multidimensional space). Each section is represented by a point in the multidimensional space. Any given variant present in multiple sections will appear as a cluster of points in that space. Each cluster can be represented by a vector, which can be determined by regression through the clusters, principal component analysis, the cluster's centroid, or the like. Thus, each variant defines a unique vector in the multidimensional space. Because the vectors are unique (and in particular have a unique orientation with respect to the coordinate system), they can be distinguished even if the number of vectors is greater than the number of dimensions. For example, four colors of probes (e.g., FAM, HEX, Cy5, ROX) can be used to plot compartments as points in a four-dimensional space, and several more numerous variants (e.g., 5, or 8, or 15...) can each be uniquely identified by their respective clusters and / or vectors in the space. The number of points in a cluster can be used to determine the abundance of that variant in the sample. In this way, the relative abundance of the variants is determined.
[0023] The multiplex dPCR assay can be used for a variety of purposes. In one example, the assay is used to detect the presence of disease after treatment. In this example (and many others exist), a tumor biopsy from a subject is subjected to genome sequencing to identify somatic mutations specific to the tumor. The sequencing can reveal tumor-specific mutations and may also reveal mutations of different abundance or MAE. Later, after the subject receives treatment, a sample from the subject is assayed by multiplex dPCR according to the disclosed method. The dPCR reveals the presence or absence of those specific tumor mutations and their relative abundance. The assay uses a relatively inexpensive and minimally invasive blood draw (sometimes referred to as a liquid biopsy). dPCR from the liquid biopsy reveals remission or residual disease.
[0024] Certain alternative embodiments of the present disclosure use variant-enriched sample preparation prior to the dPCR to preferentially increase the copy number of structural variants or mutations that may be present at low concentrations in the sample. The sample preparation process (which functions to gradually amplify the variants of interest), followed by dPCR, results in a reduction in false-negative results from the sample. As a result, the method of the present invention enhances the sensitivity of sample detection, particularly in detecting the presence of low-abundance targets in biological samples.
[0025] The method of the present invention is particularly useful for detecting low-abundance target nucleic acids in samples.Preferably, the target nucleic acid is any nucleic acid sequence whose presence is desired to be detected, including one or more target variant sequences.The target nucleic acid or variant may be, for example, a nucleic acid sequence associated with a clinical condition.In particular, the target nucleic acid sequence is a variant sequence, which is often associated with early tumorigenesis (truncal variant).
[0026] The present invention encompasses methods for detecting variant nucleic acids. An exemplary method involves compartmentalizing a sample from a subject containing a target nucleic acid into reaction compartments and performing digital PCR in the reaction compartments. The dPCR reaction can be performed using a probe set. Preferably, the probes are nucleic acid probes (e.g., hydrolysis or TaqMan probes) with optically detectable labels. The probes in each set have the same optical label, but the optical labels can differ between probes in different sets. Furthermore, in each probe set, the probes are specific to, and therefore hybridize to, different target nucleic acid variant sequences. For reference, a probe for the wild-type sequence can also be included. Using a dPCR assay, optical signals can be detected from the probes in each compartment. The signals from the wild-type and variant sequence probes are used to generate a two-color plot. Subsequently, the presence of one or more variant sequences can be identified based on the deviation of the plotted signal from the expected wild-type cluster plot. This bias can be a recognized pattern characteristic of a particular variant sequence, whether in a subject or historically in association with a particular disease or condition.
[0027] Figure 1 provides an overview of an exemplary probe set, including a wild-type probe, used to detect variant estrogen receptor 1 (ESR1) sequences in samples obtained from subjects. The ESR1 protein regulates the transcription of many estrogen-induced genes known to contribute to growth, metabolism, sexual development, pregnancy, and other reproductive functions. Crucially, ESR1 transcript variants are characteristic of certain forms of breast cancer. Among these variants are recurrent variant ESR1 sequences, which are frequently found in newly diagnosed patients with metastatic and locoregional recurrence of endocrine-treated breast cancer.
[0028] Figure 2 summarizes the probe sets from Figure 1. As shown in Figure 2, 10 unique probes were generated for variant ESR1 sequences in addition to the labeled wild-type probe. The column labeled "Target" identifies the specific ESR1 variant sequence targeted by each probe (which is represented by the amino acid mutations generated relative to the wild-type ESR1 sequence). In this case, the relative frequencies of the variant sequences are known, but the methods of the present invention can be used to elucidate the variant sequence frequencies, or include steps for doing so. "Quantification channel" refers to the optical label and, therefore, the optical channel used to detect the presence of the labeled probe in the dPCR chamber. "Reference channel" refers to the label and corresponding optical channel used to detect the wild-type sequence probe in the reaction chamber. "Discrimination channel" refers to a probe for one of the variant sequences in which an additional optical label is used to further distinguish the variant sequence, as described in more detail. Finally, as shown, the variants are separated into four different sections, each using the same optical label. Segment 1 contains the most prevalent variant sequence, and the remaining segments contain variant sequences with fairly similar frequencies and / or corresponding genomic locations of the mutations that give rise to the variants.
[0029] Figure 12 shows the probe reaction concentrations of an exemplary assay design: targets are detected in channels with target-specific probes at higher concentrations (e.g., 250 nM), and discrimination is facilitated with target-specific probes at lower concentrations (e.g., 50 nM).
[0030] In practice, the method of the present invention may include the steps of: (i) providing a sample containing one or more target nucleic acids (in this case, ESR1 sequences); (ii) optionally performing a pre-amplification, such as a variant-enrichment sample preparation reaction, using primers for each amplicon as described in Figure 1 to selectively increase the copy number of the target nucleic acid in the sample; (iii) aliquoting the sample into multiple subsamples, each in its own compartment with the required probe set; (iv) performing a dPCR reaction on the subsamples; and (v) detecting the resulting optical signals from the probes (fluorescent probes, such as hydrolysis probes, that indicate the presence of wild-type or variant sequences generated in each compartment). The optical signals from each compartment are then detected and plotted to identify the presence or absence of the variant sequence being queried.
[0031] In the exemplary assay outlined in Figures 1-2, the most common mutation, p.D538G, is detected based on its optical label using a dedicated channel (FAM / Green). The next three most common mutations (p.Y537S, p.E380Q, and p.Y537N) are detected based on the fluorophores for their sections of the probe / variant using the CY5 / Crimson channel. In certain methods of the invention, and as exemplified for variant p.Y537C, a second probe is used with a different optical label (in this case, FAM, detected by the Green channel). A discriminating probe is not necessarily required, but if used, the variants can still be simultaneously identified using fewer optical labels / channels than the variants being detected. Subsequently, in the ROX / Red channel, three (3) less frequent mutations (p.Y537N, p.Y537H, and p.Y537D) are detected in the section, and in the fourth section, three (3) less frequent mutations are detected in the ATTO550 / Orange channel (p.L536H, p.L536P, and p.L536R).
[0032] Figure 3 provides the resulting two-color plot of the p.D538G mutation (using the D538G FAM / Green channel) versus the wild-type probe signal (Hex / Yellow channel) for the dPCR reaction. In the plot, the Y-axis is the FAM / Green channel (variant) probe fluorescence intensity, and the X-axis is the HEX / Yellow channel (wild-type) fluorescence intensity. The two-color plot is divided into quadrants, with the upper left representing a positive variant probe signal only; the lower right representing a positive wild-type signal only; the upper right representing a double positive (variant and wild-type); and the lower left representing a double negative (variant and wild-type). In certain aspects, as done here, it is advantageous to dedicate a channel to the most frequently occurring variant to maximize sensitivity and provide internal validation of the assay. As shown, the signal from the FAM / Green channel from the D538G probe produces a clearly identifiable pattern when plotted against the HEX / Yellow channel, indicating the presence of the variant in the sample.
[0033] Figure 4 provides similar plots for the intercepts of variants including p.Y537S, p.E380Q, and p.Y537N, detected using their CY5 probes in the Crimson channel. In the plots on the left for each variant, the X-axis displays the wild-type probe signal (HEX / Yellow), and the Y-axis displays the variant probe signal (CY5 / Crimson). As shown, when plotted, each variant produces a signature pattern in its bias relative to the wild-type plot—vertical for p.E380Q, slightly tilted for Y537C, and significantly tilted for Y537S. As described in Figures 1-2, a small amount of Y537H probe (which had a FAM label detectable using the Green channel) was included in this assay. The plots on the right for each variant show the results of plotting the Green channel signal against the Crimson channel signal. As shown, the deviation from the wild-type plot changes when these channels are interrogated, which helps distinguish the presence of the Y537C and Y537H probes.
[0034] Figure 5 provides a synthesis of both the right and left plots of Figure 4. As shown, each mutation provides a characteristic pattern or fingerprint relative to the wild-type signal, which can be used to identify the presence or absence of each variant sequence in the sample.
[0035] Figure 6 provides an illustration of the distinct fingerprint / pattern of each variant when interrogated using the probes as described herein. Thus, the methods of the invention can be used not only to detect variant sequences based on their characteristic dPCR patterns / fingerprints relative to the wild-type probe signal, but also to distinguish between those patterns / fingerprints, whether for individual subjects or for assays of entire populations.
[0036] Figure 7 provides a composite of dPCR signals plotted for variants in sections containing Y537D, Y537N, and Y537H using the ROX / Red channel corresponding to the variant probe and HEX / Yellow for the wild type. The combined frequency of these three variants is estimated to be 8.7%. The signal quality of the two more common variants (Y537N and Y537H, combined 8.2%) is easily detected. Furthermore, the signal amplitude for Y537D is extremely rare (0.47%) and low, but it is also detectable.
[0037] Figure 8 provides individual plots for the Y537D, Y537N, and Y537H variants, which were used to discern the patterns / fingerprints for each variant.
[0038] Figure 9 provides a composite of dPCR signals plotted for variants in sections containing L536R, L536H, and L536P using the ATTO550 / Orange channel corresponding to the variant probes and HEX / Yellow for the wild type. As shown, patterns / fingerprints are clearly evident for each of these rare variant sequences when plotted against the wild type channel.
[0039] Figure 10 provides individual plots for the L536R, L536H, and L536P variants, which were used to discern the patterns / fingerprints for each variant.
[0040] The methods of the invention are useful for detecting whether a variant sequence is present in a sample, which may contain a mixture of target nucleic acids, where only a portion of the target nucleic acids may contain the variant sequence. In particular, the methods are useful for detecting the presence of a variant sequence in a sample containing target nucleic acids, only a small proportion of which may potentially contain the variant sequence.
[0041] The target nucleic acid sequence can be any target nucleic acid sequence that is desired to be detected, however, in preferred aspects, the target nucleic acid sequence is a variant sequence associated with a clinical condition (e.g., cancer).
[0042] In certain methods of the present invention for detecting the presence of a variant sequence, the method may include providing a sample containing one or more target nucleic acids and performing pre-amplification to increase the copy number of the target nucleic acid in the sample. Subsequently, the pre-amplified sample is aliquoted into a plurality of subsamples. The method may then include performing PCR (preferably dPCR) on the subsamples and detecting the target nucleic acid (e.g., a variant sequence).
[0043] Figure 11 provides an exemplary workflow showing sample acquisition, bioinformatics workflow, and sample processing for detecting and / or identifying variant sequences. This exemplary workflow shows low-pass whole genome sequencing of the tumor as the first step, followed by filtering and selection of targets (which may be known to be locations of sequence variants, such as structural variants or mutations (or combinations thereof)).
[0044] The methods of the present invention may include the use of stepwise nucleic acid amplification steps as targeted preamplification, i.e., to increase the relative abundance of a target within a sample. In some embodiments, the preamplification is exponential, increasing the abundance of a selected target. The use of targeted preamplification increases the copy number of the target nucleic acid in the sample. The preamplification may be symmetric or asymmetric, or a combination of the two. In asymmetric or stepwise amplification, a single strand of nucleic acid is preferentially copied linearly without copying other strands within the sample. Some embodiments perform stepwise amplification by providing only one primer that copies the target (instead of a forward and reverse primer pair in PCR). Certain embodiments discussed herein perform asymmetric amplification using primer-H and primer-L. See also U.S. Patent No. 11,066,707 and U.S. Publication No. 2022 / 0056533 A1 (both incorporated by reference).
[0045] The methods of the present invention have very low detection limits, which allow for the detection of target variant sequences that are often present at very low levels, and / or the detection of the presence of variant sequences that are potentially present at very low levels in mixtures containing other target nucleic acid sequences.
[0046] Sample and target nucleic acid The sample can be any sample in which it is desired to detect the presence or absence of the variant sequence. For example, if the variant sequence indicates a clinical condition, the sample can be a sample from an individual at risk of acquiring the clinical condition. The variant sequence can differ from the wild-type sequence by substitution, deletion, and / or insertion.
[0047] The method of the present invention further provides that one or more targeting nucleic acid molecules are associated with variant sequences, and wherein said variant sequences can be variants of wild-type nucleic acid sequences.Above-mentioned variant sequences can be selected from the group consisting of single nucleotide variants (SNV), insertions and deletions (indels), duplications, copy number variants (CNV), inversions and translocations.The method of the present invention can also be used to analyze cfDNA / ctDNA samples from cancer patients, and optionally, include mutations or structural variants from cfDNA / ctDNA.
[0048] In some embodiments, the methods of the invention include one or more steps of diagnosing a subject based on the detection of one or more variant sequences. Such a diagnosing step can, for example, diagnose a subject with a disease associated with the detected variant sequence, report the probability that the patient has or will develop such a disease based on the identification of one or more variant sequences, or assess the relative frequency of identified variants to assess disease progression or disease response to treatment.
[0049] The variant sequences identified using the methods of the present invention may be variants associated with a particular type or stage of cancer, or with cancers having particular characteristics (e.g., metastatic, drug-resistant, and drug-responsive). For example, certain mutations are known to be associated with patient outcomes and / or particular conditions. In certain aspects, the methods of the present invention may provide information used in treatment decisions, guidance, and monitoring, as well as in the development and clinical trials of treatments for particular conditions. For example, the effectiveness of treatment may be monitored by comparing variants and / or variant frequencies detected before, during, and after treatment. Longitudinal monitoring may be used to assess increases or decreases in variant sequences or frequencies, to identify new or absent variant sequences after treatment, which may, for example, guide subsequent treatment decisions. In certain aspects, diagnosing a subject includes diagnosing the subject with a particular stage or type of cancer associated with the detected sequence variant.
[0050] For polymorphisms or small indels (typically less than about 50 bases), it may be preferable to amplify both the wild-type and variant sequences in a polymerase reaction using a pair of primers capable of specifically amplifying the target nucleic acid sequence. It may also be preferable that the wild-type and variant sequences do not differ excessively in length from each other. For structural rearrangements (typically at least about 50 bases), the variants are amplified in a stepwise manner prior to digital PCR, and the assay also includes a parallel reference assay targeting a stable region of the genome to quantify the signal. Structural variants are described in Mahmoud, 2019, Structural variant calling: the long and the short of it, Genome Biology 20:a246 (incorporated by reference).
[0051] The sample can be any biological sample, including bodily fluid samples including bile, blood, plasma, serum, sweat, saliva, urine, feces, phlegm, mucus, sputum, tears, cerebrospinal fluid, synovial fluid, pericardial fluid, lymphatic fluid, semen, vaginal secretions, lactation or menstrual products, amniotic fluid, pleural fluid, rheum, or vomit.
[0052] Amplification and Preamplification The methods of the invention may involve the use of an amplification reaction. Amplification of a nucleic acid is the production of copies of said nucleic acid.
[0053] The method of the present invention also includes a preamplification step. This preamplification can be described as a variant-enriched sample preparation reaction step. Such preamplification can amplify only selected targets by a limited, controlled, or known amount. For example, 10 cycles of linear preamplification using primers specific to a variant increases the copy abundance of that variant in the sample by approximately 10-fold. Because the abundance is increased by a known amount, when the sequences in the reaction mixture are subsequently quantified (e.g., by digital PCR), the amount of those sequences in the original sample can be determined (using the known increased abundance of the selected variant). Furthermore, this preamplification aids in the detection of very rare sequences in a sample, greatly increasing the probability of detection of those variants. For this reason, it may be preferable for preamplification to specifically increase the copy abundance of selected variants in the sample by a known amount. According to the present invention, the preamplification step is sample preparation to preferentially increase the copy number of the selected target through asymmetric stepwise amplification, symmetric exponential amplification, or a combination of the two. Compared to the exponential production of amplicons in PCR, stepwise preamplification can be expected to proceed without more than linear copy amplification. The stepwise amplification can proceed using unpaired primers (e.g., single primers) that extend and copy through the target of interest. In some embodiments, stepwise preamplification involves the use of at least several primers, where only one primer is active in preamplification. The active primer can be active due to differences in melting temperature (Tm) and the annealing temperature of the primer. The use of these primers for asymmetric stepwise amplification is disclosed in U.S. Patent No. 11,066,707, which is incorporated herein by reference in its entirety.
[0054] The methods of the present invention can use stepwise preamplification to increase the abundance of a target of interest in a sample, even when that target is present in very low amounts. This can alleviate problems associated with molecular detection assays that rely on PCR to amplify targets, where targets present in very low abundance may go undetected due to the stochastic nature of PCR and where PCR reactions can be plagued by "dead volumes" where no detection occurs. These problems are addressed by selectively increasing the abundance of the target of interest in what is described herein as a preamplification step. In one embodiment, the preamplification step is non-exponential, i.e., not PCR. Instead, the preamplification step is preferably stepwise, which can be understood to mean that the extension products from one round of preamplification are not substrates for copying by any reverse primer. Rare targets of interest are increased in abundance by this preamplification step. The increase in abundance is approximately linear (not exponential) over the cycles of preamplification.
[0055] It may be preferable to perform stepwise (non-exponential) preamplification so that other materials in the sample are still accessible to subsequent amplification steps. A preferred embodiment uses preamplification specific to gene sequences selected for clinical significance (e.g., structural variants specific to tumor-derived nucleic acids). Indeed, the methods of the present invention can identify tumor mutations and select one or more structural variants for clinical significance (e.g., structural variants that may persist even after cancer treatments such as chemotherapy).
[0056] Preferably, in the methods of the present invention, sequences containing potential variants are selected and then pre-amplified. This pre-amplification specifically increases the abundance of copies of the selected variant in the sample. The sample can then be assayed for the selected variant, as well as any other variants potentially present in the sample, using a digital PCR (dPCR) assay, as described above. Indeed, embodiments of the present invention can be multiplexed, for example, using differentially labeled fluorescent hydrolysis probes, to simultaneously interrogate the sample for multiple variants while always using fewer distinct fluorescent reporters than the variants being detected. Because the selected variant has been increased in abundance by stepwise pre-amplification, it is not lost due to dead volume or the stochastic nature of PCR. Thus, the assays of the present invention are useful for multiplexed detection of very rare targets in a sample and may be particularly useful for detecting cancer-specific variant sequences, such as in circulating tumor DNA (ctDNA) in blood or plasma samples (e.g., in liquid biopsies). Moreover, as discussed further herein, stepwise preamplification and exponential amplification for dPCR can proceed in the presence of the same set of reagents (primers, dNTPs, polymerase, ions, probes, etc.) without the need for cleanup steps, reagent changes, or the addition of reagents as the assay progresses. Indeed, embodiments discussed herein use primer pairs that function to stepwise preamplify a target under one thermocycle and exponentially amplify the target for dPCR under a different thermocycle.
[0057] In the methods of the present invention, the nucleic acid polymerase enzymes used can have different extension temperatures. The extension temperature is the temperature that allows the enzymatic activity of the nucleic acid polymerase after primer annealing. Typically, nucleic acid polymerases are active over a certain temperature range, so the extension temperature can be any temperature within that range. Most nucleic acid polymerases have an optimum temperature, but retain activity at temperatures other than the optimum temperature. In such cases, the extension temperature can be any temperature that allows the primer to anneal, and the nucleic acid polymerase is active even if the temperature is not the optimum temperature. At the extension temperature of the nucleic acid polymerase, the enzyme is capable of catalyzing the synthesis of a new nucleic acid strand complementary to the template strand at the extension temperature. In certain methods, the extension temperature is close to the melting temperature of Primer-H. Therefore, a nucleic acid polymerase with polymerase activity at a temperature close to the melting temperature of Primer-H can be selected, and / or Primer-H can be designed to have a melting temperature close to the extension temperature.
[0058] Amplicons or products generated during a preamplification step (e.g., a variant-enrichment sample preparation reaction step) can be used as direct sample input for PCR. That is, after preamplification, reagents for PCR can be added to the reaction mixture from the preamplification. For example, if the preamplification is performed in a tube, the tube can be refilled (without cleanup) after preamplification to add any additional reagents useful in PCR. In some embodiments, the primers used in the present invention are part of multiple primer pairs capable of amplifying different target nucleic acid sequences. Preamplification can be performed under a first temperature control, and then PCR can be performed under a second temperature control after adding any additional reagents to the tube, if necessary, without any cleanup. The methods of the present invention can be performed without any cleanup step, which can result in loss of analyte material.
[0059] The methods of the present invention may involve the use of PCR reagents for sample preparation reactions and dPCR. The sample preparation reaction and / or dPCR may include at least a portion of the sample, a set of primers, and sufficient PCR reagents to enable a polymerase reaction. Methods and reagents useful for performing PCR reactions are well known to those skilled in the art. For example, the PCR reaction may include any of the nucleic acid polymerase and PCR reagents described in the "PCR Reagents" section below. Depending on the method for detecting whether the sample preparation product contains the variant sequence, the sample preparation reaction may also include a detection reagent. PCR reagents are reagents added to PCR in addition to the nucleic acid polymerase, sample, and primer set. The PCR reagents include at least nucleotides. In addition, the PCR reagents may include other compounds such as salts and buffers.
[0060] For most purposes, the PCR reagents will include nucleotides, and thus may include deoxynucleoside triphosphates (dNTPs), particularly all four naturally occurring deoxynucleoside triphosphates (dNTPs).
[0061] The PCR reagents frequently contain deoxyribonucleoside triphosphate molecules, including dATP, dCTP, dGTP, and dTTP. In some cases, dUTP is added.
[0062] The PCR reagents may also contain compounds useful for supporting the activity of nucleic acid polymerase. Thus, the PCR reagents may contain divalent cations, such as magnesium ions. The magnesium ions may be added, for example, in the form of magnesium chloride (MgCl), or magnesium acetate or magnesium sulfate may be used.
[0063] The PCR reagents may also include one or more of the following: Non-specific blocking agents, such as BSA or gelatin from bovine skin, beta-lactoglobulin, casein, milk powder, or other common blocking agents, Non-specific background / blocking nucleic acids (e.g., salmon sperm DNA), biopreservatives (e.g., sodium azide), PCR enhancing factors (e.g., betaine, trehalose, etc.), Inhibitors (e.g., RNAse inhibitors).
[0064] The PCR reagents may also contain other additives, such as dimethyl sulfoxide (DMSO), glycerol, betaine (mono)hydrate (N,N,N-trimethylglycine = [carboxymethyl] trimethylammonium), trehalose, 7-deaza-2'-deoxyguanosine triphosphate (dC7GTP or 7-deaza-2'-dGTP), formamide (methanamide), tetramethylammonium chloride (TMAC), other tetraalkylammonium derivatives (e.g., tetraethylammonium chloride (TEA-Cl) and tetrapropylammonium chloride (TPrA-Cl)), non-ionic detergents (e.g., Triton® X-100, Tween® 20, Nonidet P-40 (NP-40)), or PREXCEL-Q.
[0065] The PCR reagents may include a buffering agent.
[0066] In some cases, a nonionic ethylene oxide / propylene oxide block copolymer is added to the aqueous phase at a concentration of about 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, or 1.0%. Common biosurfactants include nonionic surfactants (e.g., Pluronic® F-68, Tetronics, Zonyl FSN). Pluronic® F-68 may be present at a concentration of about 0.5% w / v.
[0067] A wide range of common commercially available PCR buffers from various vendors can be used in place of the buffered solutions described above.
[0068] The methods of the invention generally involve the use of a nucleic acid polymerase, which can be any nucleic acid polymerase, such as a DNA polymerase, that has activity at the extension temperature.
[0069] The nucleic acid polymerase may be a DNA polymerase with 5' to 3' exonuclease activity, particularly in the methods of the invention, where the method or kit involves the use of a detection probe, such as a Taqman detection probe.
[0070] Any DNA polymerase can be used, for example, the DNA polymerase that has the 5'→3' exonuclease activity of catalyzing primer extension.For example, thermostable DNA polymerase can be used.Preferably, the nucleic acid polymerase is Taq polymerase, which allows the Taqman probe to be used to identify variant sequences as described herein.
[0071] The method of the present invention includes a sample preparation reaction step, which can be either asymmetric stepwise amplification or symmetric exponential amplification, or a combination of the two.Asymmetric stepwise amplification can proceed using unpaired, for example, a single primer (even if multiple primers are used, each can be "single" in the sense that it is unpaired with the reverse primer that anneals to the extension product of that single primer).In some embodiments, stepwise preamplification can include the use of at least two primers, where only one primer is active in the sample preparation reaction.For example, the active primer is active due to the difference in the melting temperature (Tm) and annealing temperature of the provided primers.
[0072] More preferably, the step of asymmetric stepwise amplification includes (i) providing a pair of primers capable of amplifying a target nucleic acid, wherein the pair of primers comprises Primer-H and Primer-L. According to the method of the present invention, the melting temperature of Primer-H may be about 10°C to about 22°C higher than the melting temperature of Primer-L, Primer-L comprises a sequence complementary to a fragment of an extension product of Primer-H, and the sample preparation reaction is carried out. The sample preparation reaction also comprises a nucleic acid polymerase having polymerase activity, primers, and PCR reagents.
[0073] In embodiments involving Primer-H and Primer-L, the stepwise preamplification proceeds using an annealing temperature at which Primer-H anneals (but Primer-L does not). During the stepwise amplification, Primer-H functions as a "single" primer (despite the presence of Primer-L). Primer-L does not anneal to any significant degree because the reaction mixture is not lowered below the annealing temperature of Primer-L. The stepwise preamplification can be carried out for a predetermined number of cycles (e.g., 1, 5, etc.), or for a fixed amount of time, or until the sample exhibits a result (such as a change in optical density, cleavage of a fluorescent probe, etc.). After the stepwise preamplification, the reaction mixture can then be subjected to exponential amplification. For exponential amplification, in the annealing step, the temperature is lowered to the annealing temperature of Primer-L, which promotes annealing of both Primer-H and Primer-L.
[0074] It is important to note that the above paragraphs describe the function of Primer-H and Primer-L, because those primers can function among many other primers. For example, tens, hundreds, thousands, or more loci can be probed in parallel using a corresponding number of primer pairs. For example, in a particular assay, 24 loci are probed in parallel (although the number 24 is arbitrary and can be 1, 2, 3, 6, 17, 96, 99, 384, 1,000, 1,536, or any integer multiple of those numbers). The assay of the present invention can use a primer pair for each locus, for example, 24 primer pairs. Any one or any number of those primer pairs can fit the description of Primer-H and Primer-L. However, and this is important, because it may be desirable to preamplify only certain loci, any number of the primer pairs can include forward and reverse primers, each with essentially the same annealing temperature as Primer-L.
[0075] In other embodiments, the preamplification proceeds with an unpaired "single" primer (e.g., operating at the primer-H annealing temperature). The reaction mixture (original sample, plus added reagents, plus preamplification product) can then be subjected to conditions for exponential amplification. The exponential amplification may use a primer pair (any one of which may "match" all or part of the single primer or anneal to a target within the length of the extension product of the unpaired single primer). The reactions may proceed at different temperatures, and the paired primer may not be available until after stepwise preamplification. For example, the paired primer may be added (e.g., by microfluidic handling) or released from entrapment or binding (e.g., by chemical, thermal, or photolysis of hydrogel beads).
[0076] In certain methods of the present invention, asymmetric stepwise amplification may be used, which includes the steps of: providing a primer pair capable of amplifying a target nucleic acid, the primer pair including Primer-H and Primer-L, where the melting temperature of Primer-H is about 10°C to about 22°C higher than that of Primer-L, and Primer-L contains a sequence complementary to a fragment of an extension product of Primer-H; providing a nucleic acid polymerase having polymerase activity at the extension temperature; and preparing sample preparation reactions, each containing a portion of the sample, the primer set, the nucleic acid polymerase, and PCR reagents. The melting temperature of Primer-H may be about 10°C higher than that of Primer-L. Preferably, the melting temperature of Primer-H is about 12°C higher than that of Primer-L. More preferably, the melting temperature of Primer-H is about 16°C higher than that of Primer-L.
[0077] In certain methods, the asymmetric stepwise amplification can also include a primer set, where at least one primer can specifically amplify only one strand of the target nucleic acid sequence, and a sample preparation reaction is performed in a solution containing a nucleic acid polymerase having polymerase activity, and the sample preparation reaction is performed. The sample preparation reaction also includes a nucleic acid polymerase having polymerase activity, primers, and PCR reagents.
[0078] The present invention includes a process of symmetric exponential amplification, wherein the symmetric exponential amplification comprises: (i) providing a set of primers capable of specifically amplifying the target nucleic acid sequence; (ii) providing a nucleic acid polymerase having polymerase activity at an extension temperature; (iii) preparing a sample preparation reaction each containing a portion of the sample, the primer set, the nucleic acid polymerase, and PCR reagents; and performing the sample preparation reaction.
[0079] Certain methods of the invention may involve asymmetric stepwise amplification followed by symmetric exponential amplification. The steps involved in asymmetric stepwise amplification and symmetric exponential amplification are outlined above.
[0080] The methods of the invention may also involve symmetric exponential amplification followed by asymmetric stepwise amplification. The steps involved in asymmetric stepwise amplification and symmetric exponential amplification are outlined above.
[0081] The methods of the present invention can also include asymmetric stepwise amplification, and the symmetric exponential amplification can be performed in the same reaction volume. In certain methods, the asymmetric stepwise amplification and the symmetric exponential amplification can be performed using the same primer set.
[0082] In certain methods of the invention, asymmetric stepwise amplification can be activated at higher temperatures compared to symmetric exponential amplification, e.g., a thermocycler can be programmed to go down to a higher annealing temperature for the stepwise amplification but a lower annealing temperature for the exponential amplification.
[0083] The method of the present invention may include the steps of: performing multiple PCR reactions on the sample that has undergone the sample preparation reaction, the steps including a primer pair capable of specific amplification of a target nucleic acid, and a nucleic acid polymerase having polymerase activity at an extension temperature; preparing PCR reactions, each PCR reaction including a portion of the sample, the primer set, the nucleic acid polymerase, and PCR reagents; and performing symmetric exponential amplification. Preferably, these PCR reactions are conventional PCR reactions. Preferably, the PCR is digital PCR (dPCR) and uses a variant sequence discrimination probe as described. The sample containing the product after the sample preparation reaction step can be used as a direct sample input for the PCR. Optionally, the primer pair used in the present invention is part of multiple primer pairs capable of amplifying different target nucleic acid sequences.
[0084] The methods of the invention may further involve the use of multiple primer pairs capable of amplifying different target nucleic acid sequences. Preferably, the methods of the invention involve the use of multiplex PCR.
[0085] The methods of the present invention may use quantitative PCR, quantitative fluorescent PCR (QF-PCR), multiplex fluorescent PCR (MF-PCR), real-time PCR (RT-PCR), single-cell PCR, restriction fragment length polymorphism PCR (PCR-RFLP), PCR-RFLP / RT-PCR-RFLP, hot-start PCR, nested PCR, in situ polony PCR, in situ rolling circle amplification (RCA), digital PCR (dPCR), droplet digital PCR (ddPCR), bridge PCR, picotiter PCR, and emulsion PCR.
[0086] Primer-H and Primer-L The method of the present invention may include the use of several primers, including primers designated as Primer-H and Primer-L. Primer-H is a primer with a high melting temperature, while Primer-L is a primer with a low melting temperature. The melting temperature of a primer is the temperature at which 50% of the primers form a stable double helix with their complementary sequence, and the other 50% separate into single-stranded molecules. The melting temperature can also be referred to as Tm or Tm. Preferably, Tm as used herein is calculated using the nearest neighbor method based on the method described in Breslauer, 1986, "Predicting DNA duplex stability from the base sequence," PNAS 83:3746-50 (incorporated by reference), using a salt concentration parameter of 50 mM and a primer concentration of 900 nM. For example, the method is implemented by the software "Multiple Primer Analyzer" from Life Technologies / Thermo Fisher Scientific Inc.
[0087] Certain methods of the present invention may use a primer set comprising Primer-H and Primer-L, wherein the melting temperature of Primer-H is about 10°C to about 22°C, preferably at least 10°C, more preferably at least 15°C higher than the melting temperature of Primer-L, and wherein Primer-L comprises a sequence complementary to the extension product of Primer-H.
[0088] The primer-H is preferably designed as a primer for amplifying the target sequence or a sequence complementary to the target sequence. Therefore, the primer-H is preferably capable of annealing to either the target nucleic acid sequence or a sequence complementary to the target nucleic acid sequence. For example, the primer-H may be capable of annealing to the complementary strand of the target nucleic acid sequence at or near the 5' end of the target nucleic acid sequence, or the primer-H may be capable of annealing to the target nucleic acid sequence at or near the 3' end of the target nucleic acid sequence. Therefore, the primer-H may comprise a sequence identical to the 5' end of the target nucleic acid sequence. The primer-H may further comprise a sequence identical to the 5' end of the target nucleic acid sequence. The primer-H may also comprise a sequence identical to the target nucleic acid sequence. Therefore, the primer-H may comprise a sequence complementary to the 3' end of the target nucleic acid sequence. The primer-H may further comprise a sequence complementary to the 3' end of the target nucleic acid sequence.
[0089] Similarly, Primer-L is preferably designed as a primer for amplifying the target sequence or a sequence complementary to the target sequence. If Primer-H is designed for amplifying the target sequence, Primer-L is preferably designed for amplifying a sequence complementary to the target sequence, and vice versa. Thus, Primer-L is preferably capable of annealing to either the target nucleic acid sequence or a sequence complementary to the target nucleic acid sequence. If Primer-H is capable of annealing to the target nucleic acid sequence, Primer-L is preferably capable of annealing to a sequence complementary to the target nucleic acid sequence, and vice versa. For example, Primer-L may be capable of annealing to the complementary strand of the target nucleic acid sequence at or near the 5' end of the target nucleic acid sequence, or Primer-L may be capable of annealing to the target nucleic acid sequence at or near the 3' end of the target nucleic acid sequence. Thus, Primer-L may contain a sequence identical to the 5' end of the target nucleic acid sequence. The primer-L may further comprise a sequence identical to the 5' end of the target nucleic acid sequence. The primer-L may also comprise a sequence identical to the target nucleic acid sequence. Thus, the primer-L may comprise a sequence complementary to the 3' end of the target nucleic acid sequence. The primer-L may further comprise a sequence complementary to the 3' end of the target nucleic acid sequence.
[0090] Primer-H may have a nucleotide sequence identical to the sequence at the 5' end of the target nucleic acid sequence, and Primer-L comprises or consists of a sequence identical to the complementary sequence at the 3' end of the target nucleic acid sequence.
[0091] Primer-L may have a nucleotide sequence identical to the sequence at the 5' end of the target nucleic acid sequence, and Primer-H comprises or consists of a sequence identical to the complementary sequence at the 3' end of the target nucleic acid sequence.
[0092] Primer-H and Primer-L are designed to have melting temperatures as described herein. Those skilled in the art can design Primer-H and Primer-L to have desired melting temperatures by adjusting the sequence of the primers, the length of the primers, and, if necessary, by incorporating nucleotide analogs as described in the "Primer Set" section herein above.
[0093] Primer-H is designed to have an annealing temperature significantly higher than that of Primer-L, for example, at least 10° C. higher. Thus, the melting temperature of Primer-H may be at least 12° C. higher than that of Primer-L, for example, at least 15° C. higher, preferably at least 14° C. higher, even more preferably at least 16° C. higher, even more preferably 18° C. higher (for example, at least 20° C. higher, for example, in the range of 15 to 50° C., for example, in the range of 15 to 40° C., for example, in the range of 15 to 25° C.).
[0094] Generally, it may be preferable that the melting temperature of Primer-H is as high as possible, but not higher than the highest functional extension temperature of at least one nucleic acid polymerase. The extension temperature does not need to be the optimum temperature for the nucleic acid polymerase, but it is preferable that at least one nucleic acid polymerase has activity at the melting temperature of Primer-H. Thus, the melting temperature of Primer-H may be close to or even exceed 80°C.
[0095] The melting temperature of Primer-L is sufficiently high to ensure specific annealing of Primer-L to the target nucleic acid sequence / complementary sequence of the target nucleic acid sequence, and it is also preferred that the melting temperature of Primer-H be significantly higher than that of Primer-H, so that the melting temperature of Primer-H is frequently at least 60°C. The melting temperature of Primer-H may also frequently be at least 70°C. The melting temperature of Primer-H may be, for example, in the range of 60 to 90°C, for example, in the range of 60 to 85°C, for example, in the range of 70 to 85°C, for example, in the range of 70 to 80°C.
[0096] The melting temperature of Primer-L is preferably sufficiently high to ensure specific annealing of Primer-L to the target nucleic acid sequence / the complementary sequence of the target nucleic acid sequence, but is also significantly lower than the melting temperature of Primer-H. Frequently, the melting temperature of Primer-L is in the range of 30 to 55°C, for example, in the range of 35 to 55°C, preferably in the range of 40 to 50°C.
[0097] The method of the present invention can also include the use of a primer set or a plurality of primers. The primer set or a plurality of primers includes two or more different primers. A primer set including at least one pair of primers can specifically amplify a target nucleic acid. Furthermore, the primer set according to the present invention includes at least primer-H and primer-L. Therefore, when a primer set includes only two different primers, the primer set includes primer-H and primer-L, where the primer-H and primer-L are capable of amplifying the target nucleic acid.
[0098] detection The present invention generally involves detecting whether the sample contains a variant sequence. The detection can be achieved in any suitable manner known to those skilled in the art. For example, many useful detection methods are known in the art and can be employed in conjunction with the methods of the present invention.
[0099] The detection step can include the presence of a detection reagent in the PCR reaction, which can be any detectable reagent, such as a compound containing a detectable label, where the detectable label can be, for example, a dye, radioactivity, a fluorophore, a heavy metal, or any other detectable label.
[0100] Frequently, the detection reagent comprises a fluorescent compound associated with a variant sequence-specific probe as described herein.
[0101] The detection reagent may include a detection probe. The detection probe may include a nucleotide oligomer or polymer, which may optionally include a nucleotide analog. Frequently, the detection probe may be a DNA oligomer. Typically, the detection probe is linked to a detectable label, for example, by covalent bonding. The detectable label may be any of the detectable labels described above, but is preferably an optically detectable label (e.g., a fluorophore).
[0102] The detection probe generally has the ability to specifically bind to the target nucleic acid sequence. For example, the detection probe may have the ability to specifically bind to the target nucleic acid including a variant sequence. Thus, the detection probe may be capable of annealing to the target nucleic acid sequence or a sequence complementary to the target nucleic acid sequence. Thus, the detection probe may comprise a sequence identical to the target nucleic acid sequence or a fragment of the sequence complementary to the target nucleic acid sequence. It is generally preferred that the detection probe comprises a sequence different from the sequence of any of the primers in the primer set.
[0103] Quantification The method of the present invention can include the steps of: providing a sample containing one or more target nucleic acids; performing a sample preparation reaction to increase the copy number of the target nucleic acid in the sample; aliquoting the sample into a plurality of aliquots; performing polymerase chain reaction (PCR) on the aliquots; and detecting variant nucleic acid sequences. Using targeted preamplification, the increased copy number of the target can be used for a subsequent PCR reaction used to identify the variant sequence. The advantage of preamplification is that it significantly reduces or eliminates false negatives, which are often a problem in samples with low target concentrations, by reducing the problems of dead volume and stochastic sampling error.
[0104] The present invention also provides a reference assay for detecting structural variants or mutations in a patient's cfDNA.In particular, the present invention provides a method for detecting wild-type cfDNA and any variant sequence of interest contained in the cfDNA.Therefore, the present invention provides an assay for detecting and / or quantifying the wild-type (reference) sequence of cfDNA, comprising the use of a pair of primers and a probe.The present invention further provides an assay for calculating the variant allele fraction (VAF) for the target nucleic acid in the sample.In certain aspects, the target nucleic acid can be a variant of the wild-type nucleic acid sequence in the sample.
[0105] Certain methods of the present invention may further provide for the use of publicly available information to determine genomic regions that are stable for amplification and design the assay accordingly. As an example, one skilled in the art may determine that chromosome 2, band 13 (2p13) is stable and, as a result, is not susceptible to copy number changes. The present invention further provides that this information can be used to design the assay, particularly to minimize copy number variations in the sample.
[0106] The methods of the present invention may include quantifying the amount of target nucleic acid present in the sample. Replication assays have variability in the preamplification step. In particular, the efficiency of preamplification may be less than 100%. Therefore, a correction factor is applied to back-calculate the original sample concentration. As an example of this problem, the preamplification step may result in a 50-fold amplification in one instance, but only a 30-fold amplification in another instance (yielding 50-fold or 30-fold copies of the target sample, respectively). Furthermore, the denaturing conditions of the sample also affect how many copies are measured by dPCR. As a result, the same sample may have a two-fold difference in dPCR concentration depending on the denaturing conditions (one intact double-stranded DNA molecule in one compartment is measured as one copy, while two single-stranded DNA molecules in two compartments are measured as two copies). The methods of the present invention may include the use of two positive control reactions during the preamplification. The first positive control contains positive control DNA and the assay, resulting in preamplification. The second positive control contains an equal amount of positive control DNA without the primers. The first and second positive controls are technical replicates, except that the first positive control receives the primers.
[0107] The concentration of the replicates with primers is then divided by the replicates without primers to estimate the efficiency of the assay. The calculated efficiency can be applied to measurements performed on actual samples. Indeed, the present invention provides a method for quantifying the amount of target nucleic acid in the sample. [Example]
[0108] Example Using the exemplary method of the present invention, the ESR1 variants (D538G, Y537S, Y537C, Y537N, Y537H, Y537D, L536H, L536P, and L536R) described in Figures 1-10 were simultaneously detected. Two ESR1 target sequences were amplified to generate first and second amplicons using primer pairs for each target sequence. The first amplicon contained nucleotides of the wild-type ESR1 nucleotide sequence encoding amino acids 529-552 of the ESR1 protein sequence. The second amplicon contained nucleotides of the wild-type ESR1 nucleotide sequence encoding amino acids 372-395 of the ESR1 protein sequence. Thus, the genomic location of each variant was included between the two amplicons.
[0109] Figures 1-2 show the probe sequences, the ESR1 variants they interrogate, and the associated optical labels for each probe. Figure 12 shows the probe reaction concentrations for an exemplary assay design. Targets are detected in channels with higher concentrations (e.g., 250 nM) of target-specific probes, and discrimination is facilitated with lower concentrations (e.g., 50 nM) of target-specific probes.
[0110] In the above assay, the most common mutation, p.D538G, was detected based on the optical label of its corresponding probe using a dedicated channel (FAM / Green). The next three most common mutations (p.Y537S, p.E380Q, and p.Y537N) were detected based on the fluorophore for that section of the probe / variant using the CY5 / Crimson channel. For the variant p.Y537C, a second probe was used with a different optical label (in this case, FAM, detected by the Green channel). As shown in Figure 1, the concentration of this second probe for p.Y537C was provided at a lower concentration than the other probes. Subsequently, sections with three less frequent mutations (p.Y537N, p.Y537H, and p.Y537D) were detected in the ROX / Red channel, and in a fourth section, three less frequent mutations were detected in the ATTO550 / Orange channel (p.L536H, p.L536P, and p.L536R).
[0111] All variant sequences were detected as described above, and for each variant, a characteristic dPCR signal was distinguished relative to the expected wild-type sequence dPCR signal.
[0112] References References and citations to other documents (e.g., patents, patent applications, patent publications, journals, books, articles, web content, publicly accessible databases) are made throughout this disclosure. All such documents are incorporated herein by reference in their entirety for all purposes.
[0113] equivalent Various modifications of the invention and many further embodiments thereof, in addition to those shown and described herein, will become apparent to those skilled in the art from the entire contents of this document, including references to the scientific and patent literature cited herein. The subject matter herein contains important information, exemplification and guidance that can be adapted to the practice of this invention in its various embodiments and equivalents thereof.
Claims
1. 1. A method for detecting a variant nucleic acid, the method comprising: compartmentalizing a sample containing target nucleic acids into a plurality of compartments; amplifying the nucleic acid in the compartment in the presence of a set of variant-specific probes, each of the variant-specific probes annealing to a respective variant, wherein each variant-specific probe comprises a detectable label, and wherein the set of probes has a number of distinct detectable labels that is less than the number of the respective variants; detecting a signal from said compartment; generating a plot of points representing said signal; and identifying the presence of each variant in the sample from the presence of a corresponding cluster of points in the plot; The method includes:
2. 2. The method of claim 1, wherein the generating step comprises mapping the detected signal onto a space defined by the number of distinct detectable labels.
3. 10. The method of claim 1, further comprising assigning a vector to each cluster of points, wherein each vector uniquely and specifically identifies one of the variants in the sample.
4. 2. The method of claim 1, wherein at least one of the probes detects a number of the variants in the sample, and at least a second of the probes is specific for fewer than the number of variants, is present at a different concentration than the first of the probes, and is used to distinguish between the variants by causing points in the plot to form separate clusters.
5. 2. The method of claim 1, wherein two or more of the different variant sequences are present at positions on the target nucleic acid that are amplified by one primer pair in the amplifying step.
6. The method of claim 1 , wherein the detectable label is an optical label.
7. 10. The method of claim 1, wherein the presence of one or more variant sequences is indicative of a pathological condition, optionally wherein said pathological condition is cancer.
8. 8. The method of claim 7, wherein the presence of the one or more variant sequences is indicative of minimal residual disease in the subject.
9. 10. The method of claim 1, further comprising, prior to said compartmentalizing step, obtaining an estimate of the relative abundance of said variants, and designing said variant-specific probes based on said estimate.
10. 9. The method of claim 8, wherein identifying the presence or absence of one or more of said variant sequences indicates progression or regression of said pathological condition.
11. 8. The method of claim 7, wherein the variant comprises a tumor mutation determined by sequencing tumor nucleic acid from a tumor sample.
12. 7. The method of claim 6, wherein the step of performing digital PCR further comprises using a probe specific for the wild-type sequence.
13. 2. The method of claim 1, further comprising assigning the variants to sections, and providing, for at least one section, at least one probe that detects multiple variants in the section and at least one probe that distinguishes between the multiple variants in the section.
14. 14. The method of claim 13, wherein the intercept is determined based on information about the genomic location or probable relative frequency of the variants in the sample.
15. 14. The method of claim 13, wherein each segment is defined as a set of variants that can be amplified by one primer pair.
16. 2. The method of claim 1, wherein the optical label is selected from FAM, HEX, SUN, VIC, TAMRA, ATTO550, Cy5, ROX, ATTO700, Cy5.5, Yakima Yellow, ABY, and JUN.
17. 10. The method of claim 1, wherein the sample comprises cell-free DNA (cfDNA).
18. prior to said compartmentalizing step, identifying a first pair of first and second variants among said variants that form overlapping clusters on a 2D dPCR plot; and detection probes for both variants of the pair, wherein the detection probes both have detectable optical labels of a first color; and At least one identification probe having a second color identification optical label. designing a probe set comprising: The method of claim 1 further comprising:
19. 19. The method of claim 18, wherein the amplifying step is performed with the detection probe and the discrimination probe present at different concentrations.
20. 19. The method of claim 18, wherein the detection probe further comprises a third probe specific for a third variant of the unpaired pair, the third probe having an optical label of the first color.
21. 21. The method of claim 20, wherein the plot includes a first cluster, a second cluster, and a third cluster from the first variant, the second variant, and the third variant, respectively, and vectors passing from the origin of the plot through the centroids of the clusters are non-orthogonal.
22. 21. The method of claim 20, wherein the first variant, the second variant, and the third variant are all located at positions on the target nucleic acid such that they are amplified by a single primer pair during the amplifying step.
23. The probe is a first probe specific for the first variant and having a first optical label; a second probe specific for the second variant and bearing said first optical label; a third probe specific for a third variant and bearing said first optical label; a WT probe specific for the wild-type sequence and having a second optical label; and a discrimination probe specific for said first variant and having a third optical label; The method of claim 1 , comprising:
24. 24. The method of claim 23, further comprising generating a first bicolor plot from the first optical label and the second optical label, and identifying the presence of at least the first variant and the second variant from the first bicolor plot.
25. 25. The method of claim 24, further comprising generating a second two-color plot from the first optical label and the third optical label, and distinguishing a first variant sequence from at least the second variant based on deviation of the second two-color plot from an expected dPCR two-color plot.