A method for detecting reaction volume deviations in digital polymerase chain reaction

By using convolution kernel functions and morphological image processing in digital PCR to identify effective partitions, the signal confusion caused by reaction volume deviation was resolved, thus improving the accuracy and precision of nucleic acid quantification.

CN115605611BActive Publication Date: 2026-01-06F HOFFMANN LA ROCHE & CO AG
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202180031752.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-30
Filing Date
2021-04-28
Publication Date
2026-01-06
Estimated Expiration
2041-04-28

AI Technical Summary

Technical Problem

Current digital PCR technology fails to effectively account for reaction volume deviations, making it difficult for signals to distinguish between empty and filled regions, thus affecting the accuracy and precision of nucleic acid quantification.

Method used

Fluorescence signals are processed by convolution kernel functions, combined with morphological image processing and clustering operations to identify valid and invalid partitions, correct reaction volume deviations, use kernel functions to perform convolution to combine light signals across (x, y) coordinates, set thresholds to identify valid partitions, and perform clustering and morphological processing to remove invalid partitions.

Benefits of technology

It improves the accuracy and precision of nucleic acid quantification in digital PCR, reduces false positive and false negative counts by identifying and correcting reaction volume deviations, and improves the reliability of nucleic acid concentration determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115605611B_ABST
    Figure CN115605611B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method for detecting reaction volume bias in digital polymerase chain reaction (dPCR) and to a method for determining the amount or concentration of a target nucleic acid in a sample using dPCR. The method uses a convolution with a kernel to analyze an optical image, wherein each partition is assigned a convolution value that is compared to a threshold convolution value.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit and priority of U.S. Application Serial No. 63 / 018183, filed April 30, 2020. The disclosure of the referenced application is incorporated herein by reference. Technical Field

[0003] This disclosure relates to a method for detecting reaction volume deviation in digital polymerase chain reaction (dPCR), and to a dPCR method for determining the amount or concentration of target nucleic acid in a sample, taking into account reaction volume deviation. Background Technology

[0004] For many biological, biochemical, diagnostic, or therapeutic purposes, it is necessary to accurately and precisely determine the amount or concentration of nucleic acids in a sample. Digital PCR (dPCR) provides an alternative to conventional real-time quantitative PCR for the absolute quantification of nucleic acids and the detection of rare alleles. dPCR works by partitioning the nucleic acid sample into many separate, parallel PCR reactions; some of these reactions contain the target molecule (positive), while others do not (negative). After PCR analysis, the fraction of negative reactions is used to generate an absolute count of the number of target molecules in the sample. One of the main advantages of dPCR over real-time PCR is its superior quantitative accuracy. This advantage relies on the inherent characteristics of dPCR, as quantification only requires correctly counting the positive partitions (or reaction volumes) and knowing the theoretical partition volume (the count is not very sensitive to PCR efficiency). No quantification standard is required. This eliminates potential quantification errors caused by the standard itself.

[0005] Existing technologies provide methods for identifying erroneous positive or negative counts in dropwise assays and for calibrating or normalizing signals (US 2013 / 0302792 A1). This normalization improves the separation between positive and negative counts. Therefore, normalization reduces the risk of false positive or false negative counts. The ultimate goal is to improve the accuracy and precision of nucleic acid concentration determination by correcting the signal obtained for nucleic acids.

[0006] However, existing methods do not account for errors in PCR caused by situations where the actual volume of the dPCR partition differs from the expected or desired volume.

[0007] Therefore, a method for quantifying target nucleic acids via dPCR that takes into account reaction volume deviation is needed. Summary of the Invention

[0008] This disclosure provides a method for detecting reaction volume deviation in dPCR assays, wherein the dPCR assay is used to quantify the amount or concentration of target nucleic acids in a partitioned array. The method includes the following steps:

[0009] (a) Combining optical signals across (x, y) coordinates within the array using convolutions performed with kernel functions, wherein each partition is assigned a convolution value; and

[0010] (b) Valid and invalid partitions are identified by comparing the convolution value of each partition with a threshold convolution value.

[0011] Furthermore, the method may also include (c) subjecting the data collected in step (b) to one or more additional steps, said one or more additional steps including: clustering operations and morphological image processing operations. Such morphological image processing operations may include dilation and / or erosion, while clustering may include effective trimming and / or ineffective trimming.

[0012] This disclosure also envisions a method for determining the amount or concentration of a target nucleic acid in a sample, the method comprising the steps of: (a) providing a sample suspected of containing the target nucleic acid; (b) performing dPCR on the sample in a dPCR plate comprising a partition array; (c) identifying one or more valid partitions in the partition array; and (d) calculating the amount or concentration of the target nucleic acid as the number of nucleic acids determined in step (b) per volume of a valid partition. Furthermore, the method may further include: determining the copy number N of the target nucleic acid in the one or more valid partitions identified in step (c). c And N c Divide by the effective partition volume.

[0013] In certain embodiments, this disclosure provides: a laboratory instrument adapted to perform the steps of the methods described herein; a computer program product comprising instructions that cause the laboratory instrument to perform the steps of the methods; and a computer-readable medium having the computer program product described herein stored thereon. Attached Figure Description

[0014] Figures 1A to 1B The method described herein for detecting reaction volume deviation in dPCR assays is illustrated schematically. Figure 1A A method for analyzing dPCR plates after nucleic acid amplification is shown, and Figure 1B This demonstrates a complete method from the preparation of dPCR plates to analysis using the methods described herein.

[0015] Figure 2 This is a schematic diagram of the laboratory equipment described in this article. Detailed Implementation

[0016] As detailed above, methods for reliably determining the amount or concentration of nucleic acids are particularly relevant in several industrial applications, such as in the medical field. Such applications may require precise and accurate determination of the amount or concentration of nucleic acids in a sample (e.g., a sample obtained from a patient or product). This may be of interest, for example, in the diagnosis of disease severity, in environmental technologies, or as a means of determining product quality, for example, to define contaminants or impurities.

[0017] Digital PCR (dPCR) is a biotechnological improvement on conventional polymerase chain reaction (PCR) methods, used for the direct quantification of nucleic acids (including DNA, cDNA, RNA, or mixtures thereof) and optionally, the clonal amplification of nucleic acids. The main difference between dPCR and traditional PCR (e.g., qPCR) lies in the method of measuring nucleic acid levels, with the former being more precise and accurate than PCR, but also more prone to errors by inexperienced users. The smaller dynamic range of dPCR may require sample dilution. dPCR also performs a single reaction within the sample; however, the sample is isolated into numerous partitions or zones, and the reaction is performed individually in each partition or zone. This isolation enables more reliable collection of nucleic acids and more sensitive measurement of nucleic acid levels. Furthermore, this method allows for accurate quantification.

[0018] Detailed descriptions of dPCR apparatus and methods can be found, for example, in U.S. Patent Application No. 20080160525; Vogelstein, et al., Proc. Natl. Acad. Sci. USA, Vol. 96, 9236-9241, August 1999; McCaughan, et al., J. Pathol. 2010; 220: 297-306; Mao, et al., Am. J. Transl. Res. 2019; 11(12): 7209-7222; U.S. Patent No. 10,564,102; U.S. Patent Publication No. 20180045641A1; European Patent No. 3299471B1; U.S. Patent Publication No. 20180147574A1; and U.S. Patent Publication No. 20180087090A1. The disclosure of each of these publications is incorporated herein by reference in its entirety.

[0019] In one specific embodiment, dPCR samples are isolated or partitioned in an array comprising multiple partitions (alternatively also referred to as reaction volumes or reaction wells), such that individual nucleic acid molecules within the sample are localized and concentrated within numerous isolated regions within the array. Partitioning the sample allows one to estimate the number of nucleic acids by assuming that the target molecule count within each partition follows a consistent Poisson distribution. After PCR amplification, each partition is identified as a negative or positive reaction ("0" molecules or "1 or more" molecules, respectively). The target molecules can be quantified by counting the number of positive and negative partitions and then using maximum likelihood estimation to estimate the underlying Poisson distribution. In conventional quantitative PCR, the quantification result may depend on the amplification efficiency of the PCR process. However, dPCR does not rely on the number of amplification cycles to determine the initial sample amount, thus eliminating the dependence on uncertainty index data to quantify target nucleic acids and therefore providing absolute quantification.

[0020] Because the sample is partitioned in the array, the partitioning process may result in partially filled or unfilled partitions, i.e., reaction volume deviations, and these partially filled or unfilled partitions are alternatively referred to as empty regions or filled empty regions. Fluorescence signals from empty regions may be difficult to distinguish from signals observed from otherwise filled partitions containing low- or no target molecules. Furthermore, bright spots may be observed in invalid partitions, which could be mistaken for positive signals from filled partitions.

[0021] These problems can be used Figure 1A The method of this disclosure is illustrated schematically. In short, empty regions in the signal data can be identified and distinguished by analyzing regions of a dPCR plate where the signal is consistently low across all channels. The analysis of these low-signal regions includes:

[0022] a) Feature calculation 101, which uses convolution with kernel function to combine optical signals across (x, y) coordinates in the array, thereby assigning convolution values ​​to each partition;

[0023] b) Identification of valid or invalid partitions 102, by comparing the convolution value of each partition with a threshold convolution value;

[0024] c) Optionally, additional cleaning step 103 may be performed, which includes subjecting the data collected in step (b) to one or more additional steps, including clustering operations and morphological image processing operations.

[0025] Each step of the method is described in more detail below and in Figure 1B The example dPCR plate is shown in the figure, with specific reference to it.

[0026] In short, in one specific embodiment, the dPCR plate may include one to eight reaction mixtures (samples) within a standard microplate format (SBS format, not shown in the figure). Each of the one to eight sample locations within the dPCR plate may include, for example, an inlet port at location A1 (a microstructured portion between location A1 and location A12) and an outlet port at location A12. Once a volume of fluid sample has been added to each inlet port, a partitioning fluid is added to each inlet. This can be done using a single-channel or 8-channel pipette or manually using an automated dispensing station. The separation or partitioning fluid is a hydrophobic liquid that is immiscible and non-reactive with respect to the reaction mixture (e.g., long-chain fluorinated hydrocarbons or silicone oil). Separation (or partitioning) of individual partitions containing the reaction mixture may be done passively or by applying overpressure at the inlet port or by applying negative pressure at the outlet port. Monitoring sensors may be used to ensure that the process stops when separation is complete (i.e., the separation fluid has reached the outlet port). Sample preparation steps are as follows: Figure 1B As shown in Figure 104.

[0027] After partitioning, the dPCR plate undergoes a thermal cycling process 105, followed by image analysis to detect fluorescence signals associated with each individual partition in the array. Signal data from the preliminary image analysis step 106 are collected. Table 1 summarizes the data inputs used in the method described herein:

[0028] Table 1

[0029] .

[0030] The first step in detecting reaction volume deviation is feature computation 107. The goal of feature computation is to combine signals in a way that makes the separation between invalid and valid partitions more apparent. The signals are first combined across channels and then convolved using a kernel based on (x, y) coordinates. The signals are convolved using the kernel to refine or "smooth" the image. In one embodiment, convolution can be performed on channels individually, or convolution can be performed on the sum of normalized signals across all channels.

[0031] In a specific embodiment, the signal is summed to highlight the partitions when they are dim across all channels, rather than just one. To ensure that all partitions contribute equally, the signal values ​​are normalized within each channel. In one embodiment, only those channels with a majority of positive signals are used in the calculation. Channels are identified by the useChannel tag (i.e., Standard remember 分区 = 1 (Valid) The signal is selected for use. The sum of the signals is calculated as follows:

[0032] signalChannel = Fluorescence (channel)分区

[0033]

[0034] Convolution is used to distinguish between dim regions and dim partitions. Convolution involves several steps: First, a distance function and a kernel function are established. Standard L2 distance and a custom exponential kernel can be used.

[0035]

[0036]

[0037] For ease of representation, z represents the (x, y) coordinates of a given partition, and the convolution of z is represented as:

[0038]

[0039] Please note that in it {isvalidPartition(z') ^ Dist(z, z') <= radius} In the case of an empty set, this value is not well defined. In this case, the result is set to an impossible default value, such as -1. Typically, the focus of this method is on convolutions targeting valid partitions, so the use of invalid default values ​​such as -1 is acceptable.

[0040] To initially distinguish between invalid and valid partitions, a threshold of 108 is set for the convolution values. Any partition with a convolution value less than the threshold is marked as invalid, and any partition with a convolution value exceeding the threshold remains valid. The method used to determine the threshold includes: analyzing candidate reference regions in the dPCR plate to select a region representing the filled portion of the plate, and then setting the threshold based on the convolution values ​​in that region.

[0041] There are many potential reference regions that can be plotted on a dPCR plate. Analyzing each potential reference region would be slow, so a subset is chosen, in this example, six reference regions. The ideal number of candidate reference regions can vary depending on the plate size and partition density. For illustrative purposes, six regions are used in a representative dPCR plate: ( maxx and maxy The largest partition x Coordinates and maximum y coordinate)

[0042] for

[0043] .

[0044] Please note that for i = 0 In this case, y The lower boundary was moved from 0 to 11 to avoid using the area located at the plate entrance. The last two reference areas are horizontal instead of vertical:

[0045] for

[0046]

[0047] For each reference region, the average of the effective convolutions is calculated for that region (the value is not equal to the default -1). The region with the second-highest average convolution is used as a reference. In this specific embodiment, the second-highest average convolution is chosen instead of the first, because some image artifacts may amplify the convolutions in some regions. The second-highest brightest region is less likely to have these problems, but it is likely to include fully filled partitions.

[0048] To set the threshold, the expected bias of the convolution for the filled partition is determined. First, the median absolute difference (MAD) is used:

[0049]

[0050] The default convolution value is excluded from this expression, resulting in the following standard threshold:

[0051]

[0052] Next, we assess the need for alternative thresholds. When MAD is very small, for example when λ is very high or very low, the standard threshold may not be suitable. While the possible remedies available when λ is very low are limited, simple solutions can be used to restore an appropriate threshold when λ is very high. In this case, if mean(ref) > highConvolutionThreshold and MAD(ref) < lowVarianceThreshold If so, then use an alternative threshold.

[0053]

[0054] With a threshold set, if the convolution value is less than the threshold, the partition is classified as invalid, and if the convolution value reaches or exceeds the threshold, the partition is considered valid, 109.

[0055] If the signal is noisy and / or otherwise effectively partitioned into dim or invalid partitions containing a large amount of signal, a final cleanup step may be desired. The cleanup step may include one or more of the following steps: invalid trimming 110, expansion 111, and effective trimming 112.

[0056] (i) Ineffective trimming

[0057] Empty regions can be trimmed by clustering using path connectivity. In short, the neighbors of a given partition include those partitions that share walls with that partition. For example, if the partitions in the device are square, rectangular, or hexagonal in shape, then the neighbors of a given partition are 4 or 6 partitions that share walls with that partition, respectively. In this example, a path connecting one partition to the next can be formed by starting with one partition and moving to one of its neighbors, then to one of the neighbors of the new partition, and so on. If there is a path between two invalid partitions that passes only through the invalid partition, then the two partitions are path-connected. Invalid partitions can be clustered using path connectivity by forming groups of all partitions that are path-connected to each other. voidNoise Clusters that are too small in size are reclassified as valid.

[0058] (ii) Expansion

[0059] After erroneous empty regions are removed, the remaining empty regions are dilated to ensure that any boundary empty regions are removed. Dilation can be performed using any suitable brush (e.g., a "diamond" brush with a radius of cleanupRadius), i.e., for any valid partition with coordinate z, dilation is performed if and only if an invalid partition exists. z' At that time, among them Only then was the partition reclassified as invalid.

[0060] (iii) Effective trimming

[0061] In the final step, valid partitions are trimmed by removing small groups. As described in this article, this step is performed for invalid trimming but for valid partitions. First, valid partitions are clustered by path connectivity, then those with smaller than [a certain number of groups] are trimmed. goodNoise Clustering based on size is reclassified as invalid.

[0062] Once empty zones are properly identified, the labels for these zones are changed to invalid.

[0063] The algorithm's output is summarized in Table 2:

[0064] Table 2.

[0065] .

[0066] Once invalid partitions are identified using the methods described herein, the dPCR system can quantify the amount or concentration of target nucleic acids in valid partitions of the array, regardless of signal data collected from invalid partitions. In a specific embodiment, the concentration of target nucleic acids in one or more valid partitions is calculated as the number of nucleic acid molecules per valid partition volume.

[0067] Therefore, this concentration can be determined by counting the target molecules (also known as the copy number).N c Calculated by dividing by the sampled liquid volume. Copy number N c The following is the output: First, the dPCR system identifies which valid partitions are positive and negative for the target molecules, and calculates the total for each.

[0068] For example, in a dPCR plate with 2000 valid negative partitions and 8000 valid positive partitions, the maximum likelihood estimate is used to calculate the latent parameter λ in the Poisson distribution. The estimated probability of a partition being negative can be calculated as follows:

[0069] P(negative) = number of negative cases / total number of cases.

[0070] For example, in the exemplary dPCR plate described above:

[0071] P(negative) = 2000 / (2000 + 8000) = 0.2.

[0072] If X is a Poisson random variable when modeling the number of molecules in a partition, then

[0073] P(negative) = P(X=0) = exp(-λ)

[0074] This is used to estimate λ, which is the only parameter in the Poisson distribution. Applying this principle to an example dPCR plate:

[0075] exp(-λ) = P(negative) = 0.2, that is, λ = -log(0.2) = 1.61 (rounded).

[0076] λ is an important value because it is also the expected value or mean of the Poisson distribution. In other words, λ is the average number of target molecules per partition based on the Poisson estimate. The λ estimate is used to determine N. c :

[0077] N c = λ * Number of partitions filled.

[0078] The number of partitions filled is equal to the number of non-invalid partitions: valid partitions + invalid partitions with non-empty areas. For an example dPCR plate, assuming there are 1000 additional partitions with invalid and non-empty areas, the number of partitions filled = 2000 + 8000 + 1000 = 11000, and the copy number is... N c = 1.61 * 11000 = 17710.

[0079] Concentration calculations include additional correction factors. If the sample has been treated, such as diluted, before use in dPCR, these treatment and dilution steps should be included in the calculations to obtain the amount or concentration of the target nucleic acid in the analyzed sample. Therefore, in an exemplary dPCR plate, assuming a partition volume of 1 mL and the sample is diluted to one-tenth of its original concentration before dPCR, the pre-amplification concentration would be:

[0080] N c / (Partition volume * Number of partitions filled) = 17710 / (11000 * 0.001L) =

[0081] 1610 copies per liter.

[0082] Therefore, before dilution, there are 10 * 1610 = 16100 copies per liter.

[0083] The methods described herein are used to determine the amount or concentration of nucleic acids in a sample using dPCR analysis. In this context, a sample is a quantity of material suspected of containing one or more nucleic acids to be detected, measured, and quantified. As used herein, the term includes, but is not limited to, a sample (e.g., a biopsy or medical sample), cell or tissue culture, blood, serum, plasma, needle aspirate, urine, sperm, seminal fluid, seminal plasma, prostatic fluid, excrement, tears, saliva, sweat, biopsy, ascites, cerebrospinal fluid, pleural fluid, amniotic fluid, peritoneal fluid, interstitial fluid, sputum, breast milk, lymph, bronchial lavage samples, or tissue extracts. The sample source can be solid tissue such as fresh, frozen, and / or preserved organ or tissue samples or biopsy or aspirate; or cells from any stage of pregnancy or development in the subject. Samples may contain compounds that are not naturally mixed with the sample source, such as preservatives, anticoagulants, buffers, fixatives, nutrients, antibiotics, etc.

[0084] As detailed above, the sample contains a target nucleic acid whose amount or concentration is to be determined in the methods of this disclosure. Nucleic acids are biopolymers essential for all known life forms. Therefore, nucleic acids can be used as indicators for specific organisms, but also as indicators for diseases, for example, in the case of mutations or naturally occurring variants. The target nucleic acid can be selected from the group consisting of DNA, cDNA, RNA, and mixtures thereof, or any other type of nucleic acid. Nucleic acids may contain non-nucleic acid components. It can be naturally occurring, chemically synthesized, or bioengineered. Specifically, the nucleic acid is selected from the group consisting of DNA, cDNA, RNA, and mixtures thereof.

[0085] Nucleic acids can indicate microorganisms (such as pathogens) and can be used to diagnose diseases, such as infections. Infections can be caused by bacteria, viruses, fungi, and parasites, or other objects containing nucleic acids. Pathogens can be exogenous (from environmental or animal sources or acquired from other people) or endogenous (acquired from normal flora). Samples can be selected based on signs and symptoms, should represent the disease process, and should be collected before the application of antimicrobial agents. The amount of nucleic acids in untreated samples can indicate the severity of the disease.

[0086] Alternatively, nucleic acids can indicate genetic disorders. Genetic disorders are genetic problems caused by one or more abnormalities in the genome, especially those present at birth (congenital). Most genetic disorders are very rare, affecting only one in thousands or millions of people. Genetic disorders may or may not be heritable, i.e., inherited from parents' genes. In non-heritable genetic disorders, the defect may be caused by a new mutation or alteration in the DNA. In this case, the defect is only heritable if it occurs in the germline. The same disease, such as some forms of cancer, may be caused by inherited genetic conditions in some populations, by new mutations in others, and primarily by environmental factors in still others. Clearly, the amount of mutated nucleic acids can indicate a disease state.

[0087] In a specific embodiment, the sample is a biofluid derived from a pregnant mammal containing both maternally and fetal nucleic acids (e.g., RNA or DNA). In this embodiment, nucleic acids from the maternal sample can be used to detect chromosomal doses due to fetal aneuploidy. In addition to empirically determining the frequency of nucleic acids from a specific chromosome, the proportion of fetal nucleic acids in the maternal sample can also be used to determine the risk of fetal aneuploidy based on chromosomal dose, as it affects the level of variation that is statistically significant in risk calculations. Utilizing such information when calculating the risk of aneuploidy in one or more fetal chromosomes allows for more accurate results that reflect biological differences between samples. The proportion of fetal DNA in the maternal sample is used as part of the risk calculation because the fetal proportion provides important information about the expected statistical presence of chromosomal doses. Differences from the expected statistical presence can indicate fetal aneuploidy, particularly fetal trisomy or monosomy of a specific chromosome.

[0088] In the method disclosed herein, the amount or concentration of nucleic acid is determined. The amount of substance is a standard-defined quantity. The International System of Units (SI) defines the amount of substance as being proportional to the number of basic units present, with the reciprocal of Avogadro's constant as the proportionality constant (in moles). The SI unit for the amount of substance is the mole. A mole is defined as the amount of substance containing the same number of basic units as the atoms present in 12 grams of the isotope carbon-12. Therefore, the amount of substance of a sample is calculated as the sample mass divided by the molar mass of the substance. In this context, "amount" generally refers to the copy number of the target nucleic acid sequence.

[0089] In dPCR, partitions can be miniaturized chambers of microarrays or nanoarrays, chambers of microfluidic devices, or micropores or nanopores on a chip, in a capillary, on a nucleic acid binding surface, or on beads (especially in microarrays or on a chip). The methods described herein are particularly suitable for use with array-based systems, including but not limited to commercial digital PCR platforms such as BioMark® dPCR, a microwell-based chip from Fluidigm, and QuantStudio 12k flex dPCR and 3D dPCR, both through-well based from Life Technologies. Microfluidic chip-based dPCR can have up to hundreds of partitions per panel. QuantStudio 12k dPCR performs digital PCR analysis on an OpenArray® board containing 64 partitions per subarray and a total of 48 subarrays, equivalent to a total of 3072 partitions per array.

[0090] Typically, the accuracy, and more importantly, precision, of determinations made by dPCR can be improved by using a larger number of partitions. Approximately 100 to 200, 200 to 300, 300 to 400, 700, or more partitions can be used to determine the amount or concentration in question by PCR. In a particular embodiment, dPCR is performed identically in at least 100 partitions, particularly at least 1,000 partitions, especially at least 5,000 partitions. In a specific embodiment, dPCR is performed identically in at least 10,000 partitions, particularly at least 50,000 partitions, especially at least 100,000 partitions.

[0091] For example, dPCR is performed equally in an array having between 100 and 100,000 partitions (e.g., between 1,000 and 100,000 reaction sites, or between 10,000 and 100,000 reaction sites).

[0092] The methods described herein are performed in a laboratory instrument or system configured to perform digital nucleic acid amplification reactions. As used herein, the term "nucleic acid amplification reaction" refers to a method or reaction in molecular biology for amplifying a single copy or several copies of a target DNA fragment (analyte) to a detectable amount of DNA fragment copy, comprising repeated cycles of a temperature-dependent reaction with a polymerase. Each cycle may include at least a denaturation phase (e.g., 95°C for 30 seconds), an annealing phase (e.g., 65°C for 30 seconds), and an extension phase (e.g., 72°C for 2 minutes). The dPCR plate may be in thermal contact with a thermoelectric element to heat and / or cool the sample holder to predetermined temperatures for the different phases. Typically, the nucleic acid amplification reaction comprises 20 to 40 repeated cycles, and after the nucleic acid amplification reaction is completed, the intensity of the light signal emitted from the reaction volume in the dPCR plate is measured by a detector. Based on the measured light signal intensity, the presence of nucleic acids in the sample can be determined.

[0093] Laboratory instruments used for performing nucleic acid amplification reactions are well known in the art and may include one or more of the following components (representative laboratory instrument 200 in...). Figure 2 (shown schematically in the image)

[0094] i. Sample preparation module 201, which may be a component of a laboratory instrument or a separate system, includes a pipetting device for pipetting samples and / or reagents into dPCR plate 202 and partitioning samples into one or more reaction volumes in array 203 in the dPCR plate;

[0095] ii. dPCR plate support / processing module 204, which transports dPCR plates within the laboratory instrument from one module to another;

[0096] iii. Thermal cycling module 205, which includes thermoelectric elements for heating and / or cooling the dPCR plate during the amplification reaction;

[0097] iv. Detection module 206, comprising a light source configured to emit light toward a dPCR plate (or a sub-segment thereof), and a photodetector configured to measure the signal light intensity of the light emitted from the dPCR plate (or a sub-segment thereof); and

[0098] v. Control device 207, for example, any physical or virtual processing device including processor 208, configured to control laboratory instruments and their components in such a way that sample analysis steps are performed through the laboratory instruments.

[0099] The sample preparation module may be housed within the laboratory instrument housing, or it may be a separate, independent device not housed within the laboratory instrument housing. In embodiments where the sample preparation module is a separate device, a dPCR plate is prepared in the sample preparation module, and then the plate is transported (automatically or manually) to the dPCR plate support / processing module 204 within the laboratory instrument.

[0100] Optionally, the control device may receive information from the data management unit regarding which steps need to be performed on a particular sample. For example, the processor of the control device may be embodied as a programmable logic controller adapted to execute a computer-readable program containing instructions for performing operations of laboratory instruments. As described herein, one operation is a method for detecting reaction volume deviations in a dPCR system.

[0101] One or more components of the aforementioned laboratory instrument are shown in, for example, U.S. Patent Application No. 20080160525; U.S. Patent No. 10,564,102; U.S. Patent Publication No. 20180045641A1; European Patent No. 3299471B1; U.S. Patent Publication No. 20180147574A1; and U.S. Patent Publication No. 20180087090A1. The disclosure of each of these publications is incorporated herein by reference in its entirety.

[0102] Furthermore, this disclosure envisions a computer program product including instructions that cause the laboratory instruments described herein to perform the steps of the method for detecting reaction volume deviation in a dPCR plate as described herein. Additionally, this disclosure provides a computer-readable medium having a computer program product stored thereon containing instructions that cause the laboratory instruments described herein to perform the steps of the method for detecting reaction volume deviation as described herein.

[0103] The embodiments of the subject matter and operation described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and similar structures, or one or more combinations thereof. Embodiments of the subject matter described herein can be implemented as one or more computer programs, i.e., as one or more computer program instruction modules for execution by a data processing device or for controlling the operation of a data processing device, said one or more computer program instruction modules may be encoded on a computer storage medium. Modules may include logic executed by a processor. As used herein, "logic" refers to information having any form of instruction signals and / or data that can affect the operation of a processor. Software is an example of logic.

[0104] Computer storage media may be or may be contained in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or one or more combinations thereof. Furthermore, although a computer storage medium is not a propagating signal, it may be a source or destination of computer program instructions encoded in an artificially generated propagating signal. The computer storage medium may also be or may be contained in one or more separate physical components or media (such as multiple CDs, disks, or other storage devices). The operations described herein can be implemented by a data processing device on data stored on one or more computer-readable storage devices or received from other sources.

[0105] The term "programmable processor" encompasses a wide range of devices, apparatuses, and machines that process data, including, for example, programmable microprocessors, computers, systems-on-a-chip, or a combination thereof. These devices may include special-purpose logic circuitry such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the devices may also include code that creates an execution environment for the relevant computer program, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or one or more combinations thereof. These devices and execution environments can implement various different computing model infrastructures, such as network services, distributed computing, and grid computing infrastructures.

[0106] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative languages, or procedural languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, objects, or other units suitable for use in a computing environment. Computer programs may, but do not necessarily, correspond to files in a file system. Programs can be stored in portions of files that hold other programs or data (such as one or more scripts stored in a markup language file), in a single file dedicated to the program, or in multiple coordinating files (such as files storing one or more modules, subroutines, or code portions). Computer programs can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected via a communication network.

[0107] The processes and logic flows described in this specification can be executed by one or more programmable processors that execute one or more computer programs to perform operations by processing input data and producing output results. The processes and logic flows can also be executed by special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the device can also be implemented as a special-purpose logic circuit.

[0108] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in a digital computer. Generally, a processor receives instructions and data from read-only memory or random access memory, or both. The basic elements of a computer are a processor that executes operations according to instructions and one or more storage devices that store instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, to receive data from, transfer data to, or receive data from and transfer data to them. However, a computer does not require such devices. Suitable devices for storing computer program instructions and data include various forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices like EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0109] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes backend components, such as a data server, or middleware components, such as an application server, or frontend components, such as a client computer with a graphical user interface or a web browser through which a user can interact with embodiments of the subject matter described in this specification, or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected via any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), interconnected networks (such as the Internet) and peer-to-peer networks (such as dedicated peer-to-peer networks).

[0110] The computing system may include any number of clients and servers. Typically, clients and servers are remotely configured and generally interact via a communication network. Relationships between clients and servers arise from computer programs running on their respective computers and the client-server relationships between them. In some embodiments, the server transmits data (such as HTML pages) to a client device (e.g., for displaying data to a user interacting with the client device and receiving the user's input). Data generated on the client device (such as the result of user interaction) can be received from the client device on the server.

[0111] Unless otherwise defined, all technical and scientific terms and any abbreviations used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Definitions of commonly used terms in molecular biology can be found in Benjamin Lewin, Genes V, Oxford University Press, 1994 (ISBN 0-19-854287-9); Kendrew et al. (eds.), The Encyclopedia of Molecular Biology, Blackwell Science Ltd., 1994 (ISBN 0-632-02182-9); and Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, VCH Publishers, Inc., 1995 (ISBN 1-56081-569-8).

[0112] This disclosure is not limited to the specific methodologies, schemes, and reagents described herein, as they can vary. While any methods and materials similar to or equivalent to those described herein may be used in practice, specific methods and materials are described herein. Furthermore, the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure.

[0113] As used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly indicates otherwise. Similarly, the words “comprising,” “including,” and “covering” should be interpreted as inclusive rather than exclusive. Likewise, unless the context clearly indicates otherwise, the word “or” is intended to include “and.” The term “multiple” means two or more.

[0114] The foregoing description is intended to illustrate various embodiments of this disclosure. Therefore, the specific modifications discussed should not be construed as limiting the scope of this disclosure. It will be apparent to those skilled in the art that various equivalents, changes, and modifications can be made without departing from the scope of this disclosure, and it should therefore be understood that such equivalent embodiments should be included herein.

[0115] This article cites various publications, and their public content is incorporated by citing them in their entirety.

Claims

1. A method for detecting reaction volume bias in a digital polymerase chain reaction (dPCR) assay, wherein the dPCR assay comprises quantifying an amount or concentration of a target nucleic acid in an array of partitions, the method comprising: (a) combining light signals across (x,y) coordinate combinations within the array using a convolution with a kernel function, wherein each partition is assigned a convolution value, and a kernel function of a distance function is applied, which comprises: (b) identifying valid partitions and invalid partitions by comparing the convolution value of each partition to a threshold convolution value; and optionally, (c) subjecting the data collected in step (b) to one or more additional steps comprising: a clustering operation and a morphological image processing operation.

2. The method of claim 1, further comprising: subjecting the data collected in step (b) to a morphological image processing operation comprises dilation, erosion, and combinations thereof.

3. The method of any of the preceding claims, further comprising: subjecting the data collected in step (b) to a clustering comprising valid trimming and / or invalid trimming.

4. The method of any of the preceding claims, wherein invalid partitions have a convolution value below the threshold convolution value, and valid partitions have a convolution value above the threshold convolution value.

5. The method of any of the preceding claims, wherein the array comprises a plurality of channels, and step (a) further comprises: determining which channel or channels of the plurality of channels to use in the method by a useChannel flag, 6. The method of claim 1, wherein z represents a set of (x,y) coordinates of a first partition, and the convolution of z is:

7. The method of claim 6, wherein if {isValidPartition(z') A Dist(z,z') <= radius} is an empty set, then the output for the empty set is set to a default value outside the range of the convolution.

8. The method of any of the preceding claims, wherein the convolution threshold is based on a set of convolution values within a selected reference region of the array.

9. The method of claim 8, wherein the selected reference region is selected from a vertical reference region i, a horizontal reference region j, and combinations thereof, wherein max x and max y are the maximum x coordinate and maximum y coordinate of partitions in the vertical reference region and / or the horizontal reference region, (a) the vertical reference region i is represented as (b) the horizontal reference region j is represented as and For each vertical reference region and / or each horizontal reference region, the method further comprises: calculating the mean of valid convolution values of the vertical reference region and / or the horizontal reference region, and identifying the vertical reference region and / or the horizontal reference region with the second highest mean convolution value as the selected reference region.

10. The method of claim 9, further comprising: The absolute median deviation (MAD) is calculated, expressed as and excluding the default convolution value to produce a standard threshold: voidThresh = mean(ref) - diffOff x MAD(ref).

11. The method of claim 10, wherein the method further comprises: if mean(ref) > highConvolutionThreshold and MAD(ref) < lowVarianceThreshold, then using an alternative threshold, wherein the alternative threshold is: voidThresh = mean(ref) * thresholdAdjustmentFrac where thresholdAdjustmentFrac is a fraction of the mean value used as the threshold for the substitution.

12. The method of any one of the preceding claims, wherein clustering includes path connectivity, the path connectivity including: (a) grouping the partitions in the array that are all pairwise connected to each other by contiguous paths, wherein the grouping is clusters, (b) identifying one or more clusters having a size less than an invalid noise threshold, and (c) designating the clusters identified in step (b) as invalid.

13. The method of any of the preceding claims, further comprising inflating to remove boundary voids.

14. The method of claim 13, wherein for an active partition having a coordinate z, if there exists an inactive partition z' having |z x -z' x |+|z y -z' y |≤ cleanupRadius, then the active partition is designated as an inactive partition.

15. The method of any one of the preceding claims, wherein clustering includes path connectivity, the path connectivity including: (a) grouping the partitions in the array that are all pairwise connected to each other by contiguous paths, wherein the grouping is clusters, (b) identifying one or more clusters having a size less than an invalid noise threshold, and (c) designating the clusters identified in step (b) as invalid.

16. The method of any of the preceding claims, further comprising flagging partitions identified as invalid.

17. A method for determining the amount or concentration of a target nucleic acid in a sample, the method comprising the steps of: (a) providing a sample suspected of containing the target nucleic acid; (b) performing dPCR with the sample in a dPCR plate comprising an array of partitions; (c) identifying one or more valid partitions in the array of partitions; (d) calculating the amount or concentration of the target nucleic acid as the number of nucleic acids determined in step (b) per valid partition volume; and (e) subjecting the data collected in step (c) to one or more additional steps comprising: a clustering operation and a morphological image processing operation, wherein clustering comprises path connectivity comprising: (a) grouping the partitions in the array that are all pairwise connected to each other by contiguous paths, wherein the grouping is clusters, (b) identifying one or more clusters having a size less than an invalid noise threshold, and (c) designating the clusters identified in step (b) as valid.

18. The method of claim 17, wherein the method further comprises: determining the copy number N of the target nucleic acid in the one or more effective partitions identified in step (c) c and dividing N c by the effective partition volume.

19. A laboratory instrument adapted to perform the steps of the method of any of the preceding claims.

20. A computer program product comprising instructions causing a laboratory instrument to perform the steps of the method of any of claims 1-18.

Citation Information

Patent Citations

  • Optics for analysis of microwells

    US10564102B2

  • Method and device for detecting the presence of a single target nucleic acid in a sample

    US20080160525A1

  • Calibrations and controls for droplet-based assays

    US20130302792A1

  • Plate with wells for chemical or biological reactions, and method for multiple imaging of such a plate by means of an imaging system

    US20180045641A1

  • Method for reducing quantification errors caused by reaction volume deviations in digital polymerase chain reaction

    US20180087090A1