A method for absolute quantification of nucleic acids based on bayesian inference
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU PROGRAMMABLE DIGITAL CORE BIOTECHNOLOGY CO LTD
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]然而,ddPCR所需的独立微反应单元数量极为庞大,使得整个检测系统必须依赖精密的、高稳定性的液滴生成装置以及高通量的荧光信号读取设备,导致仪器专业化程度高、购置与运行成本昂贵
传统的核酸绝对定量技术通常采用固定体积进行微滴生成,这导致检测范围受限(高浓度易饱和,低浓度信号弱)。本发明突破在于引入了贝叶斯特定优化算法作为“大脑”,通过实时分析微滴的阳性/阴性分布,动态预测样本的真实浓度范围,并据此智能调整后续微滴的生成体积。这种“边检测边调整”策略,使得系统能够自适应地覆盖从极低拷贝到极高拷贝的宽动态范围,无需进行繁琐的样本预稀释等操作。
Smart Images

Figure CN122521831A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics technology, specifically to a method for absolute quantification of nucleic acids based on Bayesian inference and its application. Technical Background Nucleic acid quantification is a fundamental analytical technique in molecular biology, clinical diagnostics, genomics, and the biotechnology industry. Its core objective is to accurately determine the concentration and copy number of deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) in a sample. Reliable quantitative results are a prerequisite for ensuring the success, reproducibility, and biological comparability of downstream experiments (such as polymerase chain reaction, high-throughput sequencing, gene cloning, and transgenic detection). Existing nucleic acid quantification techniques are mainly based on different physicochemical principles, and their development reflects the relentless pursuit of higher precision, sensitivity, and throughput.
[0002] Traditional nucleic acid quantification methods mainly include ultraviolet (UV) absorption spectroscopy and fluorescent dye methods. UV absorption spectroscopy is based on the characteristic light absorption of nucleic acid bases at a wavelength of 260 nm, calculating concentration according to the Beer-Lambert law. It has the advantages of being simple and rapid, but its specificity is limited, it is easily interfered with by common co-extracted contaminants (such as proteins, phenols, and salt ions), and it cannot distinguish between DNA and RNA. Fluorescent dye methods utilize dyes that can specifically intercalate with nucleic acids and emit enhanced fluorescence signals (such as PicoGreen and SYBR Green). The signal intensity is proportional to the nucleic acid concentration, significantly improving detection sensitivity and specificity, reducing contaminant interference, and making it suitable for trace nucleic acid analysis. Real-time quantitative PCR (qPCR) quantifies by monitoring the number of cycles (Ct value) in which the fluorescence signal reaches a set threshold during amplification, combined with a standard curve, offering higher sensitivity. However, all of the above methods fall into the category of relative quantification based on standard curves. Their accuracy heavily depends on the purity, accuracy, and linearity range of the standard, and they cannot directly obtain the absolute copy number of the target nucleic acid in the sample.
[0003] To achieve accurate quantification without relying on external standard curves, absolute nucleic acid quantification methods have emerged. Droplet Digital PCR (ddPCR), as a third-generation PCR technology, has become a key tool in molecular diagnostics and life science research due to its absolute quantification capability, ultra-high sensitivity, and excellent anti-interference performance. Its core principle is to physically divide the PCR reaction system into tens of thousands to hundreds of thousands of nano-level droplets, each droplet acting as an independent reaction unit. After amplification, flow cytometry is used to analyze the fluorescence signal of each droplet individually, and the proportion of positive droplets is statistically analyzed based on the Poisson distribution principle, thereby directly calculating the absolute copy number concentration of the target nucleic acid without relying on a standard curve. Compared to traditional quantitative PCR, this technology's significant advantage lies in its ultra-high sensitivity, effectively detecting rare mutations and subtle expression differences. Its dynamic range covers five orders of magnitude, and the detection accuracy deviation can be controlled within ±10%. Because the reaction system is segmented, ddPCR has high tolerance to PCR inhibitors, making it particularly suitable for detecting complex background samples (such as blood and soil). In clinical diagnostics, ddPCR is a core technology in liquid biopsy, enabling early cancer screening and monitoring of treatment efficacy by detecting trace amounts of cell-free tumor DNA in the blood, with sensitivity far exceeding traditional methods. In pathogen detection, it excels at detecting low viral loads (such as coronaviruses and HBV), effectively reducing the risk of false negatives. In scientific research, ddPCR is widely used in copy number variation (CNV) analysis, gene expression studies, quantitative detection of transgenic components, and CRISPR editing validation.
[0004] However, ddPCR requires an extremely large number of independent microreaction units, necessitating a sophisticated and highly stable droplet generation device and high-throughput fluorescence signal reading equipment for the entire detection system. This results in a high degree of instrument specialization and high purchase and operating costs. This strong dependence on specialized heavy equipment significantly limits the widespread adoption and application of this technology in resource-constrained scenarios (such as primary care laboratories, point-of-care diagnostics, and rapid on-site testing). Theoretically, if the threshold for the number of droplets or partitions required to achieve reliable absolute quantification can be significantly reduced, it may be possible to achieve equally reliable quantitative analysis using more general-purpose equipment (such as conventional quantitative PCR instruments and microfluidic chips) while simplifying or even eliminating some complex hardware. This would not only significantly lower the technical threshold and cost but also open up broad new application scenarios for nucleic acid absolute quantification technology, including but not limited to high-throughput screening of multiple targets, microsample processing in single-cell omics analysis, and distributed, real-time on-site molecular diagnostics.
[0005] Bayesian inference methods offer a promising computational framework for optimizing absolute nucleic acid quantification. Their core value lies in their ability to more efficiently utilize limited data, theoretically allowing for a significant reduction in the number of droplets required for reliable quantification. Traditional ddPCR relies on a Poisson statistical model, where the confidence interval width is inversely proportional to the total number of detected droplets. Typically, tens of thousands of effective droplets are needed to obtain a sufficiently narrow confidence interval to meet high-precision quantification requirements. In contrast, Bayesian methods combine prior knowledge with the limited observed positive / negative droplet data, directly providing a probability estimate of the target nucleic acid copy number and its confidence interval through a posterior probability distribution. This information fusion mechanism allows the Bayesian framework to potentially require fewer effective observations (i.e., fewer droplets) for the same quantification accuracy. This provides a new computational path for achieving high-reliability absolute quantification based on simplified droplet manipulation systems (e.g., reducing the total number of partitions) or utilizing existing general-purpose devices (e.g., microfluidic chips), potentially driving the application of this technology in portable, low-cost scenarios.
[0006] In view of this, the present invention is proposed. Summary of the Invention
[0007] To address the aforementioned issues, this patent proposes a Bayesian inference-based absolute nucleic acid quantification method. This method initially generates a small number of microdroplets and acquires their fluorescence signals to establish an initial probability distribution model. Then, using a Bayesian update mechanism, it dynamically optimizes the generation volume of subsequent microdroplets based on the results of previous detections, gradually approximating the true concentration. This approach can maintain high quantitative accuracy and confidence levels while significantly reducing the total number of microdroplets generated (by at least two orders of magnitude).
[0008] To achieve the above objectives, the present invention proposes the following specific technical solutions.
[0009] This invention first provides an absolute quantification method for nucleic acids, which is based on Bayesian inference and achieves absolute quantification of nucleic acids by adjusting the detection process simultaneously.
[0010] Furthermore, the absolute quantification method specifically includes the following steps: 1) Let X be the total copy number of nucleic acid molecules in the original nucleic acid sample. Mix the nucleic acid sample thoroughly, take a single droplet and perform nucleic acid detection to obtain the positive or negative result of the droplet. 2) If the detected droplet is positive, calculate the estimated total copy number n under the current conditions based on the following formula: If the detected droplet is negative, the estimated total copy number n under the current conditions is calculated based on the following formula: Where m is the total number of droplets detected, Ni is the ratio of the original sample volume to the volume of the i-th droplet, α is the confidence level parameter, and e is the natural constant; Calculate the droplet volume to be detected in the next round based on the estimated total copy number n under the current conditions; 3) Based on the droplet volume determined in step 2), a single droplet is extracted from the original nucleic acid sample to continue nucleic acid detection. Step 2) is repeated for iterative detection. The iteration stops when both positive and negative results are observed in the cumulative number of detected droplets. The estimated value of the total copy number n under the current conditions is then calculated based on the following formula: Where, if the detection result of the i-th droplet is positive, then bi=0; if the detection result of the i-th droplet is negative, then bi=1. Calculate the droplet volume to be detected in the next round based on the estimated total copy number n under the current conditions; 4) Based on the droplet volume determined in step 3), a single droplet is extracted from the original nucleic acid sample to continue nucleic acid detection. Step 3) is repeated for iterative detection. When the total number of droplets detected reaches the preset termination standard, the iteration stops. At this time, n, which is the solution obtained in step 3), is the final estimate of the total copy number X of the original sample.
[0011] Furthermore, in step 1), the volume of the microdroplet is 0.04-0.12 μL.
[0012] Furthermore, in step 2), the volume of droplets to be detected in the next round is 1.5-1.7v / n; where v is the total detection volume of all droplets in the current step (or under the current conditions), and n is the estimated value of the total copy number in the current step (or under the current conditions).
[0013] Furthermore, in step 3), the volume of droplets to be detected in the next round is 1.5-1.7v / n; where v is the total detection volume of all droplets at present, and n is the estimated value of the total copy number under the current conditions.
[0014] Furthermore, in step 4), the standard is a total number of detected droplets ≥ 350 droplets; preferably ≥ 650 droplets. Specifically, for ordinary precision detection standards (estimated error within ±15%), the total number of detected droplets needs to be ≥ 350 droplets; for high precision detection standards (estimated error within ±10%), the total number of detected droplets needs to be ≥ 650 droplets.
[0015] In some respects, the nucleic acid concentration in the original nucleic acid sample in step 1) is 25 copies / μL-200,000 copies / μL (corresponding to 500 copies / 20μL-4,000,000 copies / 20μL), preferably 2,000 copies / μL-200,000 copies / μL (corresponding to 40,000 copies / 20μL-4,000,000 copies / 20μL).
[0016] In some respects, the nucleic acid detection in steps 1) and 3) is PCR-based nucleic acid detection.
[0017] This invention also provides the application of the absolute quantification method described above in the absolute quantification of nucleic acids; In some preferred aspects, the absolute quantification of nucleic acid is absolute quantification during the nucleic acid PCR detection process.
[0018] The present invention also provides an electronic device, comprising: a processor and a memory; the processor and the memory are connected together, wherein the memory is used to store a computer program, and the processor is used to invoke the computer program to execute the method described in any of the preceding claims.
[0019] The present invention also provides a computer storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, perform the method described in any of the preceding claims.
[0020] Beneficial technical effects of the present invention: Traditional absolute nucleic acid quantification techniques typically employ fixed-volume droplet generation, which limits the detection range (high concentrations easily saturate, low concentrations produce weak signals). This invention breaks through this limitation by introducing a Bayesian-specific optimization algorithm as the "brain." By analyzing the positive / negative distribution of droplets in real time, it dynamically predicts the true concentration range of the sample and intelligently adjusts the generation volume of subsequent droplets accordingly. This "detect-and-adjust" strategy enables the system to adaptively cover a wide dynamic range from extremely low to extremely high copy numbers, eliminating the need for cumbersome sample pre-dilution procedures.
[0021] When a low-probability event or judgment error occurs during the detection process, causing the concentration estimation curve to deviate, the model can gradually converge the estimated value to the true concentration through iterative updates under the Bayesian inference framework. This demonstrates strong anti-interference ability and stability, ensuring that reliable detection performance can still be maintained under non-ideal conditions. Attached Figure Description
[0022] Figure 1 A schematic diagram of a Bayesian iterative feedback model; Figure 2 The iterative convergence of the copy number estimate (the initial droplet is positive); Figure 3 The iterative convergence of the copy number estimate (initial droplet was negative); Figure 4 Iterative convergence under normal conditions (the initial droplet exhibits a high probability of success); Figure 5 Iterative convergence under unconventional conditions (the initial droplet exhibits a low-probability outcome); Figure 6 The trend of the 95% confidence interval for estimating the copy number as the number of iterations increases; Figure 7 The trend of the 95% confidence interval of the estimated copy number as the number of iterations increases (total copy number 4,000,000). Detailed Implementation
[0023] The embodiments of the present invention will be described in detail below with reference to examples. However, those skilled in the art will understand that the following examples are for illustrative purposes only and should not be considered as limiting the scope of the invention. Unless otherwise specified in the examples, conventional conditions or conditions recommended by the manufacturer are followed. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased on the market.
[0024] Definitions of some terms Unless otherwise defined below, all technical and scientific terms used in the specific embodiments of this invention are intended to have the same meaning as commonly understood by those skilled in the art. While it is believed that the following terms will be well understood by those skilled in the art, the following definitions are set forth to better explain the invention.
[0025] As used in this invention, the indefinite or definite articles used when referring to singular nouns, such as “a” or “a kind”, “the”, include the plural form of the noun.
[0026] As used in this invention, the terms "comprising," "including," "having," "containing," or "involving" are inclusive or open-ended and do not exclude other unlisted elements or method steps. The term "consisting of" is considered a preferred embodiment of the term "comprising." If a group is defined below as including at least a certain number of embodiments, this should also be understood to disclose a group that preferably consists only of those embodiments.
[0027] The term "approximately" in this invention refers to an accuracy range that, as would be understood by those skilled in the art, still guarantees the technical effects of the features in question. This term typically indicates a deviation from the indicated value of ±10%, preferably ±5%.
[0028] Furthermore, the terms first, second, third, (a), (b), (c), and similar terms used in the specification and claims are for distinguishing similar elements and are not necessary for the order of description or chronological sequence. It should be understood that such terms are interchangeable in appropriate contexts, and the embodiments described in this invention can be implemented in a different order than that described or illustrated in this invention.
[0029] The terms and definitions above are provided merely to aid in understanding the invention. These definitions should not be construed as having a scope less than that understood by those skilled in the art.
[0030] The following is a detailed explanation of the contents of this invention.
[0031] The present invention provides an absolute quantification method for nucleic acid absolute quantification based on Bayesian inference.
[0032] Generally, the method of the present invention includes the following four steps: 1) Let X be the total copy number of nucleic acid molecules in the original sample. Mix the nucleic acid sample thoroughly, take a single droplet and perform nucleic acid detection to obtain the positive or negative result of the droplet. 2) If the detected droplet is positive, calculate the estimated total copy number n under the current conditions based on the following formula: If the detected droplet is negative, the estimated total copy number n under the current conditions is calculated based on the following formula: Where m is the total number of droplets detected, Ni is the ratio of the original sample volume to the volume of the i-th droplet, α is the confidence level parameter, and e is the natural constant; Calculate the droplet volume to be detected in the next round based on the estimated total copy number n under the current conditions; 3) Based on the droplet volume determined in step 2), supplement the original nucleic acid sample with single droplets and continue nucleic acid testing. Repeat step 2) for iterative testing. Stop the iteration when both positive and negative results are found in the cumulative number of tested droplets. Then, calculate the estimated value n of the total copy number under the current conditions based on the following formula: Where, if the detection result of the i-th droplet is positive, then bi=0; if the detection result of the i-th droplet is negative, then bi=1. Calculate the droplet volume to be detected in the next round based on the estimated total copy number n under the current conditions; 4) Based on the droplet volume determined in step 3), a single droplet is extracted from the original nucleic acid sample to continue nucleic acid detection. Step 3) is repeated for iterative detection. When the total number of droplets detected reaches the preset termination standard, the iteration stops. At this time, n, which is the solution obtained in step 3), is the final estimate of the total copy number X of the original sample.
[0033] According to some aspects of the invention, the volume of the droplets in step 1) is 0.04-0.12 μL.
[0034] According to some aspects of the present invention, in step 2), the volume of droplets to be detected in the next round is 1.5-1.7v / n; where v is the total detection volume of all droplets at present, and n is an estimated value of the total copy number under the current conditions.
[0035] According to some aspects of the present invention, in step 3), the volume of droplets to be detected in the next round is 1.5-1.7v / n; where v is the total detection volume of all droplets at present, and n is the estimated value of the total copy number under the current conditions.
[0036] According to some aspects of the present invention, in step 4), the standard is a total number of droplets detected ≥ 350 droplets; preferably ≥ 650 droplets. Specifically, when using a standard of ordinary precision detection (estimated error within ±15%), the total number of droplets detected needs to be ≥ 350 droplets; when using a standard of high precision detection (estimated error within ±10%), the total number of droplets detected needs to be ≥ 650 droplets.
[0037] It is understood that the method of the present invention is particularly applicable to the PCR nucleic acid detection process. Therefore, the present invention also provides the application of this absolute quantification method in nucleic acid detection, especially in the PCR nucleic acid detection process.
[0038] It is understood that, based on the core method of the present invention, the present invention may also include models, apparatuses, devices, or storage media prepared according to this method. Therefore, according to some aspects of the present invention, the present invention also provides an electronic device comprising: a processor and a memory; the processor and the memory are connected, wherein the memory is used to store a computer program, and the processor is used to invoke the computer program to perform the method described in any of the preceding claims. According to other aspects of the present invention, the present invention also provides a computer storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, perform the method described in any of the preceding claims.
[0039] The present invention will now be described in conjunction with specific embodiments.
[0040] Example 1: Exploration of the Evaluation Model of the Invention Traditional absolute quantification of nucleic acids follows the assumption that, after thorough mixing, nucleic acid molecules are independently and randomly distributed within the original sample. When the system is divided into N equal-volume droplets, each molecule has an equal probability of falling into any droplet. In this case, the nucleic acid copy number in a single droplet follows a Poisson distribution with parameter λ, where λ = n / N (n is the total copy number in the system). If the number of negative droplets detected is b, then the negative proportion q = b / N, and the total copy number in the system can be estimated using the formula n = -N·ln(q). Furthermore, the standard deviation of the likelihood estimate ln(λ) can be expressed as... This reflects the statistical uncertainty of the quantitative results. The more microdroplets N are divided, the smaller the standard deviation, and the more accurate the estimation results.
[0041] The inventors sought to maintain high quantitative accuracy and confidence levels while significantly reducing the total number of droplets generated. Their core idea was to shift from passive, fixed-scale Poisson statistics to active, sequential decision-making Bayesian inference. To this end, the following hypotheses were proposed: Relying on Bayesian inference to achieve "dynamic sampling instead of massive sampling," the absolute quantification process is viewed as a sequential closed loop of "detection-inference-decision-re-detection." First, a very small initial droplet is aspirated from a homogenized nucleic acid sample for PCR detection, obtaining a positive or negative result and establishing the posterior probability distribution of the total copy number n. The system then enters a core iterative loop: if all droplet results are consistent (all positive or all negative), it indicates insufficient data. The algorithm dynamically calculates the optimal detection volume for the next step based on the principle of maximizing information gain and supplements the sample with that volume of droplets for the next round of detection. If positive and negative droplets coexist, it indicates that the data contains sufficient information. The algorithm updates the posterior distribution of n using all historical data and assesses whether the termination condition has been met. This process is repeated until the error range of the posterior estimate converges to the target interval (e.g., ±10%). Finally, the system outputs the maximum likelihood estimate of the posterior distribution as the absolute copy number of the nucleic acid.
[0042] The following detection model is established: 1. Prior distribution Within the Bayesian inference framework, the prior distribution is set based on the estimated initial nucleic acid concentration of the sample. To reduce the average number of convergence steps in subsequent iterative optimization processes, this method sets the median value of the ddPCR instrument detection range (approximately 10-30 copies / μL) as the initial concentration prior estimate, thereby reducing the average number of convergence steps in subsequent iterative optimization processes. The method for calculating the droplet detection volume from the concentration is detailed in the subsequent acquisition function section, which shows that a reasonable initial detection droplet volume corresponding to 10-30 copies / μL is approximately 0.04-0.12 μL.
[0043] 2. Narrow the detection range The detection result of a single droplet is only either positive or negative. To achieve accurate estimation of the target nucleic acid copy number, the detection volume should be controlled to induce a statistical state where positive and negative droplets coexist in the detection result, thus providing sufficient information for parameter estimation. When the initial detection droplet is positive, it indicates a high concentration of the target nucleic acid. In this case, the droplet volume of the next round of detection should be reduced to increase the probability of obtaining a negative result. If the next round of results is also positive, the droplet volume should be iteratively reduced until a negative result appears. Conversely, if all initial droplets are negative, it indicates a low concentration of the target nucleic acid. A symmetrical strategy should be used to gradually increase the droplet volume until a positive result appears.
[0044] However, in all-positive or all-negative scenarios, the lack of observable statistical changes makes it impossible to effectively constrain the true range of nucleic acid copy numbers. Therefore, in the all-positive scenario, the total copy number estimate is set to the minimum theoretical value that can produce an all-positive result with a 95% confidence probability (α=0.05), ensuring that when reducing the detection volume, the adjustment step size will not be too large, directly crossing the optimal detection interval. Similarly, in the all-negative scenario, the total copy number estimate is set to the maximum theoretical value that can produce an all-negative result with a 95% confidence probability (α=0.05).
[0045] When all droplets are positive, then: When all droplets are negative, then: The calculated value of n is an estimate of the current total copy number assuming all droplets are positive or all are negative. This estimate guides the decision on the next round of detection volume, with the strategy being: the next round of detection volume should be 1.5-1.7v / n (where v is the sum of the volumes of the currently detected droplets; the theoretical basis for this coefficient range is detailed in the subsequent acquisition function section). This iterative decision-making process will continue until a state of coexistence of positive and negative droplets appears in the detection results, thus providing a statistical basis for subsequent accurate estimation.
[0046] 3. Posterior distribution Each detection result (positive or negative for a droplet) is considered a "Bayesian update" of the total copy number n, forming the posterior probability distribution of the current understanding of n. When the detection result includes both positive and negative droplets, the concentration of nucleic acid molecules in the sample can be estimated by maximizing the likelihood function. According to the Poisson distribution, the negative probability of a single droplet, based on the Poisson distribution assumption, satisfies the following relationship: (Where Ni is the ratio of the original sample volume to the volume of the i-th droplet).
[0047] At this point, the likelihood function of the posterior distribution can be expressed in the following form (m is the total number of droplets detected, α is the confidence level parameter, e is the natural constant, bi=1 indicates that the droplet is negative, and bi=0 indicates that it is positive): This formula can be used to find the maximum likelihood estimate under ln, then we have: That is, solve the following equation (where n is the unknown) using the posterior distribution: The calculated 'n' represents the estimated total copy number of nucleic acid molecules in the original sample under the current conditions. During the iterative phase before model convergence, this value deviates significantly from the true copy number. Its core function is to guide the decision-making process for the next round of detection volume based on preset acquisition function rules (see the acquisition function section for details). Based on the decision results, droplets of the corresponding volume are collected for further detection, continuously updating the posterior distribution until the model convergence criterion is met (see the model convergence section for details).
[0048] 4. Acquisition Function The next step, "how much volume to detect," is not a fixed value, but a decision function based on the current state of knowledge. The accuracy of the posterior distribution in estimating sample concentration is quantified by the standard deviation of the likelihood function; the smaller the standard deviation, the more accurate the estimate. The standard deviation of the likelihood function is... To obtain the optimal estimation accuracy, the standard deviation needs to be minimized. Let its first derivative... Zero: Solving the equation yields This indicates that, given a total number of droplets N, the variance of the sample concentration estimate is minimized when each droplet contains an average of approximately 1.59 copies, i.e., the estimation error is minimized. Therefore, in the dynamic adjustment strategy, the volume ratio of subsequent droplets should be adjusted based on the maximum likelihood estimate of the current sample copy number. That is, the volume of the next droplet to be tested should be 1.5-1.7v / n, and for a more accurate estimate, it should be 1.59v / n, where v is the total detection volume of m droplets.
[0049] 5. Model convergence When the 95% confidence interval of the copy number estimate falls within ±10% of the true copy number, the estimation error of the model can be considered to be no higher than 10%, achieving quantitative accuracy comparable to ddPCR. This is based on the standard deviation formula of ln(λ). The 95% confidence interval expression for λ can be obtained as follows: Maximum: Lower limit: The corresponding error level can be expressed as: When the sample size m is large, P can be approximated as: This formula is in The minimum value is reached at this time. That is, satisfying .
[0050] To control the error within ±15%, substituting P=0.15 yields a theoretical lower bound of m≥264, meaning the number of detected droplets should be at least 264; to control the error within ±10%, the number of detected droplets should be at least 593. However, in actual sequential inference, the target parameter λ itself is the parameter to be estimated, and the process of its estimated value asymptotically approaching 1.59 from the initial point consumes additional samples. Especially in the early stages of iteration, due to the large posterior variance, the information gain from a single observation is limited, causing the algorithm's convergence path to be less than the theoretically optimal path, resulting in unavoidable dynamic learning costs. Therefore, this method adds 60-80 droplets as a safety margin on top of the theoretical lower bound. Thus, the stopping conditions are established as follows: with coarse estimation accuracy (estimation error within ±15%), the cumulative number of detected droplets must reach 350; with precise estimation accuracy (estimation error within ±10%), the cumulative number of detected droplets must reach 650. This setting provides sufficient statistical redundancy for the estimation process under finite sample conditions, ensuring that the model can stably converge to the target accuracy level and terminate the iteration in practice. Solve for the maximum likelihood estimate of the current posterior distribution: The estimated value of parameter n is the total copy number of the original nucleic acid sample. The detailed dynamic adjustment process of this scheme is as follows: Figure 1 As shown.
[0051] Example 2: Establishment of the detection method of the present invention Based on the exploration in Example 1, this invention establishes a method for absolute quantification of nucleic acids based on Bayesian inference, specifically including the following steps: 1) Thoroughly mix the nucleic acid sample, and take a single droplet with a volume of 0.04-0.12 μL to carry out PCR detection to obtain the positive or negative result of the droplet; the absolute quantification method is based on Bayesian inference for absolute quantification of nucleic acid.
[0052] 2) Based on the current droplet detection results, an estimated value n of the total copy number under the current conditions is obtained, which is used to calculate the droplet volume for the next round of detection, so as to achieve a state where positive and negative results coexist as soon as possible. The value of n follows the following formula: If the detected droplet is positive, the estimated value of the total copy number n under the current conditions is calculated based on the following formula: If the detected droplet is negative, the estimated total copy number n under the current conditions is calculated based on the following formula: Where m is the total number of droplets detected, Ni is the ratio of the original sample volume to the volume of the i-th droplet, α is the confidence level parameter, and e is the natural constant; Calculate the droplet volume to be detected in the next round based on the estimated total copy number n under current conditions; the droplet volume calculation formula is: 1.5 - 1.7v / n; where v is the total detection volume of all droplets currently, and n is the estimated total copy number under current conditions. 3) Based on the previously determined droplet volume, supplement the original nucleic acid sample with single droplets and continue nucleic acid testing, repeating step 2) for iterative testing. When the cumulative tested droplets show both positive and negative results, calculate the estimated value n of the total copy number under the current conditions based on the following formula: Where, if the detection result of the i-th droplet is positive, then bi=0; if the detection result of the i-th droplet is negative, then bi=1. At this point, the volume of droplets to be detected in the next round is calculated based on the estimated total copy number n under the current conditions, using the formula 1.5-1.7v / n; where v is the total detection volume of all droplets and n is the estimated total copy number under the current conditions.
[0053] 4) Based on the droplet volume determined in step 3), supplement the original nucleic acid sample with single droplets and continue nucleic acid detection. Repeat step 3) for iterative detection. The iteration terminates when the total number of droplets detected reaches ≥350 under low-precision estimation (error within ±15%), or when the total number of droplets detected reaches ≥650 under high-precision estimation (error within ±10%). The n value calculated based on the estimation formula in step 2) is then the final estimated value of the total copy number of nucleic acid in the sample.
[0054] Example 3: Verification of the Method of the Invention To verify the model's feasibility, this study employed a Poisson distribution-based random simulation to generate experimental data streams. This method is widely accepted in computational experiments and biostatistical validation and can replace real physical experiments for methodological verification. The specific simulation process is as follows: based on the Poisson distribution parameter λ = total copy number / total volume × droplet volume, the actual nucleic acid copy number in each droplet is randomly generated. When the simulated copy number is 0, the droplet is considered negative; when it is greater than 0, it is considered positive. This simulation strategy can rigorously reproduce the core characteristic of the random distribution of nucleic acid molecules within droplets in real ddPCR detection, while completely eliminating the influence of non-model factors such as sample preparation, reagent fluctuations, and instrument noise, thus achieving a "pure" evaluation of the algorithm's statistical performance.
[0055] The total nucleic acid copy number of the original sample was set to 20,000, the total sample volume to 20 μL, the initial detection droplet volume to 0.08 μL, and the number of detection droplets per iteration to 1. The theoretical sampling process and model iterative optimization process based on this simulated data flow are shown in Figure 2. Since the initial droplet is large and positive, the model gradually reduces the volume of subsequent droplets until the fifth droplet is generated, at which point a negative result is obtained for the first time. Afterward, the volume of subsequent droplets is determined based on the maximum likelihood of all measured droplets. During the iteration process, the following pattern is observed: if the current droplet test is positive, the volume of the next droplet will decrease accordingly; if it is negative, the volume will increase. As the number of iterations increases, the oscillation amplitude of the volume adjustment gradually converges, eventually fluctuating within a small range near the true copy number, indicating that the estimation result tends to stabilize. The above simulated convergence behavior is highly consistent with the optimization process of "approaching the true concentration by adjusting the volume" in actual physical experiments, fully demonstrating the theoretical self-consistency and convergence reliability of this method. The core algorithm verification and parameter optimization can be completed before being put into real experiments.
[0056] Similarly, when the initial copy number of the sample is low (total nucleic acid copy number is 500), the model of this invention is also used for simulated detection, and the results are as follows. Figure 3 As shown, since the initial droplet detection result is negative, the model will dynamically increase the droplet volume in subsequent iterations. This adaptive adjustment mechanism enables the model to quickly locate the approximate order of magnitude of the original copy number and gradually converge to its precise range through continuous iteration.
[0057] In the dynamic detection process of ddPCR, the determination of whether each droplet is positive or negative is essentially a random event, and its statistical characteristics conform to the probability law described by the Poisson distribution. If a low-probability event occurs in the initial sampling stage (such as a droplet with a high probability of being positive being determined as negative), the concentration estimation curve may deviate from the true value, showing a shift in the opposite direction to the expectation. However, the Bayesian dynamic optimization model proposed in this paper can gradually correct the estimation bias through iterative updates based on the detection results of subsequent droplets, so that the curve eventually converges to the true concentration level, but this correction process usually requires more iterations. Taking a total nucleic acid copy number of 500 as an example, when using 0.08 μL as the initial droplet volume, the positive probability calculated according to the Poisson distribution is 86%, and the negative probability is 14%. The following figure compares the convergence behavior of the estimation curve under two conditions: a high-probability event (the initial droplet is positive) and a low-probability event (the initial droplet is negative). The results are as follows: Figure 4 and Figure 5 As shown, when starting with a high-probability event, only about 30 microdrops are needed to bring the estimated value to a stable convergence around 500 copies when the initial state conforms to the high-probability event (at this point, the estimation error is still relatively large and does not meet the invention droplet standard); while when starting with a low-probability event, about 60 microdrops are needed to achieve the same convergence accuracy, demonstrating the significant impact of initial sampling fluctuations on estimation efficiency. Despite the above efficiency differences, this model still demonstrates robust correction capabilities after the initial sampling falls into a low-probability event, ensuring the consistency and reliability of the final quantitative results, rather than systematic estimation errors caused by initial random bias.
[0058] Using the Bayesian inference model established in this invention, with total nucleic acid copy numbers of 40,000 and 4,000,000 as detection baselines, a total sample volume of 20 μL, an initial detection droplet volume of 0.08 μL, and a detection droplet count of 1 per iteration, 1000 independent simulation experiments were conducted. Based on this, the 95% confidence interval for the model's estimate of the target copy number was calculated. The statistical stability and uncertainty range of the model's estimation results were evaluated through multiple repeated experiments. Specific results are as follows: Figure 6 and 7 As shown.
[0059] The results show that there is a clear correlation between the number of samples required to obtain copy number estimates with specific accuracy: when the total copy number is 40,000, 139 microdroplets are needed to control the error within ±25%, 180 microdroplets within ±20%, 325 microdroplets within ±15%, and 628 microdroplets within ±10%. When the total copy number is 4,000,000, the number of microdroplets required to reach the same error threshold are 138 (±25%), 185 (±20%), 331 (±15%), and 641 (±10%), respectively. These results fully verify that the lower limit of the number of microdroplets established by the theoretical derivation (e.g., at least 264 microdroplets are needed to control the error within ±15%, and at least 593 microdroplets are needed to control the error within ±10%) has a rigorous mathematical basis. The actual test values closely approximate the theoretical lower bound, confirming that this Bayesian sequential inference framework maintains near-optimal statistical efficiency even in finite sample scenarios. The additional cost introduced by its dynamic learning path is kept to an extremely low level (approximately 40-70 additional droplets). Furthermore, it demonstrates that the 350 (rough estimate) and 650 (precise estimate) droplet termination conditions set by this method effectively buffer the uncertainties that may arise in actual operation due to parameter estimation fluctuations, random sampling variations, and computational convergence stability, thus achieving an excellent balance between statistical optimality and operational robustness. Notably, even with a 100-fold increase in copy number, maintaining a constant initial detection volume of 0.08 μL and achieving the same number of droplets to reach each error threshold indicates that the initial detection volume is not a critical factor. This is mainly due to the Bayesian sequential inference strategy's ability to dynamically adjust the droplet volume based on real-time observations, rapidly guiding the sampling process into a favorable inference state that yields both positive and negative results.
[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for absolute quantification of nucleic acids, characterized in that, The absolute quantification method is based on Bayesian inference and achieves absolute quantification of nucleic acids by adjusting the detection process simultaneously.
2. The absolute quantification method according to claim 1, characterized in that, The absolute quantification method includes the following steps: 1) Let X be the total copy number of nucleic acid molecules in the original nucleic acid sample. Mix the sample thoroughly, take a single droplet and perform nucleic acid detection to obtain the positive or negative result of the droplet. 2) If the detected droplet is positive, calculate the estimated total copy number n under the current conditions based on the following formula: If the detected droplet is negative, the estimated total copy number n under the current conditions is calculated based on the following formula: Where m is the total number of droplets detected, Ni is the ratio of the original sample volume to the volume of the i-th droplet, α is the confidence level parameter, and e is the natural constant; Calculate the droplet volume to be detected in the next round based on the estimated total copy number n under the current conditions; 3) Based on the droplet volume determined in step 2), a single droplet is extracted from the original nucleic acid sample to continue nucleic acid detection. Step 2) is repeated iteratively. The iteration stops when both positive and negative results are observed in the cumulative number of detected droplets. The estimated value of the total copy number n under the current conditions is then calculated based on the following formula: Where, if the detection result of the i-th droplet is positive, then bi=0; if the detection result of the i-th droplet is negative, then bi=1. Calculate the droplet volume to be detected in the next round based on the estimated total copy number n under the current conditions; 4) Based on the droplet volume determined in step 3), a single droplet is extracted from the original nucleic acid sample to continue nucleic acid detection. Step 3) is repeated for iterative detection. When the total number of droplets detected reaches the preset termination standard, the iteration stops. At this time, n, which is the solution obtained in step 3), is the final estimate of the total copy number X of the original sample.
3. The absolute quantification method according to claim 2, characterized in that, In step 1), the volume of the microdroplet is 0.04-0.12 μL.
4. The absolute quantification method according to any one of claims 2-3, characterized in that, In step 2), the volume of droplets to be detected in the next round is 1.5-1.7v / n; where v is the total detection volume of all droplets and n is the estimated total copy number under the current conditions.
5. The absolute quantification method according to any one of claims 2-4, characterized in that, In step 3), the volume of droplets to be detected in the next round is 1.5-1.7v / n; where v is the total detection volume of all droplets and n is the estimated total copy number under the current conditions.
6. The absolute quantification method according to any one of claims 2-5, characterized in that, In step 4), the standard is that the total number of droplets detected is ≥350 droplets; preferably, the standard is that the total number of droplets detected is ≥650 droplets.
7. The absolute quantification method according to any one of claims 2-6, characterized in that, The nucleic acid test mentioned is a PCR nucleic acid test.
8. The application of the absolute quantification method according to any one of claims 1-7 in nucleic acid detection; preferably, its application in PCR nucleic acid detection.
9. An electronic device, characterized in that, include: Processor and memory; The processor is connected to a memory, wherein the memory is used to store a computer program, and the processor is used to invoke the computer program to execute the absolute quantification method as described in any one of claims 1-7.
10. A computer storage medium, characterized in that, The computer storage medium stores a computer program, which includes program instructions that, when executed by a processor, perform the absolute quantification method as described in any one of claims 1-7.