A PCR result analysis system and method based on statistical distribution determination

By constructing a probability density function and calculating event probabilities, the problems of individual differences and platform stability in the qualitative interpretation of qPCR were solved, enabling rapid and accurate sample type determination and improving the stability and interpretability of qPCR.

CN122090935APending Publication Date: 2026-05-26SUZHOU HAIMIAO BIOTECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU HAIMIAO BIOTECH CO LTD
Filing Date
2026-01-08
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing qualitative interpretation methods for qPCR ignore individual differences, batch effects, and platform differences, lack probabilistic significance and uncertainty assessment, are susceptible to experimental errors, have insufficient stability across batches or platforms, and neural network models rely on large-scale training data and lack traceability of results.

Method used

By constructing a probability density function, the PCR results data of unknown samples are determined based on statistical distribution. The probability density function is constructed using methods such as parameterization, kernel density estimation, and Monte Carlo simulation. The event probability of unknown samples is calculated, and the sample type is determined by combining one-tailed or two-tailed probabilities.

Benefits of technology

It enables rapid and accurate sample type determination, improves the stability and interpretability of qPCR qualitative interpretation, reduces dependence on training data, adapts to different sample types and detection platforms, and enhances the accuracy and consistency of interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_46
    Figure SMS_46
  • Figure SMS_47
    Figure SMS_47
  • Figure SMS_85
    Figure SMS_85
Patent Text Reader

Abstract

This invention discloses a PCR result analysis system and method based on statistical distribution determination. The invention utilizes statistical methods to obtain the probability distribution of PCR result data of the target gene in real samples of various sample types, constructs a probability density function, and obtains the event probability that the PCR result data of the target gene in an unknown sample belongs to the PCR result data distribution of real samples by integrating the probability density function over a specific domain. The event probability is compared with an event probability threshold, and the result output is whether the sample type of the unknown sample is consistent with that of the real sample. This invention explores the statistical distribution characteristics of PCR result data of real samples, and combined with probability analysis, can quickly and accurately determine the sample type of unknown samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of molecular biology technology, specifically relating to a PCR result analysis system and method based on statistical distribution determination, which is suitable for qualitative analysis of PCR results and ultimately applied to disease screening, diagnosis and classification. Background Technology

[0002] Quantitative real-time polymerase chain reaction (qPCR) is a widely used nucleic acid amplification and quantitative analysis technique in molecular biology research and clinical testing. Its basic principle is to achieve exponential amplification of the target nucleic acid sequence through polymerase chain reaction (PCR), and to quantitatively analyze the amplified products by real-time monitoring of the fluorescence signal during the amplification process.

[0003] In qPCR reaction systems, the generation of fluorescence signals typically depends on fluorescent dyes or specific probes. The intensity of the fluorescence signal gradually increases with the number of amplification cycles. By setting a fluorescence threshold and recording the number of cycles in which the fluorescence signal first exceeds the threshold (i.e., the threshold cycle number, Ct value), the quantitative estimation of the target sequence can be achieved. The smaller the Ct value, the higher the initial template amount.

[0004] Current qualitative interpretation of qPCR typically relies on the Ct values ​​of one or more marker genes, comparing them to a fixed threshold to determine whether a sample is positive or negative. While this method is simple to operate, it has significant shortcomings: First, it does not consider the overall distribution characteristics of the normal and positive groups, ignoring the influence of individual differences, batch effects, and platform differences on the Ct values; second, it relies solely on single-point thresholds, lacking probabilistic significance and uncertainty assessment, making critical samples prone to misjudgment; third, it is sensitive to abnormal amplification or atypical signals, easily affected by experimental errors; and fourth, the thresholds are often set based on limited sample experience, resulting in poor model stability and cross-sample adaptability. Chinese patent document CN120412744A discloses a neural network model for qPCR result analysis. This neural network model is a single neural network model or a stacked neural network model combining multiple neural network models, taking at least three types of qPCR result data as input and the sample type of the target gene corresponding to the qPCR result data as the result output. While existing neural network methods can learn complex nonlinear relationships, their model training relies on a large amount of high-quality sample data, and the internal parameters of the model are difficult to interpret, resulting in a lack of traceability in the judgment results. At the same time, such models are sensitive to changes in the distribution of input data, lack stability across batches or platforms, and have complex calculation and deployment processes, which is not conducive to their widespread application in conventional detection environments. Summary of the Invention

[0005] To address the shortcomings of existing technologies in the qualitative interpretation of qPCR results, this invention provides a PCR result analysis system based on statistical distribution determination. A standard dataset is constructed using real samples, and a probability density function for the target gene PCR result data is established. By integrating the probability density function, the probability of an unknown sample's target gene PCR result data belonging to the distribution of PCR result data for the same target gene in real samples is calculated, thus determining whether the sample type of the unknown sample is consistent with that of the real samples. This invention explores the statistical distribution characteristics of real sample PCR result data and, combined with probability analysis, can quickly and accurately output the sample type result of unknown samples.

[0006] The specific technical solution of this invention is as follows:

[0007] A PCR result analysis system based on statistical distribution determination, wherein the system takes PCR result data of unknown sample target gene as input, analyzes the statistical distribution of PCR result data of target gene of real sample known sample type, constructs the probability density function corresponding to the PCR result data, calculates the event probability that the PCR result data of unknown sample target gene belongs to the distribution of PCR result data of real sample target gene by integrating over a specific domain of the probability density function, compares the event probability with the event probability threshold, and outputs the result as whether the sample type of unknown sample is consistent with the sample type of real sample;

[0008] The sample type refers to the source characteristics of the target gene, selected from diseases, races, ages, sexes, or other biological or environmental factors that influence gene characteristics. Sample types can be different types, or further subdivisions of the same type of factor. For example, when the sample type is a disease, the different sample types described in this invention represent different diseases (such as lung cancer, liver cancer, stomach cancer, colorectal cancer, etc.); when the sample type is a race, the different sample types described in this invention represent different races (such as Asians, Caucasians, Blacks, and Browns).

[0009] When the sample type is disease, the real sample is a disease-positive sample, and the output result is whether the unknown sample is the disease.

[0010] When the sample type is race, age, sex, or other biological or environmental factors that affect genetic characteristics, the real sample can be a specific population, such as a specific race, a specific age or age group, or a male / female sample. The output result is whether the unknown sample belongs to the population with that characteristic.

[0011] The PCR results data include one or more of the following types: the difference in CT values ​​between the target gene and the internal reference gene, the ratio of the maximum fluorescence signal amplification of the target gene and the internal reference gene, the ratio of the curvature of the fluorescence amplification of the target gene and the internal reference gene at the final amplification cycle, the ratio of the slope of the fluorescence amplification of the target gene and the internal reference gene at the final amplification cycle, and the ratio of the maximum fluorescence signal values ​​of the target gene and the internal reference gene.

[0012] Preferably, the PCR result data is selected from the difference in CT values ​​between the target gene and the internal reference gene, the ratio of the maximum fluorescence signal amplification values ​​between the target gene and the internal reference gene, and the ratio of the maximum fluorescence signal values ​​between the target gene and the internal reference gene.

[0013] When using multiple PCR result data, calculate the event probability of each PCR result data belonging to a certain sample type. The unknown sample is considered to belong to the current sample type if and only if the event probability of all data types exceeds the threshold (the sample type of the unknown sample is consistent with the sample type of the real sample).

[0014] Unknown samples and real samples were processed in the same way, the same target genes were amplified using the same primers and fluorescent probes, and the same PCR conditions were used to obtain PCR results data.

[0015] The probability density function of the analysis system described in this invention is constructed using the following method:

[0016] (1) Collect several real samples according to at least one sample type of the target gene, obtain the PCR result data of the target gene of all real samples, and clean and standardize the data to form a standard dataset. The cleaning and standardization process includes removing missing values, outliers and duplicate data, unifying sample number, gene name and label format, checking logical consistency, normalizing or Z-score standardizing to adjust the data scale, correcting batch effect when necessary, normalizing signal intensity or logarithmic transformation to eliminate experimental differences and skewed distribution, and finally forming a standard dataset that can be directly used for analysis and modeling.

[0017] (2) Based on the standard dataset, construct the probability density function using parameterization, empirical distribution function, kernel density estimation, Monte Carlo simulation, mixed distribution or Bootstrap method.

[0018] Specifically, the probability density function is selected from the following:

[0019] Based on normal distribution: , where x is the value of a random variable, i.e., the PCR result data of an unknown sample; σ is the probability density function value corresponding to x; μ is the mean, i.e., the center of the normal distribution; σ is the standard deviation, used to measure the dispersion and fluctuation of the distribution; exp() is the exponential function with the natural constant e as the base.

[0020] Based on the exponential distribution: , where x is the value of a random variable, i.e., the PCR result data of an unknown sample; Let x be the probability density function value; λ be the rate parameter, representing the average number of times the event occurs per unit time; and e be the natural constant.

[0021] Based on gamma distribution: Where Γ(k) is the gamma function, , where x is the value of a random variable, i.e., the PCR result data of an unknown sample; x represents the probability density function value; k is a shape parameter that controls the shape of the distribution curve. This is a scale parameter that controls the scaling of the distribution along the x-axis; The gamma function is used for normalization to make the sum of integrals equal to 1; e is the natural constant.

[0022] Based on the Beta distribution: in , where x is the value of a random variable, i.e., the PCR result data of an unknown sample, with a value range of [0, 1]; α is a shape parameter used to control the shape of the distribution near the 0 side; β is a shape parameter used to control the shape of the distribution near the 1 side; is the beta function, is the normalization constant, and is the coefficient whose integral is 1;

[0023] Based on the log-normal distribution: , where x is the value of the random variable, i.e., the PCR result data of the unknown sample; lnx is the natural logarithm function, i.e., the logarithmic transformation of x; μ is the logarithmic mean, i.e. the mathematical expectation of lnx; σ is the logarithmic standard deviation, i.e. the standard deviation of lnx; This is a normalization constant, used to ensure that the integral equals 1.

[0024] Based on empirical distribution: ,in , where x is the value of a random variable, i.e., the PCR result data of an unknown sample; This represents the i-th observation; n is the number of samples. This is an indicator function; it takes the value 1 if the condition is true, and 0 otherwise. The cumulative distribution proportion of the sample is the empirical distribution function; Δ represents an extremely small change. The value is the probability density function value.

[0025] Based on kernel density estimation: , where x is the value of the random variable, i.e., the PCR result data of the unknown sample; n is the number of samples; Here are the PCR results for the i-th observed sample; K is the normal distribution kernel function, symmetric and with an integral of 1; h>0 is the bandwidth, used to control the smoothing degree; e is the natural constant.

[0026] Based on mixed distribution: , where x is the value of the random variable, i.e., the PCR result data of the unknown sample; K is the number of sub-distributions, i.e. the number of components in the mixture distribution; The weight is the proportion of the k-th sub-distribution; Let be the probability density function of the k-th sub-distribution, selected from any of the aforementioned probability density functions based on the normal distribution, exponential distribution, gamma distribution, beta distribution, log-normal distribution, empirical distribution, or kernel density estimation.

[0027] Based on Monte Carlo simulation: it does not define a function itself, but relies on a large number of samples { To approximate the target distribution, kernel density estimation is used to obtain an approximate probability density function. , where x is the value of the random variable, i.e., the PCR result data of the unknown sample; n is the number of samples; represents the PCR results of the i-th observed sample; K is the normal distribution kernel function, which is symmetric and has an integral of 1; h>0 is the bandwidth, used to control the smoothing degree; e is the natural constant;

[0028] Bootstrap-based approach: Where x is the value of a random variable, i.e., the PCR result data of an unknown sample; B is the number of bootstraps; b is the resampling index, i.e., the b-th bootstrap; each Kernel density estimation from a Bootstrap sample.

[0029] Preferably, the smoothness of the kernel density estimation is controlled using the Silverman rule, calculated as follows: min() is the minimum value function, σ is the standard deviation of the x dataset, and IQR is the interquartile range.

[0030] In a specific example of the present invention, the probability density function is based on different distribution types, namely, the probability density function based on the normal distribution, the Beta distribution, the empirical distribution, the mixed distribution, or the Bootstrap method.

[0031] In the analysis system described in this invention, the event probability can be selected from one-tailed probability (left-tailed probability or right-tailed probability) or two-tailed probability.

[0032] Different methods are used to calculate the probability of different events, and the integral formula is as follows:

[0033] Right tail probability: , where X represents the random variable of the sample PCR results in the constructed distribution; This is the corresponding probability density function; The threshold is the starting point for calculating the right-tail probability, specifically the measured value of the PCR result of the sample to be evaluated. The value of X is a function of the right-tail probability, i.e., X is greater than... The probability of;

[0034] Left tail probability: , where X represents the random variable of the sample PCR results in the constructed distribution; This is the corresponding probability density function; The threshold is the endpoint for calculating the left-tail probability, specifically the measured value of the PCR result of the sample to be evaluated. The function value of the left-tail probability, i.e., X is less than The probability of;

[0035] Two-tailed probability: , where X represents the random variable of the sample PCR results in the constructed distribution; Here, μ is the probability density function, used to describe the probability density of X at each point; μ is the mean, i.e., the mathematical expectation of the random variable X; and k is the magnitude of the deviation of the data to be evaluated from the mean, specifically... , These are the measured values ​​of the PCR results for the samples to be evaluated.

[0036] In the analysis system described in this invention, when the event probability is a one-tailed probability, the event probability threshold is 0.05~0.95. When the event probability is greater than 0.05 and less than 0.95, it is determined that the sample type of the unknown sample is consistent with the sample type of the real sample; otherwise, they are inconsistent. When the event probability is a two-tailed probability, the event probability threshold is 0.05. When the event probability is greater than 0.05, it is determined that the sample type of the unknown sample is consistent with the sample type of the real sample; otherwise, they are inconsistent.

[0037] The analysis system described in this invention includes:

[0038] (1) Data processing module: Processes the raw PCR results of unknown samples or real samples to obtain PCR result data;

[0039] (2) Standard dataset construction module: Clean and standardize the target gene PCR results data of real samples to construct a standard dataset;

[0040] (3) Probability density function construction module: Based on the standard dataset, the probability density function is constructed using parameterization methods, empirical distribution functions, kernel density estimation, Monte Carlo simulation, mixture distribution or Bootstrap method;

[0041] (4) Event probability calculation module: Integrate the probability density function to calculate the event probability that the PCR result data of the target gene of the unknown sample belongs to the PCR result data of the target gene of the real sample;

[0042] (5) Result output module: compare the event probability with the event probability threshold, and output whether the sample type of the unknown sample is consistent with the sample type of the real sample.

[0043] Preferably, the system of the present invention further includes a validation dataset, which is real sample target gene PCR result data other than the standard dataset; different probability density functions and event probability integral formulas are selected for combination, and the event probability of the PCR result data of the target gene in the validation dataset belonging to the PCR result data distribution of the real sample target gene is calculated under each combination. The event probability is compared with the event probability threshold. Based on whether the judgment result is consistent with the real sample, the accuracy of the judgment result of each combination on the sample in the validation dataset is calculated. The combination of probability density function and event probability integral formula corresponding to the highest judgment accuracy of the sample in the validation dataset is selected as the optimal combination, which is used to calculate the event probability of the PCR result data of the unknown sample target gene belonging to the PCR result data distribution of the real sample target gene.

[0044] Preferably, the analysis system of the present invention further includes updating the event probability threshold based on the optimal combination, using the updated event probability threshold to determine whether the sample type of the unknown sample is consistent with the sample type of the true sample, and updating the event probability threshold using the following method: when selecting the one-tailed probability, the updated event probability threshold is less than min(P max +0.05, 0.95) and greater than max(P min -0.05, 0.05), where P max To verify the maximum probability of events in the dataset, P min To verify the minimum event probability in the dataset, `max()` is the function to maximize the probability, and `min()` is the function to minimize the probability. When choosing a two-tailed probability, the updated event probability threshold is greater than `max(P)`. min -0.05, 0.05), where P min The minimum probability of the event is given by , and max() is the function to find the maximum probability.

[0045] The analysis system described in this invention specifically includes:

[0046] Data processing module: processes the raw PCR results of unknown or real samples to obtain PCR result data;

[0047] Standard dataset construction module: Select target gene PCR results data from a portion of real samples, clean and standardize them to construct a standard dataset;

[0048] Validation dataset construction module: Cleans and standardizes the target gene PCR results data from real samples outside the standard dataset to construct the validation dataset;

[0049] Probability density function set construction module: Based on the standard dataset, several probability density functions are constructed using parametric methods, empirical distribution functions, kernel density estimation, Monte Carlo simulation, mixture distribution, or Bootstrap methods to obtain the probability density function set;

[0050] Event probability set construction module: Combines different event probability integral formulas with probability density functions in the probability density function set, calculates and outputs the event probability that the PCR result data of the target gene in the validation dataset belongs to the PCR result data distribution of the target gene in the real sample under different combinations, and obtains the event probability set;

[0051] Event probability calculation module: compares the event probability with the event probability threshold, calculates the accuracy of the judgment result of each combination pair of the sample judgment result in the verification dataset based on whether the judgment result is consistent with the real sample, selects the combination of probability density function and event probability integral formula corresponding to the highest judgment accuracy of the sample in the verification dataset as the optimal combination, and is used to calculate the event probability of the PCR result data of the target gene of the unknown sample belonging to the PCR result data distribution of the target gene of the real sample.

[0052] The results output module compares the event probability of the unknown sample with the event probability threshold and outputs whether the sample type of the unknown sample is consistent with the real sample.

[0053] Preferably, the result output module updates the event probability threshold based on the optimal combination before comparison: when a one-tailed probability is selected, the updated event probability threshold is less than min(P). max +0.05, 0.95) and greater than max(P min -0.05, 0.05), where P max To verify the maximum probability of events in the dataset, P min To verify the minimum event probability in the dataset, `max()` is the function to maximize the probability, and `min()` is the function to minimize the probability. When choosing a two-tailed probability, the updated event probability threshold is greater than `max(P)`. min -0.05, 0.05), where P minThe minimum event probability is given by 'max()', and the maximum event probability is given by 'max()'. The event probability of the unknown sample is compared with the updated event probability threshold, and the output shows whether the sample type of the unknown sample is consistent with the sample type of the real sample.

[0054] Another objective of this invention is to disclose a PCR result analysis method based on statistical distribution determination, which is achieved by using the analysis system described in this invention.

[0055] The analytical method of this invention collects several real samples according to at least one sample type to which the target gene belongs, uses statistical methods to obtain the probability distribution of PCR result data of the target gene in the real samples, constructs a probability density function, integrates over a specific domain of the probability density function, and calculates the probability of an event that the PCR result data of the target gene in the unknown sample belongs to the distribution of the QPCR result data of the target gene in the real sample. The sample type is the source characteristic of the target gene, selected from disease, race, age, sex or other biological or environmental factors that affect gene characteristics.

[0056] The PCR results data are selected from one or more of the following types: the difference in CT values ​​between the target gene and the internal reference gene, the ratio of the maximum fluorescence signal amplification of the target gene and the internal reference gene, the ratio of the curvature of the fluorescence amplification of the target gene and the internal reference gene at the final amplification cycle, the ratio of the slope of the fluorescence amplification of the target gene and the internal reference gene at the final amplification cycle, and the ratio of the maximum fluorescence signal of the target gene and the internal reference gene; the unknown sample and the real sample are treated in the same way, the same amplification primers and fluorescent probes are used for the same target gene, and the same PCR conditions are used to obtain the PCR results data.

[0057] The analytical method described in this invention includes the following steps:

[0058] (1) Constructing a standard dataset: Select real samples for the sample type of the sample to be tested, with no less than 50 samples of each real sample. Use a unified experimental protocol to conduct PCR experiments and obtain PCR result data. Clean and standardize the PCR result data to construct a standard dataset.

[0059] (2) Based on the standard dataset, construct the probability density function using parameterization, empirical distribution function, kernel density estimation, Monte Carlo simulation, mixture distribution or Bootstrap method;

[0060] (3) Obtain various PCR result data of unknown samples using the same experimental methods and conditions as real samples;

[0061] (4) Integrate the probability density function to calculate the probability that the PCR result data of the target gene of the unknown sample belongs to the real sample.

[0062] Furthermore, the analytical method includes the following steps:

[0063] (1) Constructing standard dataset and validation dataset: Select real samples for the sample type of the sample to be tested, with no less than 50 samples of each real sample. Use a unified experimental protocol to conduct PCR experiments and obtain PCR result data. Clean and standardize the PCR result data, select some PCR result data to construct a standard dataset, and select PCR result data other than the standard dataset to construct a validation dataset.

[0064] (2) Based on the standard dataset, construct a set of probability density functions using several of the following methods: parameterization, empirical distribution function, kernel density estimation, Monte Carlo simulation, mixture distribution, or Bootstrap method;

[0065] (3) Combine different event probability integral formulas with probability density functions in the probability density function set, calculate and output the event probability that the PCR result data of the target gene in the verification dataset belongs to the real sample, and obtain the event probability set;

[0066] (4) Obtain various PCR result data of unknown samples using the same experimental methods and conditions as real samples;

[0067] (5) Compare the event probabilities in the event probability set with the event probability threshold. Calculate the accuracy of each combination in judging the sample in the verification dataset based on whether the judgment result is consistent with the real sample. Select the combination of probability density function and event probability integral formula corresponding to the highest accuracy of judging the sample in the verification dataset as the optimal combination. This combination is used to calculate the event probability that the PCR result data of the target gene of the unknown sample belongs to the PCR result data distribution of the target gene of the real sample.

[0068] The analysis method described in this invention further includes comparing the event probability with the event probability threshold to determine whether the sample type of the unknown sample is consistent with the sample type of the real sample.

[0069] Preferably, the event probability is updated based on the optimal combination: when a one-tailed probability is selected, the threshold for updating the event probability is less than min(P). max +0.05, 0.95) and greater than max(P min -0.05, 0.05), where P max To verify the maximum probability of events in the dataset, P min To verify the minimum event probability in the dataset, `max()` is the function to maximize the probability, and `min()` is the function to minimize the probability. When choosing a two-tailed probability, the updated event probability threshold is greater than `max(P)`. min -0.05, 0.05), where Pmin The minimum event probability is given by 'max()', which is the maximum value function. The updated event probability threshold is used to determine whether the sample type of the unknown sample is consistent with the sample type of the real sample.

[0070] The system and method described in this invention can be used to determine the sample type of a single or multiple target genes.

[0071] Advantages of this invention:

[0072] This invention establishes a probability distribution of PCR results data from real samples of known sample types, constructs a probability density function, and calculates the probability of an unknown sample belonging to a real sample by integrating the function, thus achieving statistical discrimination of the sample type of an unknown sample. This method requires no large-scale training data, is structurally simple, computationally efficient, and can reflect the confidence level of the judgment with a clear probabilistic meaning. It also maintains high consistency and interpretability across different sample types and detection platforms, thereby significantly improving the stability and practicality of qPCR qualitative interpretation.

[0073] The statistical distribution-based PCR result analysis system and method described in this invention are based on statistical distributions, abandoning the linear division of fixed boundaries relied upon by the traditional Ct threshold method. This allows for a comprehensive consideration of inter-population differences and data fluctuations, significantly improving the robustness and accuracy of the judgment. Simultaneously, this method does not rely on large-scale training samples or complex network structures, avoiding the "black box" problem and overfitting risk of neural network models. The results have clear probabilistic meaning and interpretability, and can directly quantify the confidence level and anomaly degree of each sample. The method's calculation process is simple and efficient, easily integrated into existing qPCR analysis workflows, and maintains stable performance under cross-batch, cross-platform, and multi-sample type detection conditions. In summary, the method of this invention combines statistical robustness, interpretability, cross-scenario adaptability, and practical operability, providing a reliable new approach for high-precision analysis and standardized interpretation of qPCR qualitative results. Detailed Implementation

[0074] To make the objectives, technical solutions, beneficial effects and significant advancements of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below.

[0075] Obviously, all the embodiments described are only some embodiments of the present invention, and not all embodiments; based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0076] Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0077] It should also be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0078] The technical solution of the present invention will now be described in detail with reference to specific embodiments.

[0079] Example 1: Case Study of PCR Result Analysis of Cancer Samples

[0080] 1.1 Sequence Design

[0081] To explore the optimal combination of PCR result data types, probability density functions, and event probability calculation functions, a set of primer and probe sets targeting methylation-specific CpG sites in the plasma of colorectal cancer patients, a set targeting methylation-specific CpG sites in the plasma of gastric cancer patients, and a set targeting methylation-specific CpG sites in the plasma of lung cancer patients were designed. The specific templates were all DNA sequences consisting of 100 bases upstream and downstream of the specific site, converted to DNA templates by heavy-pressure sulfate conversion. The CpG sites were assumed to be methylated and remained CpG after conversion, while the remaining cytosine C was converted to uracil U, corresponding to thymine T in DNA amplification. The corresponding primer and probe sets covered the target CpG sites, and a methylation-specific PCR amplification system was constructed to amplify the DNA methylation templates.

[0082] The primer and probe sets and template sequence information for methylation-specific CpG sites in the plasma of colorectal cancer patients are as follows:

[0083] Specific template T1: 5'-TCGAATGTGTGGTAGGTTGTGTGTTAGTATAGATATTTTTTTAATTGGTTAAGTGATATTATGTAATTGTTATCGAGAATAAAATTAGTTTTTTTTTTGCGGTTTTATTAGCGAAAAGGAGGGGAGGGTTAGAAGTTATATTTTTTGTTTTAGGTTTTTCGTTTAGTTGTATTTTTAGGTTAGGTTGGTTTTTTTTGG-3';

[0084] Specific forward primer F1: 5'-GAATGTGTGGTAGGTTGTGTGTTAGTATAGATATTT-3';

[0085] Specific reverse primer R1: 5'-CCCTCCTTTTCGCTAATAAAACCG-3';

[0086] Specific probe P1: 5'-FAM-CTAATTTTATTCTCGATAACAATTAC-MGB-3'.

[0087] The primer and probe sets and template sequence information for methylation-specific CpG sites in the plasma of gastric cancer patients are as follows:

[0088] Specific template T2: 5'-TTATTGGCGGTTGTATATATTTTAATTTTTTGTATTTTAGTTCGTAGAGGAGATGGAGGATCGTTGTAGTTTTTGTAGTTTTGTTATTTTTGTTATTCGAAATTTTATTTGCGGTTTATTTTGGGGTTTTAGTGGGTTTTTGTTTTTATGTGGTAGTTGTGAAGGAAGGAAATGGTGAGAAAGGATGAGAGGAGGGGTG-3';

[0089] Specific forward primer F2: 5'-CGTAGAGAGGAGATGGAGGATCGT-3';

[0090] Specific reverse primer R2: 5'-CCAAAATAAACACGCAAATAAATTTCG-3';

[0091] Specific probe P2: 5'-FAM-GTAGTTTTTGTAGTTTTGTTATTTTTGTTAT-MGB-3'.

[0092] The primer and probe sets and template sequence information for methylation-specific CpG sites in the plasma of lung cancer patients are as follows:

[0093] Specific template T3: 5'-GGGAGATTTTTGGTTAGTAGGGGTTTTTTAAAGATATTGAGTTTATAGTAGGAAGTGGAGGGGAAATTTTAGTGATTGGTTCGAGATTAATTTGGATTTTCGTGTAAGGTATAATGTAGATAAAGGGGATTATTGTATTTTAAGTGGGTTTTGAAGTATTTTATATTATTTAAAAGGATAAAAATTTAGTTAGGTGTTGTT-3';

[0094] Specific forward primer F3: 5'-GATATTGAGTTTATAGTAGGAAGTGGAGGGGA-3';

[0095] Specific reverse primer R3: 5'-ATAATCCCCTTTATCTACATTATACCTTACACG-3';

[0096] Specific probe P3: 5'-FAM-TCCAAATTAATCTCGAACCAATCAC-MGB-3'.

[0097] In addition, a primer and probe set suitable for detecting the internal reference gene ACTB and its corresponding DNA template sequence information are included as follows:

[0098] Specific template T4: 5'-GTTGGTTTTTATTGTTTTATTTTTTTGTAGGGTTTATTTTTTTGTAGGGTTTATTTTTTGTTGTTTTTAATTTAGTTATATTATAAAGTTATATTTGGTTTTATTTTTAAGGTGTTTATTTTTATTTAATTGGTTTTAAGTTAGTGTATAGGTAAGTTTTGGTTGTTTTTATTTATTTTTAGGGAGATTAAAAGTTTTTAT-3';

[0099] Specific forward primer F4: 5'-GGGTTTATTTTTTTGTAGGGTTTATTTTTTG-3';

[0100] Specific reverse primer R4: 5'-CACTAACTTAAAACCAATTAAATAAAAATACACACC-3';

[0101] Specific probe P4: 5'-VIC-TAAAAATAAAACCAAATATAACTTTATAATATAAC-MGB-3'.

[0102] 1.2 Collect raw qPCR data

[0103] To facilitate the application of the method described in this invention, a total of 60 plasma samples were collected from colorectal cancer, gastric cancer, and lung cancer. 3 mL of plasma was collected from each sample. All collected plasmas were subjected to the same conditions and treatments. Free nucleic acids were extracted from each sample using a free nucleic acid extraction kit. The extracted free nucleic acids were then subjected to heavy-pressure sulfate conversion using a DNA transformation kit. This converted unmethylated cytosine in all free nucleic acid molecules of each sample to uracil, while methylated cytosine remained cytosine. This yielded purified DNA templates from all samples.

[0104] Prepare the qPCR reaction system according to the formulas described in Table 1 below. The F primers, R primers, and P probes are in a one-to-one correspondence. That is, prepare three sets of qPCR reaction solutions. The primers and probes of the first set are F1, R1, and P1; the primers and probes of the second set are F2, R2, and P2; and the primers and probes of the third set are F3, R3, and P3.

[0105] Table 1: Formulation of qPCR reaction system

[0106]

[0107] Subsequently, the corresponding qPCR experiments were performed according to the qPCR reaction procedure described in Table 2 below to obtain the qPCR result data for each independent sample.

[0108] Table 2: qPCR reaction procedure table

[0109]

[0110] 1.3 Processing of raw qPCR data

[0111] After completing the qPCR experiment, obtain the raw qPCR results for all samples. For any given sample, the following data should be included: Ct value of the FAM signal channel, maximum fluorescence signal amplification value of the FAM signal channel, curvature of the fluorescence amplification of the FAM signal channel at the final amplification cycle, slope of the fluorescence amplification of the FAM signal channel at the final amplification cycle, maximum fluorescence signal value of the FAM signal channel, Ct value of the VIC signal channel, maximum fluorescence signal amplification value of the VIC signal channel, curvature of the fluorescence amplification of the VIC signal channel at the final amplification cycle, slope of the fluorescence amplification of the VIC signal channel at the final amplification cycle, and maximum fluorescence signal value of the VIC signal channel.

[0112] The differences or ratios between the data from the FAM signal channel and the corresponding data from the VIC signal channel were calculated to obtain the following Ct value data (Data1), maximum fluorescence signal amplification ratio data (Data2), maximum fluorescence signal ratio data (Data3), final amplification cycle fluorescence amplification change curvature ratio data (Data4), and final amplification cycle fluorescence amplification change slope ratio data (Data5). After normalizing all data, the dataset for each independent sample was obtained.

[0113] 1.4 Constructing the probability density function

[0114] For any one of the three sample types—colorectal cancer, gastric cancer, and lung cancer—50 independent samples are selected from the corresponding dataset as a standard dataset. Different statistical distribution methods are used to construct different distribution models for the standard data or combinations thereof. Subsequently, the probability density function of the corresponding distribution model is calculated. The remaining samples in the dataset are used as a validation dataset. For the normal distribution method, the probability density function for all data types is: , where x is the value of the standard data, i.e. the standard data of the validation dataset sample; μ is the mean of the standard data of the standard dataset sample; σ is the standard deviation of the standard data of the standard dataset sample; exp() is an exponential function with the natural constant e as the base. For cases containing multiple data types, this function is used to calculate the probability density function of each data type.

[0115] For the exponential distribution method, the probability density function for all data types is: The function value is 0 when x is less than or equal to 0; λ is the rate parameter, which is equal to the reciprocal of the mean of the standard data in the standard dataset of the corresponding data type; and e is the natural constant. For cases containing multiple data types, this function is used to calculate the probability density function for each data type.

[0116] For the gamma distribution method, the probability density function for all data types is: ,in ; x is the specific value of each standard data in the sample to be analyzed in the validation dataset; k is the shape parameter, which is the ratio of the square of the mean to the square of the variance of the corresponding standard data type in the standard dataset; is a scaling parameter, which is the ratio of the square of the variance to the mean of the data of the corresponding standard data type in the standard dataset. The gamma function is used for normalization, ensuring the sum of integrals equals 1; e is the natural constant. For cases involving multiple data types, this function is used to calculate the probability density function for each data type separately.

[0117] For the Beta distribution method, the probability density function for all data types is: ,in ; x represents the specific value of each standard data point in the validation dataset to be analyzed; α and β are shape parameters, and their values ​​are... -1), -1), where μ is the mean of the data of the corresponding standard data type in the standard dataset. This represents the variance of the data corresponding to the standard data type in the standard dataset. For datasets containing multiple data types, this function is used to calculate the probability density function for each data type.

[0118] For the log-normal distribution method, the probability density function for all data types is: , where x is the specific value of each standard data in the validation dataset to be analyzed; lnx is the natural logarithmic transformation of x; μ is the mean of the natural logarithmic transformation of the data corresponding to the standard data type in the standard dataset; and σ is the standard deviation of the natural logarithmic transformation of the data corresponding to the standard data type in the standard dataset. For cases containing multiple data types, this function is used to calculate the probability density function for each data type.

[0119] For methods using empirical distributions, the probability density function for all data types is: ,in ; x represents the specific value of each standard data in the sample to be analyzed in the validation dataset; xi represents the data of the corresponding standard data type in each standard dataset; n is 50; This is an indicator function; it takes the value 1 if the condition is true, and 0 otherwise; Δ takes the value of... For cases involving multiple data types, this function is used to calculate the probability density function for each data type separately.

[0120] For methods using kernel density estimation, the probability density function for all data types is: Where x is the specific value of each standard data in the validation dataset to be analyzed, and n is 50; xi is the data of the corresponding standard data type in each standard dataset; h is a value of min() is the minimum value function, σ is the standard deviation of the data of the corresponding standard data type in the standard dataset, IQR is the interquartile range, and n is 50. For cases containing multiple data types, this function is used to calculate the probability density function of each data type.

[0121] For the mixed distribution method, the probability density function for all data types is: , where x is the specific value of each standard data of the sample to be analyzed in the validation dataset; K and The values ​​are automatically optimized by the mclust package in the R program; Let be the probability density function of the k-th subdistribution, which is selected from the probability density function of the normal distribution. For cases involving multiple data types, this function is used to calculate the probability density function for each data type separately.

[0122] For Monte Carlo simulation methods, the probability density functions for all data types are the same as those for density estimation methods.

[0123] For the Bootstrap method, the probability density function for all data types is: Where x is the specific value of each standard data of the sample to be analyzed in the validation dataset; B is the number of bootstraps, with a value of 50; b is the resampling index, which represents the b-th bootstrap; each The probability density function of the kernel density estimate from a single Bootstrap sample. For cases involving multiple data types, this function is used to calculate the probability density function for each data type separately.

[0124] 1.5 Calculating the probability of an event

[0125] The probability density function constructed using all the methods described above is used to sequentially calculate the different event probability function values ​​for all independent samples in the validation set, including:

[0126] Right tail probability: , where X represents the random variable of the sample PCR results in the constructed distribution; This is the corresponding probability density function; The threshold is the starting point for calculating the right-tail probability, specifically the measured value of the PCR result of the sample to be evaluated. The value of X is a function of the right-tail probability, i.e., X is greater than... The probability of;

[0127] Left tail probability: , where X represents the random variable of the sample PCR results in the constructed distribution; This is the corresponding probability density function; The threshold is the endpoint for calculating the left-tail probability, specifically the measured value of the PCR result of the sample to be evaluated. The function value of the left-tail probability, i.e., X is less than The probability of;

[0128] Two-tailed probability: , where X represents the random variable of the sample PCR results in the constructed distribution; Here, μ is the probability density function, used to describe the probability density of X at each point; μ is the mean, i.e., the mathematical expectation of the random variable X; and k is the magnitude of the deviation of the data to be evaluated from the mean, specifically... , These are the measured values ​​of the PCR results for the samples to be evaluated.

[0129] Different probability density functions and event probability integral formulas were combined to calculate the event probability that the PCR results of the target gene in the validation dataset belonged to the PCR results of real samples under each combination. When using one-tailed probability, the sample type was determined to be consistent with the real sample type if the event probability was greater than 0.05 and less than 0.95; otherwise, they were inconsistent. When using two-tailed probability, the sample type was determined to be consistent with the real sample type if the event probability was greater than 0.05; otherwise, they were inconsistent. The event probabilities were compared with the event probability thresholds. Based on whether the judgment result was consistent with the real sample, the accuracy of each combination in determining the samples in the validation dataset was calculated. The combination of probability density function and event probability integral formula corresponding to the highest accuracy in determining the samples in the validation dataset was selected as the optimal combination, used to calculate the event probability that the PCR results of the target gene in the unknown sample belonged to the PCR results of the target gene in the real sample.

[0130] The event probability threshold is further updated based on the optimal combination determined by the validation set: when choosing a one-tailed probability, the updated event probability threshold is less than min(P). max +0.05, 0.95) and greater than max(P min -0.05, 0.05), where P max To verify the maximum probability of events in the dataset, P min To verify the minimum event probability in the dataset, `max()` is the function to maximize the probability, and `min()` is the function to minimize the probability. When choosing a two-tailed probability, the updated event probability threshold is greater than `max(P)`. min -0.05, 0.05), where P min The minimum probability of the event is given by , and max() is the function to find the maximum probability.

[0131] 1.6 Validation Set Results

[0132] For validation set samples of different cancer types, we analyzed the different event probability function values ​​of these samples under different distribution analyses for different qPCR results. Specific analysis results are shown in Tables 3, 4, and 5. The analysis based on the validation set results demonstrates that, regardless of whether the samples are colorectal cancer, gastric cancer, or lung cancer, the mean trend of the event probability is similar when the same data type, distribution selection, and event probability calculation are used, proving the universality of this method for different genes. Under the same probability density function and event probability calculation conditions, the means of Data1, Data2, and Data3 in the validation dataset are all closer to the distribution center; that is, for one-tailed probabilities, the mean is closer to 0.5, and for two-tailed probabilities, the mean is closer to 1. This embodiment demonstrates that probability density functions based on different distribution types provide varying mean event probabilities. Specifically, the minimum values ​​of the left-tailed, right-tailed, and two-tailed probabilities calculated using probability density functions based on the normal, Beta, empirical, mixture, and Bootstrap methods are all greater than 0.05 across all data types. This means that the validation dataset samples analyzed using probability density functions based on these distributions can achieve 100% accuracy in sample type identification across all data types. However, the minimum value of the left-tailed probability calculated using the kernel density estimation probability density function is significantly lower in Data4 and Data5, even below the threshold, indicating a slightly weaker ability to distinguish samples compared to the aforementioned five methods. Finally, the minimum values ​​of the left-tailed, right-tailed, and two-tailed probabilities calculated using probability density functions based on the exponential, gamma, and log-normal distributions are mostly below the threshold to varying degrees across different data types, thus exhibiting the worst performance in distinguishing samples in the validation dataset.

[0133] In summary, the PCR result analysis system and method described in this invention preferably uses a single data type or a combination of multiple data types, such as Data1, Data2, and Data3, to calculate the probability density function based on the normal distribution, Beta distribution, empirical distribution, mixed distribution, and Bootstrap method, and uses left-tailed, right-tailed, or two-tailed probability analysis to evaluate the sample type of unknown samples.

[0134] Table 3: Summary of Event Probability Function Values ​​for Colorectal Cancer Samples

[0135]

[0136] Table 4: Summary of Event Probability Function Values ​​for Gastric Cancer Samples

[0137]

[0138] Table 5: Summary of Event Probability Function Values ​​for Lung Cancer Samples

[0139]

[0140] Example 2: Testing of real samples using the method described in this invention.

[0141] Ten samples of colorectal cancer, ten samples of gastric cancer, and ten samples of lung cancer were collected. Using the same DNA extraction, transformation, and qPCR detection procedures as in Example 1, raw qPCR results for 30 samples were obtained. The same qPCR data preprocessing steps were used to obtain the final qPCR data for analysis. Based on the results described in Section 1.6 of Example 1, sample type analysis was performed using three datasets: Data1, Data2, and Data3. Two-tailed probabilities were calculated using a probability density function based on a normal distribution to analyze the sample type of each sample. Based on the experimental results described in Section 1.6, the following event probability threshold was defined: the threshold for all sample types was the minimum event probability of the corresponding data type in the validation set minus 0.05. A sample was determined to belong to the current sample type if and only if the two-tailed probabilities of Data1, Data2, and Data3 were all greater than the defined threshold. The analysis results for all samples are shown in Table 6. All test set samples, whether Data1, Data2, or Data3, had two-tailed probabilities exceeding the event probability threshold, demonstrating the feasibility of using the PCR result analysis system and method described in this invention for PCR result analysis.

[0142] Table 6: Summary Table of Real Sample Test Results

[0143]

Claims

1. A PCR result analysis system based on statistical distribution determination, characterized in that... The system takes the PCR result data of the target gene of the unknown sample as input, analyzes the statistical distribution of the PCR result data of the target gene of the real sample with known sample type, constructs the probability density function corresponding to the PCR result data, calculates the event probability that the PCR result data of the target gene of the unknown sample belongs to the distribution of the PCR result data of the target gene of the real sample, compares the event probability with the event probability threshold, and outputs the result as whether the sample type of the unknown sample is consistent with the sample type of the real sample. The sample type is the source characteristic of the target gene, selected from disease, race, age, sex or other biological or environmental factors that affect gene characteristics; The PCR results data include one or more of the following types: the difference in CT values ​​between the target gene and the internal reference gene, the ratio of the maximum fluorescence signal amplification of the target gene and the internal reference gene, the ratio of the curvature of the fluorescence amplification of the target gene and the internal reference gene at the final amplification cycle, the ratio of the slope of the fluorescence amplification of the target gene and the internal reference gene at the final amplification cycle, and the ratio of the maximum fluorescence signal of the target gene and the internal reference gene; the unknown sample and the real sample are processed in the same way, the same target gene is amplified using the same primers and fluorescent probes, and the same PCR conditions are used to obtain the PCR results data.

2. The analysis system as described in claim 1, characterized in that, The probability density function is constructed using the following method: (1) Collect several real samples according to at least one sample type of the target gene, obtain the PCR result data of the target gene of all real samples, and clean and standardize the data to form a standard dataset. (2) Based on the standard dataset, construct the probability density function using parameterization, empirical distribution function, kernel density estimation, Monte Carlo simulation, mixed distribution or Bootstrap method.

3. The analysis system as described in claim 1, characterized in that... The probability density function is selected from the following: Based on normal distribution: , where x is the value of a random variable, i.e., the PCR result data of an unknown sample; σ is the probability density function value corresponding to x; μ is the mean, i.e., the center of the normal distribution; σ is the standard deviation, used to measure the dispersion and fluctuation of the distribution; exp() is the exponential function with the natural constant e as the base. Based on the exponential distribution: , where x is the value of a random variable, i.e., the PCR result data of an unknown sample; Let x be the probability density function value; λ be the rate parameter, representing the average number of times the event occurs per unit time; and e be the natural constant. Based on gamma distribution: Where Γ(k) is the gamma function, , where x is the value of a random variable, i.e., the PCR result data of an unknown sample; x represents the probability density function value; k is a shape parameter that controls the shape of the distribution curve. This is a scale parameter that controls the scaling of the distribution along the x-axis; The gamma function is used for normalization, making the sum of integrals equal to 1. e is a natural constant; Based on the Beta distribution: in , where x is the value of a random variable, i.e., the PCR result data of an unknown sample, with a value range of [0, 1]; α is a shape parameter used to control the shape of the distribution near the 0 side; β is a shape parameter used to control the shape of the distribution near the 1 side; is the beta function, and is the normalization constant, which guarantees that the integral is 1. Based on the log-normal distribution: , where x is the value of the random variable, i.e., the PCR result data of the unknown sample; lnx is the natural logarithm function, i.e., the logarithmic transformation of x; μ is the logarithmic mean, i.e. the mathematical expectation of lnx; σ is the logarithmic standard deviation, i.e. the standard deviation of lnx; This is a normalization constant, used to ensure that the integral equals 1. Based on empirical distribution: ,in , where x is the value of a random variable, i.e., the PCR result data of an unknown sample; This represents the i-th observation; n is the number of samples. This is an indicator function; it takes the value 1 if the condition is true, and 0 otherwise. The cumulative distribution proportion of the sample is the empirical distribution function; Δ represents an extremely small change. The value is the probability density function value. Based on kernel density estimation: , where x is the value of the random variable, i.e., the PCR result data of the unknown sample; n is the number of samples; This represents the PCR result data for the i-th observed sample; h>0 represents the bandwidth, used to control the smoothing level; e is the natural constant. Based on mixed distribution: , where x is the value of the random variable, i.e., the PCR result data of the unknown sample; K is the number of sub-distributions, i.e. the number of components in the mixture distribution; The weight is the proportion of the k-th sub-distribution; Let be the probability density function of the k-th sub-distribution, selected from any of the aforementioned probability density functions based on the normal distribution, exponential distribution, gamma distribution, beta distribution, log-normal distribution, empirical distribution, or kernel density estimation. Based on Monte Carlo simulation: it does not define a function itself, but relies on a large number of samples { To approximate the target distribution, kernel density estimation is used to obtain an approximate probability density function. , where x is the value of the random variable, i.e., the PCR result data of the unknown sample; n is the number of samples; This represents the PCR result data for the i-th observed sample; h>0 represents the bandwidth, used to control the smoothing level; e is the natural constant. Bootstrap-based approach: Where x is the value of a random variable, i.e., the PCR result data of an unknown sample; B is the number of bootstraps; b is the resampling index, i.e., the b-th bootstrap; each Kernel density estimation from a Bootstrap sample.

4. The analysis system as described in claim 3, characterized in that, The smoothness of the kernel density estimate is controlled using the Silverman rule, calculated as follows: min() is the minimum value function, σ is the standard deviation of the x dataset, and IQR is the interquartile range.

5. The analysis system as described in claim 1, characterized in that... The event probabilities mentioned are selected from one-tailed probabilities and two-tailed probabilities.

6. The analysis system as described in claim 5, characterized in that, The integral formula for the probability of an event is selected from: Right tail probability: , where X represents the random variable of the sample PCR results in the constructed distribution; This is the corresponding probability density function; The threshold is the starting point for calculating the right-tail probability, specifically the measured value of the PCR result of the sample to be evaluated. The value of X is a function of the right-tail probability, i.e., X is greater than... The probability of; Left tail probability: , where X represents the random variable of the sample PCR results in the constructed distribution; This is the corresponding probability density function; The threshold is the endpoint for calculating the left-tail probability, specifically the measured value of the PCR result of the sample to be evaluated. The function value of the left-tail probability, i.e., X is less than The probability of; Two-tailed probability: , where X represents the random variable of the sample PCR results in the constructed distribution; Here, μ is the probability density function, used to describe the probability density of X at each point; μ is the mean, i.e., the mathematical expectation of the random variable X; and k is the magnitude of the deviation of the data to be evaluated from the mean, specifically... , These are the measured values ​​of the PCR results for the samples to be evaluated.

7. The analysis system as described in claim 5, characterized in that... When the event probability is a one-tailed probability, the event probability threshold is 0.05~0.

95. When the event probability is greater than 0.05 and less than 0.95, the sample type of the unknown sample is determined to be consistent with the sample type of the real sample; otherwise, they are inconsistent. When the event probability is a two-tailed probability, the event probability threshold is 0.

05. When the event probability is greater than 0.05, the sample type of the unknown sample is determined to be consistent with the sample type of the real sample; otherwise, they are inconsistent.

8. The analysis system according to any one of claims 1 to 7, characterized in that... The system includes: (1) Data processing module: Processes the raw PCR results of unknown samples or real samples to obtain PCR result data; (2) Standard dataset construction module: Clean and standardize the target gene PCR results data of real samples to construct a standard dataset; (3) Probability density function construction module: Based on the standard dataset, the probability density function is constructed using parameterization methods, empirical distribution functions, kernel density estimation, Monte Carlo simulation, mixture distribution or Bootstrap method; (4) Event probability calculation module: Integrate the probability density function to calculate the event probability that the PCR result data of the target gene of the unknown sample belongs to the distribution of the PCR result data of the target gene of the real sample; (5) Result output module: compare the event probability with the event probability threshold, and output whether the sample type of the unknown sample is consistent with the sample type of the real sample.

9. The analysis system according to any one of claims 1 to 7, characterized in that... It also includes a validation dataset, which consists of real sample PCR result data of target genes outside the standard dataset. Different probability density functions and event probability integral formulas are selected and combined to calculate the event probability that the PCR result data of the target gene in the validation dataset belongs to the PCR result data of the real sample under each combination. The event probability is compared with the event probability threshold. Based on whether the judgment result is consistent with the real sample, the accuracy of the judgment result of each combination on the sample in the validation dataset is calculated. The combination of probability density function and event probability integral formula corresponding to the highest judgment accuracy of the sample in the validation dataset is selected as the optimal combination, which is used to calculate the event probability that the PCR result data of the target gene in the unknown sample belongs to the PCR result data distribution of the target gene in the real sample.

10. The analysis system as described in claim 9, characterized in that... It also includes further updating the event probability threshold based on the optimal combination, using the updated event probability threshold to determine whether the sample type of the unknown sample is consistent with the sample type of the true sample. The event probability threshold is updated using the following method: when choosing the one-tailed probability, the updated event probability threshold is less than min(P). max +0.05, 0.95) and greater than max(P min -0.05, 0.05), where P max To verify the maximum probability of events in the dataset, P min To verify the minimum event probability in the dataset, `max()` is the function to maximize the probability, and `min()` is the function to minimize the probability. When choosing a two-tailed probability, the updated event probability threshold is greater than `max(P)`. min -0.05, 0.05), where P min The minimum probability of the event is given by , and max() is the function to find the maximum probability.

11. The analysis system as described in claim 9, characterized in that... The system includes: Data processing module: processes the raw PCR results of unknown or real samples to obtain PCR result data; Standard dataset construction module: Select target gene PCR results data from a portion of real samples, clean and standardize them to construct a standard dataset; Validation dataset construction module: Cleans and standardizes the target gene PCR results data from real samples outside the standard dataset to construct the validation dataset; Probability density function set construction module: Based on the standard dataset, several probability density functions are constructed using parametric methods, empirical distribution functions, kernel density estimation, Monte Carlo simulation, mixture distribution, or Bootstrap methods to obtain the probability density function set; Event probability set construction module: Combines different event probability integral formulas with probability density functions in the probability density function set, calculates and outputs the event probability that the PCR result data of the target gene in the validation dataset belongs to the PCR result data distribution of the target gene in the real sample under different combinations, and obtains the event probability set; Event probability calculation module: compares the event probability with the event probability threshold, calculates the accuracy of the judgment result of each combination pair of the sample judgment result in the verification dataset based on whether the judgment result is consistent with the real sample, selects the combination of probability density function and event probability integral formula corresponding to the highest judgment accuracy of the sample in the verification dataset as the optimal combination, and is used to calculate the event probability of the PCR result data of the target gene of the unknown sample belonging to the PCR result data distribution of the target gene of the real sample. The results output module compares the event probability of the unknown sample with the event probability threshold and outputs whether the sample type of the unknown sample is consistent with the sample type of the real sample.

12. The analysis system as described in claim 11, characterized in that... The result output module updates the event probability threshold based on the optimal combination before comparison: when a one-tailed probability is selected, the updated event probability threshold is less than min(P). max +0.05, 0.95) and greater than max(P min -0.05, 0.05), where P max To verify the maximum probability of events in the dataset, P min To verify the minimum event probability in the dataset, `max()` is the function to maximize the probability, and `min()` is the function to minimize the probability. When choosing a two-tailed probability, the updated event probability threshold is greater than `max(P)`. min -0.05, 0.05), where P min The minimum event probability is given by 'max()', and the maximum event probability is given by 'max()'. The event probability of the unknown sample is compared with the updated event probability threshold, and the output shows whether the sample type of the unknown sample is consistent with the sample type of the real sample.

13. A method for analyzing PCR results based on statistical distribution determination, characterized in that... Based on at least one sample type to which the target gene belongs, collect several corresponding real samples, use statistical methods to obtain the probability distribution of PCR result data of the target gene in the real samples, construct a probability density function, and calculate the event probability that the PCR result data of the target gene in the unknown sample belongs to the PCR result data distribution of the target gene in the real sample. The sample type is the source characteristic of the target gene, selected from disease, race, age, sex or other biological or environmental factors that affect gene characteristics. The PCR results data are selected from one or more of the following types: the difference in CT values ​​between the target gene and the internal reference gene, the ratio of the maximum fluorescence signal amplification of the target gene and the internal reference gene, the ratio of the curvature of the fluorescence amplification of the target gene and the internal reference gene at the final amplification cycle, the ratio of the slope of the fluorescence amplification of the target gene and the internal reference gene at the final amplification cycle, and the ratio of the maximum fluorescence signal of the target gene and the internal reference gene; the unknown sample and the real sample are treated in the same way, the same amplification primers and fluorescent probes are used for the same target gene, and the same PCR conditions are used to obtain the PCR results data.

14. The analytical method as described in claim 13, characterized in that... Includes the following steps: (1) Constructing a standard dataset: Select real samples for the sample type of the sample to be tested, with no less than 50 samples of each real sample. Use a unified experimental protocol to conduct PCR experiments and obtain PCR result data. Clean and standardize the PCR result data to construct a standard dataset. (2) Based on the standard dataset, construct the probability density function using parameterization, empirical distribution function, kernel density estimation, Monte Carlo simulation, mixture distribution or Bootstrap method; (3) Obtain various PCR result data of unknown samples using the same experimental methods and conditions as real samples; (4) Integrate the probability density function to calculate the probability of an event that the PCR result data of the unknown sample target gene belongs to the distribution of the PCR result data of the real sample target gene.

15. The analytical method as described in claim 14, characterized in that... Includes the following steps: (1) Constructing standard dataset and validation dataset: Select real samples for the sample type of the sample to be tested, with no less than 50 samples of each real sample. Use a unified experimental protocol to conduct PCR experiments and obtain PCR result data. Clean and standardize the PCR result data, select some PCR result data to construct a standard dataset, and select PCR result data other than the standard dataset to construct a validation dataset. (2) Based on the standard dataset, construct a set of probability density functions using several of the following methods: parameterization, empirical distribution function, kernel density estimation, Monte Carlo simulation, mixture distribution, or Bootstrap method; (3) Combine different event probability integral formulas with probability density functions in the probability density function set, calculate and output the event probability of the distribution of the PCR result data of the target gene of the verification dataset belonging to the PCR result data of the target gene of the real sample, and obtain the event probability set. (4) Obtain various PCR result data of unknown samples using the same experimental methods and conditions as real samples; (5) Compare the event probabilities in the event probability set with the event probability threshold. Calculate the accuracy of each combination in judging the sample in the verification dataset based on whether the judgment result is consistent with the real sample. Select the combination of probability density function and event probability integral formula corresponding to the highest accuracy of judging the sample in the verification dataset as the optimal combination. This combination is used to calculate the event probability that the PCR result data of the target gene of the unknown sample belongs to the PCR result data distribution of the target gene of the real sample.

16. The analytical method according to any one of claims 13 to 15, characterized in that... The method further includes comparing the event probability with the event probability threshold to determine whether the sample type of the unknown sample is consistent with the sample type of the real sample.

17. The analytical method as described in claim 16, characterized in that... The event probabilities are updated based on the optimal combination: when a one-tailed probability is selected, the threshold for updating the event probability is less than min(P). max +0.05, 0.95) and greater than max(P min -0.05, 0.05), where P max To verify the maximum probability of events in the dataset, P min To verify the minimum event probability in the dataset, `max()` is the function to maximize the probability, and `min()` is the function to minimize the probability. When choosing a two-tailed probability, the updated event probability threshold is greater than `max(P)`. min -0.05, 0.05), where P min The minimum event probability is given by 'max()', which is the maximum value function. The updated event probability threshold is used to determine whether the sample type of the unknown sample is consistent with the sample type of the real sample.