Neural network model-based multi-dimensional QPCR result analysis system

By combining neural network models with multi-dimensional qPCR results data, the limitations of the single Ct value in traditional qPCR analysis are overcome, enabling more accurate quantitative and qualitative analysis of samples and improving the accuracy of detection and classification.

WO2026157219A1PCT designated stage Publication Date: 2026-07-30SUZHOU HAIMIAO BIOTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SUZHOU HAIMIAO BIOTECH CO LTD
Filing Date
2025-08-26
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Traditional qPCR result analysis is highly dependent on Ct values, which are easily affected by experimental conditions and cannot accurately reflect the true characteristics of the sample. This is especially true when there are discrepancies in amplification efficiency between low copy number target molecules and samples, leading to result bias and inaccurate classification.

Method used

A multi-dimensional qPCR result analysis system based on a neural network model is adopted. It uses multi-dimensional data such as CT value, fluorescence signal value, fluorescence signal amplification value, curvature and slope, combined with a neural network model to predict the sample type. The detection sensitivity and classification accuracy are improved through training and validation.

Benefits of technology

It significantly improves the sensitivity of detection and classification accuracy, enabling more accurate quantitative and qualitative analysis of samples, especially capturing subtle differences between samples under conditions of weak differences, and improving the model's generalization ability and applicability to unknown samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025116933_30072026_PF_FP_ABST
    Figure CN2025116933_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a QPCR result analysis neural network model and system. At least three types of QPCR result data selected among CT values, fluorescence signal values at all different cycle numbers during PCR, fluorescence signal increment values at all the different cycle numbers, the curvature of changes in the fluorescence signal increment values at all the different cycle numbers during the PCR, and the slope of the changes in the fluorescence signal increment values at all the different cycle numbers are used as inputs, and a sample type of a target gene corresponding to QPCR result data is used as a result for output, thereby predicting a sample type of a target gene of an unknown sample. By using multi-dimensional QPCR result data as an analysis object, and in light of the neural network model, a neural network model-based multi-dimensional QPCR result analysis system is provided, which captures subtle differences in QPCR results among samples, thereby significantly improving detection sensitivity and classification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

A multi-dimensional QPCR result analysis system based on a neural network model Technical Field

[0001] This invention belongs to the field of bioinformatics technology, specifically relating to a computer algorithm based on a neural network model that can be applied to the analysis of qPCR results. It is suitable for qualitative or quantitative analysis of singleton or multiplex qPCR results and can ultimately be applied to disease screening and classification. Background Technology

[0002] Quantitative Polymerase Chain Reaction (QPCR) is a molecular biology technique widely used in gene expression analysis, genetic mutation detection, copy number variation analysis, methylation variation analysis, and microbial detection. QPCR combines the high sensitivity of traditional PCR with the advantage of real-time detection, achieving quantitative detection of target DNA or RNA through the accumulation of fluorescence signals. Its core principle is to use fluorescent dyes (such as SYBR Green) or specific fluorescent probes (such as TaqMan probes) to monitor the accumulation of amplified products in real time during each PCR cycle.

[0003] Quantitative PCR (qPCR) is characterized by high sensitivity, high specificity, and speed, enabling the detection of extremely low concentrations of nucleic acid target molecules. Furthermore, due to its relatively simple operation and high degree of automation, qPCR technology is widely used in scientific research, medical diagnostics, food safety, and environmental monitoring.

[0004] However, while traditional qPCR results analysis provides an important tool for gene expression and nucleic acid quantification, it also has many limitations, which ultimately affect the accuracy and reliability of the results.

[0005] Traditional qPCR result analysis heavily relies on the Ct value (cycle threshold) to determine the expression level of the target gene or the nucleic acid concentration, but this method has significant limitations. First, the Ct value is easily affected by experimental conditions, such as amplification efficiency, inhibitors in the reaction system, template quality, and background noise of the fluorescence signal, leading to result bias. Second, for target molecules with low copy numbers, the Ct value fluctuates more, potentially resulting in unstable results or difficulty in distinguishing background signals. Furthermore, the Ct value is inherently a relative value, reflecting the cycle number at which the fluorescence signal reaches a specific threshold, rather than the absolute amount of the target molecule; its interpretation is highly dependent on the consistency of amplification efficiency. Especially when there are slight differences in amplification efficiency between different samples, comparing only Ct values ​​can lead to misleading conclusions.

[0006] Similarly, traditional qPCR analysis methods have significant limitations when detecting DNA origin characteristics, especially when accurate classification of unknown samples is required. For example, when determining whether a sample belongs to a healthy individual or a patient by detecting the methylation rate of a gene in blood, traditional qPCR data analysis methods face multiple challenges. First, qPCR is highly sensitive to amplification efficiency and the quality of the starting template. Complex backgrounds in different samples (such as the presence of inhibitors in the blood or the degree of DNA degradation) can lead to inconsistencies in amplification efficiency, thus affecting the accuracy of quantitative results. Furthermore, because methylation levels typically exhibit a small range of variation, traditional qPCR has limited sensitivity to detect low-abundance target molecules, potentially leading to unclear distinctions between normal and patient samples. Finally, traditional qPCR analysis methods lack multi-dimensional data support; relying solely on the Ct value as a single indicator cannot comprehensively reflect the true characteristics of a sample, thereby reducing the accuracy and reliability of sample classification. Summary of the Invention

[0007] This invention addresses the shortcomings of existing technologies that rely solely on the Ct value for analysis. By using multidimensional qPCR results data as the analysis object and combining it with a neural network model, it provides a multidimensional qPCR results analysis system based on a neural network model. By capturing subtle differences in qPCR results between samples, it significantly improves the sensitivity and classification accuracy of detection, enabling accurate quantification of samples and accurate output of multidimensional information.

[0008] The specific technical solution of this invention is as follows:

[0009] A neural network model for qPCR result analysis is provided. The neural network model is a single neural network model or a stacked neural network model combining multiple neural network models. It takes at least three types of qPCR result data as input and outputs the sample type of the target gene corresponding to the qPCR result data. The qPCR result data includes the following types: CT value, fluorescence signal value at all different cycles during PCR, fluorescence signal amplification value at all different cycles, curvature of fluorescence signal amplification value change at all different cycles during PCR, and slope of fluorescence signal amplification value change at all different cycles. The neural network model is trained and validated using a standard database and is used to predict the target gene sample type of unknown samples. The sample type is a quantitative feature of the target gene or a source feature of the target gene.

[0010] The standard database includes a quantitative database, a qualitative database, or a combination of both.

[0011] The quantitative database is used for samples of the quantitative characteristics of target genes. The construction method is as follows: prepare DNA templates of target genes with different quantification levels, use target gene primers and specific probes to perform QPCR experiments, obtain QPCR result data of target genes with different quantification levels, and use the quantification value of the target gene as a label for each QPCR result data.

[0012] The qualitative database is used when the sample type is the source characteristic of the target gene. The construction method is as follows: collect several positive real samples corresponding to the source characteristics of the target gene according to at least one source characteristic. Process all positive real samples in the same way to obtain the corresponding DNA template. Perform QPCR experiments on each DNA template using the same target gene primers and probes and QPCR conditions to obtain QPCR result data corresponding to each DNA template. Each QPCR result data is tagged with the source characteristic of the target gene.

[0013] The neural network model of the present invention uses the quantification characteristics of the target gene as methylation rate, DNA copy number, DNA mutation state, or DNA concentration, and the source characteristics of the target gene are disease, race, age, sex, or other biological or environmental factors that affect gene characteristics.

[0014] When the sample type is the quantitative feature of the target gene, the output layer of the neural network model does not require an activation function, and the output result is a quantified value, such as methylation rate, DNA copy number, DNA mutation state, or DNA concentration.

[0015] When the sample type is the source feature of the target gene, the neural network model outputs the probability of the source feature of the target gene through the activation function softmax.

[0016] Preferably, the neural network model is a fully connected neural network model, consisting of an input layer, several hidden layers, and an output layer. The input layer contains all or some types of QPCR result data. The hidden layers are fully connected layers containing several neurons. The output layer is a fully connected layer containing one or n neurons. One is the output layer when applying a quantitative database, n is the number of sample types, and n is the output layer when applying a qualitative database.

[0017] Preferably, the neural network model is selected from one of the following neural network models A and D, or is a stacked neural network model E composed of at least two of them. Neural network model A: normalized CT value, normalized dynamic fluorescence signal value of all cycles, and normalized dynamic fluorescence signal amplification value of all cycles are selected as the input layer of neural network model A, and the sample type of the target gene is selected as the output layer.

[0018] Neural Network Model B: The curvature and slope of all cycles of normalized CT value and fluorescence signal amplification value change are selected as the input layer of Neural Network Model B, and the sample type of the target gene is selected as the output layer.

[0019] Neural network model C: The normalized CT value, the normalized final cycle fluorescence signal value, and the fluorescence signal amplification value are selected as the input layer of neural network model C, and the sample type of the target gene is selected as the output layer.

[0020] Neural network model D: The curvature and slope of the final cycle of the normalized CT value and the change of fluorescence signal amplification value are selected as the input layer of neural network model D, and the sample type of the target gene is selected as the output layer.

[0021] Stacked Neural Network Model E: Select the results of at least two of the neural network models A, B, C and D as the input layer of neural network model E, and the sample type of the target gene as the output layer.

[0022] Preferably, the neural network model is a stacked neural network model that combines the above-mentioned neural network models A, B, C and D.

[0023] Internal reference genes are used to correct for potential technical biases during experiments, such as differences in RNA extraction volume and PCR efficiency, thereby normalizing data. They eliminate experimental variability, improve the comparability of results between different experiments, and enhance data reliability. Simultaneously, internal reference genes can correct for differences in sample sources caused by experimental procedures (such as RNA extraction efficiency, reverse transcription efficiency, total RNA initiation volume, etc.) or other non-biological factors, making quantitative or qualitative analyses based on differences in qPCR results more accurate. Ideally, internal reference genes should be stably expressed under all experimental conditions; their selection and validation are crucial to ensuring reliable experimental results.

[0024] Preferably, the quantitative database of the present invention also includes qPCR result data of the internal reference gene. Target gene DNA templates and internal reference gene templates (e.g., DNA copy number or DNA concentration) with different quantification levels are prepared. Target gene DNA templates with different quantification levels (e.g., copy number / concentration gradient) are mixed pairwise with internal reference gene DNA templates with different quantification levels to obtain a series of DNA template samples of target genes and internal reference genes with different quantification levels. qPCR experiments are performed on each sample using primers and specific probes for the target genes and internal reference genes. Each sample yields a set of qPCR result data for the target genes and internal reference genes. Each set of qPCR result data is labeled with the quantification level of the target gene (e.g., copy number / concentration value). The probes used for the target genes and internal reference genes have different fluorescence channels. When training and validating the neural network model using the quantitative database, a set of qPCR result data for the target genes and internal reference genes is used as input, and the data types of the qPCR result data for the target genes and internal reference genes are consistent.

[0025] The qualitative database described in this invention can minimize the influence of technical variations or non-biological factors in the experiment by strictly controlling the total amount of each DNA sample loaded, or by using internal controls to correct for the influence caused by experimental operations (such as RNA extraction efficiency, reverse transcription efficiency, total RNA initiation amount, etc.) or other non-biological factors.

[0026] Preferably, the qualitative database of the present invention also contains qPCR result data of the internal reference gene. Using primers and probes of the target gene and the internal reference gene, qPCR experiments are performed on each DNA template. Each DNA template yields a set of qPCR result data for the target gene and the internal reference gene. Each set of qPCR result data is tagged with the source characteristics of the target gene. When using the qualitative database to train and validate a neural network model, a set of qPCR result data for the target gene and the internal reference gene is used as input, and the data types of the qPCR result data for the target gene and the internal reference gene are consistent.

[0027] In the neural network model of this invention, the qualitative database further uses samples with negative source characteristics of the target gene as a control, and collects several real samples with negative source characteristics of the target gene. The processing method for obtaining DNA templates for all negative real samples, as well as the primers and probes used for qPCR experiments and the qPCR conditions, are the same as those for positive real samples. QPCR result data corresponding to each DNA template is obtained, and each qPCR result data is labeled with the source characteristics of the target gene.

[0028] The target genes in the qualitative database of the neural network model described in this invention are single genes (corresponding to specific markers for their respective sample types) or combinations of several genes (corresponding to combinations of specific markers for their respective sample types). Preferably, the number of target genes is 1 to 10. When the target genes are combinations of several genes, the qPCR results of the gene combination can be obtained using the same fluorescence channel, or the qPCR results of each gene can be obtained using different fluorescence channels.

[0029] The different sample types mentioned in this invention can refer to different types of factors: differences in diseases, races, ages, sexes, or other biological or environmental factors that affect genetic characteristics. They can also be further subdivisions of the same type of factor. For example, when the sample type is a disease, the different sample types mentioned in this invention refer to different diseases (such as lung cancer, liver cancer, stomach cancer, colorectal cancer, etc.); when the sample type is a race, the different sample types mentioned in this invention refer to different races (such as Asian, Indian, Caucasian, Native American, etc.).

[0030] When the qualitative database described in this invention contains samples with characteristics from multiple sources, it can output sample type results for different types of factors for the same sample's QPCR results data. For example, it can output diverse information about the sample regarding disease, race, age, sex, or other biological or environmental factors that affect genetic characteristics. Alternatively, it can output different sample type results for further subdivision under the same type of factor for the same sample's QPCR results data, i.e., the relationship between the sample and the subdivision under the same type of factor. For example, it can output the correlation between the sample and various diseases such as lung cancer, liver cancer, stomach cancer, and colorectal cancer, or the correlation between the sample and ethnic groups such as Asian Americans, Indians, Caucasians, and Native Americans. Alternatively, it can output sample type results for different types of factors and different sample type results for further subdivision under the same type of factor for the same sample's QPCR results data.

[0031] The definition of sample types described in this invention should be based on explicit clinical or experimental criteria, rather than solely relying on molecular characteristics or gene expression levels. For example, the determination of the type of a disease sample should rely on authoritative diagnostic methods, such as imaging examinations, pathological analysis, culture results, or other directly confirmatory methods, rather than simply inferring based on gene mutations, expression levels, or molecular markers associated with the disease. Similarly, for pathogen-related samples, their type should be based on explicit isolation and identification, specific detection, or other standardized experiments, rather than simply relying on the detection results of certain specific molecular characteristics.

[0032] In order to meet the training needs of the neural network model, the standard database uses no less than 100 samples for each sample type, and the qualitative databases for different sample types may use the same or different numbers of real samples.

[0033] Another objective of this invention is to provide a qPCR result analysis system, the system comprising a data acquisition module, an analysis module, and a result display module;

[0034] The data acquisition module is used to process and acquire qPCR result data, which includes the following types: CT value, fluorescence signal value at all different cycles during PCR, fluorescence signal amplification value at all different cycles, curvature of fluorescence signal amplification value change at all different cycles during PCR, and slope of fluorescence signal amplification value change at all different cycles. The analysis module is the neural network model described in this invention, which takes at least three types of qPCR result data as input and the sample type of the target gene as the result output. The sample type is the quantitative feature of the target gene or the source feature of the target gene.

[0035] The result display module displays the output results of the analysis module.

[0036] The qPCR result analysis system of this invention directly generates the CT value, fluorescence amplification value for each cycle, and fluorescence signal value for each cycle of the qPCR result data from the qPCR program system. The data acquisition module directly exports the data through the qPCR program. The data acquisition module calls the R program and uses the smooth.spline() function to perform spline smoothing fitting on the curve formed by the cycle number and fluorescence amplification. The slope corresponding to each cycle is calculated using the formula slope = f′(x). The curvature corresponding to each cycle is calculated using the formula k(x) = abs(f″(x)) / ((f′(x)^2)+1)^3 / 2, where slope is the slope function value, f′(x) is the first derivative at the fluorescence amplification point of the corresponding cycle, k(x) is the curvature function value, abs() is the absolute value function, and f″(x) is the second derivative at the fluorescence amplification point of the corresponding cycle.

[0037] Another objective of this invention is to provide a method for analyzing qPCR results, which uses the neural network model or qPCR result analysis system described in this invention to analyze qPCR result data of unknown samples, specifically including:

[0038] (1) Construct a standard database based on the target genes to be tested and the sample types of unknown samples;

[0039] (2) Use standard databases to train and validate neural network models to build neural network models;

[0040] (3) Perform qPCR experiments on unknown samples to obtain qPCR result data. The conditions of the qPCR experiment (including the amplification primers and probes of the target gene, the amplification primers and probes of the internal control if an internal control is used, sample preparation, PCR reaction system, qPCR reaction program, qPCR instrument and reaction conditions, etc.) are consistent with the qPCR experimental conditions used to construct the standard database.

[0041] (4) Call the neural network model constructed in step (2), input the QPCR result data of the unknown sample into the neural network model and run it. The neural network model outputs the sample type of the target gene of the unknown sample as the result. The QPCR result data of the unknown sample input is consistent with the data type and quantity of the QPCR result data input by the input layer in the neural network model training process.

[0042] The preferred step (2) of the method described in this invention includes adjusting the hyperparameters of the neural network model. By adjusting the hyperparameters in the neural network model, the model validation results are optimized, and overfitting is avoided.

[0043] The neural network model, system, or method described in this invention is applicable to various qPCR programs. As a specific example, the qPCR program includes an enzyme activation and cycling program at 95°C for five minutes. The cycling program includes denaturation: 95°C for 5-30 seconds, annealing: 5-30 seconds, extension: 72°C for 20-30 seconds, and the cycling program is performed for 40-50 cycles. Fluorescence signals are collected during the extension process at 72°C.

[0044] Advantages of this invention:

[0045] (1) This invention uses multi-dimensional qPCR result data as the analysis object combined with a neural network. It takes at least three types of qPCR result data as input and the corresponding sample type as the output. The qPCR result data includes: CT value, fluorescence signal values ​​at all different cycles during PCR, fluorescence signal amplification values ​​at all different cycles, curvature of fluorescence signal amplification value changes at all different cycles during PCR, and slope of fluorescence signal amplification value changes at all or terminated cycles. This solves the problem that existing technologies relying solely on the single indicator of Ct value cannot comprehensively reflect the true characteristics of the sample. This invention uses multi-dimensional data analysis combined with a neural network model to capture more subtle differences between samples, especially when there are only slight changes in target gene expression levels or methylation rates, significantly improving detection sensitivity and classification accuracy.

[0046] (2) This invention uses a large number of real samples during the training and validation process of the neural network model, which can learn and extract the complex relationships between the features of different samples, thereby effectively covering the diversity of various unknown samples. This large-scale data-driven method not only improves the generalization ability of the model, but also reduces the bias caused by differences in experimental conditions, and is more applicable to unknown samples.

[0047] (3) The present invention uses a stacked neural network model combining multiple neural network models to further improve the accuracy of predicting the target gene sample type of unknown samples. Attached Figure Description

[0048] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly described below.

[0049] Obviously, the accompanying drawings described below are only some of the drawings of the embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort, but these other drawings are also within the scope of the drawings required for the embodiments of the present invention.

[0050] Figure 1 is a flowchart of the qPCR result analysis system of the present invention.

[0051] Figure 2 is a qPCR amplification curve of the target gene in Example 1 of the present invention.

[0052] Figure 3 is a standard curve of qPCR amplification of the target gene in Example 1 of the present invention.

[0053] Figure 4 is a qPCR amplification curve of the target gene for the detection of unknown samples in Example 1 of the present invention.

[0054] Figure 5 is a qPCR amplification curve of the internal reference gene for the detection of unknown samples in Example 1 of the present invention.

[0055] Figure 6 is a qPCR amplification curve of the internal reference gene for the detection of unknown samples in Example 2 of the present invention.

[0056] Figure 7 is a qPCR amplification curve of the target gene for the detection of unknown samples in Example 2 of the present invention. Detailed Implementation

[0057] To make the objectives, technical solutions, beneficial effects, and significant advancements of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings provided in the embodiments of the present invention.

[0058] Obviously, all the embodiments described are only some embodiments of the present invention, and not all embodiments; based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0059] Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0060] It should also be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0061] The technical solution of the present invention will now be described in detail with reference to specific embodiments.

[0062] Example 1: Comparison of the effects of conventional qPCR and the qPCR analysis method described in this invention in analyzing DNA sample characteristics

[0063] The purpose of this embodiment is to demonstrate that, when performing quantitative detection of samples with different DNA concentrations, the QPCR result analysis method described in this invention is more accurate and stable in calculating DNA sample concentration than the traditional QPCR result analysis method.

[0064] 1. The qPCR analysis method described in this invention:

[0065] 1.1 Preparation of Standard Samples

[0066] In this embodiment, the synthesized DNA template 1 sequence is used as the target gene template, and the synthesized DNA template 2 sequence is used as the internal reference gene template. The sequence information is shown below:

[0067] DNA template 1:

[0068] DNA template 2:

[0069] The synthesized DNA templates 1 and 2 were diluted to the following concentrations: 1.5E-6 ng / μL, 0.75E-6 ng / μL, 0.375E-6 ng / μL, 0.1875E-6 ng / μL, 0.9375E-7 ng / μL, 0.46875E-7 ng / μL, 0.234375E-7 ng / μL, and 0.1171875E-7 ng / μL. Then, each concentration of DNA template 1 was individually mixed with each concentration of DNA template 2 at a 1:1 ratio to obtain DNA templates 3–66. Each template from 3 to 66 was prepared in 100 replicates. A total of 6400 standard samples were prepared, containing mixed templates of target genes at 8 concentration gradients and internal reference genes at 8 target concentrations.

[0070] 1.2 Constructing a quantitative database

[0071] Primers and probes for detecting the target gene and internal reference gene were synthesized. The sequence information of the forward primer F1, reverse primer R1, and probe P1 for detecting DNA template 1 is shown below:

[0072] F1: 5'-CTCATACTACTGGCTCTACAAGGCCG-3'

[0073] R1: 5'-GTACCCCTTGGTTGCTTCCCTAG-3'

[0074] P1: 5'-FAM-CCAGAAACCTTCCCGACTGTCTCACAGTTT-BHQ-3'

[0075] The sequence information of the forward primer F2, reverse primer R2, and probe P2 used to detect DNA template 2 is shown below:

[0076] F2: 5'-CATAACCTTCCCGCGACCG-3'

[0077] R2: 5'-GTAGAAGCCCTCTTCCCTGCC-3'

[0078] P2: 5'-VIC-AAAGAGCTCATAAAGGTGG-MGB-3'

[0079] Configure the qPCR reaction system as shown in Table 1 below, and perform qPCR experiments on DNA templates 3-66. The qPCR program is shown in Table 2 below. The qPCR instrument used is an ABI 7500 qPCR instrument. Adjust the qPCR results by setting the baseline to 3-20 cycles, setting the FAM channel threshold to 50000, and setting the VIC channel threshold to 5000.

[0080] Table 1: qPCR reaction system configuration table

[0081] Table 2: qPCR reaction procedure table

[0082] All qPCR results were exported. The CT values, fluorescence amplification values ​​for each cycle, and fluorescence signal values ​​for each cycle could be directly obtained using the qPCR program. The curvature and slope of the amplification curve at each cycle were calculated by inputting the cycle number and its corresponding fluorescence amplification value into the analysis system. The `smooth.spline()` function was used to perform spline smoothing on the curve formed by the input cycle number and fluorescence amplification. The slope for each cycle was calculated using the formula `slope = f′(x)`, and the curvature for each cycle was calculated using the formula `k(x) = abs(f″(x)) / ((f′(x)^2) + 1)^3 / 2`. Here, `slope` is the slope function value, `f′(x)` is the first derivative at the fluorescence amplification point of the corresponding cycle, `k(x)` is the curvature function value, `abs()` is the absolute value function, and `f″(x)` is the second derivative at the fluorescence amplification point of the corresponding cycle. All analytical calculations were implemented using the R programming language. The computing platform was a computer running Windows 11 Pro 23H2, with a 13th Gen Intel(R) Core(TM) i7-13700KF 3.4GHz processor and 64GB of dual-channel DDR4 Hynix memory. R version 4.3.1 was used, and the compiler was Rstudio 2023.06.2+561. All data were normalized to obtain all qPCR results for each DNA template. The corresponding DNA template concentration for each data point was used as a tag to construct a quantitative database.

[0083] 1.3 Constructing a stacked neural network model

[0084] Multiple different multilayer fully connected neural network models were constructed using the established quantitative database. The qPCR results data of parallel samples for each DNA template were randomly grouped, with 80% used for model training and 20% used for model testing.

[0085] In the first neural network model A, the input features of the input layer are normalized CT values, normalized fluorescence signal values ​​of all cycles, and fluorescence signal amplification values ​​of all cycles. The first hidden layer is a fully connected layer containing 64 neurons with ReLU activation function. The second hidden layer is a fully connected layer containing 32 neurons with ReLU activation function. The output layer contains 1 neuron with no activation function.

[0086] The input features of the second neural network model B are normalized CT values, curvature and slope values ​​of all loops, normalized CT values ​​...

[0087] The input features of the input layer of the third neural network model C are normalized CT values, fluorescence signal values ​​at the end of the loop, and fluorescence amplification values ​​at the end of the loop. The hidden layer is a fully connected layer containing 8 neurons with ReLU activation function. The output layer contains 1 neuron with no activation function.

[0088] The fourth neural network model D has input features of normalized CT values, curvature and slope values ​​of the terminating loop, fully connected hidden layers containing 8 neurons with ReLU activation function, and output layers containing 1 neuron with no activation function.

[0089] The fifth neural network model E is a stacked neural network model. The input features of its input layer are the results of the output layer of the neural network model AD. It contains a fully connected hidden layer with 8 neurons and the activation function is ReLU. The output layer contains 1 neuron and has no activation function.

[0090] The learning rate of all neural network models was set to 0.0001, and the loss function was Sparse Categorical Crossentropy. L2 regularization was used in each fully connected layer to prevent overfitting, with the L2 parameter set to 0.0001. Early stopping was implemented using Keras callbacks, with the early stopping condition being that the loss function value no longer decreased after five consecutive iterations. The training epochs were set to 100, and the number of input layer samples for each parameter update was set to 50. The labels used during training corresponded to the true concentration of the DNA template. All neural network models were developed and run using Python. The high-level framework Keras was provided by TensorFlow for Python. The development environment was Python 3.11.5, TensorFlow 2.16.1, and PyCharm Community Edition 2023.2.1 IDE. The platform for model development and operation is a computer running Windows 11 Pro 23H2, with a 13th Gen Intel(R) Core(TM) i7-13700KF 3.4GHz processor and 64GB of dual-channel DDR4 Hynix memory.

[0091] 1.4 qPCR detection of test samples

[0092] The purpose of this embodiment is to demonstrate that the qPCR result data analysis of the method of the present invention is more accurate and sensitive than that of traditional methods. In this embodiment, a mixed sample of DNA template 1 and DNA template 2 with known concentrations was used for qPCR experiments. A total of three samples were designed for this experiment. The first sample was the control group, which was a mixed sample of 2.5 μL of DNA template 1 with a concentration of 1E-7 ng / μL and 2.5 μL of DNA template 2 with a concentration of 1E-7 ng / μL. The second sample was the high expression group, which was a mixed sample of 2.5 μL of DNA template 1 with a concentration of 4E-7 ng / μL and 2.5 μL of DNA template 2 with a concentration of 1E-7 ng / μL. The third sample was the low expression group, which was a mixed sample of 2.5 μL of DNA template 1 with a concentration of 1E-8 ng / μL and 2.5 μL of DNA template 2 with a concentration of 1E-7 ng / μL. All qPCR experiments were conducted using the qPCR reaction system configuration table described in Table 1 above. The reaction procedures for all qPCR experiments were consistent with those in Table 2 above. Each DNA mixture sample was tested in triplicate, and the concentration of DNA template 2, which served as an internal control, was consistent across all mixture samples.

[0093] 1.5 Analysis of qPCR results of test samples

[0094] After the qPCR test of the three DNA mixed samples was completed, the qPCR results were adjusted. The baseline was set to 3-20 cycles, the FAM channel threshold was set to 50000, the VIC channel threshold was set to 5000, and all qPCR result data were exported. The CT value of qPCR, the fluorescence amplification value of each cycle, and the fluorescence signal value of each cycle can all be obtained directly through the qPCR program. The curvature and slope of the amplification curve at each cycle are determined by inputting the cycle number and its corresponding fluorescence amplification value into the analysis system. The `smooth.spline()` function in the R program is used to perform spline smoothing fitting on the curve formed by the input cycle number and fluorescence amplification. The slope for each cycle is calculated using the formula `slope = f′(x)`, and the curvature for each cycle is calculated using the formula `k(x) = abs(f″(x)) / ((f′(x)^2) + 1)^3 / 2`. Here, `slope` is the slope function value, `f′(x)` is the first derivative at the fluorescence amplification point of the corresponding cycle, `k(x)` is the curvature function value, `abs()` is the absolute value function, and `f″(x)` is the second derivative at the fluorescence amplification point of the corresponding cycle. All analytical calculations are implemented using the R programming language. The computing platform was a computer running Windows 11 Pro 23H2, with a 13th Gen Intel(R) Core(TM) i7-13700KF 3.4GHz processor and 64GB of dual-channel DDR4 Hynix memory. R version 4.3.1 was used, and the compiler was Rstudio 2023.06.2+561. All data was normalized. Pre-built neural network models A through E were called, and QPCR results from mixed samples of the same type as those used during model training were used as input layer data for each neural network model. The neural network models were run, and the absolute concentration of the target sample in the mixed sample was output, and the relative concentration was calculated.

[0095] 2. Traditional qPCR analysis method:

[0096] 2.1 Preparation of Standard Samples

[0097] The DNA template 1 sequence synthesized in section 1.1 was used as the target gene template, and the synthesized DNA template 2 sequence was used as the internal reference gene template. Sequence information is the same as described above.

[0098] The synthesized DNA template 1 was diluted at the following concentrations: 1.5E-6 ng / μL, 0.75E-6 ng / μL, 0.375E-6 ng / μL, 0.1875E-6 ng / μL, 0.9375E-7 ng / μL, 0.46875E-7 μL, 0.234375E-7 ng / μL, and 0.1171875E-7 ng / μL to obtain DNA templates 67-74.

[0099] 2.2 Constructing the Standard Curve

[0100] Primers and probes for detecting the target gene and internal reference gene were synthesized. The sequence information of the forward primer F1, reverse primer R1 and probe P1 for detecting DNA template 1 is as described above.

[0101] Configure the qPCR reaction system as shown in Table 3 below and perform qPCR experiments. The qPCR reaction program settings are as shown in Table 2 above. The threshold line in the qPCR results is set to 50000, and the baseline is set to 3-20. Obtain the Ct values ​​corresponding to DNA template 1 at each concentration. Construct a standard curve using the logarithm of Ct value and template concentration of sample to base 10.

[0102] Table 3: qPCR reaction system used to construct the standard curve

[0103] 2.3 qPCR detection of test samples

[0104] Same as the method described in 1.4.

[0105] 2.4 Analysis of qPCR results of test samples

[0106] After the qPCR testing of the three DNA pooled samples was completed, the qPCR results were adjusted by setting the baseline to 3-20 cycles, setting the FAM channel threshold to 50000, setting the VIC channel threshold to 5000, and exporting all qPCR result data.

[0107] Traditional relative quantitative analysis methods use the formula rc = 2 -ΔΔCt The calculation is performed, where rc represents the relative concentration of DNA samples in the low expression group and high expression group relative to the control group. In ΔΔCt, the right side calculates the difference in Ct values ​​between DNA template 1 and DNA template 2 (i.e., the difference in Ct values ​​between the target gene and the internal reference gene), and the left side calculates the difference between the difference in Ct values ​​between the target gene and the internal reference gene in all samples and the mean of the difference in Ct values ​​between the target gene and the internal reference gene in all samples in the control group.

[0108] In traditional qPCR result analysis methods based on Ct values, the analysis of absolute concentrations uses a linear regression equation obtained from a standard curve, Ct Value = k × log(c) + b, to calculate the base-10 logarithmic function value log(c) of the sample concentration. Then, the exponential function ac = 10 is used. log(c) Calculate the absolute concentration ac, where Ct Value is the Ct value, k is the slope of the linear regression equation, and b is the intercept of the linear regression equation.

[0109] 3 Experimental Results

[0110] The amplification standard curves for DNA template 1 are shown in Figures 2 and 3 and Table 4. A linear relationship was established between the base-10 logarithmic function value of the DNA sample concentration corresponding to the data points in the standard curve and the Ct value. The relationship between the Ct value and the base-10 logarithmic function value of the sample concentration is: Ct Value = -3.2662 × log(c) + 6.6973, where Ct Value is the Ct value in the qPCR result of each sample, log() is the base-10 logarithmic function, and c is the sample concentration corresponding to Ct Value.

[0111] Table 4: Statistical table of qPCR results used for plotting the standard curve

[0112] The raw data and analysis results of the mixed sample qPCR are shown in Tables 5, 6-9 and Figure 4-5 below. The relative concentration error and absolute concentration error are calculated by subtracting the absolute value of the true relative concentration and absolute concentration of the sample from the analyzed relative concentration or absolute concentration, and then dividing by the true relative concentration and absolute concentration of the sample.

[0113] The results show that neural network model E has lower relative and absolute concentration errors compared to neural network models A-D. All analysis methods based on neural network models also have lower relative and absolute concentration errors compared to traditional Ct-value-based analysis methods. This demonstrates that the qPCR result data analysis method described in this invention has a more accurate detection capability for DNA sample characteristics compared to traditional Ct-value-based qPCR result data analysis methods.

[0114] Table 5: Representative data of pooled sample qPCR results

[0115] Table 6: Analysis of qPCR results in the mixed sample control group

[0116] Table 7: Analysis of results from the low-expression group in mixed-sample qPCR

[0117] Table 8: Analysis of qPCR results in the high-expression group of mixed samples

[0118] Table 9: Summary of qPCR Results Analysis for Each Group in the Mixed Sample

[0119] Example 2: Examining the effectiveness of the qPCR analysis method described in this invention in analyzing the source characteristics of DNA samples.

[0120] The purpose of this embodiment is to demonstrate that, when performing qualitative detection of DNA samples from different sample types, the QPCR result analysis method described in this invention is more accurate and stable in determining the DNA sample type than traditional QPCR result analysis methods.

[0121] 2.1 Preparation of Standard Samples

[0122] By comparing DNA sequencing data of colorectal cancer and normal tissues in the TCGA database, bioinformatics analysis was used to select the gene SLA (DNA template 3) as the marker gene to distinguish sample types in this embodiment. The methylation rate of some CpG sites of this gene is much higher in colorectal cancer patients than in normal people, corresponding to the 17th, 54th and 102nd bases of the DNA template 3. Since the content of free nucleic acids from tissues in plasma is limited, more precise analytical methods are needed to distinguish sample types.

[0123] The standard samples used in this embodiment were plasma samples from clinically confirmed colorectal cancer patients and clinically confirmed healthy individuals without colorectal cancer. 2000 plasma samples were collected from colorectal cancer patients, and 2000 plasma samples were collected from healthy individuals without colorectal cancer. Free nucleic acids were extracted from all standard samples using a free nucleic acid extraction kit. The extracted free nucleic acids were then transformed and purified using a bisulfite conversion purification kit (after transformation, all unmethylated C bases were converted to T bases, while methylated C bases remained C bases). Finally, qPCR was used to detect the colorectal cancer-related DNA template 3 and the DNA template 4 corresponding to the internal reference gene ACTB in all samples. The DNA sequences corresponding to DNA template 3 and DNA template 4 are shown below:

[0124] DNA template 3 (NC_000008.10:134072511-134072711):

[0125] DNA template 4 (NC_000007.13:5568779-5568979):

[0126] 1. Construct a qualitative database

[0127] Primers and probes for detecting the target gene SLA and the internal reference gene ACTB were synthesized. The sequence information of the forward primer F3, reverse primer R3, and probe P3 for detecting DNA template 3 is shown below:

[0128] F3: 5'-TTGTGTATTATTTTTTTTGGTAGTTTGTTTTTTG-3'

[0129] R3: 5'-CTTAAACTACACAAAAATAACCAATCTTCACC-3'

[0130] P3: 5'-ROX-GGTGGTTGTTTATATGTGAAGAT-MGB-3'

[0131] The sequence information of the forward primer F4, reverse primer R4, and probe P4 used to detect DNA template 4 is shown below:

[0132] F4: 5'-GGGTTATTTTTTTGTGGTTGGTTTTG-3'

[0133] R4: 5'-CACCTTCTACAATAAACTACATATAACTCCCAA-3'

[0134] P4: 5'-FAM-GGTTTTGGTTAGTAGTATGGGGTG-MGB-3'

[0135] The qPCR reaction system was configured as shown in Table 10 below, and qPCR experiments were performed on all transformed DNA samples (2000 plasma samples from colorectal cancer patients and 2000 plasma samples from healthy individuals confirmed to be free of colorectal cancer). The qPCR program is shown in Table 2 above. The qPCR instrument used was an ABI 7500 qPCR instrument. The qPCR results were adjusted by setting the baseline to 3-20 cycles and setting the FAM and ROX channel thresholds to 25000. All qPCR result data were exported. The CT value, fluorescence amplification value for each cycle, and fluorescence signal value for each cycle of the qPCR were all directly obtained by exporting the qPCR program. The curvature and slope of the amplification curve at each cycle are obtained by inputting the cycle number and its corresponding fluorescence amplification value into the analysis system. The `smooth.spline()` function is used to perform spline smoothing fitting on the curve formed by the input one-to-one correspondence of cycle number and fluorescence amplification. The slope for each cycle is calculated using the formula `slope = f′(x)`, and the curvature for each cycle is calculated using the formula `k(x) = abs(f″(x)) / ((f′(x)^2) + 1)^3 / 2`. Here, `slope` is the slope function value, `f′(x)` is the first derivative at the fluorescence amplification point of the corresponding cycle, `k(x)` is the curvature function value, `abs()` is the absolute value function, and `f″(x)` is the second derivative at the fluorescence amplification point of the corresponding cycle. All analytical calculations are implemented using the R programming language. The computing platform was a computer running Windows 11 Pro 23H2, with a 13th Gen Intel(R) Core(TM) i7-13700KF 3.4GHz processor and 64GB of dual-channel DDR4 Hynix memory. R version 4.3.1 was used, and the compiler was Rstudio 2023.06.2+561. All data were normalized to obtain all qPCR results for each DNA sample. Each qPCR result was labeled with a corresponding DNA sample type (colorectal cancer or normal individuals) to construct a qualitative database.

[0136] Table 10: qPCR reaction system configuration table

[0137] 2. Construct a stacked neural network model

[0138] Multiple different multilayer fully connected neural network models were constructed using the built qualitative database. The qPCR results data of DNA templates for each sample type were randomly grouped, with 80% used for model training and 20% used for model validation.

[0139] In the first neural network model A, the input features of the input layer are normalized CT values, normalized fluorescence signal values ​​of all cycles, and fluorescence signal amplification values ​​of all cycles. The first hidden layer is a fully connected layer containing 64 neurons with ReLU activation function. The second hidden layer is a fully connected layer containing 32 neurons with ReLU activation function. The output layer contains 1 neuron with softmax activation function.

[0140] The input features of the second neural network model B are normalized CT values, curvature and slope values ​​of all loops, the input layer of the input layer, the hidden layer 1 is a fully connected layer containing 64 neurons with ReLU activation function, the hidden layer 2 is a fully connected layer containing 32 neurons with ReLU activation function, and the output layer contains 2 neurons with softmax activation function.

[0141] The input features of the input layer of the third neural network model C are normalized CT values, fluorescence signal values ​​at the end of the loop, and fluorescence amplification values ​​at the end of the loop. The hidden layer is a fully connected layer containing 8 neurons with ReLU activation function. The output layer contains 2 neurons with softmax activation function.

[0142] The fourth neural network model D has input features of normalized CT values, curvature and slope values ​​of the terminating loop, fully connected hidden layers containing 8 neurons with ReLU activation function, and output layers containing 2 neurons with softmax activation function.

[0143] The fifth neural network model E is a stacked neural network model. The input features of its input layer are the results of the output layer of the neural network model AD. It contains a fully connected hidden layer with 8 neurons and the activation function is ReLU. The output layer contains 2 neurons and the activation function is softmax.

[0144] The learning rate of all neural network models was set to 0.0001, and the loss function was Sparse Categorical Crossentropy. L2 regularization was used in each fully connected layer to prevent overfitting, with the L2 parameter set to 0.0001. Early stopping was implemented using Keras callbacks, with the early stopping condition being that the loss function value no longer decreased after 10 consecutive iterations. The training epochs were set to 100, and the number of input samples for each parameter update was set to 10. The labels during training corresponded to the true type of the samples (i.e., colorectal cancer or non-colorectal cancer). All neural network models were developed and run using Python. The high-level framework Keras was provided by TensorFlow for Python. The development environment was Python 3.11.5, TensorFlow 2.16.1, and PyCharm Community Edition 2023.2.1 IDE. The platform for model development and operation is a computer running Windows 11 Pro 23H2, with a 13th Gen Intel(R) Core(TM) i7-13700KF 3.4GHz processor and 64GB of dual-channel DDR4 Hynix memory.

[0145] 3. qPCR detection of the sample to be tested

[0146] The purpose of this embodiment is to demonstrate that the qPCR result data analysis of the method of the present invention is more accurate and sensitive than that of traditional methods. In this embodiment, DNA samples extracted and transformed from plasma of known sample types were used for qPCR experiments. A total of 12 clinically known positive colorectal cancer and 12 clinically known non-colorectal cancer plasma samples were used to extract and transform DNA samples to simulate unknown samples for qPCR experiments, which were then used to evaluate the analysis results. All qPCR experiments used the qPCR reaction system configuration table described in Table 10 above, and the reaction procedures for all qPCR experiments were consistent with those in Table 2 above.

[0147] 4. Analysis of qPCR results of the test samples

[0148] After qPCR testing of all DNA samples, the qPCR results were adjusted by setting the baseline to 3-20 cycles, setting the threshold lines for the FAM and ROX channels to 25000, and exporting all qPCR result data. The CT value, fluorescence amplification value for each cycle, and fluorescence signal value for each cycle of the qPCR could all be obtained directly through the qPCR program. The curvature and slope of the amplification curve at each cycle were determined by inputting the cycle number and its corresponding fluorescence amplification value into the computer. The `smooth.spline()` function was used to perform spline smoothing on the curves formed by the input cycle numbers and fluorescence amplification values. The slope for each cycle was calculated using the formula `slope = f′(x)`, and the curvature for each cycle was calculated using the formula `k(x) = abs(f″(x)) / ((f′(x)^2) + 1)^3 / 2`. Here, `slope` is the slope function value, `f′(x)` is the first derivative at the fluorescence amplification point of the corresponding cycle, `k(x)` is the curvature function value, `abs()` is the absolute value function, and `f″(x)` is the second derivative at the fluorescence amplification point of the corresponding cycle. All analytical calculations were implemented using the R programming language. The computing platform was a computer running Windows 11 Pro 23H2, with a 13th Gen Intel(R) Core(TM) i7-13700KF 3.4GHz processor and 64GB of dual-channel DDR4 Hynix memory. R version 4.3.1 was used, and the compiler was Rstudio 2023.06.2+561. All data was normalized. Pre-built neural network models A through E were called, using QPCR result data of the same types as those used during model training as the input layer data for each neural network model. The neural network models were run, and the probability of the simulated unknown sample type was output. Sample types with a probability exceeding 50% were used as the final output sample types of the neural network model.

[0149] The raw data and analysis results of simulated unknown sample qPCR are shown in Tables 11-13 and Figures 6-7 below. In the traditional qPCR result analysis method based on Ct values, under the conditions of this embodiment, although the methylation rate of the target gene's target site is high in both colorectal cancer and normal tissues, the limited amount of tissue-derived DNA in plasma makes it difficult to determine the difference in methylation levels of this gene in the plasma of colorectal cancer patients and normal individuals solely based on Ct values. Therefore, distinguishing between colorectal cancer patients and normal individuals by detecting the target site of this gene in plasma is quite difficult and unlikely to be applied to clinical testing. As shown in Table 11, the Ct values ​​of the methylated genes in the transformed DNA extracted from plasma samples of 12 colorectal cancer patients and 12 normal individuals were similar, showing no significant difference. Therefore, it is impossible to determine the sample type solely based on Ct values. However, as shown in Tables 12 and 13, when using the QPCR result analysis system of the present invention to analyze QPCR results, the vast majority of samples can be correctly distinguished (only one negative sample failed to be predicted in the neural network model E). Therefore, the multi-dimensional QPCR result analysis system based on the neural network model of the present invention can significantly improve the accuracy and sensitivity of result analysis when analyzing QPCR test results of samples with small differences.

[0150] Table 11: qPCR results of some representative samples

[0151] Table 12: Predicted Probability of Colorectal Cancer Positiveness Based on QPCR Results of Positive Samples Using Different Neural Network Models

[0152] Table 13: Predicted Probability of Colorectal Cancer Positiveness Based on qPCR Results of Negative Samples Using Different Negative Negative Samples

Claims

1. A neural network model for qPCR result analysis, characterized in that... The neural network model is a single neural network model or a stacked neural network model combining multiple neural network models. It takes at least three types of qPCR result data as input and outputs the sample type of the target gene corresponding to the qPCR result data. The qPCR result data includes the following types: CT value, fluorescence signal value at all different cycles during PCR, fluorescence signal amplification value at all different cycles, curvature of fluorescence signal amplification value change at all different cycles during PCR, and slope of fluorescence signal amplification value change at all different cycles. The neural network model is trained and validated using a standard database for predicting the target gene sample type of unknown samples. The sample type is the quantitative feature of the target gene or the source feature of the target gene. The standard database includes a quantitative database, a qualitative database, or a combination of both. The quantitative database is constructed as follows: DNA templates of target genes with different quantification levels are prepared, QPCR experiments are performed using target gene primers and specific probes, and QPCR result data of target genes with different quantification levels are obtained. Each QPCR result data is labeled with the quantification value of the target gene. The qualitative database is constructed as follows: based on at least one source characteristic of the target gene, several positive real samples corresponding to the source characteristics are collected. All positive real samples are processed in the same way to obtain corresponding DNA templates. For each DNA template, the same target gene primers and probes and QPCR conditions are used to perform QPCR experiments to obtain QPCR result data corresponding to each DNA template. Each QPCR result data is tagged with the source characteristic of the target gene.

2. The neural network model as described in claim 1, characterized in that... The neural network model is a fully connected neural network model, consisting of an input layer, several hidden layers, and an output layer. The input layer contains all or some types of QPCR result data. The hidden layers are fully connected layers containing several neurons. The output layer is a fully connected layer containing one or n neurons. One is the output layer when applying a quantitative database, n is the number of sample types, and n is the output layer when applying a qualitative database.

3. The neural network model as described in claim 2, characterized in that... The neural network model is selected from one of the following neural network models (AD), or a stacked neural network model composed of at least two of them. Neural network model A: The normalized CT value, the normalized dynamic fluorescence signal value of all cycles, and the normalized dynamic fluorescence signal amplification value of all cycles are selected as the input layer of neural network model A, and the sample type of the target gene is selected as the output layer. Neural Network Model B: The curvature and slope of all cycles of normalized CT value and fluorescence signal amplification value change are selected as the input layer of Neural Network Model B, and the sample type of the target gene is selected as the output layer. Neural network model C: The normalized CT value, the normalized final cycle fluorescence signal value, and the fluorescence signal amplification value are selected as the input layer of neural network model C, and the sample type of the target gene is selected as the output layer. Neural network model D: The curvature and slope of the final cycle of the normalized CT value and the change of fluorescence signal amplification value are selected as the input layer of neural network model D, and the sample type of the target gene is selected as the output layer. Stacked Neural Network Model E: Select the results of at least two of the neural network models A, B, C and D as the input layer of neural network model E, and the sample type of the target gene as the output layer.

4. The neural network model as described in claim 1, characterized in that... The quantitative database also contains qPCR results data for internal reference genes. Target gene DNA templates and internal reference gene templates with different quantification levels are prepared. These templates are then mixed pairwise to obtain a series of DNA template samples with varying quantification levels. qPCR experiments are performed on each sample using primers and specific probes for both the target and internal reference genes. Each sample yields a set of qPCR results data for both the target and internal reference genes. Each set of qPCR results data is labeled with the quantification level of the target gene. The probes for the internal reference gene and the target gene have different fluorescence channels. When training and validating neural network models using the quantitative database, a set of qPCR results data for both the target and internal reference genes is used as input, with the data types for the qPCR results for both genes being consistent.

5. The neural network model as described in claim 1, characterized in that... The qualitative database also contains qPCR result data of the internal reference gene. Using primers and probes for both the target gene and the internal reference gene, qPCR experiments are performed on each DNA template. Each DNA template yields a set of qPCR result data for both the target gene and the internal reference gene. Each set of qPCR result data is tagged with the source characteristics of the target gene. When using the qualitative database to train and validate a neural network model, a set of qPCR result data for both the target gene and the internal reference gene is used as input, and the data types of the qPCR result data for both the target gene and the internal reference gene are selected to be consistent.

6. The neural network model as described in claim 1, characterized in that... The qualitative database further uses samples with negative source characteristics of the target gene as a control, and collects several real samples with negative source characteristics of the target gene. The processing method for obtaining DNA templates, the primers and probes used for qPCR experiments, and the qPCR conditions for all negative real samples are the same as those for positive real samples. QPCR result data corresponding to each DNA template is obtained, and each qPCR result data is tagged with the source characteristics of the target gene.

7. The neural network model as described in claim 1, characterized in that... In qualitative databases, target genes are single genes or combinations of several genes that are highly specific to the sample type.

8. The neural network model as described in claim 7, characterized in that... The number of target genes is 1 to 10.

9. The neural network model as described in claim 7, characterized in that... When the target gene in the qualitative database is a combination of several genes, the same fluorescence channel can be used to obtain the QPCR results of the gene combination, or different fluorescence channels can be used to obtain the QPCR results of each gene.

10. The neural network model as described in claim 1, characterized in that... The standard database uses no fewer than 100 samples for each sample type.

11. The neural network model as described in claim 1, characterized in that... The quantitative characteristics of the target gene are methylation rate, DNA copy number, DNA mutation state, or DNA concentration. The source characteristics of the target gene are disease, race, age, sex, or other biological or environmental factors that affect gene characteristics.

12. A qPCR result analysis system, characterized in that... The system includes a data acquisition module, an analysis module, and a result display module. The data acquisition module is used to process and acquire qPCR result data, which includes the following types: CT value, fluorescence signal value at all different cycles during PCR, fluorescence signal amplification value at all different cycles, curvature of fluorescence signal amplification value change at all different cycles during PCR, and slope of fluorescence signal amplification value change at all different cycles. The analysis module is a neural network model as described in any one of claims 1-11, which takes at least three types of QPCR result data as input and the sample type of the target gene as the result output, wherein the sample type is the quantitative feature of the target gene or the source feature of the target gene. The result display module displays the output results of the analysis module.

13. The qPCR result analysis system as described in claim 12, characterized in that... The CT value, fluorescence amplification value for each cycle, and fluorescence signal value for each cycle of the qPCR results data are directly generated by the qPCR program system. The data acquisition module calls the R program and uses the smooth.spline() function to perform spline smoothing fitting on the curve formed by the cycle number and fluorescence amplification data. The slope corresponding to each cycle is calculated using the formula slope=f′(x). The curvature corresponding to each cycle is calculated using the formula k(x)=abs(f″(x)) / ((f′(x)^2)+1)^3 / 2, where slope is the slope function value, f′(x) is the first derivative at the fluorescence amplification point of the corresponding cycle, k(x) is the curvature function value, abs() is the absolute value function, and f″(x) is the second derivative at the fluorescence amplification point of the corresponding cycle.

14. A method for analyzing qPCR results, characterized in that... The analysis of qPCR result data of unknown samples is performed using the neural network model according to any one of claims 1-11 or the qPCR result analysis system according to claim 12 or 13, specifically including: (1) Construct a standard database based on the target genes to be tested and the sample types of unknown samples; (2) Use standard databases to train and validate neural network models to build neural network models; (3) Perform qPCR experiments on unknown samples to obtain qPCR result data. The conditions for the qPCR experiments are consistent with the conditions for the qPCR experiments used to construct the standard database. (4) Call the neural network model constructed in step (2), input the QPCR result data of the unknown sample into the neural network model and run it. The neural network model outputs the sample type of the target gene of the unknown sample as the result. The QPCR result data of the unknown sample input is consistent with the data type and quantity of the QPCR result data input by the input layer in the neural network model training process.

15. The method as described in claim 14, characterized in that... Step (2) involves adjusting the hyperparameters of the neural network model. By adjusting the hyperparameters in the neural network model, the model validation results are optimized to avoid overfitting.

16. The method as described in claim 14, characterized in that... The qPCR program includes enzyme activation at 95°C for five minutes and a cycling program. The cycling program includes denaturation: 95°C for 5-30 seconds, annealing: 5-30 seconds, extension: 72°C for 20-30 seconds, and the cycling program is performed for 40-50 cycles. Fluorescence signals are collected during the extension process at 72°C.