Method for specifically evaluating fragmentation degree of target nucleic acid

By designing primer sets and dPCR technology with different target sequence lengths, the problems of low sensitivity, high cost and high complexity in the existing methods are solved, and efficient, accurate evaluation and concentration detection of the degree of nucleic acid fragmentation are achieved.

WO2025162424A1PCT designated stage Publication Date: 2025-08-07SICHUAN MACCURA BIOTECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/075384
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-01
Filing Date
2025-01-27
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

The existing methods for detecting the degree of nucleic acid fragmentation have problems such as low sensitivity, high cost, complex operation, and the inability to specifically evaluate the degree of fragmentation of the target nucleic acid.

Method used

At least 2 pairs of primer sets are used to design primers of different target sequence lengths, evaluate the degree of fragmentation of the target nucleic acid by nucleic acid amplification reaction and fit the calculation formula, and quantitative detection is performed using dPCR, avoiding the dependence of additional equipment and reagents.

Benefits of technology

A high sensitivity, low cost and simple overall evaluation of the degree of fragmentation of nucleic acid samples is achieved, which can accurately reflect the degree and concentration of nucleic acid fragmentation, avoid gDNA contamination, and improve detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2025075384-FTAPPB-I100001
    Figure PCTCN2025075384-FTAPPB-I100001
  • Figure PCTCN2025075384-FTAPPB-I100002
    Figure PCTCN2025075384-FTAPPB-I100002
  • Figure PCTCN2025075384-FTAPPB-I100003
    Figure PCTCN2025075384-FTAPPB-I100003
Patent Text Reader

Abstract

Provided is a method for specifically evaluating the fragmentation degree of a target nucleic acid. The method comprises mixing a biological sample with at least two pairs of target nucleic acid primer sets, a detection marker and a nucleic acid amplification reaction solution, and performing a nucleic acid amplification reaction to obtain detection copy numbers of different target sequence lengths; fitting a calculation formula on the basis of the target sequence lengths and the detection copy numbers of the corresponding length, and evaluating the fragmentation degree of the target nucleic acid by means of determining the intersection of the obtained fitting curve with the X axis.
Need to check novelty before this filing date? Find Prior Art

Description

Method for specifically evaluating the degree of fragmentation of target nucleic acids Technical Field

[0001] The present application relates to the field of molecular biology, and more particularly to a method and system for evaluating the fragmentation of nucleic acid samples, and more particularly to a method for specifically evaluating the degree of fragmentation of a target nucleic acid. Background Art

[0002] Molecular diagnostics is a critical frontier in contemporary medicine. Nucleic acid integrity, which refers to the size and degree of fragmentation of nucleic acid fragments, is a crucial quality control indicator in molecular biological genetic testing. Incomplete nucleic acids include both naturally fragmented and artificially generated fragments. Naturally generated fragmented nucleic acids typically arise from several sources, including fragmented cfDNA (cell-free DNA) produced by apoptosis of normal cells or necrosis of tumor cells; nucleic acid degradation caused by improper sample storage or repeated freeze-thaw cycles; nucleic acid breakage caused by mechanical shear forces during nucleic acid extraction; genomic DNA damage caused by improper preparation or storage of FFPE samples; and highly fragmented ancient DNA samples. Artificially generated fragmented nucleic acids include fragments generated by breaking genomic DNA during NGS library construction and simulated fragmentation samples prepared for certain scientific research purposes.

[0003] The presence of some fragmented nucleic acids affects the test results. Some nucleic acid samples have high requirements for integrity. For example, the DNA quality of FFPE samples, such as poor nucleic acid integrity, will significantly affect the subsequent fluorescence quantification results or the success rate of second-generation sequencing library construction and the accuracy of sequencing data indicators. Therefore, before performing nucleic acid testing on FFPE samples, it is very useful to evaluate the nucleic acid fragmentation of the samples. This allows different treatments to be performed on samples with different degrees of fragmentation to avoid wasting time and money (Corces MR, Granja JM, Shams S, et al. The chromatin accessibility landscape of primary human cancers. Science 2018; 362: eaav1898; Polak P, R, Koren A, et al. Cell-of-origin chromatin organization shapes the mutational landscape of cancer. Nature 2015; 518: 360-4.). Some nucleic acid samples are fragmented, and their fragment size and distribution need to be assessed. For example, many studies have been dedicated to studying the fragment model of cfDNA for early disease screening (Mouliere F, Chandrananda D, Piskorz AM, et al. Enhanced detection of circulating tumor DNA by fragment size analysis. Sci Transl Med 2018; 10: eaat4921).

[0004] Regardless, for important studies or important samples, the degree of sample fragmentation should be evaluated to ensure the quality of downstream experiments, thereby better guiding subsequent genetic testing and the judgment of test results. Summary of the Invention

[0005] Currently, the main methods for detecting the degree of nucleic acid fragmentation are the electrophoresis methods described in GB / T 40974-2021 "Nucleic Acid Sample Quality Assessment Methods", including gel electrophoresis, capillary gel electrophoresis, and microfluidic analysis. The electrophoresis method uses the "charge effect" and "molecular sieve effect" of electrophoresis to separate nucleic acid samples of different fragment sizes. However, gel electrophoresis is time-consuming and labor-intensive, prone to aerosol contamination, and has low sensitivity, resulting in poor accuracy. It is only suitable for high-safety laboratories to evaluate the quality of high-concentration fragmentation samples. It is inaccurate for high-concentration, low-fragmentation samples, and cannot evaluate the fragmentation degree of low-concentration samples. Capillary gel electrophoresis is a new type of liquid phase separation technology that uses capillaries as separation channels and a high-voltage direct current electric field as the driving force, such as Agilent's 4150 bioanalyzer and Bioptic's Qsep100 fully automatic nucleic acid protein analysis system. Microfluidics analysis refers to the science and technology involved in systems that use microchannels with sizes ranging from tens to hundreds of microns to process or manipulate tiny fluids (with volumes ranging from nanoliters to microliters), such as Agilent's 2100 system and PerkinElmer's Labchip system. Although capillary gel electrophoresis and microfluidics analysis have overcome many of the shortcomings of ordinary gel electrophoresis, are relatively simple to operate, and have greatly improved sensitivity, they can provide high-resolution DNA fragment peak distributions and are mainly used for second-generation and third-generation sequencing quality control, sample quality control, and analysis of residual DNA in host cells of biological products. However, there are also some problems, including: this type of method requires different detection reagents for different application scenarios, and each detection reagent analyzes a different range of fragments, and cannot analyze the overall fragmentation of the sample; some instruments claim to be able to achieve picogram-level nucleic acid detection sensitivity, but often when the instrument detects samples of this concentration, many fragments are below the detection threshold, resulting in missed detections and inaccurate fragment information. Indeed, scientific research or clinical practice often also has testing needs for clinical samples with lower concentrations, such as analyzing the fragmentation distribution of extremely low-concentration cfDNA. However, because the sample concentration is lower than the detection sensitivity, relevant fragment information cannot be obtained. For example, the learning model of Agilent's 2100 system is derived from the manual judgment and scoring of 1,300 RNA samples by experts. The sample size is small, and the manual scoring itself has inaccuracies. The cost of its related instruments and supporting reagents is high, the shelf life of the reagents is short, and the instruments can only be used for fragment information analysis. It cannot obtain more sample information from a single detection process, such as whether the nucleic acid sample to be tested contains the target to be tested, the specific content of the target, the fragmentation symptoms of the target, etc., and the cost-effectiveness is relatively low. In addition, the electrophoresis step and the collection of DNA after electrophoresis have the potential for contamination. Moreover, the electrophoresis method lacks specificity and cannot evaluate the degree of fragmentation of specific target nucleic acids.For example, the literature reports that circulating HPV DNA (HPV ctDNA) may serve as a residual tumor marker at the end of chemoradiotherapy or predict recurrence during follow-up (Jeannot E, Latouche A, Bonneau C, et al. Circulating HPV DNA as a Marker for Early Detection of Relapse in Patients with Cervical Cancer. Clin Cancer Res. 2021; 27(21): 5869-5877. doi: 10.1158 / 1078-0432. CCR-21-0625); there are also literatures that study the fragment abundance and size of human mitochondrial cfDNA (Burnham P, Kim MS, Agbor-Enoh S, et al. Single-stranded DNA library preparation uncovers the origin and diversity of ultrashort cell-free DNA in plasma. Sci Rep. 2016; 6: 27859. Published 2016 Jun 14.doi:10.1038 / srep27859); If the above electrophoresis methods cannot achieve the purpose of detecting the specific degree of gene fragmentation.

[0006] Mass spectrometry has also been used to analyze nucleic acid size because nucleic acid fragments of different sizes, such as those prepared by primer extension reactions, have different molecular weights (Ding and Cantor, 2003, Proc Natl Acad Sci USA, 100, 7449-7453). However, this method also requires additional equipment and supporting reagents, making the detection costly.

[0007] Methods for detecting the degree of nucleic acid fragmentation also include Ellinger, J., et al. (Cell-Free Circulating DNA: Diagnostic Value in Patients With Testicular Germ Cell Cancer, Journal of Urology, Volume 181, Issue 1, January 2009, Pages 363-371), which discloses quantification of 106bp, 193bp, and 384bp actin-β DNA fragments by real-time quantitative PCR. The integrity of DNA is expressed by the ratio of large fragments (193 or 384bp) to short fragments (106bp). Compared with healthy subjects, it was found that the level of free DNA fragments in testicular cancer patients was significantly increased. However, the qPCR method is a relative absolute quantitative method. When using this method to detect residual host nucleic acids in biological products, it is necessary to set up multiple repeated tests for each sample, and use standard products to prepare a standard curve. Then, the mass of the sample to be tested is calculated based on the conversion relationship between the ct value obtained from the standard curve and the mass of the standard product. The operation is relatively complicated and the reagent cost is high. Secondly, qPCR measures the total signal of each reaction, and the system's resistance to sample inhibition is poor, which can affect detection sensitivity and precision. Furthermore, qPCR without nucleic acid extraction can produce strong matrix effects, affecting the accuracy of test results.

[0008] US Patent Publication No. US20180105864A1 discloses a method for determining the integrity and / or quantity of cell-free DNA (cfDNA) using digital PCR (dPCR). This method uses primers designed for two target sequences, one smaller than 300 bp and the other larger than 300 bp, to determine DNA integrity based on the test results.

[0009] However, the inventors have found that existing evaluation methods have problems such as inaccurate quantification, or are not simple enough and have poor ability to explain biological significance.

[0010] Based on this, one of the purposes of this application is to provide a method for specifically evaluating the degree of fragmentation of a target nucleic acid, the method comprising:

[0011] Designing at least two pairs of primer sets based on the target nucleic acid; the target sequence lengths of the at least two pairs of primer sets are different from each other;

[0012] After mixing the biological sample with the at least two pairs of primer sets, the detection marker, and the nucleic acid amplification reaction solution, the mixture is randomly distributed into a plurality of reaction units to perform a nucleic acid amplification reaction to obtain detection copy numbers of different target sequence lengths;

[0013] A calculation formula is fitted based on the length of the target sequence and the detected copy number of the corresponding length, in which the variable Y is the detected copy number of the target sequence and the variable X is the length of the target sequence; and

[0014] The degree of fragmentation of the target nucleic acid is evaluated based on the intersection of the fitting curve of the calculation formula and the X-axis.

[0015] In some specific embodiments, the target sequence lengths of the at least two pairs of primer sets range from 30 to 540 bp, preferably from 30 to 260 bp.

[0016] In some embodiments, the at least two pairs of primer sets share one forward primer or one reverse primer.

[0017] In some specific embodiments, the at least two primer pairs are 3 to 7 primer pairs, preferably 4 to 7 primer pairs.

[0018] In some preferred embodiments, the at least two pairs of primer sets are four pairs of primer sets.

[0019] In some specific embodiments, the lengths of the target sequences of the four pairs of primer sets are 30-50 bp, 80-100 bp, 130-150 bp, and 160-180 bp.

[0020] In some embodiments, the method further comprises the step of isolating nucleic acids in the biological sample before mixing the biological sample with the at least two pairs of primer sets.

[0021] In some embodiments, in the fitting step, a linear model, a quadratic polynomial model, or a cubic polynomial model is fitted according to the length of the target sequence and the number of detected copies of the corresponding length.

[0022] In some specific embodiments, the calculation formula fitted by the linear model is Y=kX+b.

[0023] In some embodiments, the degree of fragmentation of a target nucleic acid is assessed based on the |b / k| value.

[0024] Another object of the present application is to provide a reagent for detecting a target nucleic acid, comprising:

[0025] At least two pairs of primers are used to amplify the same target nucleic acid, wherein the target sequence lengths of the at least two pairs of primers are different from each other.

[0026] In some embodiments, when unfragmented target nucleic acid is used as a template, the coefficient of variation of the quantitative values ​​between the at least two pairs of primer sets is ≤10%.

[0027] In some embodiments, the at least two pairs of primer sets share a forward primer or a reverse primer.

[0028] In some specific embodiments, the target sequence lengths of the at least two pairs of primer sets range from 30 to 540 bp, preferably from 30 to 260 bp.

[0029] In some specific embodiments, the primer set is preferably a set of 3 to 7 primer pairs, or a set of 4 to 7 primer pairs, and more preferably a set of 4 primer pairs.

[0030] In some specific embodiments, the target sequence lengths of the four pairs of primer sets are 30-50 bp, 80-100 bp, 130-150 bp, and 160-180 bp.

[0031] Another object of the present application is to provide a kit for detecting a target nucleic acid, the kit comprising the reagents of the present application; preferably, the kit further comprises a detection marker and a nucleic acid amplification reaction solution;

[0032] In some embodiments, the nucleic acid amplification reaction solution includes dNTPs, a buffer, salt ions, and a polymerase.

[0033] Another object of the present application is to provide a method for detecting a target nucleic acid, the method comprising the following steps:

[0034] Mixing a biological sample with at least two pairs of primer sets, a detection marker, and a nucleic acid amplification reaction solution to obtain a first mixed solution; wherein the at least two pairs of primer sets have target sequences of different lengths and amplify the same target nucleic acid;

[0035] Randomly distributing the first mixed solution into a plurality of reaction units to perform a nucleic acid amplification reaction to obtain detection copy numbers of different target sequence lengths;

[0036] Fitting a first calculation formula according to the length of the target sequence and the detected copy number of the corresponding length, in which the variable Y is the detected copy number of the target sequence and the variable X is the length of the target sequence; and

[0037] The fragmentation degree of the target nucleic acid is estimated based on the intersection of the fitting curve of the first calculation formula and the X-axis and / or the copy number of the target nucleic acid is estimated based on the coefficient of the first calculation formula.

[0038] In some embodiments, the method further comprises the step of isolating nucleic acids in the biological sample before mixing the biological sample with the at least two pairs of primer sets.

[0039] In some specific embodiments, in the step of fitting the first calculation formula, a linear model, a quadratic polynomial model, or a cubic polynomial model is fitted according to the length of the target sequence and the detected copy number of the corresponding length.

[0040] In some specific embodiments, the first calculation formula fitted by the linear model is Y=kX+b.

[0041] In some embodiments, the fragmentation degree of the target nucleic acid is assessed based on the |b / k| value, and / or the copy number of the target nucleic acid is assessed based on the b value.

[0042] In some specific embodiments, the target sequence lengths of the at least two pairs of primer sets range from 30 to 540 bp, preferably from 30 to 260 bp.

[0043] In some embodiments, the at least two pairs of primer sets share one forward primer or one reverse primer.

[0044] In some specific embodiments, the primer set is a set of 3 to 7 primer pairs, or a set of 4 to 7 primer pairs, preferably a set of 4 primer pairs.

[0045] In some specific embodiments, the target sequence lengths of the four pairs of primer sets are 30-50 bp, 80-100 bp, 130-150 bp, and 160-180 bp.

[0046] In some embodiments, the method further comprises:

[0047] mixing the control biological sample with the at least two pairs of primer sets, the detection marker, and the nucleic acid amplification reaction solution to obtain a second mixed solution;

[0048] Randomly distributing the second mixed solution into a plurality of reaction units to perform a nucleic acid amplification reaction to obtain the detection copy number of target sequences of different lengths;

[0049] Fitting a second calculation formula based on the length of the target sequence and the detected copy number of the corresponding length, in which the variable Y is the detected copy number of the target sequence and the variable X is the length of the target sequence; and

[0050] The intersection of the fitting curves of the first calculation formula and the second calculation formula with the X-axis is compared to evaluate the fragmentation degree of the target nucleic acid, and / or the coefficients of the first calculation formula and the second calculation formula are compared to evaluate the copy number of the target nucleic acid.

[0051] In some embodiments, the control biological sample comprises a biological sample from a normal subject or a biological sample from a subject before administration of a drug. In some embodiments, the method further comprises a step of isolating nucleic acids in the control biological sample before mixing the control biological sample with the at least two pairs of primer sets.

[0052] In some specific embodiments, in the step of fitting the second calculation formula, a linear model, a quadratic polynomial model, or a cubic polynomial model is fitted according to the length of the target sequence and the detected copy number of the corresponding length.

[0053] In some specific embodiments, the second calculation formula fitted by the linear model is Y=k1X+b1.

[0054] In some embodiments, the |b / k| value is compared to the |b1 / k1| value to assess the degree of fragmentation of the target nucleic acid, and / or the b value is compared to the b1 value to assess the copy number of the target nucleic acid.

[0055] In some specific embodiments, the target sequence lengths of the at least two pairs of primer sets range from 30 to 540 bp, preferably from 30 to 260 bp.

[0056] In some embodiments, the at least two pairs of primer sets share one forward primer or one reverse primer.

[0057] In some specific embodiments, the primer set is a set of 3 to 7 pairs or 4 to 7 pairs of primers.

[0058] In some preferred embodiments, the primer set is a set of 4 primer pairs.

[0059] In some preferred embodiments, the target sequence lengths of the four pairs of primers are 30-50 bp, 80-100 bp, 130-150 bp, and 160-180 bp.

[0060] Another object of the present application is to provide a method for accurately detecting cfDNA, the method comprising:

[0061] Isolation of cfDNA from biological samples;

[0062] Mixing the cfDNA with at least two pairs of target nucleic acid primers, a pair of gDNA primers, a detection marker, and a nucleic acid amplification reaction solution; wherein the at least two pairs of primers have target sequences of different lengths and amplify the same target nucleic acid;

[0063] After mixing, the mixed solution is randomly distributed into multiple reaction units to perform nucleic acid amplification reaction;

[0064] Obtaining the detected copy numbers of target sequences corresponding to at least two pairs of target nucleic acid primer sets and the detected copy number of a pair of gDNA primer sets; fitting a calculation formula based on the copy number obtained by deducting the detected copy number of the gDNA primer set from the detected copy number of the target sequence corresponding to each of the at least two pairs of target nucleic acid primer sets and the target sequence length of the corresponding target nucleic acid primer set, wherein the variable Y is the copy number of the target sequence after deduction, and the variable X is the length of the target sequence; and

[0065] The degree of cfDNA fragmentation is evaluated based on the intersection of the fitting curve of the calculation formula and the X-axis, and / or the copy number of cfDNA is evaluated based on the coefficient of the calculation formula.

[0066] In a specific embodiment, in the fitting step, a linear model, a quadratic polynomial model, or a cubic polynomial model is fitted based on the copy number obtained after deducting the copy number of a pair of gDNA primer sets from the copy number of each corresponding target sequence and the target sequence length of the corresponding target nucleic acid primer set.

[0067] In some specific embodiments, the calculation formula fitted by the linear model is Y=kX+b.

[0068] In some embodiments, the degree of cfDNA fragmentation is assessed based on the |b / k| value, and / or the copy number of cfDNA is assessed based on the b value. In some embodiments, the target sequence length of the at least two primer pairs ranges from 30 to 420 bp, preferably from 30 to 260 bp.

[0069] In some specific embodiments, the target sequence length of the pair of gDNA primer sets is ≥466 bp.

[0070] In some embodiments, the at least two pairs of primer sets share one forward primer or one reverse primer.

[0071] In some specific embodiments, the primer set is a set of 3 to 7 pairs or 4 to 7 pairs of primers.

[0072] In some preferred embodiments, the primer set is a set of 4 primer pairs.

[0073] In some preferred embodiments, more preferably, the target sequence lengths of the four pairs of primer sets are 30-50 bp, 80-100 bp, 130-150 bp, and 160-180 bp.

[0074] The target nucleic acid fragmentation degree assessment method of the present application is not limited by the detection reagent to the detection fragment range, and can perform an overall assessment of the fragmentation degree of the nucleic acid sample. The present application can simultaneously achieve the assessment of the fragmentation degree of the target nucleic acid and the detection of the concentration; compared with the existing electrophoresis analysis method, it has higher sensitivity, lower cost, and does not require additional equipment and instruments, and is simple to operate, saves time, and improves efficiency; in addition, the accuracy of the fragmentation assessment using the method of the present application is higher, and the calculated concentration is closer to the true value. Using the cfDNA assessment method of the present application, the contamination of gDNA is further avoided in the cfDNA fragmentation assessment process, and a more accurate cfDNA fragmentation degree and content are obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] FIG1 is a graph showing the fitting results of a linear model, a quadratic polynomial model, and a cubic polynomial model. DETAILED DESCRIPTION

[0076] Implementation Plan and Definition

[0077] Unless otherwise stated, all technical terms and scientific terms used herein have the same implication as that generally understood in the technical field to which the application belongs.The following interpretation is a supplement to those interpretations in this area and relates to the application, but is not extrapolated to any relevant or irrelevant situation, for example, any conventionally used patent or application. Although any method and material similar or equivalent to the method and material described herein can be used when practicing the application, materials and methods as herein described are preferred. Therefore, the terms used herein are intended to merely describe specific embodiments, are not intended to limit the application.

[0078] This application is based on the following findings:

[0079] For the same fragmented sample, the longer the target sequence, the lower the detection ability; the more severe the fragmentation, the faster the detection ability of long target sequences decreases. By selecting specific sequences and designing multiple pairs of primers for specific sequences, target sequences of different lengths can be obtained. With the target sequence length as the X-axis and the mean of multiple repeated detections of the corresponding detection copy number as the Y-axis, a linear model, a quadratic polynomial model, or a cubic polynomial model is fitted to the test results of each sample. The X-axis and Y-axis are interchangeable. In this case, the degree of fragmentation of the target nucleic acid needs to be evaluated based on the intersection of the target fitting curve and the Y-axis.

[0080] For example, the corresponding linear equations Y=kX+b and R are obtained. 2 This step meets the following requirements: the coefficient of variation (CV) of the quantitative value obtained by repeating the test 10 times for each sample using each primer set is ≤15%, and the R value of the linear equation is 2 >0.90.

[0081] When Y = 0, |x| = |b / k|. This means that when the target sequence length is |b / k|, the theoretical quantitative value of the sample is 0. When the target sequence length is greater than |b / k|, the theoretical quantitative value of the sample is still 0, and when the target sequence length is greater than the template length, the theoretical quantitative value of the sample is also 0. This indicates that the |b / k| value obtained by this method can, to a certain extent, reflect the fragment length of the sample being tested and is of practical significance. Furthermore, according to the modeling data, the |b / k| value is indeed close to the main peak of the ultrasonically purified simulated sample, indicating that this method can, to a certain extent, reflect the main fragment length of the fragmented sample, that is, indirectly reflect the degree of fragmentation of the sample.

[0082] Furthermore, when X = 0, Y = b. That is, when the target sequence length is 0, the theoretical quantitative value of the sample is b. In other words, when the amplified fragment length is minimal, it essentially reflects the entire length of the gene. This indicates that the b value obtained by this method can, to a certain extent, reflect the true concentration (copy number) of the sample being tested, which has another practical significance.

[0083] Quantitative values ​​include ct values ​​or copy numbers. qPCR is a relative absolute quantitative method that requires the use of standard products to prepare a standard curve, and then the mass of the sample to be tested is calculated based on the conversion relationship between the ct value obtained from the standard curve and the mass of the standard product. The operation is relatively complicated and the reagent cost is high. In addition, a large change in the fluorescence signal value is required to bring about a slight change in the ct value, so the use of dPCR for quantitative detection has better sensitivity. In a specific embodiment, the method of the present application uses dPCR for quantitative detection.

[0084] The present application provides a method for specifically assessing the degree of nucleic acid fragmentation of a target nucleic acid. Compared with capillary gel electrophoresis and microfluidic analysis, the proposed technical solution overcomes the defects of limited detection range of sample fragments, poor fragment analysis capability for low-concentration samples, low quantitative accuracy and precision, and single instrument and reagent functions; compared with existing PCR evaluation methods, it overcomes the problem of inaccuracy caused by a small number of detection targets and errors in determination when the detection target is a sequence such as a virus that is easily mutated. At the same time, a shorter target sequence is used to analyze the presence of very small fragments in the sample, which can more comprehensively and accurately reflect the degree of fragmentation of the entire sample and the true value of the concentration of the target nucleic acid.

[0085] This application discloses a method for quantifying the same sample using multiple target sequence primers of different lengths based on a nucleic acid amplification platform. The method establishes a relationship between the sample quantification values ​​obtained for different target sequences and the target sequence lengths, thereby assessing sample quality and the degree of nucleic acid fragmentation. Nucleic acid amplification, for example, is PCR, such as qPCR or dPCR.

[0086] As used herein, the terms "one or more" and "at least one" are used interchangeably.

[0087] The term "gene" refers to a DNA segment involved in producing a polypeptide chain or transcribed RNA product. It may include regions before and after the coding region (leader and trailer) as well as intervening sequences (introns) between individual coding segments (exons).

[0088] The terms "target nucleic acid," "target gene," "target nucleic acid," or "target" are used interchangeably and refer to a nucleic acid that is the target of amplification, detection, or both.

[0089] The terms "target sequence," "object sequence," or "target nucleic acid sequence" are used interchangeably and refer to a segment of a target nucleic acid to be amplified by an agent that can anneal to a primer under annealing or amplification conditions.

[0090] The term "unfragmented target nucleic acid" refers to the full-length gene that has not been interrupted. For example, taking the GAPDH gene as an example, the unfragmented GAPDH gene refers to the full-length GAPDH gene consisting of 12 introns and 13 exons, with a gene length of approximately 38 kb.

[0091] Herein, at least 2 pairs of primer sets, for example, can be 2 pairs, 3 pairs, 4 pairs, 5 pairs, 6 pairs, 7 pairs, 8 pairs, 9 pairs, 10 pairs, 11 pairs, 12 pairs, 13 pairs, 14 pairs of primer sets, etc.

[0092] The detection label can be a probe that is complementary to the target nucleic acid; or it can be a probe that is not complementary to or identical to the target nucleic acid. In a preferred embodiment, the detection label is a probe that is not complementary to or identical to the target nucleic acid. The term "probe" refers to a labeled oligonucleotide used to detect the presence of a target.

[0093] The term "biological sample" refers to a composition in which one or more target nucleic acids may be present, including patient samples, plant or animal materials, waste materials, materials used for forensic analysis, environmental samples, circulating tumor cell (CTC) free DNA (cfDNA), liquid biopsy samples, biological products, etc. Samples can be blood samples, tissue samples, urine samples, etc. Samples include any tissue, cell, or extract derived from a living or dead organism that may contain target nucleic acids, such as peripheral blood, bone marrow, plasma, serum, biopsy tissue samples including lymph nodes, respiratory tissue or exudates, gastrointestinal tissue, urine, feces, sperm, or other body fluids. Specific target samples are tissue samples (including body fluids) from humans or animals with or suspected of having a disease or condition (particularly viral infection). Other target samples include biological products, such as samples of vaccines and pharmaceuticals. Sample components may include target and non-target nucleic acids, as well as other materials (e.g., salts, acids, bases, detergents, proteins, carbohydrates, lipids, and other organic or inorganic materials). Samples may or may not be treated to purify target nucleic acids prior to amplification. Further processing may be treatment with detergents or denaturants to release nucleic acids from cells or viruses, remove or inert non-nucleic acid components, and concentrate nucleic acids. In some cases, the method of the present application does not include the step of obtaining a biological sample from a subject, i.e., the biological sample is an isolated sample.

[0094] In some embodiments, the biological sample is plasma from a pregnant woman, in which case, for example, the plasma can be analyzed for the presence of fetal circulating DNA.

[0095] In some embodiments, the biological sample is a biological product produced by a host. In this case, the method of the present application can assess the toxicity of the biological product by detecting the residual amount of host nucleic acid in the biological product.

[0096] The terms "nucleic acid," "polynucleotide," and "oligonucleotide" refer to polymers of nucleotides (e.g., ribonucleotides or deoxyribonucleotides) and include naturally occurring (e.g., adenosine, guanosine (guanidine), cytosine, uracil, and thymidine) and non-naturally occurring (artificially modified) nucleic acids. The term is not limited by the length of the polymer (e.g., the number of monomers). Nucleic acids can be single-stranded or double-stranded and typically contain a 5'-3' phosphodiester bond, although in some cases, nucleotide analogs may have other bonds. The monomers are typically referred to as nucleotides.

[0097] The four traditional nucleotide bases are A, T / U, C, and G, of which T is found in DNA and U is found in RNA. The nucleotides found in the target are usually natural nucleotides (deoxyribonucleotides or ribonucleotides), which is also the case for the nucleotides that form the primer.

[0098] The term "amplicon" refers to the product of an amplification reaction, specifically the amplification product obtained by performing a nucleic acid amplification reaction using primers in the presence of a target nucleic acid. "Amplicon" refers to the amplification product. The 5' and 3' boundaries of the amplicon are defined by the forward and reverse primers.

[0099] The term "primer set" refers to a combination of a forward primer and a reverse primer that can amplify a target nucleic acid to produce an amplicon.

[0100] As used herein, the term "forward primer," also known as the upstream primer, is an oligonucleotide that extends uninterruptedly along the negative strand; the term "reverse primer," also known as the downstream primer, is an oligonucleotide that extends uninterruptedly along the positive strand. The positive strand, also known as the sense strand or the positive strand, is generally located at the top of the double-stranded DNA, oriented 5'-3' from left to right, and has a base sequence that is essentially the same as the mRNA of the gene in which the oligonucleotide is located; the primer that binds to this strand is the reverse primer; the negative strand, also known as the nonsense strand or the non-coding strand, is complementary to the positive strand, and the primer that binds to this strand is the forward primer. It should be understood that when the designations of the sense strand and the antisense strand are interchanged, the corresponding forward and reverse primer nomenclature may also be interchanged.

[0101] The term "fragmentation degree" refers to the fragmentation pattern exhibited by nucleic acid molecules, including the locations where breaks and fragmentation occur or the distribution of fragments. For example, nucleases in cells can break DNA into fragments of varying lengths, resulting in different fragmentation patterns. For another example, genomic regions with open chromatin arrangement are more susceptible to binding nucleases and being broken than tight regions. For another example, naked DNA that is not bound to any protein is more easily broken by nucleases, while regions protected by nucleosomes, transcription factors, etc. are less likely to be broken. By analyzing this fragmentation pattern information, it can be used to indicate different nucleic acid molecules.

[0102] The term "comprising", when preceding the recitation of steps or elements, means that the addition of further steps or elements is optional and does not exclude.

[0103] In this article, the term "calculation formula" includes formulas obtained by fitting different fitting models. Exemplarily, the fitting models include a linear model, a quadratic polynomial model, and a cubic polynomial model.

[0104] In some embodiments, the fitted model comprises a linear model.

[0105] In some specific embodiments, the calculation formula fitted by the linear model is Y=kX+b.

[0106] The term "coefficients of a calculation formula" includes parameters used to describe the relationship between different variables. In some cases, the "coefficients of a calculation formula" further include parameters obtained after operation.

[0107] In some specific embodiments, the target sequences of the at least two pairs of primer sets range in length from 30 to 550 bp.

[0108] Preferably, the target sequence lengths of the at least two pairs of primer sets range from 30 to 260 bp, such as 30 to 220 bp, and further such as 30 to 200 bp.

[0109] In some embodiments, the at least two pairs of primer sets share a forward primer or a downstream primer. Using a shared primer approach can reduce development costs.

[0110] In some specific embodiments, the primer set is a set of 3 to 7 pairs or 4 to 7 pairs of primers;

[0111] In some preferred embodiments, the primer set is a set of 4 primer pairs.

[0112] In some specific embodiments, the target sequence length of the primer set is selected from 30-50bp, 50-60bp, 60-70bp, 70-80bp, 80-90bp, 90-100bp, 100-110bp, 110-120bp, 120-130bp, 130-140bp, 140-160bp, 160bp-180bp, 180-200bp, 200-220bp, 220-240bp any two or more of 1-80 bp, 240-260 bp, 260-280 bp, 280-300 bp, 300-320 bp, 320-340 bp, 340-360 bp, 360-380 bp, 380-400 bp, 400-420 bp, 420-440 bp, 440-460 bp, 460-480 bp, 480-500 bp, 500-520 bp, and 520-540 bp.

[0113] In some specific embodiments, the target sequence lengths of the four pairs of primer sets are 30-50 bp, 80-100 bp, 130-150 bp, and 160-180 bp.

[0114] In some specific embodiments, the method of the present application further comprises a step of isolating the nucleic acid in the biological sample before mixing the biological sample with the at least two pairs of primer sets.

[0115] On the other hand, a reagent for detecting a target nucleic acid is provided, the reagent comprising:

[0116] At least two pairs of primers are used to amplify the same target nucleic acid, and the target sequence lengths of the at least two pairs of primers are different from each other.

[0117] In some specific embodiments, when unfragmented target nucleic acid is used as a template, the coefficient of variation (CV value) of the quantitative values ​​between primer sets of the present application is ≤10%.

[0118] In this context, a coefficient of variation of quantitative values ​​between groups of primers of ≤10% means that, using an unfragmented target nucleic acid as a template, the amplification efficiencies of the primers used in each primer set are substantially consistent, i.e., the amplification efficiencies of each primer set are between 90% and 110%. For example, in the presence of three target nucleic acid primer sets, the coefficient of variation of quantitative values ​​between the three primer sets (a first target nucleic acid primer set, a second target nucleic acid primer set, and a third target nucleic acid primer set) is ≤10%.

[0119] In this article, "reaction unit" refers to a basic unit that can react independently in a chemical reaction or a biological reaction. It can be a container, a tiny space, or a system that can accommodate reactants and react. For example, the reaction unit can be a droplet, a reaction chamber of a microfluidic chip, etc. There is no particular limitation on the number of reaction units in this application. It can be appropriately selected according to needs and / or detection objects. For example, the mixed solution can be distributed to 5,000 to 100,000 reaction units, such as 10,000, 20,000, 50,000, or 80,000.

[0120] In some embodiments, the at least two pairs of primer sets share a forward primer or a reverse primer.

[0121] In some specific embodiments, the primer set is preferably a set of 3 to 7 primer pairs, or a set of 4 to 7 primer pairs, and more preferably a set of 4 primer pairs.

[0122] In some specific embodiments, the target sequence lengths of the four pairs of primer sets are 30-50 bp, 80-100 bp, 130-150 bp, and 160-180 bp.

[0123] In some embodiments, the reagent further comprises a detection label.

[0124] In some preferred embodiments, the detection marker is labeled with a detection group.

[0125] On the other hand, a kit for detecting a target nucleic acid is provided, wherein the kit comprises the reagent of the present application.

[0126] Preferably, the kit further comprises a detection marker and a nucleic acid amplification reaction solution.

[0127] In some embodiments, the nucleic acid amplification reaction solution includes dNTPs, a buffer, salt ions, and a polymerase.

[0128] The term "kit" refers to any article of manufacture (eg, a package or container) that includes at least one reagent (such as a primer pair), eg, a reagent described herein for specifically amplifying a target nucleic acid.

[0129] Deoxynucleoside triphosphates, i.e., dNTPs, can be used, for example, dATP, dCTP, dGTP, dTTP, dITP, dUTP, α-thio-dNTP, biotin-dUTP, fluorescein-dUTP, digoxigenin-dUTP, 7-deaza-dGTP. dNTPs are well known in the art and are commercially available.

[0130] The buffers and salts used in the present application provide suitable stable pH and ionic conditions for nucleic acid synthesis, such as reverse transcriptase and DNA polymerase activity. A variety of buffers and salt solutions and modified buffers that can be used in the present application are known in the art, including reagents not specifically disclosed herein. Preferred buffers include, but are not limited to, TRIS, TRICINE, BIS-TRICINE, HEPES, MOPS, TES, TAPS, PIPES, and CAPS. In preferred embodiments, the kits provided include TRIS. In some embodiments, the buffer used in the methods of the present application comprises TRIS at a pH of about 8 to about 9. In certain preferred embodiments, TRIS is TRIS-HCl (pH 8.0-9.0). Salt solutions corresponding to salt ions include, but are not limited to, ammonium sulfate, magnesium chloride, potassium acetate, potassium sulfate, potassium chloride, ammonium chloride, ammonium acetate, magnesium acetate, magnesium sulfate, manganese chloride, manganese acetate, manganese sulfate, sodium solution, chloride, sodium acetate, lithium chloride, and lithium acetate. In preferred embodiments, the compositions provided include ammonium sulfate and magnesium chloride. In certain preferred embodiments, the compositions provided include TRIS-HCl, ammonium sulfate, and magnesium chloride.

[0131] For polymerase, the application has no particular restrictions. Common polymerases in the field of diagnostic reagents can be used, and exemplary polymerases can use existing known polymerases derived from thermotolerant bacteria. In a specific example, the DNA polymerase (US Patent Nos. 4,889,818 and 5,079,352) (trade name Taq polymerase) derived from aquatic thermophilic bacterium (Thermus aquaticus), the DNA polymerase (WO 91 / 09950) (rTth DNA polymerase) derived from thermophilic thermophilic bacterium (Thermus thermophilus), the DNA polymerase (WO 92 / 9689) (Pfu DNA polymerase, manufactured by Stratagenes) derived from extreme thermophilic bacterium (Pyrococcus furiosus), the DNA polymerase (EP-A455430 (trademark Vent) derived from the shore thermococcus (Thermococcus litoralis): manufactured by New England Biolabs) etc. can be commercially available, wherein the heat-resistant polymerase derived from aquatic thermophilic bacterium (Thermus aquaticus) is preferred.

[0132] For the type of detection group, the present application has no particular restrictions, and any common fluorescent group, quenching group and modifying group commonly used in the field of diagnostic reagents can be used. Exemplary fluorescent groups can be selected from various fluorescent labels, such as ALEX-350, FAM, VIC, TET, CAL FluorGold 540, JOE, HEX, CAL Fluor Orange 560, TAMRA, CAL Fluor Red 590, ROX, CAL Fluor Red 610, TEXAS RED, CAL Fluor Red 635, Quasar 670, CY3, CY5, CY5.5, Quasar 705, one or more; exemplary quenching groups can be selected from various quenchers, such as DABCYL, BHQ class (such as BHQ-1 or BHQ-2), ECLIPSE and / or TAMRA, one or more. Exemplary modifying groups are, for example, 3'MGB.

[0133] The evaluation of the degree of fragmentation of the target nucleic acid can be used to evaluate the quality of biological samples. For example, when testing FFPE samples, it is often necessary to perform second-generation sequencing on the samples, but second-generation sequencing has extremely high requirements for the degree of fragmentation of the samples. By using the evaluation method of the present application, the degree of fragmentation of biological samples can be effectively, accurately and sensitively evaluated, which is conducive to guiding subsequent experiments. Using the method of the present application, the degree of fragmentation can be evaluated by the |b / k| value. For another example, when conducting a safety assessment of biological products, such as when preparing a vaccine, it is often necessary to evaluate the host cell residues in the vaccine, and there are strict requirements on the DNA fragmentation size of the host cells. Using the method of the present application, the size of the DNA fragments of the host cells in the vaccine can be evaluated with high sensitivity by the |b / k| value. In the method of the present application, evaluating the degree of fragmentation of the target nucleic acid based on the |b / k| value means that the |b / k| value reflects the length of the main peak of the main fragment in the sample.

[0134] In addition, some samples not only need to understand the degree of nucleic acid fragmentation, but also need to consider relevant information such as the quality and concentration of the nucleic acid from multiple dimensions. For example, when performing nucleic acid testing on tumors, characterizing the degree of fragmentation and content of tumor genes can often guide disease diagnosis. For another example, when evaluating drug efficacy, it may be necessary to comprehensively consider the degree of viral nucleic acid fragmentation and content before and after medication. Based on this, the inventors found that after using primers that amplify target sequences of different lengths to detect specific target nucleic acids, the formula obtained by fitting the target sequence length and the detection copy number of the corresponding length can effectively solve the difficulties of the prior art. In the formula, the obtained |b / k| value can evaluate the degree of fragmentation of the target nucleic acid, and the b value can obtain data that is extremely close to the true content value, and this method can obtain two data at the same time, which can effectively ensure the correlation between the data.

[0135] In another aspect, a method for detecting a target nucleic acid is provided, the method comprising:

[0136] Mixing a biological sample with at least two pairs of primer sets, a detection marker, and a nucleic acid amplification reaction solution to obtain a first mixed solution; wherein the at least two pairs of primer sets have target sequences of different lengths and amplify the same target nucleic acid;

[0137] Randomly distributing the first mixed solution into a plurality of reaction units to perform a nucleic acid amplification reaction to obtain detection copy numbers of different target sequence lengths;

[0138] Fitting a first calculation formula according to the length of the target sequence and the detected copy number of the corresponding length, in which the variable Y is the detected copy number of the target sequence and the variable X is the length of the target sequence; and

[0139] The fragmentation degree of the target nucleic acid is estimated based on the intersection of the fitting curve of the first calculation formula and the X-axis, and / or the copy number of the target nucleic acid is estimated based on the coefficient of the first calculation formula.

[0140] In some embodiments, the method further comprises the step of isolating nucleic acids in the biological sample before mixing the biological sample with the at least two pairs of primer sets.

[0141] In some specific embodiments, in the step of fitting the first calculation formula, a linear model, a quadratic polynomial model, or a cubic polynomial model is fitted according to the length of the target sequence and the detected copy number of the corresponding length.

[0142] In some specific embodiments, the first calculation formula fitted by the linear model is Y=kX+b.

[0143] In some embodiments, the fragmentation degree of the target nucleic acid is assessed based on the |b / k| value, and / or the copy number of the target nucleic acid is assessed based on the b value.

[0144] In some embodiments, when unfragmented target nucleic acid is used as a template, the coefficient of variation of the quantitative values ​​between the at least two pairs of primer sets is ≤10%.

[0145] In some specific embodiments, the target sequence lengths of the at least two pairs of primer sets range from 30 to 540 bp, preferably from 30 to 260 bp.

[0146] Preferably, the at least two pairs of primer sets share a forward primer or a downstream primer.

[0147] In some specific embodiments, the primer set is 3 to 7 pairs, 4 to 7 pairs of primers, preferably 4 pairs of primers; more preferably, the target sequence length of the 4 pairs of primers is 30 to 50 bp, 80 to 100 bp, 130 to 150 bp, or 160 to 180 bp.

[0148] In some embodiments, the method further comprises:

[0149] mixing a control biological sample with at least two pairs of primer sets, a detection marker, and a nucleic acid amplification reaction solution to obtain a second mixed solution;

[0150] Randomly distributing the second mixed solution into a plurality of reaction units to perform nucleic acid amplification reaction to obtain detection copy numbers of different target sequence lengths;

[0151] Fitting a second calculation formula based on the length of the target sequence and the detected copy number of the corresponding length, in which the variable Y is the detected copy number of the target sequence and the variable X is the length of the target sequence; and

[0152] The intersection of the fitting curves of the first calculation formula and the second calculation formula with the X-axis is compared to evaluate the fragmentation degree of the target nucleic acid, and / or the coefficients of the first calculation formula and the second calculation formula are compared to evaluate the copy number of the target nucleic acid.

[0153] In some specific embodiments, in the step of fitting the second calculation formula, a linear model, a quadratic polynomial model, or a cubic polynomial model is fitted according to the length of the target sequence and the detected copy number of the corresponding length.

[0154] In some specific embodiments, the second calculation formula fitted by the linear model is Y=k1X+b1.

[0155] In some embodiments, the |b / k| value is compared to the |b1 / k1| value to assess the degree of fragmentation of the target nucleic acid, and / or the b value is compared to the b1 value to assess the copy number of the target nucleic acid.

[0156] In some specific embodiments, the |b / k| value reflects the length of the main peak of the major fragment in the biological sample, and the |b1 / k1| value reflects the length of the main peak of the major fragment in the control biological sample.

[0157] In some specific embodiments, a |b / k| value less than a |b1 / k1| value indicates that the main peak of the main fragment of the biological sample is smaller relative to the control biological sample; a |b / k| value greater than a |b1 / k1| value indicates that the main peak of the main fragment of the biological sample is larger relative to the control biological sample; or a |b / k| value equal to a |b1 / k1| value indicates that the main peak of the main fragment of the biological sample has not changed relative to the control biological sample.

[0158] In some specific embodiments, a b value less than a b1 value indicates that the number of copies of the target nucleic acid in the biological sample is less than that in the control biological sample; a b value greater than a b1 value indicates that the number of copies of the target nucleic acid in the biological sample is greater than that in the control biological sample; or a b value equal to a b1 value indicates that the number of copies of the target nucleic acid in the biological sample is equal to that in the control biological sample. In some specific embodiments, the control biological sample includes a biological sample from a normal subject or a biological sample of a subject before administration of a drug.

[0159] In some embodiments, the method further comprises the step of isolating nucleic acid in the control biological sample before mixing the control biological sample with the at least two pairs of primer sets.

[0160] In some specific embodiments, the target sequence lengths of the at least two pairs of primer sets range from 30 to 540 bp, preferably from 30 to 260 bp.

[0161] In some embodiments, the at least two pairs of primer sets share a forward primer or a downstream primer.

[0162] In some specific embodiments, the primer set is a set of 3 to 7 pairs or 4 to 7 pairs of primers.

[0163] In some preferred embodiments, the primer set is a set of 4 primer pairs.

[0164] In some preferred embodiments, the target sequence lengths of the four pairs of primers are 30-50 bp, 80-100 bp, 130-150 bp, and 160-180 bp.

[0165] Another object of the present application is to provide a method for accurately detecting cfDNA, the method comprising the following steps:

[0166] Isolation of cfDNA from biological samples;

[0167] Mixing the cfDNA with at least two pairs of target nucleic acid primers, a pair of gDNA primers, a detection marker, and a nucleic acid amplification reaction solution; wherein the at least two pairs of primers have target sequences of different lengths and amplify the same target nucleic acid;

[0168] After mixing, the mixed solution is randomly distributed into a plurality of reaction units to perform a nucleic acid amplification reaction; the detection copy number of the target sequence corresponding to at least two pairs of target nucleic acid primer sets and the detection copy number of a pair of gDNA primer sets are obtained;

[0169] A calculation formula is provided for fitting the copy number obtained by deducting the copy number detected by a pair of gDNA primer sets from the detected copy number of the target sequence of each pair of at least two pairs of target nucleic acid primer sets and the target sequence length of the corresponding target nucleic acid primer set, wherein the variable Y is the copy number of the target sequence after deduction, and the variable X is the length of the target sequence; and

[0170] The degree of cfDNA fragmentation is evaluated based on the intersection of the fitting curve of the calculation formula and the X-axis, and / or the copy number of cfDNA is evaluated based on the coefficient of the calculation formula.

[0171] In a specific embodiment, in the fitting step, a linear model, a quadratic polynomial model, or a cubic polynomial model is fitted based on the copy number obtained after deducting the copy number of a pair of gDNA primer sets from the copy number of each corresponding target sequence and the target sequence length of the corresponding target nucleic acid primer set.

[0172] In some specific embodiments, the calculation formula fitted by the linear model is Y=kX+b.

[0173] In some embodiments, the degree of fragmentation of cfDNA is assessed based on the |b / k| value, and / or the copy number of cfDNA is assessed based on the b value.

[0174] In some embodiments, when unfragmented target nucleic acid is used as a template, the coefficient of variation of the quantitative values ​​between the at least two pairs of primer sets is ≤10%.

[0175] In some specific embodiments, the target sequence lengths of the at least two pairs of primer sets range from 30 to 420 bp, preferably from 30 to 260 bp.

[0176] In some specific embodiments, the target sequence length of the pair of gDNA primers is ≥466 bp; preferably, the target sequence length of the pair of gDNA primers is ≥498 bp.

[0177] In some embodiments, the at least two pairs of primer sets share a forward primer or a downstream primer.

[0178] In some specific embodiments, the primer set is a set of 3 to 7 pairs or 4 to 7 pairs of primers.

[0179] In some preferred embodiments, the primer set is a set of 4 primer pairs.

[0180] In some preferred embodiments, more preferably, the target sequence lengths of the four pairs of primer sets are 30-50 bp, 80-100 bp, 130-150 bp, and 160-180 bp.

[0181] In various methods of the present application, the target sequences of at least two primer pairs can be named, in order of length from short to long, as a first target sequence, a second target sequence, ..., an Nth target sequence. The first target sequence is the shortest, the Nth target sequence is the longest, and the intermediate target sequences are arranged in order of length.

[0182] In some specific embodiments, the length difference between adjacent target sequences is 8 to 90 bp, such as 9 bp, 10 bp, 11 bp, 12 bp, 13 bp, 14 bp, 15 bp, 20 bp, 25 bp, 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, or 90 bp. For example, in the presence of four pairs of primer sets, the second target sequence and the first target sequence are adjacent target sequences, the second target sequence and the third target sequence are adjacent target sequences, and the third target sequence and the fourth target sequence are adjacent target sequences.

[0183] In some embodiments of the various methods of the present application, two or more of the at least two pairs of primer sets share a forward primer or a downstream primer.

[0184] In some embodiments of the various methods of the present application, the method is ex vivo or in vitro. In some embodiments of the various methods of the present application, the method is for non-diagnostic purposes.

[0185] In another aspect, an electronic device is provided, comprising:

[0186] a memory for storing executable instructions; and

[0187] The processor is used to execute the executable instructions stored in the memory to implement the method of the present application.

[0188] In another aspect, a computer-readable storage medium is provided, storing executable instructions for causing a processor to execute the instructions to implement the method of the present application.

[0189] On the other hand, a computer program product is provided, comprising a computer program or computer executable instructions, wherein when the computer program or computer instructions are executed by a processor, the method of the present application is implemented.

[0190] Unless otherwise specified, the terms "first", "second", "third" and "fourth" in the embodiments of the present application are only used to distinguish multiple similar elements, and are not intended to indicate any differences in importance or order between the elements.

[0191] The following are preferred embodiments of the present invention, and the present invention is not limited to the following preferred embodiments. It should be noted that, for those skilled in the art, any modifications and improvements made based on this inventive concept fall within the scope of protection of the present invention. The reagents used, for which the manufacturer is not indicated, are all commercially available conventional products.

[0192] In this application, the plasmids, primers and probe sets used were synthesized and provided by General Biotech (Anhui) Co., Ltd. Digital PCR detection was performed using a MacBio D600 digital PCR instrument.

[0193] Example 1 Evaluation of the degree of nucleic acid fragmentation of the human GAPDH gene

[0194] The human GAPDH gene was selected as the detection target, and four pairs of specific primers were designed. The forward primer was fixed and the reverse primer was mobile, thereby obtaining target sequences of different lengths (39, 86, 137, and 166 bp).

[0195] Digital PCR was performed on genomic DNA from 293T cells using four specific primer pairs. Each primer pair was tested 10 times, with repeatability coefficients (CVs) of 1.21%, 5.09%, 1.11%, and 5.05%, respectively. The intergroup CV for the four primer pairs was 3.68%. The sequences of the four primer pairs and their detection probes are shown in Table 1.

[0196] Table 1

[0197] Among them, the 5' end of GAPDH-P1 is modified with FAM and the 3' end is modified with BHQ1; the 5' end of GAPDH-P2 is modified with CY5 and the 3' end is modified with BHQ2; the 5' end of P is modified with ROX and the 3' end is modified with BHQ2.

[0198] Genomic DNA from 293T cells was sheared for 200 seconds using a Covaris M220 ultrasonic disruptor, following the sonication protocol in Table 2. The sheared samples were purified using Beckman Coulter magnetic beads (Agencourt AMPure XP Nucleic Acid Purification Kit, 60 mL Kit). Four fragments of varying sizes were obtained, each exhibiting a normal distribution, with major peaks of 171, 274, 364, and 466 bp, respectively. The purified nucleic acid samples were analyzed using an Agilent 4150 Fragment Analyzer.

[0199] Table 2

[0200] Digital PCR detection was performed on the nucleic acid samples that were ultrasonically sheared and purified using magnetic beads using the four primer pairs in Table 1. The detection system is shown in Table 3, the reaction procedure is shown in Table 4, and the detection results are shown in Table 5.

[0201] Table 3

[0202] Table 4

[0203] Table 5

[0204] For each nucleic acid fragment sample that was ultrasonically sheared and purified, the relationship between the target sequence length and the corresponding quantitative value of four pairs of primers was established, and it was found that they all showed a linear relationship y=kx+b, R 2 They are 0.9974, 0.9997, 0.9823 and 0.9906 respectively. The specific data are shown in Table 6.

[0205] Table 6

[0206] Analysis of the above data revealed that for the same ultrasonically fragmented and purified fragmented sample, the longer the target sequence, the lower the corresponding quantitative value, and a good linear decline relationship between the target sequence and the quantitative value. For different ultrasonically fragmented and purified fragmented samples, the shorter the purified sample fragment length, the more obvious the linear decline relationship between the target sequence and the quantitative value, and the larger the |k|.

[0207] For the linear equation y = kx + b, when y = 0, |x| = |b / k|. This point is the intersection of the line and the x-axis. This means that for ultrasonically purified fragmented samples, as the target sequence length increases, the quantitative value decreases. When the target sequence length ≥ |b / k|, the theoretical quantitative value is 0, and when the target sequence length exceeds the template length, the theoretical quantitative value is definitely 0. This indicates that the |b / k| obtained by this method can reflect the fragment length of the sample being tested and is of practical significance. Analysis of data from ultrasonically purified fragmented samples with main peaks of 171, 274, 364, and 466 bp yielded |b / k| values ​​of 178, 290, 381, and 400, respectively. These |b / k| values ​​are close to the sizes of the main peaks in the samples, demonstrating the feasibility of this method in reflecting the length of the main fragments in the samples.

[0208] Example 2: Evaluating the effect of the number of target sequences on the |b / k| value

[0209] The GAPDH gene was selected as the detection target, and seven specific primer sets were designed. The forward primer was fixed, and the reverse primer was mobile, resulting in target sequences of 39, 50, 67, 86, 104, 137, and 166 bp in length. These seven specific primer pairs were used for digital PCR detection of genomic DNA from 293T cell lines. Each primer pair was tested 10 times. The repeatability CV values ​​for each primer pair were 10.62%, 2.70%, 3.18%, 7.36%, 6.23%, 12.33%, and 1.58%, respectively. The inter-group CV for the quantitative values ​​of the seven primer pairs was 6.58%, indicating good consistency in the quantitative values ​​of the seven primer pairs. The sequences of the seven primer pairs are shown in Table 7.

[0210] Table 7

[0211] Among them, the 5' end of GAPDH-P1 is modified with FAM and the 3' end is modified with BHQ1; the 5' end of GAPDH-P2 is modified with CY5 and the 3' end is modified with BHQ2; the 5' end of P is modified with ROX and the 3' end is modified with BHQ2; the 5' end of GAPDH-P4 is modified with VIC and the 3' end is modified with BHQ1.

[0212] Digital PCR was performed using these seven pairs of specific primers and the four pairs of specific primers from Example 1 on genomic nucleic acid samples from 293T cell lines that had been sonicated for 200 seconds. Specific test results are shown in Table 8. Simultaneously, an Agilent 4150 fragment analyzer was used to analyze the samples after 200 seconds of sonication, revealing a major peak at 274 bp.

[0213] Table 8

[0214] The relationship between the target sequence length and the corresponding quantitative value mean was established for 4 primer pairs (target sequence length 39, 86, 137, 166 bp) and 7 primer pairs (target sequence length 39, 50, 67, 86, 104, 137, 166 bp), and the linear relationship established for the 4 primer pairs was y = -4.6704x + 1339.5, R 2 =0.9905, calculated |b / k| value = 287, the linear relationship established by 7 primer pairs is y = -4.8585x + 1362.5, R 2 = 0.9832, and the calculated |b / k| value is 280. These results indicate that for nucleic acid samples of genomic DNA from 293T cells sonicated for 200 seconds, there is no difference in the calculated |b / k| values ​​between 4 and 7 points. This suggests that for the same fragmented sample, the effect of the number of target sequences in the range of 4 to 7 on the |b / k| value is essentially the same.

[0215] Example 3: Evaluating the Effect of Different Concentrations of the Same Sample on |b / k|

[0216] The GAPDH gene was selected as the detection target, and four pairs of specific primers were used to amplify target sequences of 39, 86, 137, and 166 bp, respectively. The primer sets and detection probes were the same as those in Example 1.

[0217] Nucleic acid samples extracted from plasma and genomic DNA from 293T cells sonicated for 200 seconds were serially diluted three-fold to prepare four concentrations: stock solution, 3x, 9x, and 27x. Digital PCR was performed on these eight samples using four pairs of primers specific for the GAPDH gene. The results are shown in Tables 9 and 10.

[0218] Table 9 Nucleic acid samples of DNA sonication for 200s

[0219] Table 10 Nucleic acid samples extracted from plasma

[0220] For nucleic acid samples extracted from genomic DNA from 293T cells sonicated for 200 seconds (the main peak at 274 bp was detected on an Agilent 4150 fragment analyzer for the original solution, 3x, and 9x dilutions, and the 27x dilution was undetectable), the |b / k| values ​​for the four sample concentrations were 288, 299, 289, and 288, respectively, with a calculated CV of 2.07%. For nucleic acid samples extracted from plasma (the main peak at 220 bp was detected on an Agilent 4150 fragment analyzer for the original solution, 3x, and 9x dilutions, and the 27x dilution was undetectable), the |b / k| values ​​for the four sample concentrations were 222, 221, 217, and 222, respectively, with a calculated CV of 1.11%. These results indicate that there is no difference in the calculated |b / k| values ​​for different concentrations of fragmented nucleic acid samples, indicating that sample concentration has no effect on |b / k| values.

[0221] Example 4 Evaluation of primer sets at different positions and their effects on |b / k|

[0222] The GAPDH gene was selected as the detection target. Primer design scheme 1: the primer in one direction was fixed and the primer in the other direction was moved to obtain target sequences of different lengths (39, 86, 137, and 166 bp). The specific experimental process and results were the same as those in Example 1.

[0223] Primer Design Option 2: No fixed primers were used. Target sequences of varying lengths (50, 88, 145, and 167 bp) were obtained through primer design. Primer sequences and probes are shown in Table 11. Digital PCR was performed on genomic DNA from 293T cells using the four specific primer pairs from Primer Design Option 2 for the GAPDH gene. The repeatability coefficients (CVs) for each primer pair, after 10 replicates, were 3.56%, 3.51%, 1.10%, and 2.65%, respectively.

[0224] Table 11

[0225] The 5' end of P was modified with ROX, and the 3' end was modified with BHQ2. Using the four pairs of specific primers in Primer Design Scheme 2, digital PCR detection and formula fitting were performed on four ultrasonically purified fragmented samples (prepared in Example 1, with main peaks of 171, 274, 364, and 466 bp, respectively). The detection results are shown in Table 12.

[0226] Table 12

[0227] For primer design scheme 2, the |b / k| values ​​obtained after testing and formula fitting of ultrasonically purified fragment samples with main peaks of 171, 274, 364, and 466 bp were 174, 301, 377, and 414, respectively. The CV values ​​for each main peak were 1.74%, 2.75%, 0.72%, and 2.39%, respectively. For primer design scheme 1 of Example 1, the |b / k| values ​​obtained after testing and formula fitting of ultrasonically purified fragment samples with main peaks of 171, 274, 364, and 466 bp were 178, 290, 381, and 400, respectively. These results indicate that different primer design systems yield very similar detection results for samples with the same main peak.

[0228] Example 5: Evaluating the Effect of Different Sample Extractions on |b / k|

[0229] The Tiangen Extraction Kit (Magnetic Bead Large Volume Free Nucleic Acid Extraction Kit, Article No.: DP710-02) was selected. According to the instructions, a human plasma sample was divided into 5 equal parts, free DNA was extracted in parallel, and the sample concentration was determined using qubit3.0. The digital PCR detection system in Example 1 was used to detect free DNA. Each sample was tested 10 times using 4 pairs of primers. The target sequence lengths of the 4 pairs of primers and the corresponding quantitative values ​​were used to establish a relationship for linear fitting. The |b / k| of each sample was calculated. The test data are shown in Table 13. At the same time, the Guangding qsep fragment analyzer was used to perform fragment analysis on the nucleic acid samples, and the average fragment size from the lower marker to the upper marker was calculated. The calculation formula is average fragment size = avg.bp1 size * avg.bp1 proportion + avg.bp2 size * avg.bp2 proportion + ... + avg.bp n size * avg.bp n proportion. The results are shown in the following table.

[0230] Table 13

[0231] Data analysis showed that when the same sample was extracted in parallel with the same extraction reagent and tested, the obtained |b / k| value CV was <5%, that is, different sample extractions had no effect on the |b / k| value.

[0232] Example 6 Evaluation of the degree of fragmentation of HBV samples

[0233] The HBV-S region was selected as the detection target, and four pairs of specific primers were designed from conserved regions to amplify target sequences of 48, 87, 131, and 175 bp, respectively. The primer sequences and detection probes are shown in Table 14. These four pairs of specific primers were used for digital PCR detection of HBV-B plasmids. Each primer pair was tested 10 times, and the repeatability CVs of the quantitative values ​​for each primer pair were calculated to be 4.1%, 4.3%, 3.5%, and 0.2%, respectively. The total CV of all quantitative values ​​for the four primer pairs was 5.32%. This result demonstrates good consistency in the quantitative values ​​of the four primer pairs.

[0234] Table 14

[0235] The 5' end of P was modified with ROX, and the 3' end was modified with BHQ2. The HBV-B plasmid was ultrasonically fragmented for 420 s using a Covaris M220 ultrasonic disruptor according to the ultrasonic program in Table 2.

[0236] Ultrasonic fragmentation samples were purified using Beckman Coulter magnetic beads (Agencourt AMPure XP Nucleic Acid Purification Kit, 60 mL Kit). The purified nucleic acid samples were analyzed using an Agilent 4150 Fragment Analyzer. The resulting fragments showed major peaks of 238, 320, 460, and 718 bp, respectively, with a normal distribution. Digital PCR was performed on these four fragmented samples using four pairs of specific primers. Each primer pair was tested 10 times, and the repeatability (CV) of the quantitative values ​​for each primer pair was calculated. The results are shown in Table 15.

[0237] Table 15

[0238] For each nucleic acid fragment sample that was ultrasonically sheared and purified, the relationship between the target sequence length and the corresponding quantitative value of the four pairs of primers was established and the |b / k| value was calculated. It was found that all of them showed a linear relationship y=kx+b, R 2 They are 0.993, 0.9959, 0.9997 and 0.9921 respectively.

[0239] The linear trend obtained by HBV modeling was consistent with that of GAPDH, both showing that the longer the target sequence length, the lower the corresponding quantitative value, and there was a good linear decreasing relationship between the target sequence and the quantitative value. For different fragmented samples that were ultrasonically interrupted and purified, the shorter the length of the purified sample fragment, the more obvious the linear decreasing relationship between the target sequence and the quantitative value, and the larger the |k|. Analysis of the data of the fragmented samples with main peaks of 238, 320, 460, and 718 bp after ultrasonic purification showed that the obtained |b / k| were 211, 308, 400, and 547, respectively. The |b / k| was close to the size of the main peak of the simulated samples, indicating that this method can reflect the length of the main fragments of the samples to a certain extent.

[0240] Comparative Example

[0241] A viral extraction kit (Magnetic Bead-Based Large-Volume Free Nucleic Acid Extraction Kit, MacBio, Catalog No. DP710-02) was used according to the manufacturer's instructions to extract plasma samples from six HBV patients. Sample concentrations were determined using Qubit 3.0, and all concentrations were 0. An Agilent 4150 Fragment Analyzer was used to analyze these samples; no fragment information was obtained for any of them. Digital PCR was performed on these samples, with each primer pair (the same primer-probe pair as in Example 6) being used without duplication. A relationship between the target sequence length and the corresponding quantitative value for each of the four primer pairs was established, and a linear fit was performed. The |b / k| value for each sample was calculated. The test data are shown in Table 16.

[0242] Table 16

[0243] Analyzing the above 6 clinical sample data, the R value of the linear relationship established between the target sequence length and the corresponding quantitative value was 2 They are 0.97, 0.9614, 0.8616, 0.8806, 0.6439, and 0.9878 respectively. 2 =0.6439 was sequenced for the first generation, and it was found that there was a primer pair with a mismatch between the primer and the target sequence, which resulted in a linear R 2 The results of clinical sample detection suggest that when evaluating the degree of nucleic acid fragmentation, the possibility of detecting target sequence variations should be considered. In addition, the fragmentation information of nucleic acids such as viruses that infect the host cannot be distinguished from the fragment information of the host using fragment analyzers such as qubit3.0 or Agilent 4150 alone. However, digital PCR can be used to specifically detect viral targets to obtain the corresponding nucleic acid fragmentation information and quantitative information. At the same time, the detection sensitivity of digital PCR for viral clinical samples is higher than that of fragment analyzers such as qubit3.0 or Agilent 4150.

[0244] Example 7 Evaluation of the Fragmentation Degree of Escherichia coli Samples

[0245] The degree of fragmentation of Escherichia coli nucleic acid samples was assessed. The E. coli ydiJ gene was selected as the detection target, and five pairs of specific primers were designed from conserved regions to amplify target sequences of 65, 74, 91, 100, and 109 bp, respectively. The primer sequences and detection probes are shown in Table 17. Digital PCR detection of E. coli nucleic acid was performed using these five pairs of specific primers. The quantitative values ​​(CVs) for each primer pair were ≤15% after 10 replicates. The intergroup CVs for the five primer pairs were 3.39%, indicating good consistency in the quantitative values.

[0246] Table 17

[0247] Among them, EC-P 5' end is modified with T6 FAM and 3' end is modified with 3'MGB.

[0248] Using a Covaris M220 ultrasonic disruptor, the Escherichia coli nucleic acid sample was ultrasonically fragmented for 420 seconds according to the ultrasonic program in Table 2. The ultrasonically fragmented sample was purified using Beckman Coulter magnetic beads (Agencourt AMPure XP Nucleic Acid Purification Kit 60mL Kit). The purified nucleic acid sample was analyzed using an Agilent 4150 fragment analyzer to obtain a fragmented sample with a main peak of 96bp, and the fragments were normally distributed. Five pairs of specific primers were used to perform digital PCR detection on the digital PCR instruments of Qiagen, Bio-Rad, and Michael, respectively. For the three digital PCR platforms, the relationship between the target sequence length and the corresponding quantitative value of the five pairs of primers was established, and it was found that they all showed a linear relationship y=kx+b. The calculated |b / k| values ​​were 120, 117, and 117, respectively, with a CV of 1.34%. Detailed data are shown in Table 18.

[0249] Table 18

[0250] These results indicate that this fragmentation degree assessment method can obtain consistent |b / k| values ​​on different digital PCR platforms and is not affected by the instrument used.

[0251] Example 8 Assessment of gDNA Contamination in cfDNA

[0252] When separating plasma from whole blood samples, sometimes due to reasons such as sample hemolysis, improper storage, or improper operation, a certain amount of gDNA or blood cell contamination may be present in the plasma sample, resulting in a certain degree of gDNA contamination in the extracted cfDNA. This can cause inaccurate cfDNA values ​​and affect the application of cfDNA values. For example, it can cause inaccurate analysis of cfDNA content in healthy individuals and cancer patients, and affect the accuracy of calculating somatic mutation abundance, thereby affecting clinical precision medicine. Therefore, it is necessary to assess the degree of gDNA contamination in cfDNA and use certain technical methods to obtain the gDNA contamination amount and restore the true cfDNA content, thereby improving the accuracy of test results and the reliability of data analysis.

[0253] Literature indicates that 70% of cfDNA fragments are haploid, approximately 166 bp in size, with a small number of diploid and triploid fragments also present. Therefore, fragments > 498 bp can be considered gDNA contamination. This application designed the following experimental protocol to detect and subtract gDNA contamination from cfDNA.

[0254] First, cfDNA was extracted from fresh plasma from 10 healthy individuals using the Tiangen Extraction Kit (Magnetic Bead-Based Large-Volume Free Nucleic Acid Extraction Kit, Cat. No. DP710-02) according to the manufacturer's instructions. The cfDNA extracted from these 10 samples was then mixed. Different ratios of gDNA were added to these samples (simulating gDNA contamination with increasing gDNA content, with cfDNA:gDNA ratios of 1:0.1, 1:0.5, 1:1, 1:3, and 1:10, respectively) to simulate cfDNA samples with varying degrees of gDNA contamination.

[0255] A detection system of four pairs of short target sequence primers for GAPDH (39, 86, 137, and 166 bp) was used, as in Example 1;

[0256] A gDNA primer pair detection system, consisting of long target sequence primers selected from a GAPDH gene primer set with a target sequence length >498 bp. The intergroup CV of quantitative values ​​for this long target sequence primer and four pairs of short target sequence primers for the same sample was ≤10%. The reproducibility CVs for each primer pair were 8.7%, 3.8%, 8.0%, 3.3%, and 1.7%, respectively, for a total intergroup CV of 5.42%. The gDNA primer probe sequences are shown in Table 19.

[0257] Table 19

[0258] GAPDH-PC is modified with ROX at its 5' end and BHQ2 at its 3' end. Simulated contaminated cfDNA was mixed with four pairs of short target sequence primers and one gDNA primer combination detection system, and then tested by digital PCR. The results are shown in Table 20.

[0259] Table 20

[0260] The calculation formula after subtraction is fitted. Compared with the calculation formula without subtraction, the |b / k| value after subtraction is closer, as shown in the table below.

[0261] Table 21

[0262] In Table 21, for gDNA-contaminated samples, the calculated |b / k| values ​​before gDNA contamination deduction were 265, 306, 466, 764, and 2099, respectively; while the |b / k| values ​​after deduction were 226, 232, 275, 260, and 364, respectively, which were basically around 257. Among them, the |b / k| value of cfDNA:gDNA = 1:10 was relatively high. When analyzing the high content, some gDNA fragmentation was too high, affecting the overall content of small fragments. Therefore, it can be considered that the method of the present application can accurately assess gDNA contamination when gDNA contamination is relatively light; for cases with heavy gDNA contamination, gDNA deduction can further avoid inaccurate detection caused by gDNA contamination.

[0263] In addition, the method described above in this embodiment can also be used to detect the degree of nucleic acid sample fragmentation before and after sample storage, thereby further evaluating the sample quality. Specifically:

[0264] Whole blood samples were frozen and thawed three times at -20°C for two samples, frozen and thawed seven times at -20°C for two samples, and left at room temperature for one day for one sample. Plasma was then separated from the whole blood, and cfDNA was extracted using the Tiangen Extraction Kit (Magnetic Bead-Based Large-Volume Cell-Free Nucleic Acid Extraction Kit, Catalog No. DP710-02) according to the manufacturer's instructions. Digital PCR quantification was performed using the five primer pairs used for the GAPDH mock sample assay, with three replicates for each primer pair. Detailed test results are shown in Table 22. Following the above method, the |b / k| values ​​were calculated using the quantitative values ​​of four short target sequences (39, 86, 137, and 166 bp) without and with gDNA contamination removed, respectively. Detailed results are shown in Tables 22 and 23.

[0265] Table 22

[0266] Table 23

[0267] The data analysis revealed that (1) before deducting gDNA contamination, the obtained |b / k| values ​​ranged from 461 to 683, which was greater than the |b / k| values ​​of 198 to 268 for the fresh plasma of 10 healthy individuals tested; (2) after deducting gDNA contamination using 499 bp, the obtained |b / k| values ​​ranged from 175 to 241, which was close to the |b / k| values ​​of 198 to 268 for the fresh plasma of 10 healthy individuals tested.

[0268] The data results of simulated samples and whole blood separated plasma showed that the difference between the quantitative values ​​of long target sequences and short target sequences was used to represent the true quantitative value of cfDNA under each short target sequence. The |b / k| value obtained based on the true quantitative value of cfDNA fell near the |b / k| value obtained from fresh plasma of healthy people. Therefore, it was determined that this method can be used to deduct gDNA contamination from cfDNA.

[0269] Example 9 Evaluation of the distribution of residual host nucleic acid fragments in biological products

[0270] The residual fragment size distribution of residual host nucleic acids in intermediates, semi-finished products, and finished products of biological products derived from the HEK293 host cell line was evaluated. The risk gene E1A of the HEK293 host cell was selected as the detection target, and four pairs of specific primers were designed to simultaneously amplify target sequences of 83, 132, 217, and 571 bp in length in one tube. The primer sequences and detection probes are shown in Table 24, and the corresponding target sequences are shown in Table 25. Using these four pairs of specific primers, digital PCR detection was performed on 0.1, 1, and 10 ng / ul of HEK293 host cell line genomic DNA. Each concentration was tested three times, and the inter-group quantitative value CV of the four primer pairs was calculated for each test. The inter-group quantitative value CV of all four primer pairs tested was <5%. The data are shown in Table 26. This result shows that the four primer pairs have good consistency in quantitative values. Among the four primer pairs, three pairs with target sequences of 83, 132, and 217 bp were used to evaluate fragments, and the primer pair with a target sequence of 571 bp was used to determine whether the size of the residual host nucleic acid fragments met regulatory requirements.

[0271] Table 24

[0272] Among them, the 5' end of E1A-P2 is modified with ROX and the 3' end is modified with BHQ2; the 5' end of E1A-P8-FAM is modified with 6-FAM and the 3' end is modified with BHQ1; the 5' end of E1A-P7-cy5 is modified with Cy5 and the 3' end is modified with BHQ2; the 5' end of E1A-P10-VIC-MGB is modified with VIC and the 3' end is modified with MGB.

[0273] Table 25

[0274] Table 26

[0275] Using a Covaris M220 ultrasonic disruptor, according to the ultrasonic program in Table 2, the same concentration of HEK293 host cell line genomic DNA was ultrasonically fragmented for 25s, 50s, 100s, 200s, and 400s to simulate the residual host nucleic acids with different degrees of digestion in biological products. At the same time, the same concentration of HEK293 host cell line genomic DNA without ultrasonic treatment was used as a control. Using this one-tube 4-plex system, digital PCR detection was performed on the fragmented samples prepared by ultrasonication. Each sample was tested three times, and the mean of the three quantitative values ​​of each pair of primers was calculated. The specific test results are shown in Table 27. For each ultrasonically fragmented nucleic acid sample and full-length sample, a linear relationship between the target sequence length of 83, 132, and 217bp and the corresponding quantitative value mean was established, and the |b / k| value was calculated. The R of all fragmented samples 2 Both values ​​were >0.9. Furthermore, as the degree of nucleic acid fragmentation in the simulated sample increased, the |b / k| value showed a decreasing trend. These results suggest that the |b / k| value can be used to analyze the extent of digestion of residual host nucleic acids during the process; smaller |b / k| values ​​indicate more complete digestion of residual host nucleic acids. Furthermore, the values ​​of the 217 and 571 bp target sequences in the same test result can be used to assess whether the size of residual host nucleic acid fragments meets regulatory requirements.

[0276] Table 27

[0277] Example 10 The significance of the b value in the fitting formula

[0278] When using the |b / k| value to assess the degree of nucleic acid fragmentation, it was found that when X = 0 in the linear relationship, Y = b. That is, when the target sequence length is 0, the theoretical quantitative value of the sample is b. In other words, when the amplified fragment length is minimal, it can essentially reflect the total copy number of the gene. This indicates that the b value obtained by this method can, to a certain extent, reflect the true concentration (copy number) of the tested sample, allowing for sample quantity traceability.

[0279] The human GAPDH gene was selected as the detection target, and 4 pairs of specific primers were designed to amplify target sequences of 50, 88, 111, and 145 bp, respectively. The primer sequences and detection probes are shown in Table 28, and the corresponding target sequences are shown in Table 29.

[0280] Table 28

[0281] The 5' end of P is modified with ROX, and the 3' end is modified with BHQ2.

[0282] Table 29

[0283] The genomic DNA of the 293T cell line was diluted to 4E4 copies / ul, and the genomic DNA of the 293T cell line was ultrasonically fragmented for 50s, 100s, and 200s according to the ultrasonic program in Table 2 to simulate nucleic acid samples with different degrees of fragmentation. At the same time, non-sonicated samples of the same concentration were used as full-length controls. Four pairs of specific primers were used to perform digital PCR detection on full-length samples and fragmented samples. Each pair of primers was tested three times, and the mean of the three quantitative values ​​of each pair of primers was calculated. The specific test results are shown in Table 30. For each ultrasonically fragmented nucleic acid sample and full-length sample, a linear relationship between the target sequence length of 50, 88, 111, and 145 bp and the corresponding quantitative value mean was established, and the |b / k| value was calculated. The R of all fragmented samples 2 All >0.9.

[0284] Comparison of the quantitative values ​​of full-length samples and different sonication samples revealed that the quantitative values ​​of each primer showed a downward trend with increasing sonication duration, with a more pronounced downward trend observed with longer sonication duration. Comparison of the quantitative values ​​of target sequences of different lengths within the same sample revealed that the quantitative values ​​of each sonication sample showed a downward trend with increasing primer amplification length, with a faster decline observed with longer target sequences, consistent with theoretical results. Further comparison of the b values ​​of full-length samples and different sonication samples revealed that the b values ​​of each sonication sample were similar to those of the full-length sample (see Table 31 for specific data). These results suggest that the |b / k| value can be used to assess the degree of sample fragmentation and, based on the b value, to obtain a quantitative value that is as close to the true value as possible for the sample being tested. This method can be applied to trace the quantitative values ​​of nucleic acid standards, quality control products, and other materials.

[0285] Table 30

[0286] Table 31

[0287] Example 11 Establishment of different fitting models and their evaluation results

[0288] A model for assessing nucleic acid fragmentation using Escherichia coli nucleic acid samples was established. The E. coli ydiJ gene was selected as the detection target, and 22 pairs of specific primers were designed from conserved regions. Target sequence lengths ranged from 56 to 176 bp. Primer sequences are shown in Table 32. The detection probe is SEQ ID NO. 31.

[0289] Table 32

[0290] Digital PCR was performed on full-length E. coli genomic DNA samples using 22 specific primer pairs. Each primer pair was tested once, and the quantification coefficients (CVs) across all primer sets were calculated. The CVs for full-length genomic DNA across all 22 specific primer pairs were 2.6%, demonstrating good consistency across all primer pairs.

[0291] E. coli genomic DNA was ultrasonically fragmented for 420 s using a Covaris M220 ultrasonic disruptor according to the ultrasonic program in Table 2. The ultrasonically fragmented samples were purified using Beckman Coulter magnetic beads (Agencourt AMPure XP Nucleic Acid Purification Kit 60 mL Kit). The purified nucleic acid samples were analyzed using an Agilent 4150 Fragment Analyzer. This method yielded six fragmented samples with a normal distribution, with the main peaks at 96, 150, 288, 347, 414, and 538 bp, respectively. Digital PCR was performed on these fragmented samples using 22 pairs of specific primers, with each sample tested once for each primer pair. Specific test data are shown in Table 33.

[0292] Table 33

[0293] The test data in Table 33 provide the quantitative values ​​corresponding to the target sequence length (bp) and the target sequence length, which helps to select the most appropriate model to describe the relationship between them. The relationship between the target sequence length (bp) and the quantitative value corresponding to the target sequence was fitted by a linear model, a quadratic polynomial model, and a cubic polynomial model, and the mean square error (MSE) and the coefficient of determination (R) corresponding to each model were calculated. 2 ), and also evaluated different fitted models based on their ability to explain the biological significance of this relationship. The above analysis was performed on each sample, resulting in the data in Table 34 and Figure 1 below.

[0294] Table 34

[0295] Analyzing the above results, it can be seen that for fragmented nucleic acids, when establishing a relationship model between target sequence length and corresponding quantitative value, the cubic polynomial model provides the lowest MSE and the highest R 2The fitting effect is the best. Compared with other models, it can follow the trend of data points more closely. It should be the best model to describe the relationship trend between target sequence length and corresponding quantitative value. However, the parameters of the cubic polynomial are relatively large. When the amount of data is limited, it is easy to overfit. That is, the model fits the training data very well, but has poor generalization ability for new data. In addition, as the degree of the polynomial increases, the complexity of the model increases, and it is difficult to understand the correspondence between each coefficient and biological significance. The mean square error (MSE) and the coefficient of determination (R 2 ) is slightly worse than a cubic polynomial, and the model's ability to explain the biological significance of this relationship is average.

[0296] Compared with the cubic polynomial and quadratic polynomial, the mean square error (MSE) of the linear model is slightly higher, R 2 Although relatively low, its biological significance in explaining the relationship between target sequence length and corresponding quantitative values ​​is clear. The general form of a linear function is y = kx + b (k and b are constants, k ≠ 0). The function graph is a straight line, with k being the slope of the line, which determines its inclination, and b being the intercept of the line on the y-axis. The coordinates of the intersection of the line and the y-axis are (0, b). Based on the target sequence length and the number of detected copies of the corresponding length, the fitting formula is Y = kX + b, where Y is the number of detected copies of the target sequence and X is the length of the target sequence. When Y = 0, |x| = |b / k|. This means that when the target sequence length is |b / k|, the theoretical quantitative value of the sample is 0. When the target sequence length is greater than |b / k|, the theoretical quantitative value of the sample is still 0, and when the target sequence length is greater than the template length, the theoretical quantitative value of the sample is also 0. This indicates that the |b / k| value obtained by this method can, to a certain extent, reflect the fragment length of the sample being tested and is of practical significance. Based on the modeling data, the |b / k| value is indeed close to the main peak of the ultrasonically purified simulated sample. Specific data are shown in Table 35. This indicates that this method can, to a certain extent, reflect the main fragment length of the fragmented sample, which indirectly reflects the degree of fragmentation. Furthermore, when X = 0, Y = b. That is, when the target sequence length is 0, the theoretical quantitative value of the sample is b. In other words, when the amplified fragment length is minimal, it essentially reflects the entire length of the gene. This indicates that the b value obtained by this method can, to a certain extent, reflect the true concentration (copy number) of the tested sample, which has another practical significance.

[0297] Table 35

Claims

1. A method for specifically evaluating the degree of target nucleic acid fragmentation, characterized in that: The method comprises: Designing at least two pairs of primer sets based on the target nucleic acid sequence; the target sequence lengths of the at least two pairs of primer sets are different from each other; After mixing the biological sample with the at least two pairs of primer sets, the detection marker, and the nucleic acid amplification reaction solution, the mixture is randomly distributed into a plurality of reaction units to perform a nucleic acid amplification reaction to obtain detection copy numbers of different target sequence lengths; Fitting a calculation formula based on the length of the target sequence and the detected copy number of the corresponding length, in which the variable Y is the detected copy number of the target sequence and the variable X is the length of the target sequence; and The degree of fragmentation of the target nucleic acid is evaluated based on the intersection of the fitting curve of the calculation formula and the X-axis.

2. The method according to claim 1, wherein In the fitting step, a linear model, a quadratic polynomial model, or a cubic polynomial model is fitted according to the length of the target sequence and the number of detected copies of the corresponding length; preferably, the calculation formula is Y=kX+b, and the intersection of the fitting curve and the X-axis is |b / k|.

3. The method according to claim 1, wherein The X-axis and the Y-axis are interchanged, and the degree of fragmentation of the target nucleic acid is evaluated based on the intersection of the fitting curve of the calculation formula and the Y-axis.

4. The method according to claim 1, wherein The target sequence lengths of the at least two pairs of primers range from 30 to 540 bp; preferably, the at least two pairs of primers share one forward primer or one reverse primer.

5. The method according to claim 1, wherein The primer set is a set of 3 to 7 primer pairs, preferably a set of 4 primer pairs; more preferably, the target sequence length of the 4 primer pairs is 30 to 50 bp, 80 to 100 bp, 130 to 150 bp, or 160 to 180 bp.

6. A method for detecting a target nucleic acid, characterized in that: The method comprises: Mixing a biological sample with at least two pairs of primer sets, a detection marker, and a nucleic acid amplification reaction solution to obtain a first mixed solution; wherein the at least two pairs of primer sets have target sequences of different lengths and amplify the same target nucleic acid; Randomly distributing the first mixed solution into a plurality of reaction units to perform a nucleic acid amplification reaction to obtain detection copy numbers of different target sequence lengths; Fitting a first calculation formula according to the length of the target sequence and the detected copy number of the corresponding length, in which the variable Y is the detected copy number of the target sequence and the variable X is the length of the target sequence; and The fragmentation degree of the target nucleic acid is evaluated based on the intersection of the fitting curve of the first calculation formula and the X-axis, and / or the copy number of the target nucleic acid is evaluated based on the coefficient of the first calculation formula.

7. The method according to claim 6, wherein The method further comprises: mixing the control biological sample with the at least two pairs of primer sets, the detection marker, and the nucleic acid amplification reaction solution of claim 5 to obtain a second mixed solution; Randomly distributing the second mixed solution into a plurality of reaction units to perform a nucleic acid amplification reaction to obtain detection copy numbers of different target sequence lengths; Fitting a second calculation formula based on the length of the target sequence and the detected copy number of the corresponding length, in which the variable Y is the detected copy number of the target sequence and the variable X is the length of the target sequence; and The degree of fragmentation of the target nucleic acid is evaluated by comparing the intersection of the fitting curves of the first calculation formula and the second calculation formula with the X-axis, and / or the coefficients of the first calculation formula and the second calculation formula are compared to evaluate the copy number of the target nucleic acid.

8. The method according to claim 7, wherein In the step of fitting the first calculation formula and the second calculation formula, a linear model, a quadratic polynomial model, or a cubic polynomial model is fitted according to the length of the target sequence and the detected copy number of the corresponding length; preferably, the first calculation formula is Y=kX+b and the second calculation formula is Y=k1X+b1, and in the evaluation step, the b / k| value and the |b1 / k1| value are compared to evaluate the fragmentation degree of the target nucleic acid, and / or the b value and the b1 value are compared to evaluate the copy number of the target nucleic acid.

9. A method for accurately detecting cfDNA, characterized in that: The method comprises: Isolation of cfDNA from biological samples; Mixing the cfDNA with at least two pairs of target nucleic acid primers, a pair of gDNA primers, a detection marker, and a nucleic acid amplification reaction solution; wherein the at least two pairs of primers have target sequences of different lengths and amplify the same target nucleic acid; After mixing, the mixed solution is randomly distributed into multiple reaction units to perform nucleic acid amplification reaction; Obtaining the detected copy numbers of the target sequences corresponding to at least two pairs of target nucleic acid primer sets and the detected copy number of a pair of gDNA primer sets; fitting a calculation formula based on the copy number obtained by deducting the detected copy number of the gDNA primer set from the detected copy number of each target sequence in the at least two pairs of target nucleic acid primer sets and the target sequence length of the corresponding target nucleic acid primer set, wherein the variable Y is the copy number of the target sequence after deduction, and the variable X is the length of the target sequence; and The degree of fragmentation of cfDNA is evaluated based on the intersection of the fitting curve of the calculation formula and the X-axis, and / or the copy number of cfDNA is evaluated based on the coefficient of the calculation formula.

10. The method according to claim 9, wherein: The target sequence lengths of the at least two pairs of primer sets range from 30 to 420 bp, and the target sequence length of the pair of gDNA primer sets is ≥466 bp; preferably, the at least two pairs of primer sets share one forward primer or one reverse primer.

Citation Information

Patent Citations

  • Information statistical method for tumor ctDNA

    CN109112191A

  • Method of correcting amplification deviation in amplicon sequencing

    CN110741094A

  • Method for judging blood stain formation time in forensic medicine by detecting RNA degradation degree

    CN111850106A

  • Standard substance and kit for calibrating copy number concentration of digital PCR (Polymerase Chain Reaction) instrument

    CN115786477A

  • Methods of quantifying cell-free DNA

    WO2016054255A1