Determination of nucleic acid sequence concentration

The method corrects the underestimation of nucleic acid concentrations in fragmented samples by using a correction factor based on length distribution and measurement parameters, enhancing the accuracy of nucleic acid quantification.

JP2026016403APending Publication Date: 2026-02-03STILLA TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025162355
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-16
Filing Date
2025-09-29
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Current methods for quantifying nucleic acid sequences in biological samples underestimate the concentration due to fragmentation, leading to inaccurate measurements when amplification fails in PCR assays.

Method used

A method to correct the measured concentration of nucleic acid sequences in fragmented samples by using a correction factor based on the length distribution (LD) and parameters of the measurement method, allowing for the determination of the concentration in unfragmented nucleic acids.

Benefits of technology

The method provides a more accurate estimation of nucleic acid concentrations by accounting for fragmented nucleic acids, improving the reliability of quantitative analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016403000001_ABST
    Figure 2026016403000001_ABST
Patent Text Reader

Abstract

To provide a method for determining the concentration of a nucleic acid sequence without underestimating the presence of the target nucleic acid sequence by fragmentation.SOLUTION: The present invention is a method of determining the concentration of a sequence to be detected in non-fragmented nucleic acid by applying a correction factor to the concentration of said sequence to be detected measured in fragmented nucleic acid. Determining the length distribution (LD) of nucleic acids in a sample comprising fragmented nucleic acids, wherein fragmented nucleic acids are derived from said non-fragmented nucleic acids; ii. measuring the concentration of said sequence to be detected in said sample comprising fragmented nucleic acids by a measurement method; and iii. correcting the measured concentration of said sequence to be detected in the sample comprising fragmented nucleic acids by a correction factor to obtain: Obtaining the concentration of the sequence to be detected in the unfragmented nucleic acid, wherein the correction factor is based on the length distribution (LD) and at least one parameter of the measurement method.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to European Patent Application No. EP19306346, filed October 16, 2019, the contents of which are incorporated herein by reference in their entirety.

[0002] FIELD OF THE INVENTION The present invention relates to a method for determining the concentration of a nucleic acid sequence in a biological sample. [Background technology]

[0003] A normal human has two sets of 23 chromosomes in every healthy diploid cell. Under some conditions, mutations can occur in one or more of these chromosomes, leading to chromosomal abnormalities. These abnormalities may be associated with genetic diseases, cancer, and other diseases. Detecting chromosomal abnormalities can identify individuals prone to developing certain diseases or define the most recommended treatment for a given individual. In this regard, testing for chromosomal abnormalities is extremely valuable.

[0004] In addition to human health, the detection of chromosomal abnormalities is relevant to additional species, including but not limited to insects, bacteria, plants, and mixed samples containing organic matter such as soil and food, and genetically engineered genetically modified organism (GMO) detection can be performed, for example, for quality control. The detection of genetic mutations occurring in viruses, viroids, and other non-chromosomal genomes is also highly relevant.

[0005] In this context, it relates to detecting, both qualitatively and quantitatively, indications of genomic abnormalities, ie, alterations to specific nucleic acid sequences in a biological sample.

[0006] Such detection is currently efficiently achieved by amplifying the target nucleic acid in a sample of interest. Amplification can be achieved by combining oligonucleotide primers with the sample and then subjecting the sample to amplification conditions compatible with nucleic acid quantification, such as polymerase chain reaction (PCR) conditions. These amplification methods allow for the generation of multiple copies of a single target nucleic acid sequence, thus reaching the detection threshold.

[0007] However, quantitatively measuring the concentration of such specific nucleic acid sequences in a biological sample is more relevant, for example, when the aim is not only to detect but also to quantify genetic diseases, e.g., to monitor the evolution of rare mutations.

[0008] In some cases, the nucleic acid present in biological samples is damaged, and typically fragmented into short sequences.In this fragmented nucleic acid population, the sequence that should be amplified in a given PCR assay may be randomly cut.In this scenario, the amplification of this nucleic acid sequence cannot occur in PCR, so the presence of the nucleic acid sequence of interest will be underestimated.

[0009] The present invention proposes a method to correct such underestimation problems. Summary of the Invention

[0010] The present invention relates to a method for determining the concentration of a detected sequence (as further defined below) in unfragmented nucleic acid, comprising correcting the measured concentration of said detected sequence in a sample comprising fragmented nucleic acid by a correction factor, thereby obtaining the concentration of said detected sequence in unfragmented nucleic acid, wherein the correction factor is based on the length distribution (LD) and at least one parameter of the measurement method.

[0011] In some embodiments, a method for determining the concentration of a detected sequence in unfragmented nucleic acid comprises the steps of: i. measuring the concentration of the sequence to be detected in the sample containing fragmented nucleic acids by a measurement method; and ii. correcting the measured concentration of the detected sequence in the sample containing fragmented nucleic acids by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acids, wherein the correction factor is based on the length distribution (LD) and at least one parameter of the measurement method. A method is provided that includes:

[0012] In some embodiments, a method for determining the concentration of a detected sequence in unfragmented nucleic acid comprises the steps of: i. determining the length distribution (LD) of nucleic acids in a sample comprising fragmented nucleic acids, wherein the fragmented nucleic acids are derived from said non-fragmented nucleic acids; ii. measuring the concentration of the sequence to be detected in the sample containing fragmented nucleic acids by a measurement method; and iii. correcting the measured concentration of the detected sequence in the sample containing fragmented nucleic acids by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acids, wherein the correction factor is based on the length distribution (LD) and at least one parameter of the measurement method. A method is provided that includes:

[0013] In some embodiments, the method comprises: i. providing a sample of fragmented nucleic acids, wherein the fragmented nucleic acids are derived from said unfragmented nucleic acids; ii. providing a length distribution (LD) of nucleic acid fragments of the fragmented nucleic acids; iii. measuring the concentration of said sequence to be detected in said sample of fragmented nucleic acids by a measurement method; iv. determining a correction factor according to the nucleic acid fragment length distribution (LD) of the fragmented nucleic acid and parameters of the measurement method; and v. correcting the concentration of said detected sequence in the fragmented nucleic acid by said correction factor to obtain the concentration of said detected sequence in unfragmented nucleic acid. Includes:

[0014] In one embodiment, the sample of fragmented nucleic acids is any combination of the following three categories: i. Either a cell-free sample or a cell-containing sample; ii. Either naturally fragmented or artificially fragmented samples; and iii. Any type of deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0015] In one embodiment, the measurement method is an isothermal quantitative nucleic acid amplification method, preferably selected from loop-mediated isothermal amplification and nucleic acid sequence-based quantitative amplification.

[0016] In another embodiment, the measurement method is a non-isothermal quantitative nucleic acid amplification method, preferably selected from quantitative polymerase chain reaction, real-time polymerase chain reaction, digital polymerase chain reaction, multiplex polymerase chain reaction and multiplex digital polymerase chain reaction.

[0017] In one embodiment, the parameters of the measurement method include the length of the sequence to be amplified, as further defined below. In a particular configuration of this embodiment, at least 5% of the nucleic acid fragments have a length shorter than the length of the sequence to be amplified. In another particular configuration of this embodiment, at most 95% of the nucleic acid fragments have a length shorter than the length of the sequence to be amplified.

[0018] In one embodiment, the length of the sequence to be amplified is longer than 40 bp and shorter than 200 bp, preferably longer than 50 bp and shorter than 170 bp, more preferably longer than 65 bp and shorter than 150 bp, and even more preferably longer than 70 bp and shorter than 130 bp.

[0019] In one embodiment, the length distribution (LD) of the nucleic acid fragments is in the range of 25 bp to 350 bp, preferably 30 bp to 320 bp, more preferably 35 bp to 290 bp, and even more preferably 40 bp to 270 bp.

[0020] The present invention further provides a method for determining a function of a first concentration of a first detected sequence S1 and a second concentration of a second detected sequence S2 in unfragmented nucleic acid, comprising the steps of: i. determining the concentration of S1 in the unfragmented nucleic acid according to any of the determination methods described above; ii. determining the concentration of S2 in the unfragmented nucleic acid according to any of the determination methods described above; and iii. calculating the function of the S1 concentration and the S2 concentration Includes; wherein the S1 concentration and the S2 concentration are determined in the same sample; and the length of the sequence to be amplified associated with S1 is different from the length of the sequence to be amplified associated with S2.

[0021] The present invention also provides a system configured to determine the concentration of a detected sequence in unfragmented nucleic acid, comprising: i. a module configured to measure the concentration of the sequence to be detected in a sample of fragmented nucleic acids, the fragmented nucleic acids being derived from the non-fragmented nucleic acids; ii. a module configured to calculate a correction factor according to the nucleic acid fragment length distribution (LD) of the fragmented nucleic acid and parameters of the measurement; iii. a module configured to calculate the concentration of the detected sequence in unfragmented nucleic acid using the correction factor; The present invention also relates to a system comprising:

[0022] In a particular configuration, the module configured to measure the concentration of the detected sequence in the fragmented nucleic acid is an isothermal quantitative nucleic acid amplification module, preferably selected from a loop-mediated isothermal amplification module and a nucleic acid sequence-based quantitative amplification module.

[0023] In another configuration, the module configured to measure the concentration of the detected sequence in the fragmented nucleic acids is a non-isothermal quantitative nucleic acid amplification module, preferably selected from a quantitative polymerase chain reaction module, a real-time polymerase chain reaction module, a digital polymerase chain reaction module, a multiplex polymerase chain reaction module, and a multiplex digital polymerase chain reaction module.

[0024] In one embodiment, the parameters of said measurement include the length of the sequence to be amplified.

[0025] definition In the present invention, the following terms have the following meanings:

[0026] The term "amplicon" refers to the nucleic acid product of an amplification reaction. An amplicon can be single-stranded or double-stranded, or a combination thereof.

[0027] The term "amplification" refers to a reaction in which replication occurs repeatedly over time to form multiple copies of at least one segment of a template molecule. Amplification can result in an exponential or linear increase in copy number as amplification progresses. Typical amplifications result in a greater than 1,000-fold increase in copy number and / or signal. Exemplary amplification reactions for droplet-based assays disclosed herein can include polymerase chain reaction (PCR) or ligase chain reaction, each of which is driven by thermal cycling. Droplet-based assays may additionally or alternatively use other amplification reactions that can be performed isothermally. Amplification can be performed or assayed for in an amplification mixture, which is any composition that, when present in the composition, can generate multiple copies of a nucleic acid target molecule. An "amplification mixture" can include any combination of at least one primer or primer pair, at least one probe, at least one replicative enzyme (e.g., at least one polymerase, e.g., at least one DNA and / or RNA polymerase, e.g., reverse transcriptase), and deoxynucleotide (and / or nucleotide) triphosphates (dNTPs and / or NTPs), as well as a buffer containing, among other components, any components essential for replicative enzyme activity.

[0028] The term "assay" refers to a procedure and / or reaction used to characterize a sample, and any signals, values, data, and / or results obtained from that procedure and / or reaction.

[0029] The term "cleaved" is equivalent to the term "fragmented" in relation to nucleic acid sequences.

[0030] The term "detected sequence" refers to the exact sequence quantified from a sample in the method of the present invention. In the case of a point mutation, the detected sequence includes the detected mutation. In some embodiments, a fluorescent reporter is used. When a fluorescent reporter linked to a specific sequence is used (for example, as shown in Figure 1A), the "detected sequence" can be shorter than the sequence to be amplified and can be contained within the sequence (as further defined herein). When a free fluorescent reporter is used (for example, as shown in Figure 1B), the detected sequence is usually an amplicon.

[0031] The term "digital PCR" or "dPCR" refers to a PCR assay performed to determine the presence / absence, concentration, and / or copy number of a nucleic acid target in a sample based on whether aliquots of the sample support amplification of the target. The concept of digital PCR can be extended to other types of analytes in addition to nucleic acids.

[0032] The term "label" refers to an identification and / or differentiation marker or identifier attached to or incorporated into any entity, such as a compound, a biological particle (e.g., a cell, a bacterium, a spore, a virus, or an organelle), or a droplet. A label can be, for example, a dye that renders the entity optically detectable and / or optically distinguishable. Exemplary dyes used for labeling are fluorescent dyes (fluorophores) and fluorescence quenchers.

[0033] The term "length" in reference to nucleic acids refers to the number of consecutive nucleotides or bases that form a single-stranded molecule or the number of consecutive base pairs that form a double-stranded molecule. Length is measured in nucleotides and base pairs.

[0034] The term "multiplex digital PCR" refers to a digital PCR assay performed to simultaneously amplify at least two different nucleic acid sequences, particularly 2, 3, 4, 5, 6, 7, 8, or more different nucleic acid sequences (similar to performing many separate PCR reactions all together in a single pot). This method uses multiple primers to amplify nucleic acids in a sample. Specifically, "multiplex digital PCR" includes "duplex digital PCR" and "triple digital PCR." Conversely, a digital PCR assay performed to amplify one nucleic acid sequence is "single digital PCR," often abbreviated to "digital PCR."

[0035] The term "multiplex PCR" refers to a PCR assay performed to simultaneously amplify at least two different nucleic acid sequences, particularly 2, 3, 4, 5, 6, 7, 8, or more different nucleic acid sequences (similar to performing many separate PCR reactions all together in a single pot). This method uses multiple primers to amplify nucleic acids in a sample. Specifically, "multiplex PCR" includes "duplex PCR" and "triple PCR." Conversely, a PCR assay performed to amplify one nucleic acid sequence is a "single PCR," often abbreviated to "PCR."

[0036] The term "nucleic acid" refers to both deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), whether the product of amplification, produced synthetically, the product of reverse transcription of RNA, or naturally occurring. Generally, nucleic acids are single- or double-stranded molecules and are composed of naturally occurring nucleotides.

[0037] The term "nucleotide," in addition to referring to naturally occurring ribonucleotide or deoxyribonucleotide monomers, is understood herein to refer to its related structural variants, including derivatives and analogs and chemical modifications, that are functionally equivalent for the particular context in which the nucleotide is used (e.g., hybridization with complementary bases: adenine (A) pairs with thymine (T), guanine (G) pairs with cytosine (C)), unless the context clearly indicates otherwise.

[0038] The term "aliquot" refers to a separated portion of a bulk volume. The aliquots may be sample aliquots generated from a sample (such as a prepared sample) that forms the bulk volume. The aliquots generated from the bulk volume may be substantially uniform in size or may have different sizes (e.g., a set of two or more individual uniformly sized aliquots). An exemplary aliquot is a "droplet." The aliquots may also vary in size, either in a predetermined size distribution or in a random size distribution.

[0039] The term "PCR" or "polymerase chain reaction" refers to a nucleic acid amplification assay that relies on alternating cycles of heating and cooling (i.e., thermal cycling) to achieve successive rounds of replication. PCR can be performed by thermal cycling between two or more temperature set points (such as a higher melting (denaturation) temperature and a lower annealing / extension temperature), or, inter alia, between three or more temperature set points (such as a higher melting temperature, a lower annealing temperature, and an intermediate extension temperature). Other types of PCR, such as touchdown PCR, can be included in this definition, in which the annealing and / or extension temperatures can be changed during the cycling reaction.

[0040] The term "primer" refers to an oligonucleotide that can act as an initiation point for template-directed nucleic acid synthesis when placed under conditions in which polynucleotide elongation is initiated; e.g., conditions including the presence of the necessary nucleoside triphosphates (determined by the template to be copied) and polymerase in an appropriate buffer, at a suitable temperature or temperature cycle (e.g., as in the polymerase chain reaction).

[0041] The term "probe" refers to a nucleic acid that is attached to at least one label, such as at least one dye.

[0042] The term "qualitative PCR" refers to a PCR-based analysis that determines whether a target is present in a sample, generally without substantially quantifying the presence of the target. In an exemplary embodiment, qualitative digital PCR can be performed by determining whether an aliquot packet contains (positive sample) or does not contain (negative sample) at least a predefined percentage of positive droplets.

[0043] The terms "quantitative PCR," "qPCR," "real-time quantitative polymerase chain reaction," or "kinetic polymerase chain reaction" refer to a PCR-based assay that determines the concentration and / or copy number of a target in a sample. This technique uses PCR to simultaneously amplify and quantify target nucleic acids, with quantification achieved by sequence-specific probes containing intercalating fluorescent dyes that are detectable only when hybridized to the target nucleic acid or fluorescent reporter molecules that are detectable only upon sequence amplification.

[0044] The term "real-time PCR" refers to a PCR-based assay in which amplicon formation is measured during the reaction, such as after one or more thermal cycles are completed prior to the final thermal cycle of the reaction. Real-time PCR generally provides target quantitation based on the kinetics of target amplification.

[0045] The term "replication" refers to the process of forming copies (i.e., direct copies and / or complementary copies) of a nucleic acid or segment thereof. Replication generally involves enzymes such as polymerases and / or ligases, among others. The nucleic acid and / or segment being replicated is the template (and / or target) for replication.

[0046] The term "sample" refers to a compound, composition, and / or mixture of interest from any suitable source. A sample is a general object of interest for an assay that analyzes aspects of the sample (such as aspects related to at least one analyte that may be present in the sample). Samples can be analyzed in their natural state as collected and / or in an altered state after, for example, inter alia, storage, preservation, extraction, lysis, dilution, concentration, purification, filtration, mixing with one or more reagents, pre-amplification (e.g., to achieve target enrichment by performing limited cycles (e.g., <15) of PCR on the sample prior to PCR), amplicon removal (e.g., treatment with uracil-d-glycosylase (UDG (UNG, also known as uracil-N-glycosylase gene)) prior to PCR to eliminate carryover contamination from previously generated amplicons (i.e., amplicons are generated with dUTP instead of dTTP and therefore digestible with UDG)), fractionation, or any combination thereof. Clinical samples may include, among others, nasopharyngeal washings, blood, plasma, cell-free plasma, buffy coat, saliva, urine, stool, sputum, mucus, wound swabs, biopsy tissue, milk, bodily fluid aspirates, swabs (e.g., nasopharyngeal swabs), and / or tissues. Samples may be collected for diagnostic purposes (e.g., quantitative measurement of a clinical analyte such as an infectious pathogen) or for monitoring purposes (e.g., to determine whether an environmental analyte of interest, such as a biothreat agent, has exceeded a predetermined threshold).

[0047] The term "sequence to be amplified" refers to a nucleic acid containing a "sequence to be detected" (as defined above) that begins with the sequence of the forward primer and ends with the sequence complementary to the reverse primer, including any additional base pairs located between the primer sequences. The amplification product of the sequence to be amplified is an amplicon. Ultimately, amplification allows for quantification of the "sequence to be amplified" by determining its concentration. The length of the sequence to be amplified is referred to as "L a " [Brief explanation of the drawings]

[0048] [Figure 1A] FIG. 1 is a schematic diagram of a quantitative nucleic acid amplification method, by way of example and not limitation. Here, double-stranded fragmented nucleic acid is indicated by a horizontal solid black line. Two primers (forward primer FP and reverse primer RP) define the left and right boundaries of the sequence to be amplified (SA), respectively. Both strands of the sequence to be amplified (SA) are copied during amplification. In FIG. 1A, the sequence to be detected (DS) is part of the sequence to be amplified (SA). A fluorescently labeled probe (probe) specifically binds to the sequence to be detected. In FIG. 1B, the sequence to be detected (DS) and the sequence to be amplified (SA) are the same. A fluorescent dye (FD) is intercalated into the double-stranded nucleic acid produced during amplification. [Figure 1B] FIG. 1 is a schematic diagram of a quantitative nucleic acid amplification method, by way of example and not limitation. Here, double-stranded fragmented nucleic acid is indicated by a horizontal solid black line. Two primers (forward primer FP and reverse primer RP) define the left and right boundaries of the sequence to be amplified (SA), respectively. Both strands of the sequence to be amplified (SA) are copied during amplification. In FIG. 1A, the sequence to be detected (DS) is part of the sequence to be amplified (SA). A fluorescently labeled probe (probe) specifically binds to the sequence to be detected. In FIG. 1B, the sequence to be detected (DS) and the sequence to be amplified (SA) are the same. A fluorescent dye (FD) is intercalated into the double-stranded nucleic acid produced during amplification. [Figure 2]Figure 2 is a graph showing the length distribution (LD) of DNA fragments in the sample used in Example E1, centered around 150 bp, where f(i) is the probability (Y-axis—arbitrary units) that a fragment in the sample has a length of i base pairs (X-axis). [Figure 3] FIG. 3 shows the probability P of non-cleavage of the sequence to be amplified, estimated for the samples shown in FIG. 2, as a function of the length La of the sequence to be amplified (unit: base pairs bp). [Figure 4] FIG. 4 shows the predicted correction factor (PCF) estimated for the samples shown in FIG. 2 as a function of the length La (in base pairs bp) of the sequence to be amplified. DETAILED DESCRIPTION OF THE INVENTION

[0049] The present application provides methods and systems for correcting the measured concentration of a detected sequence in a nucleic acid sample containing fragmented nucleic acids. The methods and systems provided herein enable the determination of a corrected concentration of the detected sequence that more closely approximates the true concentration of the detected sequence in unfragmented nucleic acids. When the concentration of a detected sequence is measured by amplifying a sequence to be amplified that includes the detected sequence, fragments that contain a portion of the sequence to be amplified but have a length shorter than the length of the sequence to be amplified are not replicated. This failure to replicate truncated target regions (i.e., the sequence to be amplified) leads to an underestimation of the concentration. The methods provided herein can be applied to correct the resulting underestimation of the concentration of nucleic acids in unfragmented nucleic acid samples.

[0050] Thus, in some embodiments, the present invention provides methods for determining the concentration of a detected sequence in unfragmented nucleic acid (including determining the copy number of said detected sequence) (hereinafter referred to as "Basic Methods"), comprising the steps of: (a) determining the copy number of said detected sequence; (b) determining the copy number of said detected sequence; and (c) determining the copy number of said detected sequence. In some embodiments, the present invention provides methods for calibrating the concentration of a detected sequence (including the copy number of said detected sequence) in a sample containing fragmented nucleic acid molecules. Also provided are kits, software, devices, and other articles of manufacture useful for the methods described herein.

[0051] I. The present method The present invention relates to a method for determining the concentration of a detected sequence in unfragmented nucleic acid, comprising correcting the measured concentration of said detected sequence in a sample comprising fragmented nucleic acid by a correction factor, thereby obtaining the concentration of said detected sequence in unfragmented nucleic acid, wherein the correction factor is based on the length distribution and at least one parameter of the measurement method.

[0052] In some embodiments, a method for determining the concentration of a detected sequence in unfragmented nucleic acid comprises the steps of: i. measuring the concentration of the sequence to be detected in the sample containing fragmented nucleic acids by a measurement method; and ii. correcting the measured concentration of the detected sequence in the sample containing fragmented nucleic acids by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acids, wherein the correction factor is based on the length distribution (LD) and at least one parameter of the measurement method. A method is provided which includes:

[0053] In some embodiments, a method for determining the concentration of a detected sequence (as further defined below) in unfragmented nucleic acid comprises the steps of: i. amplifying a target region containing the sequence to be detected (hereinafter referred to as the "sequence to be amplified") in the sample containing fragmented nucleic acids; ii. measuring the concentration of the detected sequence; and iii. correcting the measured concentration of the detected sequence in the sample containing fragmented nucleic acids by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acids, wherein the correction factor is based on the length distribution (LD) and at least one parameter of the measurement method. A method is provided that includes:

[0054] In some embodiments, provided herein are methods for determining the concentration of a detected sequence in unfragmented nucleic acid, comprising the steps of: i. determining the length distribution (LD) of nucleic acids in a sample comprising fragmented nucleic acids, wherein the fragmented nucleic acids are derived from said non-fragmented nucleic acids; ii. measuring the concentration of the detected sequence in the sample containing fragmented nucleic acids using a measurement method; and correcting the measured concentration of the detected sequence in the sample containing fragmented nucleic acids by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acids, wherein the correction factor is based on the length distribution (LD) and at least one parameter of the measurement method. The method includes:

[0055] In some embodiments, provided herein is a method for calibrating the measured concentration of a detected sequence in a sample containing fragmented nucleic acids so that it more closely reflects the concentration of that detected sequence in unfragmented nucleic acids, the method comprising the steps of: i. determining the length distribution (LD) of nucleic acids in a sample containing fragmented nucleic acids, wherein the fragmented nucleic acids are derived from the non-fragmented nucleic acids; and ii. correcting the measured concentration by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acid, the correction factor being based on at least one parameter of the measurement method used to determine the length distribution (LD) and the concentration of the detected sequence. The method includes:

[0056] In some embodiments, provided herein are methods for calibrating a measured concentration of a detected sequence in a sample comprising fragmented nucleic acids, wherein the measured concentration of the detected sequence is an underestimate of the actual concentration of the detected sequence in the sample, comprising the steps of: i. determining the length distribution (LD) of nucleic acids in a sample comprising fragmented nucleic acids, wherein the fragmented nucleic acids are derived from said non-fragmented nucleic acids; and ii. correcting the measured concentration by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acid, wherein the correction factor is based on at least one parameter of the measurement method used to determine the length distribution (LD) and the concentration of the detected sequence. The method includes:

[0057] In some embodiments, provided herein is a method for correcting the measured concentration of a detected sequence in a sample containing fragmented nucleic acids so that it more closely reflects the concentration of the detected sequence in unfragmented nucleic acids, the method comprising the steps of: i. determining the length distribution (LD) of nucleic acids in a sample comprising fragmented nucleic acids, wherein the fragmented nucleic acids are derived from said non-fragmented nucleic acids; ii. measuring the concentration of the sequence to be detected in the sample containing fragmented nucleic acids by a measurement method; and iii. correcting the measured concentration of the detected sequence in the sample containing fragmented nucleic acids by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acids, wherein the correction factor is based on the length distribution (LD) and at least one parameter of the measurement method. The method includes:

[0058] In some embodiments, provided herein are methods for correcting a measured concentration of a detected sequence in a sample comprising fragmented nucleic acids, wherein the measured concentration of the detected sequence is an underestimate of the actual concentration of the detected sequence in the sample, the method comprising the steps of: i. determining the length distribution (LD) of nucleic acids in a sample comprising fragmented nucleic acids, wherein the fragmented nucleic acids are derived from said non-fragmented nucleic acids; ii. measuring the concentration of the sequence to be detected in the sample containing fragmented nucleic acids by a measurement method; and iii. correcting the measured concentration of the detected sequence in the sample containing fragmented nucleic acids by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acids, wherein the correction factor is based on the length distribution (LD) and at least one parameter of the measurement method. The method includes:

[0059] In some embodiments, the methods do not involve obtaining a sequence of the fragmented nucleic acid. In some embodiments, the methods do not involve obtaining or predicting genetic coordinates of the fragmented nucleic acid. In some embodiments, the methods do not involve reordering the fragmented nucleic acid into a contiguous sequence.

[0060] The fragmented nucleic acid samples in some embodiments are any combination of the following three categories, as described in the "Nucleic Acid Samples and Length Distribution" subsection below: i. Either a cell-free sample or a cell-containing sample; ii. Either naturally fragmented or artificially fragmented samples; and iii. Any kind of deoxyribonucleic acid or ribonucleic acid.

[0061] In some embodiments, the measurement method does not include sequencing the nucleic acid in the sample. In some embodiments, measuring the concentration of the sequence to be detected can include amplifying a sequence to be amplified that includes the sequence to be detected. In some embodiments, the measurement method is an isothermal quantitative nucleic acid amplification method (e.g., loop-mediated isothermal amplification or nucleic acid sequence-based quantitative amplification). In some embodiments, the measurement method is a non-isothermal quantitative nucleic acid amplification method (e.g., quantitative polymerase chain reaction, real-time polymerase chain reaction, digital polymerase chain reaction, multiplex polymerase chain reaction, and multiplex digital polymerase chain reaction).

[0062] In some embodiments of any of the foregoing methods, the sequence to be detected and the sequence to be amplified are the same, and the measuring step comprises detecting incorporation of a label (e.g., an intercalating dye, such as a fluorescent dye, including a fluorophore) into nucleic acids produced during amplification, as shown in FIG. 1B.

[0063] In some embodiments, the detected sequences are a subset of the sequences to be amplified, and / or the measuring step comprises detecting binding of a labeled probe (e.g., a fluorescently labeled probe comprising a fluorophore) to the detected sequences, as shown in Figure 1A.

[0064] The measurement method carried out in the measurement step of the basic method of the present invention essentially uses several relevant parameters to determine the correction factor. In a particular embodiment, the measurement method is an amplification method using a replication and amplification mixture.

[0065] In some embodiments, the method can further comprise determining a correction factor based on the length distribution of nucleic acids in the sample and at least one parameter of the measurement method.

[0066] In some embodiments, at least one parameter of the measurement method can include the length of the sequence to be amplified. In some embodiments, at least one parameter of the measurement method can include only the length of the sequence to be amplified.

[0067] Thus, for example, in some embodiments, provided herein are methods for determining the concentration of a detected sequence in unfragmented nucleic acid, comprising the steps of: i. determining the length distribution (LD) of nucleic acids in a sample comprising fragmented nucleic acids, wherein the fragmented nucleic acids are derived from said non-fragmented nucleic acids; ii. measuring the concentration of the detected sequence in the sample containing fragmented nucleic acids by a measurement method; and correcting the measured concentration of the detected sequence in the sample containing fragmented nucleic acids by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acids, wherein the correction factor is based on the length distribution (LD) and the length of the sequence to be amplified, including the detected sequence. The method includes:

[0068] In some embodiments, provided herein is a method for calibrating the measured concentration of a detected sequence in a sample containing fragmented nucleic acids so that it more closely reflects the concentration of the detected sequence in unfragmented nucleic acids, comprising the steps of: i. determining the length distribution (LD) of nucleic acids in a sample comprising fragmented nucleic acids, wherein the fragmented nucleic acids are derived from said non-fragmented nucleic acids; and ii. correcting the measured concentration by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acid, wherein the correction factor is based on the length distribution (LD) and length of the sequence to be amplified, including the detected sequence. The method includes:

[0069] In some embodiments, provided herein is a method for correcting the measured concentration of a detected sequence in a sample containing fragmented nucleic acids so that it more closely reflects the concentration of the detected sequence in unfragmented nucleic acids, the method comprising the steps of: i. determining the length distribution (LD) of nucleic acids in a sample comprising fragmented nucleic acids, wherein the fragmented nucleic acids are derived from said non-fragmented nucleic acids; ii. measuring the concentration of the sequence to be detected in the sample containing fragmented nucleic acids by a measurement method; and iii. correcting the measured concentration of the detected sequence in the sample containing fragmented nucleic acids by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acids, wherein the correction factor is based on the length distribution (LD) and length of the sequences to be amplified, including the detected sequences. The method includes:

[0070] The length of the sequence to be amplified is L a which is actually the sum of the lengths of the primers (forward and reverse) and any additional base pairs located between them.

[0071] In one embodiment, the length L of the sequence to be amplified a is equal to or greater than the length of the primers (forward and reverse), and in some embodiments is about 200 bp or less, e.g., about 170 bp or less, 150 bp or less, or 130 bp or less.

[0072] For replication, a primer must first bind to the nucleic acid. Binding is optimal when all bases of the primer are complementary to the nucleic acid (for DNA, adenine is complementary to thymine and guanine is complementary to cytosine; for RNA, adenine is complementary to uracil and guanine is complementary to cytosine), but binding can also be efficient if some bases of the primer are not complementary to the nucleic acid. In other words, if the nucleic acid sequence to which the primer is to bind is shortened by a few bases, the replication process can still be efficient. A relevant parameter is L a , or L a -n (n is an integer greater than 0), or rL a (r and L a and product) (the shortening coefficient r is in the range of 75% to 100%, but rL a (where n is always an integer value). In some embodiments, n is an integer between 1 and 15, between 1 and 10, or between 1 and 5. In some embodiments, r is greater than 0.75 and less than 1, greater than 0.8 and less than 1, greater than 0.85 and less than 1, greater than 0.9 and less than 1, or greater than 0.95 and less than 1.

[0073] The correction factor is the length parameter L acan be determined according to any of the embodiments described in the "Determining the Correction Factor" subsection below.

[0074] The correction factor is calculated by multiplying L by 1, as described in the "Determining the Correction Factor" subsection below, to account for shortened fragments that still result in replication. a -n (n is an integer greater than 0 (e.g., 1 to 15, 1 to 10, or 1 to 5)) or rL a (r and L a and product) (the shortening coefficient r is in the range of 75% to 100%, but rL a is always an integer value).

[0075] In some embodiments, a method for determining the concentration of a detected sequence in unfragmented nucleic acid comprises the steps of: i. determining the length distribution of nucleic acids in a sample comprising fragmented nucleic acids, wherein the fragmented nucleic acids are derived from said non-fragmented nucleic acids; ii. measuring the concentration of said detected sequence in said sample comprising fragmented nucleic acids by a measurement method; and iii. correcting the measured concentration of said detected sequence in said sample comprising fragmented nucleic acids by a correction factor to obtain the concentration of said detected sequence in non-fragmented nucleic acids, wherein the correction factor is

number

number

[0076] In some embodiments, L is L a is.

[0077] Regarding the length distribution of nucleic acid fragments, L a (or L a -n or rL a ) fragments of length shorter than 1 / 2 s will not be replicated anyway. If such a fragment contains part of the sequence to be amplified, it will not be replicated, which will lead to an underestimation of the concentration.

[0078] The correction factor described herein can also be based on the probability that the sequence to be amplified will be fragmented based on sequence fragmentation bias. In some embodiments, the relative probability that the sequence to be amplified will be fragmented is known. In some embodiments, the probability that the sequence to be amplified will be fragmented depends on the cause of fragmentation (e.g., spontaneous nucleic acid fragmentation or fragmentation by physical means such as sonication). For example, the correction factor can be adjusted by multiplying the unadjusted correction factor by the probability that the sequence to be amplified will be fragmented based on sequence fragmentation bias. Alternatively, the probability that the sequence to be amplified will be fragmented based on sequence fragmentation bias can be accounted for by modifying the fragmentation length distribution (LD) curve.

[0079] In some embodiments, the at least one parameter of the measurement method further comprises a parameter of the amplification step selected from the group consisting of the GC content of the sequence to be amplified, the GC content of the amplification primers, the length of the amplification primers, the type of polymerase used, and the temperature of the amplification cycle. In some embodiments, the additional parameter of the amplification step can be incorporated as a product of a factor and a fragment length distribution (LD) curve, or as a product of a calibration factor and a correction factor, as described in the "Calibration of the Correction Factor" subsection below.

[0080] In some embodiments, the at least one parameter of the measurement method further comprises a parameter of the measurement step selected from the group consisting of the sequence of the detection probe, the photostability of the fluorophore used, the chemical stability of the fluorophore used, the quantum yield of the fluorophore used, and the wavelength of the fluorophore used. In some embodiments, the additional parameter of the amplification step can be incorporated as a product of a factor and a fragment length distribution (LD) curve, or as a product of a calibration factor and a correction factor, as described in the "Calibration of the Correction Factor" subsection below.

[0081] In some embodiments, the correction involves multiplying the measured concentration in a sample containing fragmented nucleic acids by a correction factor. In some embodiments, the correction involves multiplying the measured concentration in a sample containing fragmented nucleic acids by the correction factor and an additional correction factor. In some embodiments, the additional correction factor is based on the probability that the sequence to be amplified will be fragmented based on sequence fragmentation bias. In some embodiments, the additional correction factor is based on at least one parameter of the measurement method, as described above. In some embodiments, the additional correction factor is based on at least one parameter of the measurement method that affects sequence amplification (e.g., the GC content of the sequence to be amplified). In some embodiments, the additional correction factor is based on at least one parameter of the measurement method that affects detection of the sequence to be detected (e.g., the photostability or chemical stability of the fluorophore used). In some embodiments, the additional correction factor is based on an experimentally determined calibration factor, as described in the "Calibration of the Correction Factor" subsection below.

[0082] In some embodiments, the correction is applied when at least 5% of the nucleic acids in the sample have a length shorter than the length of the sequence to be amplified, hi some embodiments, 95% or less of the nucleic acid fragments in the sample have a length shorter than the length of the sequence to be amplified.

[0083] The present application also provides a method for determining a function of a first concentration of a first detected sequence (detected sequence S1) and a second concentration of a second detected sequence (detected sequence S2) in the same sample of unfragmented nucleic acid, where the function can be a fraction or a ratio.

[0084] To this end, the method described herein comprises: i. measuring the concentration of a second sequence to be detected in the sample containing fragmented nucleic acids by a measurement method; and ii. correcting the concentration of said second detected sequence in those fragmented nucleic acids by a correction factor to obtain the concentration of said second detected sequence in unfragmented nucleic acids, wherein the correction factor is based on the length distribution (LD) and at least one parameter of the measurement method. It may further include:

[0085] In some embodiments, measuring the concentration of the second detected sequence comprises amplifying a second sequence to be amplified that includes the second detected sequence. In some embodiments, the measurement method comprises a multiplex amplification step. In some embodiments, the length of the first sequence to be amplified is different from the length of the second sequence to be amplified, and different correction factors are applied to the first detected sequence and the second detected sequence. In some embodiments, the lengths of the first sequence to be amplified and the second sequence to be amplified are the same. In some embodiments, other parameters of the measurement method differ between the first detected sequence and the second detected sequence (e.g., parameters related to probe binding / detection, or any parameters affecting amplification, as described in the "Determining Correction Factors" subsection below).

[0086] In some embodiments, the first sequence to be amplified comprises a mutant allele-detecting sequence or a variant allele-detecting sequence, and the second sequence to be amplified comprises a corresponding reference allele-detecting sequence. In some embodiments, the first sequence to be amplified comprises an insertion or deletion compared to the corresponding reference sequence contained in the second sequence to be amplified. In some embodiments, the method further comprises determining a corrected mutant allele frequency (MAF) of the first detected sequence by comparison with the second detected sequence. In some embodiments, the method further comprises determining a corrected variant allele frequency (VAF) of the first detected sequence by comparison with the second detected sequence.

[0087] In some embodiments, the first sequence to be amplified is amplified from a variant nucleic acid containing a copy number variation (CNV), and the second sequence to be amplified is amplified from a reference nucleic acid. In some embodiments, the method comprises: amplifying the corrected copy number variation (CNV) of the reference sequence compared to the variant sequence; 実際値 ) determining the

[0088] In any of the foregoing embodiments, the method may further include calibrating a correction factor based on the measured concentrations of the detected sequences in a first nucleic acid sample having a first fragment length distribution (LD1) and a second nucleic acid sample having a second fragment length distribution (LD2).

[0089] Nucleic acid samples and length distribution In some embodiments provided herein, a sample containing fragmented nucleic acids is provided. These fragmented nucleic acids are actually fragments from the original nucleic acids that have undergone fragmentation, i.e., non-fragmented nucleic acids. In fact, the concentration of the sequence to be detected in the non-fragmented nucleic acids is the desired measurement, but the available sample is actually fragmented. In some embodiments, the method does not include fragmenting non-fragmented nucleic acids to generate a sample containing fragmented nucleic acids.

[0090] The sample may contain any type of fragmented nucleic acid. In particular, the sample may be a cell-free sample, i.e., a biological fluid in which nucleic acids have been released from cells, such as saliva, plasma, urine, or whole blood; or a cell-containing sample, i.e., a biological sample that essentially contains cells, such as biopsy tissue. Furthermore, the sample may be a naturally fragmented sample, i.e., the non-fragmented nucleic acid is naturally degraded before sample collection, i.e., within a living organism, or after sample collection due to preservation processing or storage conditions. Finally, the sample may be an artificially fragmented sample, i.e., the non-fragmented nucleic acid is artificially degraded after sample collection according to the needs of the measurement method. Furthermore, the sample may be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0091] Examples of cell-free, naturally fragmented DNA samples are: DNA digested by caspase-activated DNases after lysosomal DNase II digestion during apoptosis and after phagocytosis of dying cells; and Circulating DNA cleaved by plasma nucleases is.

[0092] Examples of naturally fragmented RNA samples include, but are not limited to, messenger RNA cleaved by ribonucleases (such as endonucleases and exonucleases). These samples can be cell-free or cell-containing samples.

[0093] Examples of artificially fragmented DNA and RNA samples are: Nucleic acids fragmented after formalin-fixed, paraffin-embedded (FFPE) samples (examples of cell-containing samples); and Nucleic acids fragmented using acoustic shearing (example of acellular sample) is.

[0094] The measurement method carried out in the measurement step of the basic method of the present invention can be an isothermal quantitative nucleic acid amplification method or a non-isothermal quantitative nucleic acid amplification method.

[0095] The isothermal quantitative nucleic acid amplification method can be loop-mediated isothermal amplification or nucleic acid sequence-based quantitative amplification, which can be combined with a reverse transcription step to enable RNA detection.

[0096] Non-isothermal quantitative nucleic acid amplification methods can be quantitative polymerase chain reaction (PCR), real-time polymerase chain reaction (PCR), digital polymerase chain reaction (DPCR), multiplexed polymerase chain reaction (PCR), or multiplexed digital polymerase chain reaction (DPCR), which can be combined with a reverse transcription step to allow for RNA detection.

[0097] In some embodiments of the present method, a length distribution (LD) of nucleic acid fragments is provided. In the present invention, fragments refer to individual nucleic acids resulting from natural or artificial fragmentation of the original unfragmented nucleic acid. By way of example and not limitation, an original unfragmented nucleic acid having a length of 10,000 base pairs (bp) can be fragmented into, for example, 25 fragments having a length of 75 bp, 50 fragments having a length of 100 bp, and 25 fragments having a length of 125 bp, to generate a population of short-chain nucleic acids. Next, it is possible to define the length distribution of the nucleic acid fragments of this population, i.e., the number of fragments having a given length for all possible lengths. The length distribution can also be defined using a common statistical function (e.g., Gaussian or Poisson distribution) or parameters such as the mean and standard deviation.

[0098] Suitable devices for measuring the length distribution of nucleic acid fragments in a sample are, for example, the Tape Station 4200 or Bioanalyzer 2100 (both instruments from Agilent Technologies), or the LabChip GX Touch nucleic acid analyzer (from PerkinElmer).

[0099] Regarding the length distribution of nucleic acid fragments, the length of the sequence to be amplified, L a (or La -n or rL a Fragments having a length shorter than 1 / 2 (as described above) are not replicated. If such a fragment contains a portion of the sequence to be amplified, it will not be replicated, which will lead to an underestimation of the concentration. The methods provided herein can be applied to correct the resulting underestimation of the concentration of nucleic acids in unfragmented nucleic acid samples.

[0100] The correction factor is that only a few fragments a (or L a -n or rL a ) may be smaller. In practice, Applicant believes that the correction factor may be smaller if at least 5% of the nucleic acid fragments have a length shorter than L a (or L a -n or rL a ) is more appropriate when the length is shorter than L a (or L a -n or rL a ) is at least 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44% In some embodiments, determining a correction factor and correcting the concentration of the detected sequences may be 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, or 85%. ... a (or L a -n or rL a ) is applied when the length is shorter than

[0101] Furthermore, the correction factor is that almost all fragments are La (or L a -n or rL a ) can be very high. In practice, Applicant believes that the correction factor should be such that at most 95% of the nucleic acid fragments have lengths shorter than L a (or L a -n or rL a ) is more appropriate when the length is shorter than L a (or L a -n or rL a ) is at most 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, It can be 56%, 55%, 54%, 53%, 52%, 51%, 50%, 49%, 48%, 47%, 46%, 45%, 44%, 43%, 42%, 41%, 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, or 15%.

[0102] In particular, L a (or L a -n or rL a) The percentage of fragments with lengths shorter than 5% is 5%-95%, 5%-90%, 5%-85%, 5%-80%, 5%-75%, 5%-70%, 5%-65%, 5%-60%, 5%-55%, 5%-50%, 10%-95%, 10%-90%, 10%-85%, 10%-80%, 10%-75%, 10%-70%, 10%-65%, 10%-60%, 10%-55%, 10%-50%, 15%-95%, 15%-90%, 1%-10%. 5%~85%, 15%~80%, 15%~75%, 15%~70%, 15%~65%, 15%~60%, 15%~55%, 15%~50%, 20%~95%, 20%~90%, 20%~85%, 20%~80%, 20%~75%, 20%~70%, 20%~65%, 20%~60%, 20%~55%, 20%~50%, 25%~95%, 25%~90%, 25%~85%, 25%~80%, 25%~75%, 25% ~70%, 25%~65%, 25%~60%, 25%~55%, 25%~50%, 30%~95%, 30%~90%, 30%~85%, 30%~80%, 30%~75%, 30%~70%, 30%~65%, 30%~60%, 30%~55%, 30%~50%, 35%~95%, 35%~90%, 35%~85%, 35%~80%, 35%~75%, 35%~70%, 35%~65%, 35%~60%, 35%~ It may be in a range selected from 55%, 35% to 50%, 40% to 95%, 40% to 90%, 40% to 85%, 40% to 80%, 40% to 75%, 40% to 70%, 40% to 65%, 40% to 60%, 40% to 55%, 40% to 50%, 45% to 95%, 45% to 90%, 45% to 85%, 45% to 80%, 45% to 75%, 45% to 70%, 45% to 65%, 45% to 60%, 45% to 55%, and 45% to 50%.

[0103] In another embodiment, L a In some embodiments, the length L of the sequence to be amplified is longer than 40 bp and shorter than 200 bp, preferably longer than 50 bp and shorter than 170 bp, more preferably longer than 65 bp and shorter than 150 bp, and even more preferably longer than 70 bp and shorter than 130 bp. aranges from 65-70bp, 70-75bp, 75-80bp, 80-90bp, 90-100bp, 100-110bp, 110-120bp, or 120-130bp.

[0104] In another embodiment, the length distribution of the nucleic acid fragments is in the range of 10 bp to 1000 bp, preferably 25 bp to 350 bp, preferably 30 bp to 320 bp, more preferably 35 bp to 290 bp, and even more preferably 40 bp to 270 bp.

[0105] Determination of correction coefficients In some embodiments, a correction factor is determined. In fact, if the original non-fragmented nucleic acid is fragmented within the sequence to be amplified, the measurement method used cannot detect the sequence to be amplified. This lack of detection leads to an underestimation of the concentration. This correction factor depends on the length distribution of nucleic acid fragments and parameters related to defining the probability that copies of the sequence to be amplified will be cut during fragmentation, such as the length of the sequence to be amplified, and potentially parameters related to the measurement method in particular.

[0106] In some embodiments, the measured concentration of the detected sequence is corrected with a correction factor to obtain the concentration of the detected sequence in the unfragmented nucleic acid, i.e., in the original sample.

[0107] In some embodiments, the measurement method performed to measure the concentration of the sequence to be detected essentially uses several relevant parameters to determine a correction factor, hi some embodiments, the measurement method is an amplification method using a replication and amplification mixture.

[0108] A particularly relevant parameter is the length of the sequence to be amplified (hereafter referred to as L a The length of the sequence to be amplified is L a is actually the sum of the lengths of the primers (forward and reverse) and any additional base pairs located between them.

[0109] In one embodiment, the length L of the sequence to be amplified a is equal to or greater than the length of the primers (forward and reverse) and is shorter than 200 bp, preferably shorter than 170 bp, more preferably shorter than 150 bp, and even more preferably shorter than 130 bp.

[0110] For replication, a primer must first bind to the nucleic acid. Binding is optimal when all bases of the primer are complementary to the nucleic acid (for DNA, adenine is complementary to thymine and guanine is complementary to cytosine; for RNA, adenine is complementary to uracil and guanine is complementary to cytosine), but binding can also be efficient if some bases of the primer are not complementary to the nucleic acid. In other words, if the nucleic acid sequence to which the primer is to bind is shortened by a few bases, the replication process can still be efficient. A relevant parameter is L a , or L a -n (n is an integer from 1 to 15), or rL a (r and L a and product) (the shortening coefficient r is in the range of 75% to 100%, but rL a In some embodiments, r can be greater than 0.75 and less than 1, greater than 0.8 and less than 1, greater than 0.9 and less than 1, or greater than 0.95 and less than 1 (where rL a is always an integer value).

[0111] In some embodiments, the value of n or r can be predicted based on the predicted strength of the primer binding to the sequence to be amplified, shortened by several bases, to which the primer should bind. For example, the value of n or r can be predicted based on parameters of the sequence to be amplified and / or the primer (e.g., the GC content of the sequence to be amplified; the GC content of the amplification primer; the length of the amplification primer; the type of polymerase used; the temperature of the amplification cycle). In some embodiments, the value of n or r can be predicted based on (a) the predicted melting temperature (Tm) of the primer binding to the full-length sequence to be amplified, (b) the predicted melting temperature of the primer binding to the sequence to be amplified shortened by n or r (i.e., La-n or r.La) as described above, and (c) the annealing temperature used in the amplification method.

[0112] In some embodiments, the value of n or r can be determined experimentally, for example, in a calibration step as described in the "Calibrating the Correction Factor" subsection below. In some embodiments, n or r can be determined experimentally for a given sequence to be amplified and set of amplification primers for one sample, and then used to calibrate the L parameter for any sample with the same sequence to be amplified and amplification conditions.

[0113] Length parameter L a Using this, the correction factor can be determined as follows: Let P(X) be the probability of event X. L a is the length of the sequence to be amplified (unit: number of base pairs). Let f be the probability distribution of the lengths of nucleic acid fragments in a sample (f(i) is the probability that a fragment in a sample has a length of i base pairs).

number

number

number

number

[0114] By applying Bayes' rule P(A∩B)=P(A / B)P(B)

number

number

[0115] Furthermore, let N be the total number of fragments in the fragmented nucleic acid sample:

number

[0116] Therefore:

number

[0117] Also,

number

[0118] Finally, the concentration of the detected sequence in the sample of unfragmented nucleic acid (C 実際値 ) is the measured concentration (C 測定値 ) multiplied by a correction factor:

number

[0119] This correction factor is determined by the length distribution of nucleic acid fragments in the sample (f) and the parameters of the amplification method, i.e., the length of the sequence to be amplified, L a Depends on.

[0120] The correction factor is L to take into account the shortened fragments that still lead to replication. a -n (n is an integer greater than 0 (e.g., an integer between 1 and 15, between 1 and 10, or between 1 and 5)) or rL a (r and L a and product) (the shortening coefficient r is in the range of 75% to 100%, but rL a is always an integer value).

[0121] For example, in some embodiments where a set of amplification primers is still able to bind to and amplify a target region shortened by n, the correction factor is

number

number

[0122] In some embodiments, where a set of amplification primers is still able to bind to and amplify a target region that has been shortened by a shortening factor r, the correction factor is

number

number

[0123] In particular, it may be appropriate to measure the frequency of mutated nucleic acid, for example, mutant allele frequency (MAF).In this particular case, the first detected sequence (S1) is mutant (mut), and the second detected sequence (S2) is wild type (wt).If both detected sequences are related to the sequences to be amplified of different lengths, the correction factors for both concentrations are different and need to be taken into account.

[0124] It may also be appropriate to measure the ratio of the amplified variant nucleic acid (var) to the amplified reference nucleic acid (ref), i.e., copy number variation (CNV) (variant is not necessarily a mutation of the reference).

[0125] In some embodiments, the method comprises the steps of:

[0126] In the first step, the concentration of S1 in the unfragmented nucleic acid is determined according to the basic method described above, in which a correction factor is determined taking into account the length of the sequence to be amplified relative to S1.

[0127] In a second step, the concentration of S2 in the unfragmented nucleic acid is determined according to the basic method described above, in which a correction factor is determined taking into account the length of the sequence to be amplified relative to S2.

[0128] In the third step, a function of S1 and S2 concentrations is determined.

[0129] In this method, the S1 concentration and the S2 concentration are determined in the same sample. Multiplex PCR or multiplex digital PCR is particularly suitable for performing measurements of S1 and S2 concentrations in the same sample.

[0130] In this method, the length of the sequence to be amplified associated with S1 is different from the length of the sequence to be amplified associated with S2.

[0131] In the specific case of mutant allele frequency (MAF), the function is a fraction determined using the following relationship:

number

[0132] The corrected variant allele frequency (VAF) can be determined according to the same method used to determine the mutant allele frequency (MAF).

[0133] In the specific case of copy number variation (CNV), the function is a ratio determined using the following relationship:

number

[0134] Correction factor calibration In some embodiments, the method further comprises calibrating the correction factor, hi some embodiments, the calibration factor is calibrated based on parameters of the sequence to be amplified or parameters of the measurement method (e.g., parameters of the amplification step or the detection step).

[0135] In some embodiments provided herein, the correction factor can also be based on the probability that the sequence to be amplified will be fragmented based on sequence fragmentation bias. In some embodiments, the relative probability that the sequence to be amplified will be fragmented is known. In some embodiments, the probability that the sequence to be amplified will be fragmented depends on the cause of fragmentation (e.g., spontaneous nucleic acid fragmentation or fragmentation by physical means such as sonication).

[0136] In some embodiments, the fragmentation bias depends on the chromatin structure of the region containing the sequence to be amplified. In some embodiments, the fragmentation bias depends on the relative occurrence of sequences associated with fragmentation within the sequence to be amplified (e.g., sites targeted by a restriction enzyme).

[0137] For example, in some embodiments, the nucleic acid sample may contain DNA digested by caspase-activated DNase during apoptosis. In some embodiments, the correction factor may include a correction factor based on the probability that the sequence to be amplified will be fragmented during apoptosis. In most somatic tissues, apoptotic DNA cleavage results in the formation of fragments approximately 195 bp in length and multiples thereof, whereas the fragmentation pattern of neuronal chromatin is characterized by a size of approximately 165 bp. The repeatable length corresponds to the size of a single nucleosome (including degraded DNA linkers). Within the nucleosome core, DNA is protected from nucleases by histones, while the linkers are vulnerable to digestion. Thus, in some embodiments, a correction factor is applied to the correction factor that represents the probability that the sequence to be amplified will be fragmented based on its location within the chromatin structure (e.g., within the nucleosome or linker region).

[0138] In some embodiments, the sequence fragmentation bias is a positional bias. For example, in some embodiments, the nucleic acid sample is an RNA sample. RNA transcripts may be preferentially cleaved at specific positions within the transcript, for example, at the beginning and / or end of the transcript. (Tuerk A, Wiktorin G, Guler S (2017) Mixture models reveal multiple positional bias types in RNA-Seq data and lead to accurate transcript concentration estimates. PLOS Computational Biology 13(5): e1005515). Therefore, this type of bias is referred to as a positional bias. Therefore, in some embodiments, a correction factor is applied to the correction coefficient, which represents the probability that a sequence to be amplified will be fragmented based on the position of the sequence to be amplified within the RNA transcript.

[0139] In some embodiments, the nucleic acid sample can contain artificially fragmented nucleic acids, such as nucleic acids produced by mechanical shearing or enzymatic fragmentation. For nucleic acid samples containing enzymatically fragmented nucleic acids, a correction factor can be applied to account for the recognition sequence or sequence preference of the enzyme used (e.g., the sequence bias of a transposase).

[0140] In some embodiments, at least one parameter of the measurement method can further include a parameter of the amplification step selected from the group consisting of the GC content of the sequence to be amplified, the GC content of the amplification primer, the length of the amplification primer, the type of polymerase used, and the temperature of the amplification cycle. In some embodiments, at least one parameter of the measurement method can further include a parameter of the measurement step selected from the group consisting of the sequence of the detection probe, the photostability of the fluorophore used, the chemical stability of the fluorophore used, the quantum yield of the fluorophore used, and the wavelength of the fluorophore used. In some embodiments, the additional parameter of the amplification step can be incorporated as a product of a coefficient and a length distribution (LD) curve, or as a product of a calibration coefficient and a correction coefficient.

[0141] In some embodiments, the effect of additional parameters of the measurement method, such as those described above, can be determined experimentally and incorporated into the correction factor as an experimentally determined calibration factor.

[0142] In some embodiments, the method further includes calibrating a correction factor based on the measured concentrations of the detected sequences in a first nucleic acid sample having a first fragment length distribution (LD1) and a second nucleic acid sample having a second fragment length distribution (LD2). In some embodiments, the first nucleic acid sample is an unfragmented sample. In some embodiments, the first nucleic acid sample contains fragmented nucleic acids in which less than 5% of the nucleic acid fragments have a length shorter than the length of the sequence to be amplified. In some embodiments, the first nucleic acid sample is used to determine a ground truth correction factor for the second nucleic acid sample. The relative error between the ground truth correction factor and the predicted correction factor can be determined and applied as a calibration factor to calculate the corrected concentration of the detected sequence.

[0143] In some embodiments, as one skilled in the art would readily appreciate, n or r can be determined based on the value of n or r that results in the best match between the predicted correction coefficients and the ground truth correction factors.

[0144] In some methods according to any one of the embodiments provided herein, the method further comprises correcting the concentration with a calibration factor based on any one of the above parameters. In some embodiments, the calibration factor can be applied as a corrected dilution factor according to any one of the measurement methods provided herein.

[0145] II. The system of the present application Finally, the invention relates to a system adapted to determine the concentration of a sequence to be detected in unfragmented nucleic acid, said system comprising the following modules:

[0146] The first module is configured to measure the concentration of a sequence to be detected in a sample of fragmented nucleic acids, said fragmented nucleic acids being derived from said non-fragmented nucleic acids, and this module implements the measurement step of the basic method of the invention.

[0147] A suitable module includes a real-time thermocycler, which performs the amplification reaction using a reaction mixture containing primers, intercalating fluorescent dyes or probes, and a polymerase enzyme in an appropriate buffer. Alternatively, a digital PCR platform can be used instead of the real-time thermocycler. Such a digital PCR platform consists of a PCR reservoir (often a tube, plate, or microfluidic chip) partitioning system, a thermocycler, a fluorescence reader, and analysis software.

[0148] The second module is configured to calculate a correction factor according to the nucleic acid fragment length distribution of the fragmented nucleic acid and the measurement parameters. This module is generally a computer device including a display screen, at least one microprocessor, a data exchange module, and at least one computer-readable storage medium. Alternatively, this module can be connected to a remote server including at least one microprocessor, a data exchange module, and at least one computer-readable storage medium.

[0149] A computer program comprising instructions that, when executed by a computer or a remote server, can cause the computer or the remote server to automatically calculate a correction factor.

[0150] A computer-readable medium containing instructions that can be used when a computer program is executed by a computer or remote server. In one embodiment, the computer-readable medium is a non-transitory computer-readable medium.

[0151] The third module is configured to calculate the concentration of the detected sequence in the unfragmented nucleic acid by a correction factor.

[0152] The first module can be an isothermal quantitative nucleic acid amplification module or a non-isothermal quantitative nucleic acid amplification module.

[0153] The isothermal quantitative nucleic acid amplification module generally contains primers compatible with isothermal amplification, intercalating fluorescent dyes and polymerase enzymes in a suitable isothermal buffer to carry out the amplification reaction, and a temperature-controlled fluorescent scanner with compatible analysis software. Preferred modules are those that perform loop-mediated isothermal amplification, nucleic acid sequence-based quantitative amplification, signal-mediated amplification and strand displacement amplification of RNA technologies.

[0154] Non-isothermal quantitative nucleic acid amplification modules are generally quantitative polymerase chain reaction modules, real-time polymerase chain reaction modules, digital polymerase chain reaction modules, multiplex polymerase chain reaction modules and multiplex digital polymerase chain reaction modules.

[0155] The second module may calculate a correction factor using the length of the sequence to be amplified as a parameter of the measurement performed by the first module.

[0156] All these parameters can be included one by one or in combination in the calculation of the correction factor performed by the second module.

[0157] IIi. Exemplary Embodiments One aspect of the present application provides a method for correcting the concentration of a detected sequence in unfragmented nucleic acid. Another aspect of the present application provides a system configured to determine the concentration of a detected sequence in unfragmented nucleic acid.

[0158] To this end, the present application provides the following exemplary embodiments: Embodiment 1. A method for determining the concentration of a sequence to be detected in unfragmented nucleic acid, comprising the steps of: i. providing a sample of fragmented nucleic acids, said fragmented nucleic acids being derived from said unfragmented nucleic acids; ii. providing a length distribution (LD) of nucleic acid fragments of the fragmented nucleic acids; iii. measuring the concentration of said sequence to be detected in said sample of fragmented nucleic acids by a measurement method; iv. determining a correction factor according to the length distribution (LD) of the nucleic acid fragments of the fragmented nucleic acid and the parameters of the measurement method; and v. correcting the concentration of said detected sequence in those fragmented nucleic acids by said correction factor to obtain the concentration of said detected sequence in unfragmented nucleic acids. A method comprising:

[0159] Embodiment 2. A method for determining the concentration of a sequence to be detected in non-fragmented nucleic acids according to embodiment 1, wherein the sample of fragmented nucleic acids is any combination of the following three categories: i. Either a cell-free sample or a cell-containing sample; ii. Either naturally fragmented or artificially fragmented samples; and iii. Any kind of deoxyribonucleic acid or ribonucleic acid.

[0160] Embodiment 3. A method for determining the concentration of a sequence to be detected in unfragmented nucleic acid according to any of the preceding embodiments, wherein said measurement method is an isothermal quantitative nucleic acid amplification method, preferably selected from loop-mediated isothermal amplification and nucleic acid sequence-based quantitative amplification.

[0161] Embodiment 4. A method for determining the concentration of a sequence to be detected in unfragmented nucleic acid according to embodiment 1 or 2, wherein said measurement method is a non-isothermal quantitative nucleic acid amplification method, preferably selected from quantitative polymerase chain reaction, real-time polymerase chain reaction, digital polymerase chain reaction, multiplex polymerase chain reaction and multiplex digital polymerase chain reaction.

[0162] Embodiment 5. A method for determining the concentration of a sequence to be detected in non-fragmented nucleic acid according to embodiment 3 or 4, wherein the parameters of said measurement method include the length of the sequence to be amplified.

[0163] Embodiment 6. A method for determining the concentration of a sequence to be detected in a non-fragmented nucleic acid according to embodiment 5, wherein at least 5% of the nucleic acid fragments have a length shorter than the length of the sequence to be amplified.

[0164] Embodiment 7. A method for determining the concentration of a sequence to be detected in a non-fragmented nucleic acid according to embodiment 5 or 6, wherein at most 95% of the nucleic acid fragments have a length shorter than the length of the sequence to be amplified.

[0165] Embodiment 8. A method for determining the concentration of a sequence to be detected in non-fragmented nucleic acid according to any one of embodiments 5 to 7, wherein the length of the sequence to be amplified is longer than 40 bp and shorter than 200 bp, preferably longer than 50 bp and shorter than 170 bp, more preferably longer than 65 bp and shorter than 150 bp, and even more preferably longer than 70 bp and shorter than 130 bp.

[0166] Embodiment 9. A method for determining the concentration of a sequence to be detected in a non-fragmented nucleic acid according to any one of embodiments 1 to 8, wherein the length distribution of the nucleic acid fragments is within the range of 25 bp to 350 bp, preferably 30 bp to 320 bp, more preferably 35 bp to 290 bp, and even more preferably 40 bp to 270 bp.

[0167] Embodiment 10. A method for determining a function of a first concentration of a first detected sequence S1 and a second concentration of a second detected sequence S2 in unfragmented nucleic acid, comprising: iv. determining the concentration of S1 in the unfragmented nucleic acid according to any one of embodiments 5 to 9; v. Determining the concentration of S2 in the unfragmented nucleic acid according to any one of embodiments 5 to 9; and vi. calculating said function of said S1 concentration and said S2 concentration Includes; wherein the S1 concentration and the S2 concentration are determined in the same sample; and the length of the sequence to be amplified associated with S1 is different from the length of the sequence to be amplified associated with S2.

[0168] Embodiment 11. A system configured to determine the concentration of a detected sequence in unfragmented nucleic acid, comprising: i. a module configured to measure the concentration of the sequence to be detected in a sample of fragmented nucleic acids, the fragmented nucleic acids being derived from the non-fragmented nucleic acids; ii. a module configured to calculate a correction factor according to the nucleic acid fragment length distribution (LD) of the fragmented nucleic acid and parameters of the measurement; iii. a module configured to calculate the concentration of the detected sequence in unfragmented nucleic acid using the correction factor; A system comprising:

[0169] Embodiment 12. A system configured to determine the concentration of a sequence to be detected in unfragmented nucleic acids according to embodiment 11, wherein the module configured to measure the concentration of the sequence to be detected in the fragmented nucleic acids is an isothermal quantitative nucleic acid amplification module, preferably selected from a loop-mediated isothermal amplification module and a nucleic acid sequence-based quantitative amplification module.

[0170] Embodiment 13. A system configured to determine the concentration of a detected sequence in unfragmented nucleic acids according to embodiment 11, wherein the module configured to measure the concentration of the detected sequence in the fragmented nucleic acids is a non-isothermal quantitative nucleic acid amplification module, preferably selected from a quantitative polymerase chain reaction module, a real-time polymerase chain reaction module, a digital polymerase chain reaction module, a multiplex polymerase chain reaction module, and a multiplex digital polymerase chain reaction module.

[0171] Embodiment 14. A system configured to determine the concentration of a sequence to be detected in unfragmented nucleic acid according to embodiment 12 or 13, wherein the parameters of said measurement comprise the length of the sequence to be amplified.

[0172] IV. Embodiments One aspect of the present application provides a method for correcting a problem related to underestimating the presence of a nucleic acid sequence of interest in a nucleic acid sample containing fragmented nucleic acids. Another aspect of the present application provides a system configured to determine the concentration of a detected sequence in unfragmented nucleic acids.

[0173] To this end, the present application provides the following embodiments: Embodiment 1'. A method for determining the concentration of a sequence to be detected in unfragmented nucleic acid, comprising the steps of: i. determining the length distribution of nucleic acids in a sample comprising fragmented nucleic acids, wherein the fragmented nucleic acids are derived from said unfragmented nucleic acids; ii. measuring the concentration of the sequence to be detected in the sample containing fragmented nucleic acids by a measurement method; and iii. correcting the measured concentration of said detected sequence in said sample containing fragmented nucleic acids by a correction factor to obtain the concentration of said detected sequence in unfragmented nucleic acids, wherein the correction factor is based on the length distribution and at least one parameter of the measurement method. A method comprising:

[0174] Embodiment 2'. The method of embodiment 1', wherein measuring the concentration of the sequence to be detected comprises amplifying a sequence to be amplified that includes the sequence to be detected.

[0175] Embodiment 3'. The method of embodiment 2', wherein the sequence to be detected and the sequence to be amplified are the same, and the measuring step comprises detecting incorporation of a label into nucleic acids produced during amplification.

[0176] Embodiment 4'. The method of embodiment 3', wherein the label is a fluorescent dye comprising a fluorophore.

[0177] Embodiment 5'. The method of embodiment 2', wherein the sequences to be detected are a subset of the sequences to be amplified, and the measuring step comprises detecting binding of a labeled probe to the sequences to be detected.

[0178] Embodiment 6'. The method of embodiment 4', wherein the labeled probe is a fluorescently labeled probe comprising a fluorophore.

[0179] Embodiment 7'. The method of any one of the preceding embodiments, further comprising determining a correction factor based on the length distribution (LD) of nucleic acids in the sample and at least one parameter of said measurement method.

[0180] Embodiment 8'. At least one parameter of the measurement method is the length of the sequence to be amplified (L a13. The method of any one of the preceding embodiments, comprising:

[0181] Embodiment 9'. The correction coefficient is

number

number

[0182] Embodiment 10'. A set of amplification primers is capable of still binding to the sequence to be amplified shortened by n, and the correction factor is

number

number

[0183] Embodiment 11'. The method of embodiment 10', wherein n is an integer from 1 to 5.

[0184] Embodiment 12'. A set of amplification primers is capable of still binding to a sequence to be amplified that has been shortened by a shortening factor r, and the correction factor is

number

number

[0185] Embodiment 13'. The method of any one of the preceding embodiments, wherein the correction factor is also based on the probability that the sequence to be amplified will be fragmented based on a sequence fragmentation bias.

[0186] Embodiment 14'. The method of any of embodiments 2' to 13', wherein at least one parameter of the measurement method further comprises a parameter of the amplification step selected from the group consisting of: the GC content of the sequence to be amplified; the GC content of the amplification primers; the length of the amplification primers; the type of polymerase used; and the temperature of the amplification cycle.

[0187] Embodiment 15'. The method of any one of embodiments 4' or 6', wherein at least one parameter of the measuring method further comprises a parameter of the measuring step selected from the group consisting of: the sequence of the detection probe; the photostability of the fluorophore used; the chemical stability of the fluorophore used; the quantum yield of the fluorophore used; and the wavelength of the fluorophore used.

[0188] Embodiment 16'. The method of any one of the preceding embodiments, wherein the correction comprises multiplying the concentration measured in the sample containing the fragmented nucleic acid by a correction factor.

[0189] Embodiment 17'. The method of any one of the preceding embodiments, wherein the correction is applied if at least 5% of the nucleic acids in the sample have a length shorter than the length of the sequence to be amplified.

[0190] Embodiment 18'. The method of any one of the preceding embodiments, wherein no more than 95% of the nucleic acid fragments in the sample have a length shorter than the length of the sequence to be amplified.

[0191] Embodiment 19'. The method of any one of the preceding embodiments, wherein the length of the sequence to be amplified is greater than 40 bp and less than 200 bp.

[0192] Embodiment 20'. The method of embodiment 19', wherein the length of the sequence to be amplified is greater than 50 bp and less than 170 bp.

[0193] Embodiment 21'. The method of embodiment 19', wherein the length of the sequence to be amplified is longer than 65 bp and shorter than 150 bp.

[0194] Embodiment 22'. The method of embodiment 19', wherein the length of the sequence to be amplified is longer than 70 bp and shorter than 130 bp.

[0195] Embodiment 23'. The method of any one of the preceding embodiments, wherein the length distribution (LD) of the nucleic acid fragments in the sample is comprised between 25 bp and 350 bp.

[0196] Embodiment 24'. The method of embodiment 23', wherein the length distribution (LD) of the nucleic acid fragments in the sample is in the range of 30 bp to 320 bp.

[0197] Embodiment 25'. The method of embodiment 23', wherein the length distribution (LD) of the nucleic acid fragments in the sample is in the range of 35 bp to 290 bp.

[0198] Embodiment 26'. The method of embodiment 23', wherein the length distribution (LD) of the nucleic acid fragments in the sample is in the range of 40 bp to 270 bp.

[0199] Embodiment 27'. The method of any one of the preceding embodiments, which does not include obtaining the sequence of the fragmented nucleic acid.

[0200] Embodiment 28'. The method of any one of the preceding embodiments, which does not include obtaining or predicting genetic coordinates of the fragmented nucleic acids.

[0201] Embodiment 29'. The method of any one of the preceding embodiments, which does not include forming a contiguous sequence of the fragmented nucleic acids.

[0202] Embodiment 30'. The method of any one of the preceding embodiments, which does not include fragmenting unfragmented nucleic acids to produce a sample comprising fragmented nucleic acids.

[0203] Embodiment 31'. The method of any one of the preceding embodiments, wherein the measuring method does not include sequencing nucleic acids in the sample.

[0204] Embodiment 32'. The method of any one of the preceding embodiments, wherein the sample comprising fragmented nucleic acids is a cell-containing sample.

[0205] Embodiment 33'. The method of any one of embodiments 1' to 31', wherein the sample containing fragmented nucleic acids is a cell-free sample.

[0206] Embodiment 34'. The method of any one of the preceding embodiments, wherein the sample comprising fragmented nucleic acids is a naturally fragmented sample.

[0207] Embodiment 35'. The method of any one of embodiments 1' to 29' or 31' to 33', wherein the sample containing fragmented nucleic acids is an artificially fragmented sample.

[0208] Embodiment 36'. The method of any one of the preceding embodiments, wherein the sample comprising fragmented nucleic acids is a deoxyribonucleic acid sample.

[0209] Embodiment 37'. The method of any one of the preceding embodiments, wherein the sample comprising fragmented nucleic acids is a ribonucleic acid sample.

[0210] Embodiment 38'. A method for determining the concentration of a sequence to be detected in non-fragmented nucleic acid according to any one of embodiments 2' to 37', wherein the measurement method is an isothermal quantitative nucleic acid amplification method.

[0211] Embodiment 39'. The method of embodiment 38', wherein the measuring method is selected from the group consisting of loop-mediated isothermal amplification and nucleic acid sequence-based quantitative amplification.

[0212] Embodiment 40'. The method of any one of embodiments 2' to 37', wherein the measurement method is a non-isothermal quantitative nucleic acid amplification method.

[0213] Embodiment 41'. The method of embodiment 40', wherein the measuring method is selected from the group consisting of quantitative polymerase chain reaction, real-time polymerase chain reaction, digital polymerase chain reaction, multiplex polymerase chain reaction, and multiplex digital polymerase chain reaction.

[0214] Embodiment 42'. The method of any one of the preceding embodiments, further comprising calibrating a correction factor based on the measured concentrations of detected sequences in a first nucleic acid sample having a first fragment length distribution (LD1) and a second nucleic acid sample having a second fragment length distribution (LD2).

[0215] Embodiment 43'. i. measuring the concentration of a second sequence to be detected in the sample containing fragmented nucleic acids by a measurement method; and ii. correcting the measured concentration of said second detected sequence in those fragmented nucleic acids by a correction factor to obtain the concentration of said second detected sequence in unfragmented nucleic acids, wherein said correction factor is based on the length distribution (LD) and at least one parameter of the measurement method. 13. The method of any one of the preceding embodiments, further comprising:

[0216] Embodiment 44'. The method of embodiment 43', wherein measuring the concentration of the second detected sequence comprises amplifying a second sequence to be amplified that includes the second detected sequence.

[0217] Embodiment 45'. The method of embodiment 44', wherein the length of the first sequence to be amplified is different from the length of the second sequence to be amplified.

[0218] Embodiment 46'. The method of any of embodiments 43'-45', wherein the first sequence to be amplified comprises a mutant allele-detecting sequence and the second sequence to be amplified comprises a corresponding reference allele-detecting sequence.

[0219] Embodiment 47'. The method of any of embodiments 43' to 45', wherein the first sequence to be amplified contains an insertion or deletion compared to the corresponding reference sequence contained in the second sequence to be amplified.

[0220] Embodiment 48'. The method of any of embodiments 43' to 45', further comprising determining a corrected mutant allele frequency (MAF) of the first detected sequence compared to the second detected sequence.

[0221] Embodiment 49'. Corrected MAF (MAF 実際値 )but

number

[0222] Embodiment 50'. The method of any of embodiments 43' to 45', wherein the first sequence to be amplified is amplified from a variant nucleic acid comprising a copy number variation (CNV), and the second sequence to be amplified is amplified from a reference nucleic acid.

[0223] Embodiment 51'. The method of any of embodiments 43' to 45', wherein the first sequence to be amplified is amplified from a variant nucleic acid comprising a copy number variation (CNV), and the second sequence to be amplified is amplified from a reference nucleic acid.

[0224] Embodiment 52’. A method according to any of Embodiments 43’ to 45’, wherein the first sequence to be amplified is amplified from a variant nucleic acid containing a copy number polymorphism (CNV), and the second sequence to be amplified is amplified from a reference nucleic acid.

[0225] Embodiment 53’. The formula:

Number

[0226] Embodiment 54’. A method according to any one of Embodiments 43’ to 51’, wherein the measurement method includes a multiplex amplification step.

[0227] Embodiment 55’. A system configured to determine the concentration of a detected sequence in non-fragmented nucleic acid, comprising: i. a module configured to measure the concentration of the detected sequence in a sample containing fragmented nucleic acids (the fragmented nucleic acids are derived from the non-fragmented nucleic acids); ii. a module configured to calculate a correction factor according to the length distribution (LD) of the nucleic acids in the sample and the parameters of the measurement; iii. a module configured to calculate the concentration of the detected sequence in the non-fragmented nucleic acid by the correction factor A system comprising.

[0228] Embodiment 56’. A system configured to determine the concentration of a detected sequence in non-fragmented nucleic acid according to Embodiment 53’, wherein the module configured to measure the concentration of the detected sequence in the fragmented nucleic acid is an isothermal quantitative nucleic acid amplification module selected from a loop-mediated isothermal amplification module and a nucleic acid sequence-based quantitative amplification module.

[0229] Embodiment 57'. A system configured to determine the concentration of a sequence to be detected in non-fragmented nucleic acids according to embodiment 53', wherein the module configured to measure the concentration of the sequence to be detected in the fragmented nucleic acids is a non-isothermal quantitative nucleic acid amplification module selected from a quantitative polymerase chain reaction module, a real-time polymerase chain reaction module, a digital polymerase chain reaction module, a multiplex polymerase chain reaction module, and a multiplex digital polymerase chain reaction module.

[0230] Embodiment 58'. A system configured to determine the concentration of a sequence to be detected in unfragmented nucleic acid according to embodiment 54' or 55, wherein the parameters of said measurement include the length of the sequence to be amplified.

[0231] Embodiment 59'. A method for determining copy number variation (CNV) of a first concentration of a first detected sequence S1 and a second concentration of a second detected sequence S2 in unfragmented nucleic acid, comprising: i. measuring the concentrations of S1 and S2 in a nucleic acid sample containing fragmented nucleic acids; ii. correcting the measured concentrations of S1 and S2 in the sample containing fragmented nucleic acids by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acids, wherein the correction factor is based on the length distribution (LD) and at least one parameter of the measurement method; iii. calculating the CNV using the corrected concentrations of S1 and S2; Including, where the concentrations of S1 and S2 are measured in the same sample; and At least one parameter of the measurement method associated with S1 is different from at least one parameter of the measurement method associated with S2 ,method.

[0232] Embodiment 60'. The method of embodiment 59', wherein measuring the concentrations of S1 and S2 comprises amplifying a first sequence to be amplified that includes S1 and a second sequence to be amplified that includes S2.

[0233] Embodiment 61'. The method of embodiment 60', wherein at least one parameter of the measuring method comprises the length of the sequence to be amplified.

[0234] Embodiment 62'. The method of embodiment 61', wherein the length of the sequence to be amplified comprising S1 is different from the length of the sequence to be amplified comprising S2.

[0235] While various embodiments have been described and illustrated, the detailed description should not be construed as limiting the scope of the present disclosure, as various modifications may be made thereto by those skilled in the art without departing from the true spirit and scope of the present disclosure, as defined by the following claims. [Example]

[0236] Example 1 Method of Example 1: a. Preparation of sonicated DNA The starting sample was 200 ng / μl stock DNA (=6.06E+4 cp / μl) (human genomic DNA, Bio-35025, Bioline, Paris, France) (average length greater than 50 kbp according to the manufacturer).

[0237] Use Covaris® Microtubes-15 (containing 15-20 μl ± 1 μl) (according to the manufacturer, they can accommodate up to 1 μg of DNA). Therefore, a dilution step of the DNA stock solution is performed using TE:Tris-EDTA. This results in a DNA solution of 60 ng / μl (= 1.82E+4 cp / μl), which can be placed in a Covaris® Microtube-15 and sonicated.

[0238] Sonication is performed with an M220 Focused-ultrasonicator (Covaris®, Brighton, UK).

[0239] First, the smallest fragments achievable by Covaris® (length distribution centered at 150 bp) are prepared to approximate as closely as possible the distribution of human DNA fragment lengths, i.e., to the pattern found in human plasma (163, 316, and 465 bp).

[0240] Subsequently, length distributions centered around 200 bp, 350 bp and 550 bp are prepared.

[0241] All sonications were performed in Covaris® Microtubes-50 (55 μl±2.5) or snap-cap microtubes (130 μl±5), which according to the manufacturer can accommodate up to 5 μg of DNA.

[0242] b. Verification of sonicated DNA samples using TapeStation To obtain the base pair distribution necessary to calculate all theoretical correction factors, grayscale data were extracted from the electrophoresis images on a 4200 TapeStation system (Agilent Technologies, Santa Clara, CA, USA). To convert mass units to base pair units, the image intensity was inverted and then divided by the expected fragment length at the pixel location in the image.

[0243] Use the high-sensitivity D1000 ScreenTape.

[0244] c. Primers, probes and fluorophores The primers and probes used were synthesized by Eurogentec (Eurogentec, Angers, France) and purified by reversed-phase high-performance liquid chromatography (HPLC).

[0245] The fluorophores used in the TriPlex PCR experiments were FAM, HEX, and cyanine (Cy5), as follows: A blue channel FAM fluorophore for detecting the sequence BRAF V600 WT with a sequence to be amplified of 117 base pairs in length; A HEX fluorophore in the green channel to detect the sequence EGFR L858 WT with a sequence to be amplified of 78 base pairs in length; A Cy5 fluorophore in the red channel to detect the sequence ALB using a sequence to be amplified of 81 base pairs in length.

[0246] The prepared template BRAF-EGFR-ALB is shown in Table 1. The notation {} denotes a locked nucleobase.

[0247] [Table 1]

[0248] d. PCR mix PCR reactions were performed using Perfecta® Multiplex qPCR ToughMix® (Quanta Biosciences, Beverly, MA, USA) at a final concentration of 1×, supplemented with 0.1 μM fluorescein (VWR International, Fontenay-sous-Bois, France).

[0249] The PCR mix assembly is as follows: Perfecta® qPCR Multiplex ToughMix® (1x) Fluorescein (0.1 μM) Oligonucleotide BRAF V600 WT, FAM fluorophore (1x) Oligonucleotide EGFR L858 WT, HEX fluorophore (1x) Oligonucleotide ALB, Cy5 fluorophore (1x) ·water

[0250] [Table 2]

[0251] The samples are obtained by diluting the target sequences in PCR mix as described in Table 2, so that the expected final concentration of each target sequence without sonication is 3000 cp / μL.

[0252] e.dPCR experiment Samples were applied to the injection chambers of a sapphire chip (Stilla Technologies, Villejuif, France) with a volume of 27 μL per chamber. Three replicates were performed for each sonicated sample (three chambers per sample). One unsonicated sample was applied in triplicate to three separate chambers.

[0253] The Naica™ Geode (Stilla Technologies, Villejuif, France) is programmed to divide the sample.

[0254] PCR conditions are as follows: 95°C for 10 minutes, followed by 45 cycles of 95°C for 30 seconds and 58°C for 15 seconds.

[0255] The default exposure times for image acquisition with Naica™ Prism3 (Stilla Technologies, Villejuif, France) for the blue, green, and red channels are 65 ms, 250 ms, and 50 ms, respectively.

[0256] f. Calculation of correction factors The correction factor predicted according to the method of the present invention ("Prediction") is calculated from the experimentally measured fragment length distribution of the sample and from the length in base pairs (bp) of the sequence to be amplified, as shown in Figure 2 for a length distribution centered at 150 bp.

[0257] To calculate the correction factor according to the method of the present invention, the probability that the sequence to be amplified is not cut is required as a function of its length, as shown in Figure 3 for a length distribution centered at 150 bp. The correction factor for the same conditions is shown in Figure 4.

[0258] A ground truth correction factor ("ground truth") is obtained in vitro by calculating the ratio of experimentally measured concentrations of the detected sequence in unsonicated and sonicated samples.

[0259] Relative error ("relative error") is defined as the error of the predicted correction factor relative to the ground truth correction factor.

[0260] Results of Example 1: The experimental and theoretical results are compared in Table 3. The predicted and ground truth correction factors were obtained using TriPlex digital PCR experiments to measure the concentrations of sequences to be amplified with different sequence lengths (78 bp, 81 bp, 117 bp) in sonicated samples with different fragment lengths (150 bp, 200 bp, 350 bp, 550 bp). Each experimental measurement was performed in triplicate, and the values ​​shown are the average of the triplicate values.

[0261] [Table 3]

[0262] The obtained relative error values ​​range from 1% to 17%, indicating that the predicted correction factors are consistently accurate and show a slight overestimation relative to the ground truth correction factors.

[0263] From the above, it can be inferred that the method of the present invention provides results that can be directly used in practical conditions, since the length of the sequence to be amplified is representative of a standard PCR sequence and the fragment length distribution is consistent with naturally occurring fragmentation.

Claims

1. 1. A method for determining the concentration of a sequence to be detected in unfragmented nucleic acid, comprising the steps of: i. determining the length distribution (LD) of nucleic acids in a sample comprising fragmented nucleic acids, wherein the fragmented nucleic acids are derived from said non-fragmented nucleic acids; ii. measuring the concentration of the sequence to be detected in the sample containing fragmented nucleic acids by a measurement method; and iii. correcting the measured concentration of the detected sequence in a sample containing fragmented nucleic acids by a correction factor to obtain the concentration of the detected sequence in unfragmented nucleic acids, wherein the correction factor is based on the length distribution (LD) and at least one parameter of the measurement method. A method comprising:

2. The measurement of the concentration of the sequence to be detected includes amplifying a sequence to be amplified that includes the sequence to be detected, and at least one parameter of the measurement method is the length (L a 10. The method of claim 1, comprising:

3. The correction coefficient is [Equation 1] (In the formula, L is the length of the sequence to be amplified (L a ) and f(i) is the probability that a fragment in the sample has a length of i base pairs; and [Equation 2] is the average length of the nucleic acid fragment) The method of claim 2, wherein the value is determined by:

4. 3. The method of claim 2, wherein the correction is applied if at least 5% of the nucleic acids in the sample have a length shorter than the length of the sequence to be amplified.

5. 3. The method of claim 2, wherein the length of the sequence to be amplified is greater than 40 bp and less than 200 bp.

6. The method of claim 1, wherein the length distribution (LD) of the nucleic acid fragments in the sample is in the range of 25 bp to 350 bp.

7. The method of any one of claims 1 to 6, which does not comprise obtaining or predicting genetic coordinates of the fragmented nucleic acids.

8. The method of claim 1, wherein the measurement method is an isothermal quantitative nucleic acid amplification method.

9. The method according to claim 1, wherein the measurement method is a non-isothermal quantitative nucleic acid amplification method.

10. i. measuring the concentration of a second detectable sequence in the sample containing fragmented nucleic acids by a measurement method, wherein measuring the concentration of the second detectable sequence comprises amplifying a second sequence to be amplified that contains the second detectable sequence; and ii. correcting the measured concentration of said second detected sequence in the fragmented nucleic acid by a correction factor to obtain the concentration of said second detected sequence in unfragmented nucleic acid, wherein the correction factor is based on the length distribution (LD) and at least one parameter of the measurement method, and wherein the length of the first sequence to be amplified is different from the length of the second sequence to be amplified. The method of claim 1 further comprising:

11. The method of claim 10, wherein the measurement method comprises a multiplex amplification step.

12. 1. A system configured to determine the concentration of a detected sequence in unfragmented nucleic acid, comprising: i. an amplification module configured to measure the concentration of the sequence to be detected in a sample comprising fragmented nucleic acids, the fragmented nucleic acids being derived from the non-fragmented nucleic acids; ii. A module configured to calculate a correction factor depending on the length distribution (LD) of nucleic acids in the sample and at least one parameter of said measurement; iii. A module configured to calculate the concentration of the detected sequence in unfragmented nucleic acid by the correction factor; A system comprising:

13. 13. A system configured to determine the concentration of a detected sequence in unfragmented nucleic acids as described in claim 12, wherein the module configured to measure the concentration of the detected sequence in the fragmented nucleic acids is an isothermal quantitative nucleic acid amplification module selected from a loop-mediated isothermal amplification module and a nucleic acid sequence-based quantitative amplification module.

14. 13. A system configured to determine the concentration of a detected sequence in unfragmented nucleic acids according to claim 12, wherein the module configured to measure the concentration of the detected sequence in the fragmented nucleic acids is a non-isothermal quantitative nucleic acid amplification module selected from a quantitative polymerase chain reaction module, a real-time polymerase chain reaction module, a digital polymerase chain reaction module, a multiplex polymerase chain reaction module, and a multiplex digital polymerase chain reaction module.

15. 13. A system configured to determine the concentration of a sequence to be detected in unfragmented nucleic acid according to claim 12, wherein the parameters of the measurement include the length of the sequence to be amplified.