Method and apparatus for correcting data set for target analyte in sample
The method corrects real-time PCR data sets by identifying local maximum points and setting baselines to address errors in multiple amplification curve analysis, enhancing detection accuracy for both single and multiple amplification reactions.
Patent Information
- Application Number
- PCT/KR2025/009082
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-28
- Filing Date
- 2025-06-27
- Publication Date
- 2026-01-02
AI Technical Summary
Existing real-time PCR methods face challenges in accurately analyzing multiple amplification curves due to errors in determining target presence, particularly when using algorithms designed for single amplification curves, leading to incorrect analysis of samples with multiple amplification patterns.
A method for correcting a data set by differentiating the amplification function to identify local maximum points, setting a baseline, and subtracting background signals, along with techniques to smooth and standardize data sets, enabling accurate detection of target analytes in both single and multiple amplification reactions.
The method enhances the accuracy of real-time PCR analysis by effectively distinguishing between single and multiple amplification curves, reducing errors and improving the reliability of target analyte detection.
Smart Images

Figure KR2025009082_02012026_PF_FP_ABST
Abstract
Description
Method and device for correcting a data set for target analytes in a sample
[0001] The present invention relates to a method and apparatus for correcting a data set for a target analyte in a sample.
[0002] Polymerase chain reaction (PCR), the most widely used method for nucleic acid amplification, involves repeated cycles of denaturation of double-stranded DNA, annealing of oligonucleotide primers to a DNA template, and extension of the primers by DNA polymerase (Mullis et al., U.S. Pat. Nos. 4,683,195, 4,683,202, and 4,800,159; Saiki et al., Science 230:1350-1354 (1985)).
[0003] Real-time PCR (Real-time PCR) is a PCR-based technology for detecting target nucleic acid molecules in a sample in real time. To detect a specific target analyte, real-time PCR utilizes a signal-generating means that generates a detectable fluorescent signal proportional to the amount of the target molecule. A data set containing each measurement point and the signal value at the measurement point is obtained. The intensity of the fluorescent signal is proportional to the amount of the target molecule. An amplification curve or amplification profile curve, which displays the intensity of the fluorescent signal versus the measurement point, is obtained from the data set.
[0004] Typically, the amplification curve for real-time PCR is divided into a baseline region, an exponential region, and a plateau region. The exponential region is the region where the emitted fluorescent signal increases in proportion to the increase in PCR amplification product, and the plateau region is the region where the increase in PCR amplification product and the emission of fluorescent signals reach saturation, and no further increase in fluorescent signals is observed. The baseline region refers to the region where there is little change in fluorescent signal during the initial cycle of PCR. Since the baseline region does not have enough fluorescent signal emitted by the PCR reaction product to be detected, the background signal, which is the fluorescent signal of the reaction sample itself and the fluorescent signal of the measurement system itself, can account for most of the fluorescent signal in the baseline region.
[0005] Meanwhile, amplification curves are classified into single amplification curves, where amplification occurs once, and multiple amplification curves, where amplification occurs multiple times, depending on the sample type. In the case of multiple amplification curves, errors can occur when analyzing the curves and determining target presence using algorithms designed for single amplification curve analysis.
[0006] The problem to be solved in one embodiment includes, but is not limited to, providing a technique for correcting a data set for a target analyte in a sample.
[0007] A method for correcting a data set for a target analyte in a sample according to a first embodiment comprises the steps of: obtaining a data set for the target analyte; the data set is obtained from a signal generation response for the target analyte using a signal generation means, the data set includes a plurality of data points including a cycle number and a signal value, generating an amplification function including the data set; differentiating the amplification function to generate a derivative; the derivative includes at least one peak and a local maximum point located on the peak, and selecting a first local maximum point having a smallest cycle number among the at least one local maximum point; determining an amplification recognition cycle using the first local maximum point; setting a baseline fitted to a data set, wherein the amplification recognition cycle is located in a cycle less than the first local maximum point, and corresponding to a cycle less than or equal to the amplification recognition cycle; and correcting the data set by subtracting the baseline from the data set.
[0008] In addition, the step of selecting the first maximum point may include the steps of: differentiating the derivative to generate a second derivative; the second derivative includes at least one pair of positive second derivative intervals and negative second derivative intervals corresponding to peaks of the derivative; and selecting a valid peak among peaks of the derivative using the positive second derivative interval or the negative second derivative interval; and selecting a maximum point with the smallest cycle number among the maximum points located at the valid peak as the first maximum point.
[0009] In addition, the step of selecting the valid peak may select a peak among the peaks of the derivative, the number of cycles corresponding to the width of the corresponding positive second derivative section or the negative second derivative section, as a valid peak, which is equal to or greater than a predetermined standard.
[0010] In addition, the step of selecting the valid peak may select a peak among the peaks of the derivative, in which the absolute value of the function value of the corresponding positive second derivative section or negative second derivative section is equal to or greater than a predetermined standard, as the valid peak.
[0011] In addition, the step of determining the amplification recognition cycle may determine a cycle corresponding to a predetermined function value of the derivative as the amplification recognition cycle.
[0012] In addition, the step of determining the amplification recognition cycle may determine a cycle that is subtracted by a predetermined number of subtraction cycles from the first maximum point as the amplification recognition cycle.
[0013] Additionally, the step of setting the baseline may include a step of setting a baseline fitted to the data set corresponding to a predetermined baseline start cycle to the amplification recognition cycle.
[0014] Additionally, the baseline may be a quadratic function.
[0015] Additionally, the axis of symmetry of the quadratic function may be located in a cycle greater than or equal to the last cycle of the data set.
[0016] Additionally, the axis of symmetry of the quadratic function may be located in a cycle less than or equal to the last cycle of the data set.
[0017] Additionally, the baseline may be a linear function.
[0018] Additionally, the baseline may be a constant function.
[0019] Additionally, the constant of the above constant function may be the average value of the signal value of the data set corresponding to the cycle less than or equal to the amplification recognition cycle.
[0020] Additionally, the constant of the above constant function may be a signal value corresponding to the amplification recognition cycle.
[0021] Additionally, the method may further include a step of determining the presence or absence of a target analyte in the sample using the function value of the derivative.
[0022] Additionally, the step of determining the presence or absence may include a step of determining the presence or absence of the target analyte in the sample using the maximum function value of the derivative.
[0023] Additionally, the above maximum function value may be the above local maximum point.
[0024] In addition, the step of determining the presence or absence may include a step of determining the presence or absence of the target analyte in the sample by comparing the maximum function value with a predetermined threshold value for presence or absence.
[0025] Additionally, the method may further include a step of determining the presence or absence of a target analyte in the sample by using the characteristics of the corrected data set.
[0026] Additionally, the step of determining the presence or absence may include a step of determining the presence or absence of the target analyte in the sample using the maximum slope of the corrected data set.
[0027] In addition, the step of generating the amplification function may further include the steps of obtaining an n-th order change value of a signal value in each cycle from the data set (n is an integer greater than or equal to 2); calculating a product of the n-th order change values in two consecutive cycles; and counting the number of cycles in which the product of the n-th order change values is less than a first threshold value to identify a jump error and correct the jump error or generate a smoothed data set.
[0028] Additionally, the above nth change value may be a second change value.
[0029] Additionally, the target analyte may be a target nucleic acid molecule.
[0030] Additionally, the signal generation reaction may be a process of amplifying a signal value with or without amplification of the target nucleic acid molecule.
[0031] Additionally, the signal generating means can generate a signal dependent on the formation of a dimer.
[0032] Additionally, the step of correcting the data set may include a step of generating a standardization coefficient using (i) a signal value in a reference cycle or (ii) a modified value of the signal value in the reference cycle; and a step of applying the standardization coefficient to the signal values of the data set to generate the corrected data set.
[0033] In addition, the signal-generating reaction includes multiple signal-generating reactions for the same target analyte in different reaction environments, the data set is a plurality of data sets, the plurality of data sets are obtained from a plurality of sets of the plurality of signal-generating reactions, and the reference cycle or the reference cycle and the reference value can be equally applied to the plurality of data sets.
[0034] Additionally, the multiple signal-generating reactions may be performed in different reaction environments, including different devices, different reaction tubes or wells, different samples, different amounts of target analyte, or different primers or probes.
[0035] Additionally, the corrected data set may be generated by (i) applying the standardization coefficient to the data set to generate a standardized data set, or (ii) baselining the data set or the standardized data set.
[0036] Additionally, the method may further include a step of determining the presence or absence of a target analyte in the sample by comparing a cycle number of the data set corresponding to a predetermined threshold signal value with a predetermined threshold cycle number.
[0037] A computer program stored in a computer-readable recording medium according to the second embodiment may be programmed to perform each step included in the above-described method.
[0038] In a computer-readable recording medium storing a computer program according to the third embodiment, the computer program may be programmed to perform each step included in the above-described method.
[0039] A computer device according to a fourth embodiment comprises a memory storing at least one command; and a processor. Wherein, when the at least one command is executed by the processor, a data set for the target analyte is obtained; the data set is obtained from a signal generation response for the target analyte using a signal generation means, the data set includes a plurality of data points including a cycle number and a signal value, and an amplification function including the data set is generated; a derivative is generated by differentiating the amplification function; the derivative includes at least one peak and a local maximum point located on the peak, and a first local maximum point having a smallest cycle number is selected among the at least one local maximum point; an amplification recognition cycle is determined using the first local maximum point; the amplification recognition cycle is located in a cycle less than the first local maximum point, and a baseline fitted to a data set corresponding to a cycle less than or equal to the amplification recognition cycle is set; and the data set is corrected by subtracting the baseline from the data set.
[0040] In addition, by executing the at least one instruction by the processor, the derivative is differentiated to generate a second derivative, the second derivative including at least one pair of positive second derivative intervals and negative second derivative intervals corresponding to peaks of the derivative, and by using the positive second derivative interval or the negative second derivative interval, a valid peak among peaks of the derivative can be selected, and a local maximum having the smallest cycle number among local maximums located at the valid peak can be selected as a first local maximum.
[0041] In addition, by executing at least one command by the processor, a peak of the derivative, the number of cycles of which is equal to or greater than a predetermined standard corresponding to the width of the corresponding positive second derivative section or negative second derivative section, can be selected as a valid peak.
[0042] In addition, by executing at least one command by the processor, a peak of the derivative, in which the absolute value of the function value of the corresponding positive second derivative section or negative second derivative section is equal to or greater than a predetermined standard, can be selected as a valid peak.
[0043] In addition, by executing at least one command by the processor, a cycle corresponding to a predetermined function value of the derivative can be determined as the amplification recognition cycle.
[0044] In addition, by executing at least one command by the processor, a cycle deducted by a predetermined number of deduction cycles from the first maximum point can be determined as the amplification recognition cycle.
[0045] Additionally, the step of setting a baseline fitted to the data set corresponding to the amplification recognition cycle from a predetermined baseline start cycle may be included by executing at least one command by the processor.
[0046] Additionally, the baseline may be a quadratic function.
[0047] Additionally, the axis of symmetry of the quadratic function may be located in a cycle greater than or equal to the last cycle of the data set.
[0048] Additionally, the axis of symmetry of the quadratic function may be located in a cycle less than or equal to the last cycle of the data set.
[0049] Additionally, the baseline may be a linear function.
[0050] Additionally, the baseline may be a constant function.
[0051] Additionally, the constant of the above constant function may be the average value of the signal value of the data set corresponding to the cycle less than or equal to the amplification recognition cycle.
[0052] Additionally, the constant of the above constant function may be a signal value corresponding to the amplification recognition cycle.
[0053] In addition, by executing at least one command by the processor, the presence or absence of a target analyte in the sample can be determined using the function value of the derivative.
[0054] In addition, by executing at least one command by the processor, the presence or absence of a target analyte in the sample can be determined using the maximum function value of the derivative.
[0055] Additionally, the above maximum function value may be the above local maximum point.
[0056] In addition, by executing at least one command by the processor, the presence or absence of a target analyte in the sample can be determined by comparing the maximum function value with a predetermined presence / absence threshold value.
[0057] Additionally, by executing at least one command by the processor, the presence or absence of a target analyte in the sample can be determined using the characteristics of the corrected data set.
[0058] Additionally, by executing at least one command of the above by the processor, the presence or absence of a target analyte in the sample can be determined using the maximum slope of the corrected data set.
[0059] In addition, by executing at least one instruction of the above by the processor, an n-th order change value of a signal value in each cycle is obtained from the data set (n is an integer greater than or equal to 2), a product of the n-th order change values in two consecutive cycles is calculated, and a number of cycles in which the product of the n-th order change values is less than a first threshold value is counted to identify a jump error and correct the jump error or generate a smoothed data set, thereby generating the amplification function.
[0060] Additionally, the above nth change value may be a second change value.
[0061] Additionally, the target analyte may be a target nucleic acid molecule.
[0062] Additionally, the signal generation reaction may be a process of amplifying a signal value with or without amplification of the target nucleic acid molecule.
[0063] Additionally, the signal generating means can generate a signal dependent on the formation of a dimer.
[0064] In addition, the data set can be corrected by generating a standardization coefficient using (i) a signal value in a reference cycle or (ii) a modified value of the signal value in the reference cycle by executing at least one instruction of the above by the processor, and applying the standardization coefficient to the signal values of the data set to generate the corrected data set.
[0065] In addition, the signal-generating reaction includes multiple signal-generating reactions for the same target analyte in different reaction environments, the data set is a plurality of data sets, the plurality of data sets are obtained from a plurality of sets of the plurality of signal-generating reactions, and the reference cycle or the reference cycle and the reference value can be equally applied to the plurality of data sets.
[0066] Additionally, the multiple signal-generating reactions may be performed in different reaction environments, including different devices, different reaction tubes or wells, different samples, different amounts of target analyte, or different primers or probes.
[0067] Additionally, the corrected data set may be generated by (i) applying the standardization coefficient to the data set to generate a standardized data set, or (ii) baselining the data set or the standardized data set.
[0068] In addition, by executing at least one command by the processor, the presence or absence of a target analyte in the sample can be determined by comparing a cycle number of the data set corresponding to a predetermined threshold signal value with a predetermined threshold cycle number.
[0069] The method of correcting a data set for a target analyte in a sample as described above provides an analysis method for an amplification reaction, which is effectively applied to not only a single amplification reaction but also multiple amplification reactions.
[0070] In addition, an algorithm for baselining is provided to detect a background signal in an amplification reaction through an embodiment, which is an improvement over the baselining algorithm in a conventional single amplification reaction, and effectively performs baselining in a multiple amplification reaction.
[0071] FIG. 1 illustrates an example of a detection device and a data set correction device connected thereto according to one embodiment.
[0072] FIG. 2 is an exemplary block diagram of a data set correction device according to one embodiment.
[0073] Figure 3 is a data analysis process for a single amplified signal according to one embodiment.
[0074] Figure 4 illustrates errors that occur when a single amplified signal analysis method is applied to a multi-amplified signal.
[0075] Figures 5a and 5b illustrate a multi-amplification signal analysis method according to one embodiment.
[0076] Figures 6a and 6b illustrate a multi-amplification signal analysis method according to one embodiment.
[0077] FIG. 7 illustrates an example of improved signal analysis by an analysis method for a multi-amplified signal according to one embodiment.
[0078] FIG. 8 is an exemplary flowchart illustrating a method for correcting a data set for a target analyte in a sample according to one embodiment.
[0079] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined solely by the scope of the claims.
[0080] When describing embodiments of the present invention, detailed descriptions of known functions or configurations will be omitted if they are deemed to unnecessarily obscure the gist of the invention. Furthermore, the terms described below are defined in light of their functions in the embodiments of the present invention and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the overall content of this specification.
[0081] Before explaining Figure 1, let us look at the terms used herein.
[0082] The term “target analyte” as used herein encompasses a variety of substances (e.g., biological substances such as compounds and non-biological substances), specifically biological substances, and more specifically nucleic acid molecules (e.g., DNA and RNA), proteins, peptides, carbohydrates, lipids, amino acids, biological compounds, hormones, antibodies, antigens, and metabolites. Most specifically, the target analyte is a target nucleic acid molecule.
[0083] As used herein, the terms "target nucleic acid molecule," "target nucleic acid," "target nucleic acid sequence," or "target sequence" refer to a nucleic acid sequence to be analyzed, detected, or quantified. The target nucleic acid sequence may be single- or double-stranded. The target nucleic acid sequence includes not only the sequence initially present in the sample but also the sequence generated during the reaction.
[0084] The target nucleic acid sequence includes DNA (gDNA and cDNA), RNA molecules, and hybrids thereof (chimeric nucleic acids). The sequence may be double-stranded or single-stranded. If the starting nucleic acid sequence is double-stranded, the two strands can be denatured into single-stranded or partially single-stranded molecules. The denaturation process can be performed using conventional techniques, including, but not limited to, heat, alkali, formamide, urea, and glycoxal treatment, enzymatic methods (e.g., helicase action), and binding proteins. For example, denaturation can be achieved by heat treatment at a temperature of 80-105°C. General methods for such treatments are disclosed in Joseph Sambrook, et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2001).
[0085] The target nucleic acid sequence includes any naturally occurring prokaryotic nucleic acid, eukaryotic nucleic acid (e.g., protozoa and parasites, fungi, yeast, higher plants, lower animals, and higher animals including mammals and humans), viral nucleic acid (e.g., herpes virus, HIV, influenza virus, Epstein-Barr virus, hepatitis virus, poliovirus, etc.), or viroid nucleic acid. The target nucleic acid sequence includes a nucleic acid molecule that is or can be produced by recombinant methods or chemical synthesis. Accordingly, the target nucleic acid sequence may or may not exist in nature. The target nucleic acid sequence includes a known or unknown sequence.
[0086] The term “sample” as used herein includes biological samples (e.g., cells, tissues, and body fluids) and non-biological samples (e.g., food, water, and soil), and the biological samples include, for example, viruses, bacteria, tissues, cells, blood (including whole blood, plasma, and serum), lymph, bone marrow fluid, saliva, sputum, swabs, aspirations, milk, urine, stool, eye fluid, semen, brain extracts, spinal fluid, synovial fluid, thymic fluid, bronchial lavage fluid, ascites, and amniotic fluid. When the target analyte is a target nucleic acid molecule, the sample may undergo a nucleic acid extraction process (see: Sambrook, J. et al., Molecular Cloning. A Laboratory Manual, 3rd ed. Cold Spring Harbor Press (2001)). The nucleic acid extraction process may vary depending on the type of sample. Additionally, if the extracted nucleic acid is RNA, it may additionally undergo a reverse transcription process to synthesize cDNA (Reference: Sambrook, J. et al., Molecular Cloning. A Laboratory Manual, 3rd ed. Cold Spring Harbor Press (2001).
[0087] As used herein, the term "signal-generating reaction" means a reaction that generates a signal depending on the properties of a target analyte in a sample. The properties are, for example, the activity, amount, or presence (or absence) of the target analyte, and specifically, the presence or absence of the target analyte in the sample. The signal-generating reaction of the present invention includes a biological reaction and a chemical reaction. The biological reaction includes a genetic analysis process such as PCR, real-time PCR, microarray, and invader assays, an immunological analysis process, and a bacterial growth assay. The chemical reaction includes a chemical analysis process that includes the production, change, or destruction of a chemical substance. According to one embodiment of the present invention, the signal-generating reaction is a genetic analysis process.
[0088] According to one embodiment of the present invention, the signal-generating reaction is a nucleic acid amplification reaction, an enzymatic reaction, or microbial growth.
[0089] A signal-generating response is accompanied by a signal change. As used herein, the term "signal" means a measurable output.
[0090] The progress of a signal-generating reaction is assessed by measuring the signal. The signal value or change in signal serves as an indicator, either qualitatively or quantitatively indicating the characteristics of the target analyte, specifically the presence or absence of the target analyte. This signal change includes both an increase and a decrease in signal.
[0091] The term "signal-generating means" in this specification means a material used to generate a signal indicating a characteristic of a target analyte to be analyzed, specifically the presence or absence.
[0092] Various signal-generating means are known. Signal-generating means include labels. The labels include fluorescent labels, luminescent labels, chemiluminescent labels, electrochemical labels, and metal labels. A single label or an interactive dual label comprising a donor molecule and an acceptor molecule may be used in a form conjugated to one or more oligonucleotides. Alternatively, an intercalating dye itself may also serve as a signal-generating means. The signal-generating means may additionally include an enzyme having a nucleic acid cleavage activity (e.g., an enzyme having a 5' nucleic acid cleavage activity or an enzyme having a 3' nucleic acid cleavage activity) and an oligonucleotide (e.g., a primer and a probe) to generate a signal.
[0093] Examples of oligonucleotides that serve as signal-generating means include oligonucleotides that specifically hybridize to a target nucleic acid sequence (e.g., probes and primers); when a probe or primer hybridized to a target nucleic acid sequence is cleaved to release a fragment, the oligonucleotide that serves as a signal-generating means includes a capture oligonucleotide that specifically hybridizes to said fragment; when the fragment hybridized to said capture oligonucleotide is extended to produce an extended strand, the oligonucleotide that serves as a signal-generating means includes an oligonucleotide that specifically hybridizes to said extended strand; the oligonucleotide that serves as a signal-generating means includes an oligonucleotide that specifically hybridizes to said capture oligonucleotide; and the oligonucleotide that serves as a signal-generating means includes combinations thereof.
[0094] According to one embodiment of the present invention, the signal-generating means comprises a fluorescent label. More specifically, the signal-generating means comprises a fluorescent single label or an interactive dual label comprising a donor molecule and an acceptor molecule (e.g., an interactive dual label comprising a fluorescent reporter molecule and a quencher molecule).
[0095] Various methods are known for generating a signal indicating the presence of a target analyte, particularly a target nucleic acid molecule, using the above signal-generating means. These methods include: TaqMan probe method (US Pat. No. 5,210,015), molecular beacon method (Tyagi et al., Nature Biotechnology, 14 (3): 303 (1996)), Scorpion method (Whitcombe et al., Nature Biotechnology 17:804-807 (1999)), Sunrise (or Amplifluor) method (Nazarenko et al., Nucleic Acids Research, 25(12):2516-2521 (1997), and US Pat. No. 6,117,635), Lux method (US Pat. No. 7,537,886), CPT (Duck P, et al. Biotechniques, 9:142-148 (1990)), LNA method (US Pat. No. 6,977,295), Plexor method (Sherrill CB, et al., Journal of the American Chemical Society, 126:4550-4556 (2004)), Hybeacons method (DJ French, et al., Molecular and Cellular Probes (2001) 13, 363-374 and U.S. Pat. No. 7,348,141), Dual-labeled, self-quenched probe method (U.S. Pat. No. 5,876,930), Hybridization probe method (Bernard PS, et al., Clin Chem 2000, 46, 147-148), PTOCE (PTO cleavage and extension) method (WO 2012 / 096523), PCE-SH (PTO Cleavage and Extension-Dependent Signaling Oligonucleotide Hybridization) method (WO 2013 / 115442), PCE-NH (PTO Cleavage and Extension-Dependent Non-Hybridization) method (PCT / KR2013 / 012312) and CER method (WO 2011 / 037306).
[0096] According to one embodiment of the present invention, the signal-generating reaction is a reaction that amplifies a signal value with or without amplification of the target nucleic acid molecule.
[0097] As used herein, the term “amplification reaction” means a reaction that increases or decreases a signal.
[0098] According to one embodiment of the present invention, the amplification reaction refers to a signal increase (or amplification) reaction generated by the signal-generating means depending on the presence of the target analyte. This amplification reaction may or may not be accompanied by amplification of the target analyte (e.g., a nucleic acid molecule). More specifically, in the present invention, the amplification reaction refers to a signal amplification reaction accompanied by amplification of the target analyte.
[0099] According to one embodiment of the present invention, the amplification reaction for amplifying a signal indicating the presence of a target analyte (e.g., a target nucleic acid molecule) can be performed in a manner in which the signal is amplified while the target nucleic acid molecule is amplified (e.g., a real-time PCR method). Alternatively, the amplification reaction can be performed in a manner in which only the signal is amplified without amplifying the target analyte [e.g., the CPT method (Duck P, et al., Biotechniques, 9:142-148 (1990)), the Invader assay (U.S. Patent Nos. 6,358,691 and 6,194,149)].
[0100] Target analytes, particularly target nucleic acid molecules, can be amplified in a variety of ways. For example, various methods for amplification of target analytes are known, including polymerase chain reaction (PCR), ligase chain reaction (LCR) (U.S. Pat. Nos. 4,683,195 and 4,683,202; PCR Protocols: A Guide to Methods and Applications (Innis et al., eds, 1990)), strand displacement amplification (SDA) (Walker, et al. Nucleic Acids Res. 20(7):1691-6 (1992); Walker PCR Methods Appl 3(1):1-6 (1993)), and transcription-mediated amplification (Phyffer, et al., J. Clin. Microbiol. 34:834-841 (1996); Vuorinen, et al., J. Clin. Microbiol. 33:1856-1859 (1995)), nucleic acid sequence-based amplification (NASBA) (Compton, Nature 350(6313):91-2 (1991)), rolling circle amplification (RCA) (Lisby, Mol. Biotechnol. 12(1):75-99 (1999); Hatch et al., Genet. Anal. 15(2):35-40 (1999)), and Q-Beta Replicase (Lizardi et al., BiolTechnology 6:1197(1988)).
[0101] According to one embodiment of the present invention, the amplification reaction amplifies the signal while amplifying the target analyte (specifically, the target nucleic acid molecule). According to one embodiment of the present invention, the amplification reaction is performed according to PCR or real-time PCR.
[0102] According to one embodiment of the present invention, the signal-generating means generates a signal dependently on the formation of a dimer. The expression “generating a signal dependently on the formation of a dimer” used in conjunction with the signal-generating means means that the detected signal is provided dependently on the association or dissociation of two nucleic acid molecules. Specifically, the signal is generated by a dimer formed dependently on the cleavage of a mediation oligonucleotide that specifically hybridizes to the target nucleic acid sequence. As used herein, the term “mediation oligonucleotide” is an oligonucleotide that mediates the formation of a dimer that does not include the target nucleic acid sequence. According to one embodiment of the present invention, the cleavage of the mediation oligonucleotide itself does not generate a signal, and the fragment formed by the cleavage participates in successive reactions for signal generation after hybridization and cleavage of the mediation oligonucleotide. According to one embodiment of the present invention, the intermediate oligonucleotide comprises an oligonucleotide that hybridizes to a target nucleic acid sequence and is cleaved to release a fragment, thereby mediating the formation of a duplex. Specifically, the fragment mediates the formation of the duplex by extension of the fragment from a capture oligonucleotide. According to one embodiment of the present invention, the intermediate oligonucleotide comprises (i) a 3'-targeting portion comprising a hybridizing nucleotide sequence complementary to the target nucleic acid sequence and a 5'-tagging portion comprising a nucleotide sequence non-complementary to the target nucleic acid sequence. According to one embodiment of the present invention, cleavage of the intermediate oligonucleotide releases a fragment, which specifically hybridizes to the capture oligonucleotide and extends on the capture oligonucleotide.According to one embodiment of the present invention, an intermediate oligonucleotide hybridized to a target nucleic acid sequence is cleaved to release a fragment, which specifically hybridizes to a capture oligonucleotide, which extends to form an extended strand, and an extended duplex is formed between the extended strand and the capture oligonucleotide, thereby providing a signal indicating the presence of the target nucleic acid sequence. A representative example of a signal-generating means that generates a signal dependent on the formation of a duplex is the PTOCE method (WO 2012 / 096523).
[0103] According to one embodiment of the present invention, the signal-generating means generates a signal dependent on the cleavage of the detection oligonucleotide. Specifically, the signal is generated by hybridization of the detection oligonucleotide to a target nucleic acid sequence and cleavage of the detection oligonucleotide. The signal resulting from hybridization of the detection oligonucleotide to a target nucleic acid sequence and cleavage of the detection oligonucleotide can be generated by various methods, including the TaqMan probe method (U.S. Patent Nos. 5,210,015 and 5,538,848).
[0104] The data set obtained through the amplification reaction includes amplification cycles or cycle numbers.
[0105] The term "cycle" in this specification refers to a unit of change in a condition in multiple measurements involving a change in said condition. The change in said constant condition may refer to an increase or decrease in, for example, temperature, reaction time, number of reactions, concentration, pH, and / or the number of replications of a measurement target (e.g., target nucleic acid molecule). Accordingly, a cycle may be a time or process cycle, a unit operation cycle, or a reproductive cycle.
[0106] For example, when measuring the substrate degradation capacity of an enzyme according to substrate concentration, the degree of substrate degradation is measured several times by varying the substrate concentration, and the enzyme's substrate degradation capacity is analyzed from these measurements. In this case, a constant change in condition is an increase in substrate concentration, and the unit of substrate concentration increase used is set as one cycle.
[0107] As another example, in the case of isothermal amplification of nucleic acids, one sample can be measured several times with different reaction times. In this case, the reaction time is a change in conditions, and the unit of reaction time is set as one cycle.
[0108] More specifically, the term "cycle" means one unit of repetition, when a reaction is repeated in a certain process or at certain time intervals.
[0109] For example, in the case of polymerase chain reaction (PCR), one cycle refers to a reaction unit that includes a denaturation step of a target nucleic acid molecule, an annealing (hybridization) step between the target nucleic acid molecule and a primer, and an extension step of the primer. In this case, a change in a certain condition is an increase in the number of reaction repetitions, and a repeating unit of a reaction that includes the above series of steps corresponds to one cycle.
[0110] A data set obtained by a signal-generating reaction comprises multiple data points including cycle numbers and signal values.
[0111] As used herein, the term “signal value” or “signal value” means a value of a signal (e.g., signal intensity) actually measured in a cycle of a signal-generating reaction or a modified value thereof. The modified value may include a mathematically processed value of the measured signal value. Examples of mathematically processed values of the measured signal value may include a logarithm or derivative of the measured signal value. The derivative of the measured signal value may include multiple derivatives.
[0112] In this specification, the term "data point" means a single coordinate value that includes a cycle number and a signal value. The term "data" means all information that constitutes a data set. For example, each cycle and signal value of an amplification reaction is data.
[0113] Data points obtained by a signal-generating reaction, particularly an amplification reaction, can be represented as coordinate values that can be represented in a two-dimensional rectangular coordinate system. In the coordinate values, the X-axis represents the corresponding cycle number of the amplification reaction, and the Y-axis represents the signal value measured or processed in the corresponding cycle.
[0114] The term "data set" as used herein refers to a collection of data points. For example, the data set may include a raw data set, which is a collection of data points obtained directly from a signal-generating reaction (e.g., an amplification reaction) using a signal-generating means. Alternatively, the data set may be a modified data set obtained by modifying a data set including a collection of data points obtained directly from the signal-generating reaction. The data set may include some or all of the data points obtained from the signal-generating reaction or modified data points thereof.
[0115] The data set can be plotted and an amplification curve can be obtained from the data set.
[0116] According to one embodiment of the present invention, the data set is corrected using the total signal change value. As used herein, the term “total signal change value” refers to the amount of signal change (increased or decreased) in the data set. When determining the total signal change value in one section of the data set, the difference between the signal values in the first and last cycles of the data set or the difference between the maximum and minimum signal values in one section of the data set can be determined as the total signal change value.
[0117] As used herein, the term "displacement" refers to the degree of signal change in a data set, specifically, the degree of signal change from a reference signal value. Here, the reference signal value may be the smallest signal value, the signal value of the first cycle, the average signal value of the baseline region, or the signal value at the start of the amplification region. In calculating the displacement, the final signal value may be the largest signal value, the signal value of the last cycle, or the average signal value of the plateau region.
[0118] The term “determined based on the cycles of the baseline region” used in this specification when referring to a fitting interval means selecting the entire cycle or a portion of the cycles of the baseline region to use as the fitting interval.
[0119] The cycles of the baseline region as the fitting interval can be selected from among all or part of the cycles of the baseline region determined by conventional methods (e.g., U.S. Patent No. 8,560,247 and WO 2016 / 052991) or empirically.
[0120] When the first quadratic function is generated and the cycle (Cmin) when it is at its minimum is greater than or equal to a threshold value or the coefficient of x2 is greater than 0, a second quadratic function is generated in the fitting interval. The fitting interval can be determined based on a cycle in the baseline region of the data set. Alternatively, the fitting interval is determined based on the cycle (Cmin) when the first quadratic function is at its minimum. The term “determined based on the cycle (Cmin) when the first quadratic function is at its minimum” used when referring to the fitting interval means that the end cycle of the fitting interval is determined based on Cmin or a cycle around Cmin (specifically, a cycle within Cmin±5 cycles). The fitting interval is determined from the above-described start fitting cycle to the end cycle, and the end cycle can be determined as a cycle determined based on the cycle (Cmin) when the first quadratic function is at its minimum.
[0121] Second-order quadratic functions include various functions, such as linear functions, polynomial functions (e.g., quadratic and cubic functions), and exponential functions. Linear or nonlinear regression analysis methods can be used to determine the second-order quadratic function that best matches a data set.
[0122] According to one embodiment of the present invention, an exponential function may be used instead of the second quadratic function.
[0123] An exponential function fitted to signal values in the baseline region of the data set is generated, and a corrected data set is generated by subtracting the exponential function from the data set.
[0124] As used herein, the term "amplification recognition cycle" refers to the point in time at which the amplification reaction of a target analyte in a sample can be signal-recognized. Typically, in polymerase chain reaction (PCR), the amplification reaction begins from the initial cycle, but the signal is often buried in noise and difficult to clearly distinguish. Therefore, in the present invention, based on the parameters preceding or at the maximum point of the derivative, the cycle at which the amplification curve deviates from the baseline and begins to exhibit a significant increasing pattern is analytically recognized, and this is defined as the "amplification recognition cycle."
[0125] This is not the starting point of the reaction in a biological or chemical sense, but rather the first cycle where the amplification reaction forms a statistically significant upward curve from a data analysis perspective, and can be determined by various mathematical criteria depending on the algorithm (e.g., passing the derivative threshold, a certain distance before a local maximum, the shape of the second derivative, etc.).
[0126] Therefore, the “amplification recognition cycle” is not simply the point at which amplification begins (start cycle), but is used as a reference point by which the algorithm interpreting the amplification pattern recognizes the start of the amplification reaction.
[0127] With reference to the drawings below, various implementation examples of the present invention will be examined.
[0128] FIG. 1 illustrates a detection device (200) and a data set correction device (100) connected thereto according to an embodiment. These devices (100, 200) may be connected to each other via wired or wireless communication. However, the block diagram illustrated in FIG. 1 is merely exemplary, and the spirit of the present invention is not limited to what is illustrated in FIG. 1. For example, components not illustrated in FIG. 1 may be additionally connected to these devices (100, 200). Alternatively, unlike what is illustrated, the data set correction device (100) may be implemented by being included in the detection device (200). However, the following description will be made on the assumption that the aforementioned components (100, 200) are implemented or connected as illustrated in FIG. 1. Hereinafter, each component will be examined in detail.
[0129] Referring to FIG. 1, the analysis system includes a detection device (200) and a data set correction device (100). The analysis system can determine the presence or absence of a target analyte in a sample and display the results to a user.
[0130] The detection device (200) is a device that performs a nucleic acid amplification reaction, an enzymatic reaction, or microbial growth. For example, when the detection device (200) performs a nucleic acid amplification reaction, the detection device (200) can repeatedly perform an operation of increasing and decreasing the temperature of the samples. The detection device (200) can obtain a data set by measuring a signal generated from the samples for each cycle.
[0131] The detection device (200) can be connected to the data set correction device (100) via a cable (130) or can be connected wirelessly. The detection device (200) transmits the obtained data set to the data set correction device (100) wired or wirelessly.
[0132] The data set correction device (100) obtains a data set from the detection device (200). The data set correction device (100) analyzes the data set to determine whether a target analyte is present or absent in the sample. In other words, the data set correction device (100) determines whether the sample is positive or negative.
[0133] The data set correction device (100) may include a display unit. The display unit may display a data set or display a plotted data set as a graph. The display unit may display various functions, such as a sigmoid function, a step function, or a linear function. In addition, the display unit may indicate whether a target analyte is present or absent, and may display detection results for each sample.
[0134] The data set correction device (100) can read a data set contained in a recording medium. The recording medium can store a data set or a program used in the data set correction device (100). The recording medium may be a CD or USB, etc.
[0135] The detection device (200) is implemented to perform a nucleic acid amplification reaction and a nucleic acid detection operation on a sample. Depending on the embodiment, the detection device (200) may be implemented to perform a nucleic acid detection operation without performing a nucleic acid amplification reaction. However, the following description will assume that the detection device (200) is implemented to also perform a nucleic acid amplification reaction.
[0136] Meanwhile, as the nucleic acid amplification reaction described above is performed, if the sample in the reaction vessel contains a target analyte, not only will its amount be amplified, but the magnitude of the signal generated by the signal generation means described above may also be amplified as described above. Accordingly, the detection device (200) is implemented to detect the magnitude of this signal. Specifically, the detection device (200) can monitor in real time the magnitude of the signal described above, which changes as the nucleic acid amplification reaction progresses. The monitored result can be output from the detection device (200) in the form of a data set as mentioned in the definition of terms, for example, the intensity of the signal per cycle. Here, the intensity of the signal per cycle is information as a material that serves as a basis for determining the presence or absence of the target analyte in the corresponding sample.
[0137] Meanwhile, since the size of the signal generated by the signal generation means can also be amplified when the above nucleic acid amplification reaction is performed, the nucleic acid amplification reaction will be considered hereinafter as causing the signal generation reaction discussed above as a term.
[0138] Below, each functional configuration of the data set correction device (100) will be examined in more detail with reference to FIG. 2.
[0139] FIG. 2 is an exemplary block diagram of a data set correction device (100) according to one embodiment. Before referring to FIG. 2, it should be noted that the data set correction device (100) can be implemented in a laptop, server, cloud, etc., but is not limited thereto.
[0140] Referring to FIG. 2, the data set correction device (100) includes, but is not limited to, a communication unit (120), a memory (140), and a processor (160).
[0141] First, the communication unit (120) is implemented as a wired or wireless communication module. Through this communication unit (120), the data set correction device (100) can communicate with the outside world. For example, as will be described later, the data set correction device can transmit and receive setting values, threshold values, etc. input by researchers and developers to and from the outside world through the communication unit (120).
[0142] Various types of data are stored in the memory (140). The stored data may include, but are not limited to, the aforementioned setting values, threshold values, etc. For example, the memory (140) may store data (or data sets) received from the detection device (200) or data processed by the processor (160), which may include the aforementioned data sets, noise-removed data sets from which noise has been removed, and amplification curves approximated from the noise-removed data sets.
[0143] Additionally, the memory (140) may store data in units of plates or strips. Additionally, the memory (140) may store data for a plurality of plates and amplification data generated based on the data for a plurality of plates.
[0144] Here, data about the plate includes numbers or letters for identifying the plate, information recorded about the plate, or data about the reaction wells contained in the plate. Among these, information recorded about the plate includes various information such as the date, time, and method of performing the amplification reaction on the plate. In addition, data about the reaction wells includes a data set for each reaction well, a noise-removed data set, an amplification curve, a Ct value, a signal value, etc.
[0145] Meanwhile, in FIG. 2, the memory (140) is depicted as a separate configuration from the processor (160), but the memory (140) may be implemented as a single device with the processor (160). For example, the memory (140) may be a storage such as a cache included within the processor (160).
[0146] Next, the processor (160) may be implemented by a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller unit (MCU), or a dedicated processor on which the methods according to one embodiment are performed. In addition, such a processor (160) may be implemented by at least one core.
[0147] Figure 3 is a data analysis process for a single amplified signal according to one embodiment.
[0148] Referring to FIG. 3, the data analysis process for a single amplified signal includes, but is not limited to, jump error confirmation and correction (310, 320), background signal separation (330), amplified signal extraction (340), Ct value calculation (350), and fitting curve provision (360).
[0149] Meanwhile, the applicant has developed an analysis process for an amplification reaction (WO2019 / 066572). According to this, the fitting accuracy of a nonlinear function for a data set is used as a direct indicator for target analyte analysis, and a data set for a target analyte is obtained from a signal-generating reaction using a signal-generating means to analyze the target analyte without false positive and false negative results, particularly false positive results, and the obtained data set including a plurality of data points including cycle numbers and signal values is corrected, and a nonlinear function is generated for the corrected data set, and the fitting accuracy of the nonlinear function for the corrected data set is determined, and the presence or absence of the target analyte in the sample can be determined using the fitting accuracy.
[0150] Referring to the jump error confirmation process (310) of FIG. 3, the original signal (311) and the jump error (312) can be confirmed. Here, the original signal refers to a signal before processing, such as correction, is applied to the signal measured by the aforementioned signal generation means. In addition, the jump error refers to an abnormal change in the signal, and includes, for example, a sudden increase / decrease in the signal, a spike, or a step-shaped signal. Such abnormal signals can be caused by various reasons, such as changes in the annealing temperature, bubble formation in the tube, or the presence of foreign substances in the sample. In qualitative and quantitative analysis of data, abnormal signals often lead to misinterpretation, such as false positive or false negative results, which can reduce the accuracy and reliability of the analysis. Not only do abnormal signals have various causes, but even if the exact cause is identified, it is not easy to eliminate the cause in advance. Therefore, before determining the presence or absence of a target nucleic acid from a data set, the data set must be analyzed to determine whether an abnormal signal has occurred, and if an abnormal signal is identified, the abnormal signal must be corrected.
[0151] Referring to the jump error correction process (320) of FIG. 3, a correction signal (321) in which the jump error portion is smoothed can be confirmed. Regarding jump error correction, various methods have been reported, including those described in U.S. Patent Publication No. 2015 / 0186598, so further explanation is omitted.
[0152] Referring to the background signal separation process (330) and the amplified signal extraction process (340) of Fig. 3, it is possible to confirm a configuration for separating the background signal (332) and the amplified signal (341) from the signal (331) transmitted in the previous process.
[0153] In one embodiment, a given signal can be baselined to remove background signals. A baseline-subtracted data set can be obtained by various methods known in the art (see, e.g., U.S. Patent No. 8,560,247 and WO 2016 / 052991). The baseline region refers to a region where the fluorescence signal remains constant with little change during the initial cycles of a PCR reaction. Because the level of PCR product in this region is not sufficient to be detected, most of the fluorescence signal in this region is due to background signals, including the fluorescence signal of the reaction sample itself and the fluorescence signal of the measurement system itself. A baseline-subtracted data set can also be obtained using a baselined method developed by the applicant of the present invention (see, e.g., WO 2019 / 066572).
[0154] Referring to the Ct value calculation process (350) of FIG. 3, a configuration can be confirmed that provides a Ct (Cycle threshold), which is a cycle (352) corresponding to a predetermined threshold value. This Ct can be used to determine the presence or absence of a target analyte according to a predetermined standard.
[0155] Referring to the fitting curve provision process (360) of FIG. 3, a configuration can be confirmed that provides a fitted curve, such as a nonlinear function, corresponding to an amplified signal. This curve can be used to determine the presence of a target analyte using other criteria, such as parameters related to the nonlinear function, rather than the aforementioned Ct.
[0156] Meanwhile, the amplification signal analysis method described with reference to FIG. 3 may result in errors in multi-amplification signals, not single-amplification signals. Here, "single" and "multiple" are determined by the number of exponential increases in the signal value within the entire signal, approximately 45 cycles. In other words, if the signal increases exponentially once, it is classified as single-amplification, and if the signal increases exponentially twice or more, it can be classified as multiple-amplification.
[0157] Below, errors that occur when applying a conventional amplified signal analysis method to a multi-amplified signal are described with reference to Fig. 4.
[0158] Figure 4 illustrates errors that occur when a single amplified signal analysis method is applied to a multi-amplified signal.
[0159] First, referring to the first graph (410) of Fig. 4, an example of a false negative processed due to a linear pattern by multiple amplification can be confirmed. Here, a linear pattern means that when analyzing the single amplification signal described above, when the form of the amplification shows a linear increase pattern rather than an exponential function-like increase pattern, the signal increase is considered to be due to external factors such as a background signal rather than a signal increase due to the amplification of the target analyte. Therefore, when the degree of signal amplification shows a slope of a linear pattern similar to a linear function rather than a slope like an exponential function, the presence or absence of the target analyte can be processed as a negative regardless of the displacement of the signal.
[0160] However, applying the aforementioned linear pattern algorithm to a multi-amplified signal can result in errors. Let's assume that the multi-amplifications are evenly distributed throughout the entire cycle. For example, if the entire cycle is 45 cycles, the number of amplifications is 3, and the amplification points are distributed at cycles 5, 20, and 35, then considering that a single amplification is observed in a section of more than 10 cycles, the amplification curve will exhibit a pattern of continuous increase throughout the entire cycle. In this case, if the single-amplification analysis method that negatively processes the aforementioned linear pattern is followed, false negatives may occur.
[0161] Next, referring to the second graph (420) of Fig. 4, we can see an example where the Ct value is delayed due to a baseline algorithm error. Looking at the second graph (420) of Fig. 4, we can see a nonlinear function of a solid line fitted to the dotted raw data (421). As a result of obtaining Ct from the fitted nonlinear function, Ct is 27.10326. This Ct value is obtained by fitting to the fitted nonlinear function, and in this embodiment, the fitted nonlinear function is a sigmoid function. That is, assuming a single amplification, we fitted to a sigmoid function and baselined the bending part before 20 cycles. However, if the raw data is multi-amplified, the bending before 20 cycles may be a movement caused by actual amplification. Considering that the Ct value must come from the cycle value prior to the first amplification, it can be seen that the actual Ct value of the second signal (420) in FIG. 4 must be located 20 cycles prior to the illustrated 27.10326.
[0162] Next, referring to the third graph (430) of Fig. 4, we can see an example where the Ct value is brought forward due to an error in the fitting function algorithm. Looking at the third graph (430) of Fig. 4, we can see a solid nonlinear function fitted to the dotted raw data (431). As a result of obtaining Ct from the fitted nonlinear function, Ct is 14.03034. This Ct value is obtained by fitting to the fitted nonlinear function, and in this example, the fitted nonlinear function is a sigmoid function. However, if we check the raw data (431), the first amplification starts at about 20 cycles, and then the slope increase that appears to be the second amplification occurs once more. In other words, the data corresponding to multiple amplification was fitted to a sigmoid function for single amplification analysis. From the first to the last amplification of the raw data, the amplification interval is affected by the change in the slope of the sigmoid function, resulting in a longer amplification interval (unlike typical amplification graphs, where amplification typically completes in about 10 to 15 cycles). As the amplification interval lengthens, the slope of the sigmoid function becomes more gradual, leading to an earlier Ct value.
[0163] With reference to Fig. 4, we examined errors that occur when analyzing a multi-amplified signal using a single-amplified signal analysis method, and an algorithm that can improve this is described below.
[0164] Figures 5a and 5b illustrate a multi-amplification signal analysis method according to one embodiment.
[0165] Referring to the upper graph (510) of Fig. 5(a), an amplified signal (511) formed by multiple amplification can be confirmed. Referring to the lower graph (520) of Fig. 5(a), a derivative (521) obtained by differentiating the amplified signal (511) of the upper graph and a second derivative (522) obtained by differentiating the derivative (521) can be confirmed.
[0166] According to one embodiment, a method for calibrating a data set for a target analyte in a sample may set a predetermined threshold value for a function value of a derivative obtained by differentiating an amplification signal. The horizontal line (523) in the lower graph of Fig. 5(a) represents the threshold value of the predetermined derivative. This threshold value may intersect the derivative, and a cycle (524) corresponding to this intersection may be set as the starting point of amplification of the amplification signal. According to one embodiment, a method for calibrating a data set for a target analyte in a sample may perform baselineing up to the starting point of amplification.
[0167] Meanwhile, the derivative may contain multiple peaks, which represent changes in the slope of the amplified signal and, consequently, the beginning and end of amplification. Therefore, multiple peaks indicate the presence of multiple amplifications, indicating that amplification began in the cycle preceding the first peak.
[0168] The presence or absence of a target can also be determined using the maximum function value of the derivative. Since this is similarly used in single-amplification signal analysis methods, a detailed description will be omitted. The maximum function value of the derivative can be the largest peak value among multiple peaks. Here, the peak value can be a local maximum.
[0169] Referring to the upper graph (530) of Fig. 5(b), it can be confirmed that the baseline (532) was obtained by performing baseline lining up to the starting point (533) of the amplification set in Fig. 5(a).
[0170] Referring to the lower graph (540) of Fig. 5(b), it can be confirmed that the signal (541) for the amount of pure amplification obtained by subtracting the background signal from the amplified signal (531) by subtracting the baseline (532) obtained in Fig. 5(b) from the amplified signal (531) is obtained. Meanwhile, although the background signal (332) and the amplified signal (341) were previously described with reference to Fig. 3, the term “amplified signal” is used throughout this specification regardless of whether the background signal is included. In parts where the inclusion of the background signal is important in the description of the technology, the subtraction of the background signal is described once more with emphasis.
[0171] Referring to the lower graph (540) of Fig. 5(b), it can be confirmed that the amplified signal (541) from which the baseline has been subtracted has a similar shape to the amplified signal (542) before subtraction. This is because the shape of the baseline is a straight line. According to one embodiment, the baseline may be a straight line or a curve.
[0172] With reference to FIG. 5, a schematic process of a multi-amplification signal analysis method according to an embodiment was examined.
[0173] A method for correcting a data set for a target analyte in a sample according to one embodiment comprises the steps of: obtaining a data set for the target analyte; the data set is obtained from a signal generation response for the target analyte using a signal generation means, the data set includes a plurality of data points including a cycle number and a signal value, and generating an amplification function including the data set; differentiating the amplification function to generate a derivative; the derivative includes at least one peak and a local maximum point located on the peak, and selecting a first local maximum point having a smallest cycle number among the at least one local maximum point; determining an amplification recognition cycle using the first local maximum point; setting a baseline fitted to a data set, wherein the amplification recognition cycle is located in a cycle less than the first local maximum point, and corresponding to a cycle less than or equal to the amplification recognition cycle; and correcting the data set by subtracting the baseline from the data set.
[0174] In addition, the step of selecting the first maximum point may include the steps of: differentiating the derivative to generate a second derivative; the second derivative includes at least one pair of positive second derivative intervals and negative second derivative intervals corresponding to peaks of the derivative; and selecting a valid peak among peaks of the derivative using the positive second derivative interval or the negative second derivative interval; and selecting a maximum point with the smallest cycle number among the maximum points located at the valid peak as the first maximum point.
[0175] In addition, the step of selecting the valid peak may select a peak among the peaks of the derivative, the number of cycles corresponding to the width of the corresponding positive second derivative section or the negative second derivative section, as a valid peak, which is equal to or greater than a predetermined standard.
[0176] In addition, the step of selecting the valid peak may select a peak among the peaks of the derivative, in which the absolute value of the function value of the corresponding positive second derivative section or negative second derivative section is equal to or greater than a predetermined standard, as the valid peak.
[0177] In addition, the step of determining the amplification recognition cycle may determine a cycle corresponding to a predetermined function value of the derivative as the amplification recognition cycle.
[0178] In addition, the step of determining the amplification recognition cycle may determine a cycle that is subtracted by a predetermined number of subtraction cycles from the first maximum point as the amplification recognition cycle.
[0179] Additionally, the step of setting the baseline may include a step of setting a baseline fitted to the data set corresponding to a predetermined baseline start cycle to the amplification recognition cycle.
[0180] Additionally, the baseline may be a quadratic function.
[0181] Additionally, the axis of symmetry of the quadratic function may be located in a cycle greater than or equal to the last cycle of the data set.
[0182] Additionally, the axis of symmetry of the quadratic function may be located in a cycle less than or equal to the last cycle of the data set.
[0183] Additionally, the baseline may be a linear function.
[0184] Additionally, the baseline may be a constant function.
[0185] Additionally, the constant of the above constant function may be the average value of the signal value of the data set corresponding to the cycle less than or equal to the amplification recognition cycle.
[0186] Additionally, the constant of the above constant function may be a signal value corresponding to the amplification recognition cycle.
[0187] Additionally, the method may further include a step of determining the presence or absence of a target analyte in the sample using the function value of the derivative.
[0188] Additionally, the step of determining the presence or absence may include a step of determining the presence or absence of the target analyte in the sample using the maximum function value of the derivative.
[0189] Additionally, the above maximum function value may be the above local maximum point.
[0190] In addition, the step of determining the presence or absence may include a step of determining the presence or absence of the target analyte in the sample by comparing the maximum function value with a predetermined threshold value for presence or absence.
[0191] Additionally, the method may further include a step of determining the presence or absence of a target analyte in the sample by using the characteristics of the corrected data set.
[0192] Additionally, the step of determining the presence or absence may include a step of determining the presence or absence of the target analyte in the sample using the maximum slope of the corrected data set.
[0193] In addition, the step of generating the amplification function may further include the steps of obtaining an n-th order change value of a signal value in each cycle from the data set (n is an integer greater than or equal to 2); calculating a product of the n-th order change values in two consecutive cycles; and counting the number of cycles in which the product of the n-th order change values is less than a first threshold value to identify a jump error and correct the jump error or generate a smoothed data set.
[0194] Additionally, the above nth change value may be a second change value.
[0195] Additionally, the target analyte may be a target nucleic acid molecule.
[0196] Additionally, the signal generation reaction may be a process of amplifying a signal value with or without amplification of the target nucleic acid molecule.
[0197] Additionally, the signal generating means can generate a signal dependent on the formation of a dimer.
[0198] Additionally, the step of correcting the data set may include a step of generating a standardization coefficient using (i) a signal value in a reference cycle or (ii) a modified value of the signal value in the reference cycle; and a step of applying the standardization coefficient to the signal values of the data set to generate the corrected data set.
[0199] In addition, the signal-generating reaction includes multiple signal-generating reactions for the same target analyte in different reaction environments, the data set is a plurality of data sets, the plurality of data sets are obtained from a plurality of sets of the plurality of signal-generating reactions, and the reference cycle or the reference cycle and the reference value can be equally applied to the plurality of data sets.
[0200] Additionally, the multiple signal-generating reactions may be performed in different reaction environments, including different devices, different reaction tubes or wells, different samples, different amounts of target analyte, or different primers or probes.
[0201] Additionally, the corrected data set may be generated by (i) applying the standardization coefficient to the data set to generate a standardized data set, or (ii) baselining the data set or the standardized data set.
[0202] Additionally, the method may further include a step of determining the presence or absence of a target analyte in the sample by comparing a cycle number of the data set corresponding to a predetermined threshold signal value with a predetermined threshold cycle number.
[0203] Hereinafter, the multi-amplification signal analysis method will be described in more detail with reference to FIGS. 6a and 6b.
[0204] FIG. 6a illustrates a multi-amplification signal analysis method according to one embodiment.
[0205] The upper graph (610) of Fig. 6a shows a multi-amplified signal (611) according to an embodiment. The curvature before 10 cycles is noise that occurred between measurements, and it is a multi-amplified signal that starts the first amplification around 20 cycles and then undergoes one more amplification.
[0206] In the lower graph (620) of Fig. 6a, the derivative (621) and the second derivative (622) of the multi-amplified signal (611) can be confirmed. As described above with reference to Fig. 5, by setting a threshold (623) for the derivative, the cycle (624) where the derivative (621) and the threshold (623) intersect can be obtained. The cycle (624) where the threshold intersects can be set as the start of amplification, and baseline can be performed in the cycles below that.
[0207] In this way, the cycle (624) that intersects the threshold is used for baseline, and since this cycle indicates the start of amplification, it must be located before the peak of the derivative, which can be considered an indicator of amplification. In other words, a condition for the cycle that intersects the threshold may be that it is located in the cycle before the first peak of the derivative. To this end, it is necessary to determine the peak of the derivative, and the method of determining the peak of the derivative using the second derivative (622) will be explained.
[0208] The peak of the derivative is used to identify the start cycle of amplification, and as mentioned above, the maximum peak value can also be used to determine the presence or absence of a target. However, the original function, i.e., the amplification function before differentiation, may have inflections that are not due to amplification due to errors in the measuring device. If there are multiple inflections due to causes other than amplification, these inflections will appear as peaks in the derivative through the first differentiation, which can affect the determination of the start cycle of amplification. In this case, the validity of the derivative peak can be determined using the second derivative. In other words, the second derivative can be used to select a valid peak. The peak of the derivative resulting from inflections due to causes other than amplification may have a narrower or smaller peak width or size than the peak caused by amplification. Using this, the width, size, area, and sign change of the second derivative peak can be used to determine the validity of the corresponding derivative peak. Here, the validity of a derivative peak can be determined by determining whether the peak is a result of amplification. Therefore, by setting thresholds in advance for the width, size, area, and sign change of the aforementioned second-order derivative peak, these can be used as a reference to determine the validity of the corresponding derivative peak.
[0209] Referring to the lower graph (620) of Fig. 6ab, a configuration for judging the validity of a derivative can be confirmed using the width (625, 626) of the negative peak of the second derivative. At this time, an upper limit and a lower limit can be set for the width of the negative peak of the second derivative. If the width of the peak of the second derivative is below or above the standard, it can be interpreted that the signal was not derived from a signal due to target amplification.
[0210] Referring to the upper graph (630) of Fig. 6b, it is possible to confirm a configuration in which baseline lining is performed up to the start cycle of amplification (632) using the start cycle of amplification (632) obtained in the previous process.
[0211] Referring to the lower graph (640) of Fig. 6b, a signal (641) obtained by subtracting the baseline from the amplified signal (631) using the baseline (633) obtained in the previous process can be confirmed. In addition, it can be confirmed that the Ct value obtained using the preset threshold value is 19.52.
[0212] The above describes an analysis method for multiple amplified signals according to one embodiment. This method for analyzing multiple amplified signals according to one embodiment can be considered a method for correcting a data set for target analytes within a sample, and can prevent errors arising from conventional methods of analyzing data by fitting it to a nonlinear function.
[0213] Below, an example is described in which an error in the conventional method described with reference to FIG. 4 is improved when analyzed using an analysis method for a multi-amplified signal according to an embodiment.
[0214] FIG. 7 illustrates an example of improved signal analysis by an analysis method for a multi-amplified signal according to one embodiment.
[0215] Referring to Figure 7, three pairs of graphs can be confirmed. Each pair has a configuration that matches the three graphs (410, 420, 430) of Figure 4.
[0216] First, if we check the pair of graphs at the top, it is an example of a false negative with a linear pattern, which was explained with reference to the graph (410) of FIG. 4. In the analysis method for a multiple amplification signal according to an embodiment, there is no need for an algorithm for negative processing by a linear pattern. Since the linear pattern, which determines a signal increase pattern that occurs in a section exceeding 10 to 15 cycles, which is a typical PCR amplification section, as a signal error, is a signal pattern that can sufficiently appear in multiple amplification, such a pattern can be interpreted by the analysis method for a multiple amplification signal according to an embodiment. The Ct value of the upper graph interpreted by the analysis method for a multiple amplification signal according to an embodiment is confirmed to be 11.02699.
[0217] Next, if we check the graph pair of interruptions, it is an example of Ct being delayed due to a baseline algorithm error, as explained with reference to the graph (420) of Fig. 4. In the analysis method for a multi-amplification signal according to an embodiment, there is no need to perform the “sigmoid function fitting considering single amplification” that was the cause of the graph (420), but rather, the start cycle of amplification is found through peak analysis of the derivative, and the presence or absence of the function is determined. Through this, the Ct value that was delayed can also be obtained by reflecting the actual amplification signal (Ct value 12.62827).
[0218] Finally, if we look at the graph pair below, it is an example where Ct is advanced due to single-amplification nonlinear function fitting, as explained with reference to graph (430) of Fig. 4. This is also due to the same cause as the second graph (420), and in the analysis method for multiple amplification signals according to one embodiment, since the amplification data shape is interpreted as it is rather than fitting to a nonlinear function with a single shape in batches, the Ct value error can be resolved.
[0219] FIG. 8 illustrates a flowchart illustrating a method for correcting a data set for a target analyte within a sample according to one embodiment. This flowchart is merely exemplary, and the scope of the present invention is not limited thereto. For example, depending on the embodiment, each step may be performed in a different order than that illustrated in FIG. 8, or at least one step not illustrated in FIG. 8 may be additionally performed, or at least one of the steps illustrated in FIG. 8 may not be performed.
[0220] Referring to FIG. 8, a step (S100) of obtaining a data set for the target analyte is performed. At this time, the data set is obtained from a signal generation response for the target analyte using a signal generation means, and the data set includes a plurality of data points including a cycle number and a signal value.
[0221] Additionally, a step (S200) of generating an amplification function including the above data set is performed.
[0222] In addition, a step (S300) of differentiating the amplification function to generate a derivative is performed. At this time, the derivative includes at least one peak and a local maximum point located on the peak.
[0223] In addition, a step (S400) of selecting the maximum point with the smallest cycle number among the above at least one maximum point as the first maximum point is performed.
[0224] In addition, a step (S500) of determining an amplification recognition cycle using the first maximum point is performed. At this time, the amplification recognition cycle is located in a cycle less than the first maximum point.
[0225] Additionally, a step (S600) of setting a baseline fitted to a data set corresponding to a cycle less than or equal to the above amplification recognition cycle is performed.
[0226] Additionally, a step (S700) of correcting the data set by subtracting the baseline from the data set is performed.
[0227] Meanwhile, the method according to the various embodiments described above can be implemented in the form of a computer program stored in a computer-readable recording medium programmed to perform each step of the method, and can also be implemented in the form of a computer-readable recording medium storing a computer program programmed to perform each step of the method.
[0228] The above description is merely an illustrative illustration of the technical idea of the present invention, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential quality of the present invention. Therefore, the embodiments disclosed in the present invention are intended to illustrate, rather than limit, the technical idea of the present invention, and the scope of the technical idea of the present invention is not limited by these embodiments. The scope of protection of the present invention should be interpreted by the following claims, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of the rights of the present invention.
Claims
1. Method for calibrating a data set for target analytes in a sample, including the following steps: A step of obtaining a data set for the target analyte; the data set is obtained from a signal generation response for the target analyte using a signal generation means, and the data set includes a plurality of data points including a cycle number and a signal value. generating an amplification function including the above data set; A step of generating a derivative by differentiating the above amplification function; the derivative includes at least one peak and a local maximum point located on the peak, A step of selecting a maximum point having the smallest cycle number among the above at least one maximum point as a first maximum point; A step of determining an amplification recognition cycle using the first maximum point; the amplification recognition cycle is located in a cycle less than the first maximum point, A step of setting a baseline fitted to a data set less than or equal to the above amplification recognition cycle; and A step of correcting the data set by subtracting the baseline from the data set.
2. In paragraph 1, The step of selecting the first maximum point is: A step of generating a second derivative by differentiating the above derivative; the second derivative includes at least one pair of positive second derivative intervals and negative second derivative intervals corresponding to peaks of the derivative, A step of selecting at least one valid peak among the peaks of the derivative by using the positive second derivative interval or the negative second derivative interval; and A method comprising the step of selecting a local maximum having the smallest cycle number among local maximums located at at least one valid peak as a first local maximum.
3. In paragraph 2, The step of selecting at least one valid peak is a method of selecting a peak among the peaks of the derivative, the number of cycles corresponding to the width of the positive second derivative interval or the negative second derivative interval being greater than a predetermined standard, as a valid peak.
4. In paragraph 2, The step of selecting at least one valid peak is a method of selecting a peak of the derivative whose absolute value of the function value in the positive second derivative interval or the negative second derivative interval is greater than a predetermined standard as a valid peak.
5. In paragraph 1, The step of determining the above amplification recognition cycle is a method of determining a cycle in which the function value of the above derivative matches a predetermined reference value as the amplification recognition cycle.
6. In paragraph 1, The step of determining the amplification recognition cycle is a method of determining a cycle that is a predetermined number of cycles prior to the first maximum point as the amplification recognition cycle.
7. In paragraph 1, The step of setting the above baseline is a method of setting a baseline fitted to a data set from a predetermined baseline start cycle to the amplification recognition cycle.
8. In paragraph 7, The above baseline is a quadratic function.
9. In paragraph 8, A method in which the axis of symmetry of the above quadratic function is located in a cycle greater than or equal to the last cycle of the above data set.
10. In paragraph 8, A method in which the axis of symmetry of the above quadratic function is located in a cycle less than or equal to the last cycle of the above data set.
11. In paragraph 7, The above baseline is a linear function.
12. In paragraph 7, The above baseline is a constant function.
13. In paragraph 12, A method in which the constant of the above constant function is the average value of the signal value of the data set in the section less than the above amplification recognition cycle.
14. In paragraph 12, A method in which the constant of the above constant function is a signal value of the above amplification recognition cycle.
15. In paragraph 1, A method further comprising a step of determining the presence or absence of a target analyte in the sample using the function value of the derivative.
16. In paragraph 15, A method wherein the step of determining the presence or absence of the target analyte in the sample comprises a step of determining the presence or absence of the target analyte in the sample using the maximum function value of the derivative.
17. In paragraph 15, The above maximum function value is the above local maximum point.
18. In paragraph 16, A method in which the step of determining the presence or absence of the target analyte in the sample comprises a step of determining the presence or absence of the target analyte in the sample by comparing the maximum function value with a predetermined threshold value for presence or absence.
19. A computer program stored on a computer-readable recording medium that is executed to include each step included in any one of paragraphs 1 to 18.
20. A computer-readable recording medium storing a computer program that is performed to include each step included in any one of paragraphs 1 to 18.
21. Memory for storing at least one instruction; and Includes a processor, By executing at least one instruction by the processor, Obtaining a data set for the target analyte; wherein the data set is obtained from a signal generation response for the target analyte using a signal generation means, and wherein the data set includes a plurality of data points including cycle numbers and signal values, Generate an amplification function including the above data set; Differentiating the above amplification function to generate a derivative; the derivative includes at least one peak and a local maximum point located on the peak, Selecting the first maximum point having the smallest cycle number among the above at least one maximum point; Determine the amplification recognition cycle using the first maximum point; and the amplification recognition cycle is located in a cycle less than the first maximum point, Setting a baseline fitted to a data set below the above amplification recognition cycle; and Correcting the data set by subtracting the baseline from the data set Computer devices.
Citation Information
Patent Citations
Method for calibrating a data set for target analytes
KR102270771B1
Methods for analyzing samples
KR102336732B1
Real-time PCR elbow calling by equation-less algorithm
US20100070185A1
Multi-stage, regression-based PCR analysis system
US20140032126A1
Universal method to determine real-time PCR cycle threshold values
US20190095576A1
Cited By
PCR melting curve analysis method and system based on multiple verification
CN121811967A