A method and system for PCR melt curve analysis based on multiplex validation
By employing multiple validation mechanisms and adaptive baseline alignment technology, the problems of noise and multi-channel crosstalk in PCR melting curve analysis were solved, achieving high-precision Tm value measurement and reducing the false positive rate.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TAIZHOU FOOD & DRUG INSPECTION INSTITUTE
- Filing Date
- 2026-03-06
- Publication Date
- 2026-06-19
AI Technical Summary
Existing PCR melting curve analysis methods are susceptible to noise, lack multi-dimensional signal authenticity assessment, struggle to identify non-specific signals in complex samples, and have a high false positive rate.
A multi-verification mechanism is adopted, which verifies multiple dimensions such as peak height, peak-to-valley difference, half-peak width, and full-peak width. Combined with adaptive baseline alignment and point-to-point signal strength comparison, linear interpolation method is used to improve the measurement accuracy of Tm value and eliminate multi-channel crosstalk interference.
It significantly reduces the false positive rate, improves multi-peak resolution, enhances data quality and comparability, and ensures the accuracy and reliability of Tm value measurement.
Smart Images

Figure CN121811967B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of PCR detection technology, and in particular to a method and system for PCR melting curve analysis based on multiple validation. Background Technology
[0002] After the PCR amplification reaction is complete, the negative derivative of fluorescence intensity, i.e., the melting curve, is often obtained by gradually increasing the temperature to degrade the amplification products. This is used to examine the specificity of the amplification products or for SNP genotyping. During the heating process, when the temperature reaches the point where half of the melting occurs, a large amount of fluorescent dye is released, and the fluorescence intensity decreases rapidly, thus forming a peak point on the melting curve. The temperature corresponding to this peak point is T. m Value (melting temperature), while T m The number and position of the items are the key points of the examination.
[0003] Currently, the main techniques used in melting curve analysis are as follows:
[0004] 1) Cluster analysis-based methods:
[0005] The target gene category is determined through dataset correction, normalization, target peak calculation, and cluster analysis. Key technical measures include setting positive and negative thresholds, locating regions before and after melting, removing background curves, normalization, calculating the negative first derivative, peak search, and cluster analysis. This method is suitable for software analysis of multiple melting curves and can eliminate the influence of inter-instrument differences.
[0006] 2) Overlapping peak separation method based on higher-order derivatives:
[0007] The Savitzky-Golay derivative method is used to calculate the second derivative curve, and overlapping peaks are separated through superposition fitting. Key techniques include smoothing and denoising, Savitzky-Golay second derivative calculation, initial overlapping peak position determination, nonlinear least squares algorithm superposition fitting, and multi-parameter function fitting. This method is suitable for analyzing complex melting curves with significant background interference and overlapping peaks.
[0008] 3) High-resolution melting curve analysis method:
[0009] Higher-order derivatives are used to improve resolution, and cluster analysis is combined to determine T. m The main technical measures include wavelet denoising, higher-order derivative (second or fourth order) calculation, extreme point detection, peak-valley screening, and adaptive clustering analysis. This method can achieve high-resolution melting curve analysis with a detection resolution of 0.1℃-1℃.
[0010] While the aforementioned technical solutions represent innovations in their respective fields, existing methods often rely on a single threshold for judgment, making them susceptible to noise and lacking a mechanism for assessing signal authenticity from a global perspective. They also exhibit insufficient ability to identify non-specific signals. Traditional methods primarily depend on single indicators such as peak height or peak-to-valley difference, lacking comprehensive judgment based on multiple dimensions such as peak width and peak shape, and have limited ability to identify complex samples (such as continuous peaks and weak shoulder peaks).
[0011] Therefore, in view of the above-mentioned prior art, the present invention provides a PCR melting curve analysis method and system based on multiple validation. Summary of the Invention
[0012] The purpose of this invention is to address the shortcomings of existing technologies by providing a PCR melting curve analysis method and system based on multiple validation. Through a multiple peak validation mechanism, considering multiple dimensions such as peak height, peak-to-valley difference, half-peak width, and full-peak width, it significantly reduces the false positive rate. Adaptive baseline alignment and point-to-point signal intensity comparison are employed to accurately identify and eliminate crosstalk interference between multiple channels. Based on a global maximum peak determination mechanism, the authenticity of the signal is judged from a holistic perspective, effectively eliminating non-specific signals. Data quality and comparability are improved through data replacement at the initial heating stage and integration time normalization. A linear interpolation method is used to overcome sampling interval limitations, achieving higher accuracy in Tm analysis. m Value measurement.
[0013] To achieve the above objectives, the present invention adopts the following technical solution:
[0014] A PCR melting curve analysis method based on multiple validation includes:
[0015] Step S1. Obtain data information for each well in each channel of the PCR reaction plate, preprocess the obtained data information, and generate the initial melting curve for each well in each channel; the data information includes raw fluorescence data, temperature data, and integration time data;
[0016] Step S2. The initial melting curve is smoothed using a bidirectional low-pass filter, and the first derivative of the smoothed melting curve is calculated to generate a negative derivative curve.
[0017] Step S3. Calculate the second derivative of the obtained negative derivative curve, and identify the initial candidate peak set by detecting the zero crossings of the second derivative;
[0018] Step S4. Perform a quadruple validation on each candidate peak in the initial candidate peak set based on peak height threshold, peak-to-valley difference threshold, half-width at half-maximum, and full-width at half-maximum. Calculate the melting temperature T for candidate peaks that pass the quadruple validation using linear interpolation. m The value is used to obtain the first valid T. m Value set;
[0019] Step S5. In the first valid T m Identify continuous peak groups from the value set, and filter the peak values within the continuous peak groups based on an adaptive relative peak height ratio threshold to obtain the second effective T. m Value set;
[0020] Step S6. Place the second valid T m The peak value in the value set is compared with the global maximum peak value. If there is no T corresponding to the global maximum peak value... m If the value is less than a certain threshold, it is considered a false positive sample, and all T samples are removed. m Value; if there exists a T corresponding to the global maximum peak value. m If the value is positive, it is determined to be a true positive sample, and step S7 is executed;
[0021] Step S7. Normalize and display the melting curve and negative derivative curve of the samples determined to be true positive samples.
[0022] Furthermore, the preprocessing of the acquired data information in step S1 specifically includes: invalid data processing and integration time normalization processing of the original fluorescence data.
[0023] Furthermore, the smoothing of the initial melting curve in step S2 is achieved by a bidirectional exponentially weighted moving average filtering method.
[0024] Furthermore, step S4 specifically includes:
[0025] Step S41. Determine whether the negative derivative value corresponding to the peak point is greater than the preset peak height threshold. If not, exclude the peaks that do not meet the conditions; if yes, retain the peaks that meet the conditions and proceed to step S42.
[0026] Step S42. Calculate the difference between the first negative derivative and the difference between the second negative derivative of the peak point and the valley point to its left and right respectively. Determine whether the difference between the first negative derivative and the difference between the second negative derivative are both greater than the preset peak-valley difference threshold. If not, exclude the peak points that do not meet the conditions; if so, retain the peak points that meet the conditions and proceed to step S43.
[0027] Step S43. Calculate the absolute value of the first difference and the absolute value of the second difference between the peak point and the temperature corresponding to the valley point to its left and the valley point to its right, respectively. Determine whether the absolute value of the first difference and the absolute value of the second difference are both greater than the preset half-peak width threshold. If not, exclude the peak points that do not meet the conditions; if so, retain the peak points that meet the conditions and proceed to step S44.
[0028] Step S44. Calculate the temperature difference between the valley point to the right of the peak point and the valley point to the left of the peak point, and determine whether the temperature difference is greater than the preset full peak width threshold. If not, exclude the peak that does not meet the condition; if yes, retain the peak that meets the condition and proceed to step S45.
[0029] Step S45. Calculate the melting temperature T using linear interpolation on the retained peak values. m The value is used to obtain the first valid T. m Value set.
[0030] Furthermore, in step S45, the melting temperature T is calculated using linear interpolation. m The values specifically include:
[0031] Obtain the second derivative values at the peak position and the next subsequent position. Calculate the interpolation ratio based on the second derivative value at the peak position and the obtained second derivative values. Then, based on the interpolation ratio, the temperature difference between two adjacent temperature sampling points, and the temperature at the peak position, obtain the first effective T. m Value set.
[0032] Furthermore, step S5 specifically includes:
[0033] Step S51. Traverse the first valid T m The set of values is used to determine whether the negative derivative of each peak is positive and whether the smooth curve value is greater than the smooth curve value corresponding to the next peak. If so, the current peak and the next peak are grouped into a continuous peak group, and the process continues until the condition is no longer met. The number of continuous peaks is then counted.
[0034] Step S52. Traverse all peaks within a continuous peak group and select the peak with the largest negative derivative value;
[0035] Step S53. Calculate the ratio of the negative derivative value of each peak in the continuous peak group to the negative derivative value of the largest peak, as the relative peak height ratio, and determine whether there are peaks with a relative peak height ratio greater than the preset peak height ratio threshold. If so, re-execute the four-fold judgment verification in step S4 for the peaks that meet the conditions to obtain the second valid T. m Value set.
[0036] Furthermore, step S6 specifically includes:
[0037] Step S61. Identify the peak with the largest negative derivative value in the candidate peak set from step S3 as the global maximum peak;
[0038] Step S62. Traverse the second valid T m If any valid peak in the value set has the same index as the global maximum peak, the sample is considered a true positive sample; otherwise, it is considered a false positive sample, and all T values of the sample are removed.m value.
[0039] Furthermore, the procedure before step S7 includes:
[0040] When the detection channel in a true positive sample meets the crosstalk triggering condition, the adaptive baseline alignment algorithm is used to eliminate the false positive samples corresponding to T caused by inter-channel crosstalk. m value.
[0041] Furthermore, the T corresponding to the elimination of false positive samples caused by inter-channel crosstalk through the adaptive baseline alignment algorithm... m The specific value is:
[0042] Obtain the negative derivative curves of the detection channel and the reference channel, and scale the negative derivative curves of the detection channel and the reference channel proportionally.
[0043] The minimum difference between the negative derivative curve of the scaled detection channel and the negative derivative curve of the reference channel is calculated and used as the baseline offset to perform baseline alignment on the negative derivative curve of the reference channel.
[0044] Compare the T values corresponding to the negative derivative curves of the detection channel and the reference channel. m At the specified value location, the signal intensity of the negative derivative curve of the detection channel is compared with the signal intensity of the negative derivative curve of the reference channel after baseline alignment. Based on the comparison result, it is determined whether to clear the T value of the detection channel. m value.
[0045] Correspondingly, a PCR melting curve analysis system based on multiple validation is also provided, including:
[0046] The acquisition module is used to acquire data information from each well of each channel of the PCR reaction plate, preprocess the acquired data information, and generate the initial melting curve for each well of each channel; the data information includes raw fluorescence data, temperature data, and integration time data;
[0047] The first calculation module is used to smooth the initial melting curve using a bidirectional low-pass filtering method, and to calculate the first derivative of the smoothed melting curve to generate a negative derivative curve.
[0048] The second calculation module is used to calculate the second derivative of the obtained negative derivative curve and identify the initial candidate peak set by detecting the zero crossing point of the second derivative.
[0049] The verification module performs a four-fold verification on each candidate peak in the initial candidate peak set based on peak height threshold, peak-to-valley difference threshold, half-width, and full-width. For candidate peaks that pass the four-fold verification, the melting temperature T is calculated using linear interpolation. m The value is used to obtain the first valid T.m Value set;
[0050] The filtering module is used for the first valid T m Identify continuous peak groups from the value set, and filter the peak values within the continuous peak groups based on an adaptive relative peak height ratio threshold to obtain the second effective T. m Value set;
[0051] The judgment module is used to determine the second valid T. m The peak value in the value set is compared with the global maximum peak value. If there is no T corresponding to the global maximum peak value... m If the value is less than a certain threshold, it is considered a false positive sample, and all T samples are removed. m Value; if there exists a T corresponding to the global maximum peak value. m If the value is positive, it is determined to be a true positive sample;
[0052] The display module is used to normalize and display the melting curve and negative derivative curve of samples determined to be true positive samples.
[0053] Compared with the prior art, the present invention has the following main advantages:
[0054] 1. A multi-level peak quality control mechanism effectively reduces the false positive rate:
[0055] Traditional single threshold methods are sensitive to baseline drift (when the baseline rises slowly, the derivative curve will form a wide and low "peak", which cannot be effectively distinguished by the peak height threshold alone), are extremely fragile to data (instantaneous fluctuations or drops during data acquisition will form sharp spurious peaks on the derivative curve), and lack biological constraints (the physical characteristics of the DNA melting process, such as melting temperature range and peak shape symmetry, are not considered). This invention verifies the authenticity of peak values from multiple dimensions through a four-layer progressive peak screening mechanism: The first layer is peak height threshold verification (≥0.016), which uses three times the standard deviation of noise levels from multiple negative control samples in the laboratory as the discrimination threshold to filter out noise peaks below the threshold; the second layer is peak-valley difference verification (≥0.0169), which sets the peak-valley difference distribution of known positive samples to five times the standard deviation of noise fluctuation to ensure that the peak value is significantly higher than the surrounding baseline; the third layer is half-peak width verification (left and right half-peak width difference ≥0.99℃), which verifies the symmetry of the peak based on the principle that the difference between the left and right half-peak widths of a normal melting peak in DNA melting kinetics should be less than 1℃, excluding asymmetric sharp peaks caused by data mutations; the fourth layer is peak width verification (≥4.5℃), which, based on the thermodynamic theory of DNA melting and laboratory measured data (the melting peak width range of typical PCR products is 4.5-8℃), ensures that the peak width conforms to the physical process characteristics of DNA melting.
[0056] 2. Improve the resolution of multiple peaks:
[0057] Traditional first-order derivative methods have limitations in multi-peak identification: using a fixed Savitzky-Golay window (e.g., 21-point) may lead to over-smoothing, causing adjacent peaks to be merged; direct differentiation after a single smoothing can easily miss minor peaks with small amplitudes but real existence; and peak detection is easily interfered with in areas with strong high-frequency noise. This invention improves multi-peak resolution through bidirectional low-pass filtering, dynamic windowing, and second-order derivative zero-crossing detection. The bidirectional low-pass filtering first performs forward filtering (from low temperature to high temperature) and then reverse filtering (from high temperature to low temperature), and the results of the two filterings are fused. The mathematical principle is: forward filtering y[i] = y[i-1] + η(x[i] - y[i-1]) and reverse filtering z[i] = z[i+1] + η(y[i] - z[i+1]), where η is the filtering coefficient. Bidirectional filtering is equivalent to applying a zero-phase filter, which can effectively suppress the phase shift introduced by unidirectional filtering, preserve the true position of the peak, and has a stronger suppression effect on noise while better preserving the true signal. Dynamic windowing technology adaptively adjusts the Savitzky-Golay window size based on data length (an odd-numbered window, typically 11-41 points, representing about 2% of the data points for short data). Maintaining the window size to data point ratio within an appropriate range ensures smoothness without excessive loss of detail. Short windows better preserve peak details and are suitable for identifying closely spaced multi-peaked peaks, while longer windows offer stronger noise suppression. Dynamic adjustment achieves optimal peak resolution under varying data quality conditions. Second derivative zero-crossing detection locates peak positions by finding the extreme points of the first derivative (i.e., the zero-crossing points of the second derivative), more accurately capturing subtle fluctuations compared to direct peak detection methods. Attached Figure Description
[0058] Figure 1 This is a flowchart of a PCR melting curve analysis method based on multiple validation provided in Example 1;
[0059] Figure 2 This is a schematic diagram of the four-fold judgment verification provided in Example 1;
[0060] Figure 3 This is a schematic diagram of the original data provided in Example 1;
[0061] Figure 4 This is a schematic diagram of the original data smoothing process provided in Example 1;
[0062] Figure 5 This is a schematic diagram of the curve after processing the first negative derivative of the melting data provided in Example 1;
[0063] Figure 6 This is a schematic diagram of a pseudo-peak sample caused by abnormal data drop provided in Example 2;
[0064] Figure 7 This is a schematic diagram of the identification of the three-peak sample provided in Example 3. Detailed Implementation
[0065] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.
[0066] The purpose of this invention is to address the shortcomings of existing technologies by providing a method and system for PCR melting curve analysis based on multiple validation.
[0067] Example 1
[0068] This embodiment provides a PCR melting curve analysis method based on multiple validation, such as... Figure 1 As shown, it includes:
[0069] Step S1. Obtain data information for each well in each channel of the PCR reaction plate, preprocess the obtained data information, and generate the initial melting curve for each well in each channel; the data information includes raw fluorescence data, temperature data, and integration time data;
[0070] Step S2. The initial melting curve is smoothed using a bidirectional low-pass filter, and the first derivative of the smoothed melting curve is calculated to generate a negative derivative curve.
[0071] Step S3. Calculate the second derivative of the obtained negative derivative curve, and identify the initial candidate peak set by detecting the zero crossings of the second derivative;
[0072] Step S4. Perform a quadruple validation on each candidate peak in the initial candidate peak set based on peak height threshold, peak-to-valley difference threshold, half-width at half-maximum, and full-width at half-maximum. Calculate the melting temperature T for candidate peaks that pass the quadruple validation using linear interpolation. m The value is used to obtain the first valid T. m Value set;
[0073] Step S5. In the first valid T m Identify continuous peak groups from the value set, and filter the peak values within the continuous peak groups based on an adaptive relative peak height ratio threshold to obtain the second effective T. m Value set;
[0074] Step S6. Place the second valid T mThe peak value in the value set is compared with the global maximum peak value. If there is no T corresponding to the global maximum peak value... m If the value is less than a certain threshold, it is considered a false positive sample, and all T samples are removed. m Value; if there exists a T corresponding to the global maximum peak value. m If the value is positive, it is determined to be a true positive sample, and step S7 is executed;
[0075] Step S7. Normalize and display the melting curve and negative derivative curve of the samples determined to be true positive samples.
[0076] In step S1, data information of each well in each channel of the PCR reaction plate is obtained, and the obtained data information is preprocessed to generate the initial melting curve of each well in each channel; wherein the data information includes raw fluorescence data, temperature data and integration time data.
[0077] like Figure 3 The image shows a schematic diagram of the acquired raw data, where blue represents the first-channel curve, green represents the second-channel curve, yellow represents the third-channel curve, and red represents the fourth-channel curve.
[0078] Read data from a complete melting process from the PCR instrument software interface or data file. The data includes raw fluorescence data, temperature data, and integration time data; raw fluorescence data is stored in three dimensions: channel, well position, and sampling point; temperature data is stored in two dimensions: channel and sampling point; and integration time data is stored in one channel.
[0079] The acquired data is processed for invalid data and the integration time is normalized to generate the initial melting curves for each channel and each orifice.
[0080] In the initial stage of melting and heating, the temperature and fluorescence signals may be unstable. Therefore, it is necessary to calculate the length of invalid data. At this point, 10% of the total data length is taken and rounded to the nearest integer. Then, for each channel, the temperature and fluorescence data within the first few invalid data lengths are replaced with the corresponding data value of the first valid position (i.e., the 5% position), thus eliminating the influence of unstable data in the initial heating stage. For example, if the total length is 1000 points, the first 100 points are replaced with the value of the 50th point, thereby eliminating the influence of unstable data in the initial stage.
[0081] Since different fluorescence channels may employ different integration times to achieve the optimal signal-to-noise ratio, directly comparing fluorescence values is not comparable. Therefore, the raw fluorescence data is normalized based on the integration time of each channel. Specifically, the raw fluorescence data of each sampling point at each well in each channel is divided by the integration time of that channel to ensure the comparability of data from different channels.
[0082] This embodiment is specifically optimized for PCR heating characteristics. At the start of the PCR melting experiment, the temperature control system gradually increases the temperature from the initial temperature, resulting in initial temperature fluctuations and unstable fluorescence signals. Analysis of the experimental data revealed that the first 5% of the data contained significant noise. Replacing this noise with the data from the first stable position effectively eliminated these interferences without affecting the identification of the true melting peak, as the true melting peak typically appears in the higher temperature region.
[0083] Integration time normalization solves a crucial problem in multi-channel detection. Different channels may use different integration times to optimize the signal-to-noise ratio, resulting in significant differences in the original fluorescence intensity values. By normalizing by dividing by the integration time, it ensures that the data from different channels are compared under the same dimensions, laying the foundation for subsequent crosstalk cancellation and other cross-channel analyses.
[0084] This method eliminates the effects of unstable data during the initial heating phase and differences between different channels, thus improving data quality. In repeated experiments with different batches of reagents and different operators, the consistency of results is significantly improved, with the coefficient of variation decreasing from over 10% in traditional methods to below 3%.
[0085] In step S2, a bidirectional low-pass filter is used to smooth the initial melting curve, and the first derivative of the smoothed melting curve is calculated to generate a negative derivative curve.
[0086] To smooth the melting curve, suppress high-frequency noise, and avoid peak position shifts caused by traditional unidirectional filtering, this embodiment employs a bidirectional exponentially weighted moving average (EWMA) filtering method, specifically:
[0087] The smoothing coefficient alpha is calculated as: alpha = 100 / total_len, where total_len represents the total data length. The smaller the alpha value, the higher the smoothness.
[0088] Starting from the first fluorescence data point, the filtering result of the current data point is set as the filtering result of the previous data point plus the smoothing coefficient multiplied by the difference between the original value of the current data point and the filtering result of the previous data point, expressed as:
[0089] smoothed_forward[i]=smoothed_forward[i-1]+alpha*(raw_data[i]-smoothed_forward[i-1])
[0090] Where smoothed_forward[i] represents the forward filtering result of the current data point; smoothed_forward[i-1] represents the filtering result of the previous data point; and raw_data[i] represents the original value of the current data point.
[0091] Following the above method, all fluorescence data points are processed sequentially to complete the forward filtering of all fluorescence data points.
[0092] Starting from the last fluorescence data point, the filtering result of the current point is set as the filtering result of the next point plus the smoothing coefficient multiplied by the difference between the forward filtering value of the current point and the filtering result of the next point, expressed as:
[0093] smoothed_bidirectional[i]=smoothed_bidirectional[i+1]+alpha*(smoothed_forward[i] - smoothed_bidirectional[i+1])
[0094] Where smoothed_bidirectional[i] represents the backward filtering result of the current data point; smoothed_bidirectional[i+1] represents the filtering result of the next data point; and smoothed_forward[i] represents the forward filtering result of the current data point.
[0095] All data points are processed sequentially using the above method to complete the backward filtering of all fluorescence data points.
[0096] After bidirectional filtering, the final smoothed_bidirectional curve is both smooth and retains the accurate position of the original peak, while also eliminating the phase delay caused by unidirectional filtering, maintaining the accuracy of the peak position, and ensuring T. m Value positioning is not distorted.
[0097] like Figure 4 The diagram shows the smoothed data after processing the original data. Blue represents the curve of the first channel, green represents the curve of the second channel, yellow represents the curve of the third channel, and red represents the curve of the fourth channel.
[0098] Compared to traditional unidirectional filtering, bidirectional filtering in this embodiment eliminates the phase delay problem. Unidirectional filtering causes a time shift in the output signal relative to the input signal, resulting in a deviation in the peak position and affecting T. m Value accuracy. This invention combines forward and backward filtering to ensure that the output signal is in phase with the input signal, maintaining the time position of signal characteristics without distortion and ensuring T. m The accuracy of value positioning.
[0099] In this embodiment, the first derivative of the smoothed melting curve is calculated to generate a negative derivative curve, specifically as follows:
[0100] (1) Dynamic window Savitzky-Golay filtering: Calculate the window size window_size, expressed as: window_size = total_len * 0.05 and round up. If window_size is even, add 1 to make it odd to ensure the filter coefficients are symmetrical. Use a first-order polynomial for fitting and directly output the first-order derivative value.
[0101] (2) First derivative calculation process: The output first derivative value is low-pass filtered to further smooth it. The filter coefficient is 0.65 times the smoothing coefficient alpha. Then, it is scaled inversely, which is expressed as: negative_derivative = -1*filtered_derivative, where negative_derivative represents the negative derivative curve; filtered_derivative represents the filtered derivative value; and 1 represents the scaling factor.
[0102] The peak of the negative derivative curve corresponds to the inflection point of the original melting curve, i.e., the location of the melting temperature, which is the subsequent temperature T. m Value recognition provides feature curves, such as Figure 5 The schematic diagram of the curves after processing the first negative derivative of the melting data shows that blue represents the first-channel curve, green represents the second-channel curve, yellow represents the third-channel curve, and red represents the fourth-channel curve.
[0103] In step S3, the second derivative is calculated on the obtained negative derivative curve, and the initial candidate peak set is identified by detecting the zero-crossing points of the second derivative.
[0104] To accurately locate the peak of the negative derivative curve, it is necessary to find the zero-crossing point of its first derivative (i.e., the second derivative of the original melting curve).
[0105] (1) Second derivative calculation: The second derivative curve second_derivative is obtained by using the negative derivative curve negative_derivative of the Savitzky-Golay filter.
[0106] (2) Zero-crossing detection: Traverse the second derivative curve's `second_derivative` array and detect the sign change between adjacent points. When the current point is positive and the next point is negative, mark that position as a maximum (peak); when the current point is negative and the next point is positive, mark that position as a minimum (valley), that is:
[0107] If second_derivative[i] > 0 and second_derivative[i+1] < 0, then mark a maximum point (peak) at index i.
[0108] If second_derivative[i] < 0 and second_derivative[i+1] > 0, then mark a minimum point (valley) at index i.
[0109] (3) Initial Peak Identification: Traverse all marked points, collect the indices of points marked as maxima into a peak list, and collect the indices of points marked as minima into a valley list, preparing a candidate peak set for subsequent peak verification. The points in the peak list are the initially identified candidate peaks (T). m Value position.
[0110] This embodiment adapts to different experimental conditions and sampling densities while maintaining algorithm stability. The fixed window size method exhibits significant performance fluctuations when data length varies, potentially leading to over-smoothing and loss of detail on short data sets, and insufficient smoothing and retention of excessive noise on long data sets. By using a relative window size, the algorithm's performance remains stable with fluctuations of less than 5% over a wide range from 100 to 3056 points.
[0111] In step S4, each candidate peak in the initial candidate peak set is subjected to a four-fold judgment verification based on the peak height threshold, peak-to-valley difference threshold, half-width, and full-width. The melting temperature T is calculated for the candidate peaks that pass the four-fold judgment verification using linear interpolation. m The value is used to obtain the first valid T. m Value set.
[0112] like Figure 2 As shown, specifically:
[0113] Step S41. Determine whether the negative derivative value corresponding to the peak point in the negative derivative curve is greater than the preset peak height threshold H. th (e.g., 0.2), if not, it means that the peak is regarded as noise and the peak that does not meet the condition is excluded; if yes, the peak that meets the condition is retained and step S42 is executed.
[0114] Step S42. For a candidate peak, find its nearest left and right valley points in the valley list. Calculate the difference ΔL between the first negative derivative of the peak point that meets the condition and its left valley point, and calculate the difference ΔR between the peak point that meets the condition and its right valley point. Determine whether both the difference ΔL and ΔR are greater than the preset peak-valley difference threshold Δth (e.g., 0.0169). If not, exclude peaks that do not meet the condition; if so, retain peaks that meet the condition and proceed to step S43. This condition ensures that the peak has sufficient prominence.
[0115] Step S43. Calculate the absolute value of the first temperature difference between the peak point and the valley point to its left (i.e., the left half-peak width W). L), and calculate the absolute value of the second difference between the temperature corresponding to the peak point and the valley point to its right (i.e., the right half-peak width W). R Determine the width W of the left half-peak. L , right half-peak width W R Are all values greater than the preset half-peak width threshold W? half_th (e.g., 0.99℃), if not, then exclude peak values that do not meet the conditions; if yes, then retain peak values that meet the conditions and execute step S44; this condition can exclude excessively sharp noise spikes.
[0116] Step S44. Calculate the temperature difference between the valley points to the right and left of the peak point (i.e., the full peak width W). F Determine the full peak width W F Is it greater than the preset full peak width threshold W? full_th (e.g., 6.5℃). If not, exclude peaks that do not meet the conditions; if yes, retain peaks that meet the conditions and proceed to step S45. This condition ensures that it is a real melting peak with a certain width, rather than a local fluctuation.
[0117] In this embodiment, only peak values that simultaneously meet the above four conditions are considered valid melting point peaks.
[0118] Step S45. Calculate the melting temperature T using linear interpolation on the retained peak values. m The value is used to obtain the first valid T. m Value set.
[0119] For each valid peak, linear interpolation is used to accurately locate it near the zero-crossing point. Let the peak index be p, and obtain its corresponding second derivative value y0, and the second derivative value y1 corresponding to the next point index p+1. Ideally, the zero-crossing point is located between p and p+1. Then, the interpolation ratio fraction between p and p+1 is calculated, expressed as: fraction = y0 / |y0-y1|. Finally, the temperature at the peak position is added to the interpolation ratio and multiplied by the temperature difference between two adjacent temperature sampling points to obtain the accurate T. m The value is represented as:
[0120] T m = temperature[p] + fraction * (temperature[p+1]- temperature[p])
[0121] Where temperature[p] represents the temperature with peak index p; temperature[p+1] represents the temperature with peak index p+1.
[0122] The first effective T can be obtained using the above method. mValue set. This embodiment achieves sub-sampling level precision positioning through linear interpolation, improving positioning accuracy by approximately 10 times compared to directly using peak index, thus overcoming the limitation of sampling interval on measurement accuracy.
[0123] This embodiment employs multi-dimensional feature fusion for judgment, significantly reducing the false positive rate compared to single threshold judgment methods. Traditional methods rely solely on peak height or peak-to-valley difference as a single indicator, making them susceptible to noise interference and resulting in misjudgments. In contrast, this method uses quadruple verification to comprehensively judge from multiple perspectives, including peak height, depth, and width, forming a mutually corroborating judgment system that greatly improves the reliability of peak identification.
[0124] In step S5, in the first effective T m Identify continuous peak groups from the value set, and filter the peak values within the continuous peak groups based on an adaptive relative peak height ratio threshold to obtain the second effective T. m Value set.
[0125] For samples with multiple melting peaks (such as mixed templates or heterozygotes), their negative derivative curves may show multiple consecutive peaks.
[0126] (1) Continuous peak determination: Traverse the first valid T m The set of values for the current peak i If its negative derivative curve has a positive negative_derivative value, and its corresponding smoothed melting curve value is greater than the next peak, then... j The corresponding smooth curve value will be the peak. i and peak j Grouped into the same continuous peak group, and continue to use peak... j Compare with subsequent peaks until the condition is no longer met, and finally count the number of peaks in the group.
[0127] (2) Adaptive threshold filtering: For the identified continuous peak group (e.g., containing 3 peaks), first find the peak with the largest negative derivative curve negative_derivative value in the group as the largest peak in the group. Then, calculate the ratio of the negative derivative value of each peak in the group to the negative derivative value of the largest peak as the relative peak height ratio;
[0128] Preset a peak-to-height ratio threshold R th (e.g., 0.78), determine whether the calculated relative peak height ratio is greater than the peak height ratio threshold R. th If a peak value exists, then select those with a relative peak height ratio greater than R. th For the peak value, the four-fold judgment verification in step S4 is re-executed on the selected peak value to obtain the second valid T. m Value set.
[0129] The relative threshold in this embodiment can automatically adapt to the signal intensity of different samples, eliminating the need for a fixed threshold and supporting mixed samples with multiple T. m Value detection, capable of identifying up to 6 different T values. m value.
[0130] This embodiment eliminates the need for a fixed threshold and automatically adapts to different signal intensities. Traditional methods using fixed thresholds for filtering consecutive peaks are ineffective when there are significant differences in signal intensity among different samples. This embodiment employs a relative threshold, calculating the ratio of each peak to the largest peak within the group. The threshold of 0.65 was determined through extensive experimental optimization using heterozygous samples, effectively identifying genuine minor peaks while eliminating weak noise peaks.
[0131] In step S6, the second valid T m The peak value in the value set is compared with the global maximum peak value. If there is no T corresponding to the global maximum peak value... m If the value is less than a certain threshold, it is considered a false positive sample, and all T samples are removed. m Value; if there exists a T corresponding to the global maximum peak value. m If the value is positive, it is determined to be a true positive sample, and step S7 is executed.
[0132] To determine whether a sample is a true positive, the melting peak generated by the actual amplification product should be one of the most prominent features of the entire melting curve.
[0133] (1) Identify the global maximum peak: Traverse all the initial candidate peak sets obtained in step S3, find the peak with the largest negative derivative curve negative_derivative value, record its negative derivative value and index position, and use it as the global maximum peak.
[0134] (2) False positive determination logic: Initialize a Boolean variable. Assuming it is a false positive, then iterate through all the second valid T variables that have passed the four-fold verification in step S5. m For each valid peak in the set of values, check if its index matches the global maximum peak index. If a valid peak index equals the global maximum peak index, set the false positive flag to false, indicating the sample is a true positive. If the false positive flag remains true after the iteration, it means none of the valid peaks are the global maximum peak, and the sample is judged as a false positive. This implies that none of the peaks identified as valid are the most significant peaks on the curve, which does not conform to the logic of true amplification. Therefore, the sample is judged as a false positive, and all T values for that sample are cleared. m value.
[0135] This embodiment determines signal authenticity from a global perspective, rather than focusing on local features. Existing technologies often use local thresholds for judgment, making it difficult to distinguish between real and non-specific signals. This method is based on biological principles; the true melting point peak should be the most prominent feature in the curve. If none of the verified peaks is the largest peak, it indicates potential background interference or non-specific amplification, while the largest peak may be noise or other interference signals.
[0136] This embodiment effectively eliminates non-specific signals and prevents misjudgments. In experimental verification, this mechanism successfully identified and eliminated false positive signals caused by factors such as reagent contamination and temperature fluctuations, reducing the false positive rate from 3-5% in traditional methods to below 1%.
[0137] Before step S7, the method further includes: when the detection channel in a true positive sample meets the crosstalk triggering condition, using an adaptive baseline alignment algorithm to eliminate the T corresponding to the false positive sample caused by inter-channel crosstalk. m value.
[0138] In multi-channel fluorescence detection, the emission spectrum of one channel may be partially received by the detectors of adjacent channels, leading to crosstalk. This embodiment performs adaptive cancellation for specific channel pairs (taking channel 0-FAM as the reference channel and channel 1-HEX as the detection channels as an example).
[0139] (1) Crosstalk detection triggering conditions: Crosstalk detection is performed on channel 1 of the current hole position when the following three conditions are met simultaneously. The three conditions are:
[0140] The current process is working on a specific channel (e.g., channel 1, HEX channel);
[0141] A reference channel (e.g., channel 0, FAM channel) has at least one valid T at this aperture location. m value;
[0142] T of two channels m The absolute value of the difference is less than 1.1 degrees Celsius.
[0143] (2) Adaptive baseline alignment algorithm:
[0144] Curve scaling: Multiply the negative derivative curve of the detection channel (e.g., channel 1) by a scaling factor (e.g., 1.0) to obtain the scaled negative derivative curve of channel 1. Multiply the negative derivative curve of the reference channel (e.g., channel 0) by a crosstalk scaling factor (e.g., 0.08) to obtain the scaled negative derivative curve of channel 0.
[0145] Baseline alignment: Calculate the minimum value of the negative derivative curve of the scaled detection channel and the minimum value of the negative derivative curve of the scaled reference channel. Subtract the minimum value of the negative derivative curve of the reference channel from the minimum value of the negative derivative curve of the detection channel to obtain the offset. Add this offset to all points of the negative derivative curve of the reference channel to obtain the reference curve after baseline alignment, thus eliminating possible baseline differences between the two channels.
[0146] Point-to-point signal strength comparison: Find the detection channel T respectively m The temperature location corresponding to the value and the reference channel T m For each temperature location corresponding to a given value, find the index of the nearest sampling point and compare the detection channel's T value with the index of the nearest sampling point. m The signal strength at the value position is the same as that of the aligned reference channel at its T value. m If the signal strength at a given location is weaker than the aligned reference channel signal, it is considered crosstalk, and the T signal at that location of the detection channel is adjusted. m The value is reset to zero.
[0147] The adaptive scaling factor in this embodiment can adapt to the actual crosstalk ratio, baseline alignment eliminates the influence of baseline drift, and point-to-point comparison at T m The signal strength is directly compared with the value location, making the determination more accurate.
[0148] This embodiment can adapt to the actual crosstalk ratios of different instruments without manual parameter tuning. Baseline alignment eliminates the influence of baseline drift by calculating the minimum difference, ensuring that the two curves are compared on the same reference. Point-to-point comparison is performed directly on T... m By comparing signal strength at different value locations, the limitations of the overall threshold method are avoided, resulting in more accurate judgment.
[0149] In step S7, the melting curve and negative derivative curve of the samples determined to be true positives are normalized and displayed.
[0150] This embodiment only normalizes positive samples. The criterion for determining a positive sample is that the sample has at least one non-zero T. m value.
[0151] (1) Feature Scaling (Min-Max Normalization): Find the maximum and minimum values of the melting curve data or negative derivative curve to be displayed. For each point in the curve, calculate the normalized value, expressed as:
[0152] normalized_value=(value - data_min) / (data_max - data_min) *display_factor
[0153] Wherein, normalized_value represents the normalized value of the data point, and value represents the numerical value of the data point; data_min represents the minimum value; data_max represents the maximum value; display_factor represents the normalization factor. In the display stage, the normalization factor is set to 100, which maps the data to the range of [0, 100], making it easier to compare different samples under the same scale.
[0154] (2) Reverse scaling: For the negative derivative curve, in order to make the peak of the negative derivative curve face upward when displayed, each data point is multiplied by a negative display scaling factor (such as 3056) to obtain a numerical range suitable for the interface display.
[0155] This embodiment uses a phased processing approach to calculate T. m The original scale is maintained during calculation to ensure accuracy, while scaling is performed during display to optimize the visual effect. This separates data processing from result display, ensuring calculation accuracy while meeting the display requirements of the user interface.
[0156] The first stage of this embodiment calculates T. m The scaling factor is set to 1 during the initial calculation, and to 3056 during the second stage of normalization display, with a normalization factor of 100. This optimizes the display while maintaining calculation accuracy, achieving separation between data processing and result presentation. The original scale is used during the calculation stage to avoid accuracy loss due to scaling, ensuring T... m The accuracy of value calculation is improved. During the display phase, positive samples are normalized and scaled to map the values to a range suitable for interface display, enhancing the visual contrast and readability of the curves.
[0157] Example 2
[0158] The difference between the PCR melting curve analysis method based on multiple validation in this embodiment and that in Embodiment 1 is as follows:
[0159] This approach effectively reduces the false positive rate through a multi-level peak quality control mechanism, specifically compared to traditional methods:
[0160] In the comparison of typical cases Figure 6 This is a pseudo-peak sample caused by an abnormal data drop, specifically the fluorescence change trend of an actual pseudo-peak data case; among which... Figure 6 The first curve in the diagram is a schematic diagram of the original curve smoothed by this scheme, and the second curve is a schematic diagram of the melted data processed by the first negative derivative of this scheme.
[0161] The spurious peak sample exhibited abnormal data acquisition at approximately 50.3℃, with the fluorescence curve momentarily dropping before recovering. Traditional methods identify the first "peak" (at 50.4℃) on the derivative curve and determine it as positive. However, although this method's "peak" passed the first two layers of verification, the third layer of verification showed a left half-peak width of 0.3℃, a right half-peak width of 1.8℃, and a half-peak width difference of 1.5℃ (>0.99℃). This indicates severe peak asymmetry, and the peak width was only 2.1℃ (<4.5℃). Therefore, it was correctly excluded. After visual inspection and biological verification by professionals, the sample was indeed a false positive.
[0162] Example 3
[0163] The difference between the PCR melting curve analysis method based on multiple validation in this embodiment and that in Embodiment 1 is as follows:
[0164] This solution can improve the resolution of multi-peaks, specifically compared to traditional methods:
[0165] In the comparison of typical cases Figure 7 It involves identifying three-peaked samples, specifically the fluorescence change trend in actual three-peaked data cases; among which... Figure 7 The first curve in the diagram is a schematic diagram of the original curve after smoothing using this method, and the second curve is a schematic diagram of the melting data after processing with the first negative derivative of this method.
[0166] Theoretically, this three-peak sample should contain three target sequences and produce three melting peaks (temperatures of approximately 50℃, 60℃, and 78℃, respectively). Traditional methods failed to identify all three peaks due to nearby noise interference. This method successfully identified the three peaks (peak 1: 49.3℃; peak 2: 60℃; peak 3: 78.6℃; with better noise suppression, the peaks were correctly detected) through bidirectional filtering and dynamic windowing techniques. Sequence analysis confirmed that the sample does indeed contain three target fragments, and the identification results of this method are consistent with theoretical expectations.
[0167] Example 4
[0168] This embodiment provides a PCR melting curve analysis system based on multiple validation, including:
[0169] The acquisition module is used to acquire data information from each well of each channel of the PCR reaction plate, preprocess the acquired data information, and generate the initial melting curve for each well of each channel; the data information includes raw fluorescence data, temperature data, and integration time data;
[0170] The first calculation module is used to smooth the initial melting curve using a bidirectional low-pass filtering method, and to calculate the first derivative of the smoothed melting curve to generate a negative derivative curve.
[0171] The second calculation module is used to calculate the second derivative of the obtained negative derivative curve and identify the initial candidate peak set by detecting the zero crossing point of the second derivative.
[0172] The verification module performs a four-fold verification on each candidate peak in the initial candidate peak set based on peak height threshold, peak-to-valley difference threshold, half-width, and full-width. For candidate peaks that pass the four-fold verification, the melting temperature T is calculated using linear interpolation. m The value is used to obtain the first valid T. m Value set;
[0173] The filtering module is used for the first valid T m Identify continuous peak groups from the value set, and filter the peak values within the continuous peak groups based on an adaptive relative peak height ratio threshold to obtain the second effective T. m Value set;
[0174] The judgment module is used to determine the second valid T. m The peak value in the value set is compared with the global maximum peak value. If there is no T corresponding to the global maximum peak value... m If the value is less than a certain threshold, it is considered a false positive sample, and all T samples are removed. m Value; if there exists a T corresponding to the global maximum peak value. m If the value is positive, it is determined to be a true positive sample;
[0175] The display module is used to normalize and display the melting curve and negative derivative curve of samples determined to be true positive samples.
[0176] It should be noted that the PCR melting curve analysis system based on multiple validation provided in this embodiment is similar to that in Embodiment 1, and will not be described in detail here.
[0177] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A multiplexed validation based PCR melt curve analysis method, characterized in that, include: Step S1. Obtain data information for each well in each channel of the PCR reaction plate, preprocess the obtained data information, and generate the initial melting curve for each well in each channel; the data information includes raw fluorescence data, temperature data, and integration time data; Step S2. The initial melting curve is smoothed using a bidirectional low-pass filter, and the first derivative of the smoothed melting curve is calculated to generate a negative derivative curve. Step S3. Calculate the second derivative of the obtained negative derivative curve, and identify the initial candidate peak set by detecting the zero crossings of the second derivative; Step S4. Perform a quadruple validation on each candidate peak in the initial candidate peak set based on peak height threshold, peak-to-valley difference threshold, half-width at half-maximum, and full-width at half-maximum. Calculate the melting temperature T for candidate peaks that pass the quadruple validation using linear interpolation. m The value is used to obtain the first valid T. m Value set; Step S5. In the first valid T m Identify continuous peak groups from the value set, and filter the peak values within the continuous peak groups based on an adaptive relative peak height ratio threshold to obtain the second effective T. m Value set; Step S6. Place the second valid T m The peak value in the value set is compared with the global maximum peak value. If there is no T corresponding to the global maximum peak value... m If the value is less than a certain threshold, it is considered a false positive sample, and all T samples are removed. m value; If there is a T m value corresponding to the global maximum peak, it is determined as a true positive sample, and step S7 is executed. Step S7. Normalize and display the melting curve and negative derivative curve of the samples determined to be true positive samples; Step S4 specifically involves: Step S41. Determine whether the negative derivative value corresponding to the peak point is greater than the preset peak height threshold. If not, exclude the peaks that do not meet the conditions; if yes, retain the peaks that meet the conditions and proceed to step S42. Step S42. Calculate the difference between the first negative derivative and the difference between the second negative derivative of the peak point and the valley point to its left and right respectively. Determine whether the difference between the first negative derivative and the difference between the second negative derivative are both greater than the preset peak-valley difference threshold. If not, exclude the peak points that do not meet the conditions; if so, retain the peak points that meet the conditions and proceed to step S43. Step S43. Calculate the absolute value of the first difference and the absolute value of the second difference between the peak point and the temperature corresponding to the valley point to its left and the valley point to its right, respectively. Determine whether the absolute value of the first difference and the absolute value of the second difference are both greater than the preset half-peak width threshold. If not, exclude the peak points that do not meet the conditions; if so, retain the peak points that meet the conditions and proceed to step S44. Step S44. Calculate the temperature difference between the valley point to the right of the peak point and the valley point to the left of the peak point, and determine whether the temperature difference is greater than the preset full peak width threshold. If not, exclude the peak that does not meet the condition; if yes, retain the peak that meets the condition and proceed to step S45. Step S45. Calculate the melting temperature T using linear interpolation on the retained peak values. m The value is used to obtain the first valid T. m Value set; The melting temperature T is calculated in the step S45 using linear interpolation m Values include in particular: The second derivative value of the peak position and a position after the peak position is obtained, an interpolation ratio is calculated according to the second derivative value of the peak position and the obtained second derivative value, and a first effective T m value set is obtained according to the interpolation ratio, a temperature difference of the adjacent two temperature sampling points and the temperature of the peak position. Step S5 specifically involves: Step S51. Traverse the first valid T m The set of values is used to determine whether the negative derivative of each peak is positive and whether the smooth curve value is greater than the smooth curve value corresponding to the next peak. If so, the current peak and the next peak are grouped into a continuous peak group, and the process continues until the condition is no longer met. The number of continuous peaks is then counted. Step S52. Traverse all peaks within a continuous peak group and select the peak with the largest negative derivative value; Step S53. Calculate the ratio of the negative derivative value of each peak in the continuous peak group to the negative derivative value of the largest peak, as the relative peak height ratio, and determine whether there are peaks with a relative peak height ratio greater than the preset peak height ratio threshold. If so, re-execute the four-fold judgment verification in step S4 for the peaks that meet the conditions to obtain the second valid T. m Value set; Step S6 specifically involves: Step S61. Identify the peak with the largest negative derivative value in the candidate peak set from step S3 as the global maximum peak; Step S62. Traversing all valid peaks in the second valid T m value set, if there is a valid peak whose index is the same as the index of the global maximum peak, the sample is determined as a true positive sample; if not, it is determined as a false positive sample, and all T m values of the sample are cleared.
2. The PCR melting curve analysis method based on multiple validation according to claim 1, characterized in that, The preprocessing of the acquired data information in step S1 specifically includes: invalid data processing and integration time normalization processing of the original fluorescence data.
3. The PCR melting curve analysis method based on multiple validation according to claim 1, characterized in that, In step S2, the initial melting curve is smoothed by a two-way exponentially weighted moving average filtering method.
4. The PCR melting curve analysis method based on multiple validation according to claim 1, characterized in that, Before step S7, the following also includes: When the detection channel in a true positive sample meets the crosstalk triggering condition, the adaptive baseline alignment algorithm is used to eliminate the false positive samples corresponding to T caused by inter-channel crosstalk. m value.
5. The PCR melting curve analysis method based on multiple validation according to claim 4, characterized in that, The false positive sample corresponding to T is eliminated by the adaptive baseline alignment algorithm m The value is specifically: Obtain the negative derivative curves of the detection channel and the reference channel, and scale the negative derivative curves of the detection channel and the reference channel proportionally. The minimum difference between the negative derivative curve of the scaled detection channel and the negative derivative curve of the reference channel is calculated and used as the baseline offset to perform baseline alignment on the negative derivative curve of the reference channel. comparing the signal intensity of the negative derivative curve of the detection channel at the position corresponding to the T m value of the negative derivative curve of the reference channel after alignment with the baseline with the signal intensity of the negative derivative curve of the detection channel, and determining whether to clear the T m value of the detection channel based on the comparison result.
6. An analytical system based on the PCR melting curve analysis method based on multiple validation as described in any one of claims 1-5, characterized in that, include: The acquisition module is used to acquire data information from each well of each channel of the PCR reaction plate, preprocess the acquired data information, and generate the initial melting curve for each well of each channel. The data includes raw fluorescence data, temperature data, and integration time data; The first calculation module is used to smooth the initial melting curve using a bidirectional low-pass filtering method, and to calculate the first derivative of the smoothed melting curve to generate a negative derivative curve. The second calculation module is used to calculate the second derivative of the obtained negative derivative curve and identify the initial candidate peak set by detecting the zero crossing point of the second derivative. The verification module performs a four-fold verification on each candidate peak in the initial candidate peak set based on peak height threshold, peak-to-valley difference threshold, half-width, and full-width. For candidate peaks that pass the four-fold verification, the melting temperature T is calculated using linear interpolation. m The value is used to obtain the first valid T. m Value set; a screening module for identifying a group of consecutive peaks in the first set of effective T m values, screening the peaks within the group of consecutive peaks based on an adaptive relative peak height ratio threshold to obtain a second set of effective T m values; The judgment module is used to determine the second valid T. m The peak value in the value set is compared with the global maximum peak value. If there is no T corresponding to the global maximum peak value... m If the value is less than a certain threshold, it is considered a false positive sample, and all T samples are removed. m value; If there is a T m value corresponding to the global maximum peak, then it is determined as a true positive sample; The display module is used to normalize and display the melting curve and negative derivative curve of samples determined to be true positive samples.
Citation Information
Patent Citations
Real-time fluorescent quantitative PCR data processing method and device
CN114882948A
Method and device for determining Tm value of high-resolution melting curve and electronic equipment
CN116543840A
High-precision spectral signal peak detection method and system
CN120354155A