Multiplexed fluorescent quantitative fusion peak detection automated recognition analysis system

By smoothing the melting curve using the information entropy empirical mode decomposition algorithm and the first-order difference method, and identifying the Tm value, the problem of identifying the fusion peak in multiplex quantitative PCR with low resolution and many curve clutters is solved. This enables automated analysis and evaluation of instrument and reagent stability, improving detection efficiency and optimization capabilities.

CN118861601BActive Publication Date: 2026-04-07AUTOBIO LABTEC INSTR CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify fusion peaks in multiplex quantitative PCR under conditions of low resolution, high curve clutter, and unstable data. Furthermore, there is a lack of automated analysis systems to assess the stability and performance of instruments and reagents.

Method used

The Empirical Mode Decomposition (EMD) algorithm of information entropy is used to smooth the melting curve. The Tm value is identified by combining first-order difference and multiple threshold screening methods. Stability distribution map is generated by combining instrument and reagent information to provide automated analysis.

Benefits of technology

It improves the efficiency and accuracy of fusion peak detection, provides feedback on the stability of instrument reagents, and offers data support for product optimization. It is suitable for environments with low resolution and high curve clutter.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118861601B_ABST
    Figure CN118861601B_ABST
Patent Text Reader

Abstract

This invention discloses an automated identification and analysis system for multiple fluorescence quantitative fusion peak detection, comprising an information acquisition module, an information processing module, and a statistical analysis module. A simple first-order difference is performed on the original melting curve to obtain the melting peak curve. An information entropy adaptive EMD filter is then used to obtain a smoothed melting peak curve. The fusion peak on the melting peak curve is identified using a peak lookup method. Simultaneously, this invention collects relevant information from instruments and reagents, combining it with identified Tm values ​​and time, to analyze the stability of reagents or instruments from different perspectives, providing feedback on the stability of different batches of instruments and reagents, and offering reference data resources for product upgrades and optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multiplex fluorescence quantitative technology, and more particularly to an automated identification and analysis system for the detection of multiplex fluorescence quantitative fusion peaks. Background Technology

[0002] Multiplex fluorescence quantitative PCR (MPCR) is a technique based on quantitative real-time PCR (qPCR). It utilizes a combination of different fluorescent groups and the instrument's ability to detect fluorescence in different channels to achieve real-time quantitative detection of multiple targets. Multiplex PCR is mainly used for the simultaneous detection or identification of multiple pathogens, certain genetic diseases, and oncogene typing. For example, the simultaneous detection or identification of multiple pathogens involves adding specific primers for multiple pathogens to the same PCR reaction tube, performing PCR amplification, and simultaneously detecting multiple pathogens or identifying one or more pathogen types.

[0003] The melting curve is obtained by gradually increasing the temperature of the reaction solution after the PCR reaction and monitoring the subtle changes in the fluorescence signal of the fluorescent dye in real time. Before heating, the PCR product is a double-stranded structure with a high fluorescence intensity signal. As the temperature is gradually increased, the DNA double strands gradually open, the amount of dye embedded in the double strands decreases, and the fluorescence signal gradually decreases. After obtaining the melting curve, a first-order derivative is performed to obtain the melting peak curve. The Tm value is the peak point formed on the melting peak curve. The Tm value represents the specificity of the corresponding target location on the corresponding channel, which can prove that a specific gene fragment at the relevant target location has been detected, and thus can identify the specific type of related disease.

[0004] For data output from fluorescence instruments with high resolution and a large number of sampling points, the melting peak curve can be directly generated by simulating the first derivative using the differential method. Alternatively, the Tm value can be identified by setting a peak pre-value and using a basic peak height identification method. However, these methods place high demands on the design and cost of the detection instruments and cannot obtain accurate Tm values ​​for the current common problems of fewer sampling points, more curve clutter, unstable data, and closely spaced target peaks.

[0005] At present, there are no systems or products for automated analysis, identification and statistics of multiplex fluorescence quantitative detection. It is impossible to provide data comparison and analysis on the performance of different batches of instruments and reagents, to provide reagent manufacturers and users with feedback on the stability of instruments and reagents, and to provide reference data resources for product upgrades and optimization. Summary of the Invention

[0006] The purpose of this invention is to provide an automated identification and analysis system for multiple fluorescence quantitative fusion peak detection. This system is suitable for identifying, statistically analyzing, and analyzing multiple fluorescence quantitative fusion peaks that are characterized by a small number of sampling points, numerous curve clutters, unstable data, and closely spaced target peaks. It improves the single-pass detection efficiency of multiple fluorescence quantitative fusion peak detection, provides feedback on the stability of different batches of instrument reagents, and offers reference data resources for product upgrades and optimization.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] The automated identification and analysis system for multiple fluorescence quantitative fusion peak detection of the present invention includes an information acquisition module, an information processing module, and a statistical analysis module. The information acquisition module acquires instrument, reagent, and reaction information. The information processing module is mainly used to smooth the raw melting curve data obtained from the reaction information, obtain the melting peak curve, and identify the corresponding Tm value. The statistical analysis module, based on the instrument information, reagent information, reaction information, and the smoothness and Tm value of the melting peak curve, provides a stability distribution diagram of the instrument within a preset time and the Tm value distribution of different batches of reagents in different channels, and reports the stability of the instrument, reagents, and different target locations.

[0009] Furthermore, the instrument information includes the instrument manufacturer, instrument model, and instrument production batch; the reagent information includes the reagent manufacturer and reagent production batch; and the reaction information includes the original melting curve data and channel information.

[0010] Furthermore, the smoothing process includes performing a first-order difference operation on the original melting curve data to obtain a non-stationary melting peak curve; and using an information entropy-based empirical mode EMD filtering algorithm to smooth the non-stationary melting peak curve.

[0011] Furthermore, the information entropy-based empirical mode EMD filtering algorithm includes using the empirical mode EMD decomposition algorithm to calculate multiple intrinsic function component data of the non-stationary melting peak curve, calculating the information entropy of the intrinsic function component data, and realizing the smoothing processing of the non-stationary melting peak curve according to the relationship between the information entropy and the preset step threshold.

[0012] Furthermore, the identified Tm value includes ordinary peak identification and fused peak identification.

[0013] Furthermore, the ordinary peak identification, based on the melting peak curve, filters out the set of peak coordinates whose vertical coordinate values ​​are greater than those on the left and right sides; removes peak coordinates that do not meet the preset peak height threshold, preset peak height steepness threshold, preset peak spacing threshold, or whose coordinates are located on the left and right boundaries of the melting peak curve; removes peak coordinates that do not meet the thresholds for the proportion of the highest peak, the thresholds for upward and downward trends, or the thresholds for the minimum number of upward and downward movements on the left and right sides, and determines the final set of peaks.

[0014] Furthermore, in the fusion peak identification, several points to the left and right of each peak coordinate in the final peak set are found, and the peak coordinates are connected to the farthest left and right points with straight lines; the distance from each left and right point of the peak coordinate to the straight line is calculated one by one, the convex point with the farthest distance on the left and right sides is found, and then it is determined whether the convex point meets the target peak threshold. If it does, it is confirmed as the Tm value.

[0015] Furthermore, the statistical analysis module provides stable information entropy distribution diagrams of melting peak curves for one year, six months, and three months, stable information entropy distribution diagrams of melting peak curves in the instrument under different batches of reagents and different channels; distribution of Tm values ​​under different batches of reagents and different channels, and distribution of Rm values ​​calculated from the corresponding Tm value positions.

[0016] Furthermore, by arbitrarily combining channels, reagent batches, and experimental times, a certain number of original melting peak curve data are selected for first-order difference operations. The information entropy of the first-order difference value is calculated, and the variances of the information entropy at the 5th, 25th, 75th, and 95th quantiles are statistically analyzed. A stable information entropy distribution map of the melting peak curve is plotted. The sum of the products of each quantile weight and each quantile variance is calculated, and the sum of the products is compared with the total variance threshold, and the variances of each quantile are compared with the respective quantile variance thresholds to determine the stability of the instrument or reagent.

[0017] The advantages of this invention lie in performing a simple first-order difference on the original melting curve to obtain the melting peak curve, then using an information entropy adaptive EMD filter to obtain a smooth melting peak curve, and finally identifying the fusion peak on the melting peak curve through peak lookup. Simultaneously, this invention collects relevant information about instruments and reagents, combining it with identified Tm values, time, etc., to analyze the stability of reagents or instruments from different perspectives, providing feedback on the stability of different batches of instruments and reagents, and offering reference data resources for product upgrades and optimization. Attached Figure Description

[0018] Figure 1 This is a framework diagram of the system described in this invention.

[0019] Figure 2 This is a flowchart of the information processing module in the system described in this invention.

[0020] Figure 3 This is a flowchart of the ordinary peak identification process in the system described in this invention.

[0021] Figure 4 This is a flowchart of the peak identification process in the system described in this invention. Detailed Implementation

[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0023] like Figure 1 The automated identification and analysis system for multiple fluorescence quantitative fusion peak detection described in this invention includes an information acquisition module, an information processing module, and a statistical analysis module.

[0024] The information acquisition module mainly includes acquiring instrument, reagent, and reaction information. Instrument information includes the instrument manufacturer, model, and batch number; reagent information includes the reagent manufacturer and batch number; reaction information includes raw melting curve data and channel information.

[0025] like Figure 2 As shown, the information processing module is mainly used to smooth the original melting curve data obtained from the reaction information, and then identify the corresponding Tm value after obtaining the melting peak curve. The smoothing process includes performing a first-order difference operation on the original melting curve data to obtain a non-stationary melting peak curve; and using an information entropy-based empirical mode EMD filtering algorithm to smooth the non-stationary melting peak curve.

[0026] The Empirical Mode Decomposition (EMD) algorithm based on information entropy primarily applies varying degrees of filtering to the melting curve after simple first-order differencing, resulting in a smoothed melting peak curve. The EMD algorithm is used to process the non-stationary melting peak curve after simple differencing, yielding multiple intrinsic function component data. Information entropy is then calculated for these intrinsic function component data. According to the definition of information entropy, a smaller information entropy value indicates more important information, requiring a lower order of rejection. Conversely, a larger information entropy value indicates more white noise components, requiring a higher order of rejection. Based on the relationship between information entropy and a preset stepped threshold, adaptive filtering is applied to the simple non-stationary melting peak curve to achieve smoothing.

[0027] After obtaining a smooth melting peak curve, the output Tm value needs to be identified based on this smooth melting peak curve. Specifically, Tm value identification includes ordinary peak identification and fused peak identification.

[0028] like Figure 3 The diagram shows the flowchart for ordinary peak identification. Based on the smoothed melting peak curve, the set of peak coordinates with ordinate values ​​greater than those on the left and right sides is selected. Then, each peak is judged one by one. Coordinates that do not meet the preset peak height threshold are removed. Based on the steepness of the rise and fall on the left and right sides of a peak coordinate, it is determined whether it is just a particularly gentle convex peak, and coordinates whose steepness of the rise and fall on the left and right sides does not meet the preset peak height steepness are removed.

[0029] Steepness = Mean of the first-order difference to the left of the current peak coordinate / Negative mean of the first-order difference to the right of the current peak coordinate.

[0030] In addition, when designing targets, a certain distance is usually maintained between peaks to facilitate the identification of Tm values. Therefore, peak coordinates that do not meet the preset peak spacing can be eliminated based on whether the distance between two peak coordinates and the x-axis meets the preset peak spacing. Furthermore, some peak coordinates with insignificant increases should be eliminated based on a preset threshold for the proportion of the highest peak.

[0031] The threshold for the proportion of the highest peak is calculated as: Current peak height / Highest peak height. Peak height is the ordinate of the peak value coordinate.

[0032] Based on the ratio of rising to falling trends, remove the peak coordinates where only one side shows a significant rise or fall.

[0033] The ratio of rising to falling trends generally includes the ratio of rising to falling trends to the left of the current peak coordinate and the ratio of rising to falling trends to the right of the current peak coordinate.

[0034] The ratio of rising to falling trends on the left side = minimum steepness of the left side / maximum steepness of the left side.

[0035] Ratio of rising to falling trend on the right side = minimum steepness of the right side / maximum steepness of the right side.

[0036] Based on the proportion of rising and falling points on the left and right sides, peak coordinates that do not meet the minimum threshold for the number of rising and falling points on the left and right sides are removed. Rising and falling points are points whose first-order difference value is greater than 0.

[0037] The percentage of rising and falling points on the left = minimum number of rising and falling points on the left / maximum number of rising and falling points on the left.

[0038] The percentage of right-side rising and falling points = minimum number of right-side rising and falling points / maximum number of right-side rising and falling points.

[0039] Finally, since the initial target position is designed to avoid positions too close to the edge, this principle can be used to eliminate the peak coordinates located at the left and right boundaries of the melting peak curve, thereby determining the final peak set.

[0040] like Figure 4 The diagram shows the fusion peak identification flowchart. After finding a set of easily identifiable peak positions using a simple peak threshold setting and peak search method, the fusion peaks in the melting curve possess an inherent characteristic: the TM values ​​of the two peaks should be within a certain target range. Based on the target range setting, the fusion peaks of potential peak values ​​(TM values) can be determined by the left and right proximity points. A straight line is formed by connecting the left and right proximity points to the peak point. Based on the distance from each point to the straight line, the points on both sides with the largest distance to the line are selected to identify the most convex points on both sides.

[0041] Specifically, first, find several points to the left and right of each peak coordinate in the final peak set, typically 25 points. If the requirement of 25 points is not met, cover the point farthest from the peak coordinate. Connect the peak coordinate with the farthest points to the left and right. Calculate the distance from each point to the left and right of the peak coordinate to the line, and find the two points farthest from the line on both sides. Then, determine whether these two points meet the target peak threshold. If they do, confirm the point as a Tm value. Otherwise, delete the point.

[0042] Finally, the statistical analysis module of the automated identification and analysis system for multiple fluorescence quantitative fusion peak detection described in this invention can provide a stability distribution map of the instrument within a preset time period and the distribution of Tm values ​​of different batches of reagents in different channels based on instrument information, reagent information, reaction information, and the smoothness and Tm value of the melting peak curve. It also reports the stability of the instrument, reagents, and different target locations. For example, it provides stability entropy distribution maps of the melting peak curve stability information for one year, six months, and three months.

[0043] The method for plotting the stable information entropy distribution diagram of the melting peak curve is as follows:

[0044] Arbitrarily combine channels, reagent batches, and experimental times, then select a certain number of original melting peak curve data under the combined conditions to perform first-order difference operations, calculate the information entropy of the first-order difference values, and statistically analyze the variances of the information entropy at the 5th, 25th, 75th, and 95th quantiles, respectively, to plot a stable information entropy distribution map of the melting peak curves. Then, calculate the sum of the products of each quantile weight and each quantile variance, compare the sum of these products with the total variance threshold, and compare each quantile variance with its respective quantile variance threshold to determine the stability of the instrument or equipment.

[0045] When determining instrument stability, a certain number of original melting peak curve data from the same reagent batch, the same test time, and the same channel can be selected to plot the stability information entropy distribution of the melting peak curve. The sum of the products of each quantile weight and each quantile variance and the total variance threshold, as well as the sum of each quantile variance and each quantile variance threshold, can be compared to determine the instrument stability.

[0046] When determining reagent stability, a certain number of original melting peak curve data from different reagent batches, under the same test time and the same channel can be selected to plot the stability information entropy distribution of the melting peak curve. The stability of the reagent can be determined by comparing the sum of the products of each quantile weight and each quantile variance with the total variance threshold, and the sum of the variances of each quantile with each quantile variance threshold.

[0047] The distribution of stable information entropy of melting peak curves in the instrument under different batches of reagents and different channels (four-channel or six-channel), including the mean, variance, upper and lower 75th percentiles, and upper and lower 95th percentiles. The distribution of Tm values ​​under different batches and channels of reagents, such as the range of mean, variance, and Rm values ​​calculated from the corresponding Tm value positions, including the mean, variance, upper and lower 75th percentiles, and upper and lower 95th percentiles.

[0048] Based on the analysis report provided by the statistical analysis module, instrument optimization and reagent optimization can be achieved. During the instrument optimization phase, a lower information entropy of the curve indicates more stable data, while a higher information entropy indicates greater data fluctuation. Based on this principle, the controlled variable method can be used, employing stable reagents, adjusting the heating time, and fine-tuning the sampling frequency of the instrument during reagent compatibility testing.

[0049] When optimizing reagents, controlling the instrument and reagent concentrations allows for the verification of the rationality of the reagent target design based on the degree of overlap between two Tm value locations in the data. Simultaneously, the effectiveness of the target primers and other reagent components at the same Tm value is evaluated based on the distribution of Rm values.

[0050] When instruments and reagents are used together, if there are abnormal fluctuations in the curve or unclear peaks, the main causes of the abnormalities can be analyzed by controlling variables, and then individual analysis and optimization can be performed.

Claims

1. An automated identification and analysis system for the detection of multiple fluorescence quantitative fusion peaks, characterized in that... The system includes an information acquisition module, an information processing module, and a statistical analysis module. The information acquisition module acquires instrument, reagent, and reaction information. The information processing module is mainly used to smooth the raw melting curve data obtained from the reaction information, obtain the melting peak curve, and identify the corresponding Tm value. The statistical analysis module, based on the instrument information, reagent information, reaction information, and the smoothness and Tm value of the melting peak curve, provides a stability distribution diagram of the instrument within a preset time and the Tm value distribution of different batches of reagents in different channels, and reports the stability of the instrument, reagents, and different target locations. The smoothing process includes performing a first-order difference operation on the original melting curve data to obtain a non-stationary melting peak curve; and using an information entropy-based empirical mode EMD filtering algorithm to smooth the non-stationary melting peak curve. The information entropy-based empirical mode EMD filtering algorithm includes using the empirical mode EMD decomposition algorithm to calculate multiple intrinsic function component data of the non-stationary melting peak curve, calculating the information entropy of the intrinsic function component data, and realizing the smoothing of the non-stationary melting peak curve according to the relationship between the information entropy and the preset step threshold.

2. The automated identification and analysis system for multiple fluorescence quantitative fusion peak detection according to claim 1, characterized in that: The instrument information includes the instrument manufacturer, instrument model, and instrument production batch; the reagent information includes the reagent manufacturer and reagent production batch; and the reaction information includes the original melting curve data and channel information.

3. The automated identification and analysis system for multiple fluorescence quantitative fusion peak detection according to claim 1, characterized in that: Identifying the Tm value includes ordinary peak identification and fused peak identification.

4. The automated identification and analysis system for multiple fluorescence quantitative fusion peak detection according to claim 3, characterized in that: The ordinary peak identification process involves filtering out peak coordinates whose vertical coordinate values ​​are greater than those on the left and right sides based on the melting peak curve; removing peak coordinates that do not meet preset peak height thresholds, preset peak height steepness thresholds, preset peak spacing thresholds, or whose coordinates are located on the left and right boundaries of the melting peak curve; and removing peak coordinates that do not meet thresholds for the proportion of the highest peak, upward and downward trends, or the minimum number of upward and downward movements on the left and right sides, thus determining the final peak set.

5. The automated identification and analysis system for multiple fluorescence quantitative fusion peak detection according to claim 4, characterized in that: The fusion peak identification process involves finding several points to the left and right of each peak coordinate in the final peak set, connecting the peak coordinates with the farthest points on the left and right with straight lines, calculating the distance from each point to the straight line, finding the convex point with the farthest distance on both the left and right sides, and then determining whether the convex point meets the target peak threshold. If it does, it is confirmed as the Tm value.

6. The automated identification and analysis system for multiple fluorescence quantitative fusion peak detection according to claim 1, characterized in that: The statistical analysis module provides stable information entropy distribution maps of melting peak curves for one year, six months, and three months based on the original melting peak curve data under the selected channel, reagent batch, and test time, to analyze the stability of reagents and instruments; and provides the distribution of Tm values ​​under different reagent batches and different channels, as well as the distribution of Rm values ​​calculated from the corresponding Tm value positions.

7. The automated identification and analysis system for multiple fluorescence quantitative fusion peak detection according to claim 6, characterized in that: Using arbitrary combinations of channels, reagent batches, and experimental times, a certain number of original melting peak curve data are selected for first-order difference operations. The information entropy of the first-order difference values ​​is calculated, and the variances of the information entropy at the 5th, 25th, 75th, and 95th quantiles are statistically analyzed. A stable information entropy distribution map of the melting peak curve is plotted. The sum of the products of each quantile weight and each quantile variance is calculated, and the sum of the products is compared with the total variance threshold, and the variances of each quantile are compared with their respective quantile variance thresholds to determine the stability of the instrument or reagent.

Citation Information

Patent Citations

  • Method for online monitoring surface particle pollutants of large-diameter reflector

    CN110389088A

  • Devices, systems, and methods for high-resolution melt analysis

    US20170323051A1