Biological Sample Classification via DNA Melting Curve Derivatives
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for classifying biological samples into groups face challenges in accuracy, automation, and manual selection of variables, especially when the number of groups is large or when samples do not belong to known groups.
Innovation Solution
A method involving the acquisition of DNA melting curves, analysis of descriptors such as second derivatives, and the use of a random forest method for classification, which includes learning from reference curves, descriptor elimination, and calculation of a confidence index to determine group membership.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If manual variable selection is used for classification, then the process can be customized for specific research questions, but the process cannot be automated and requires manual intervention for each new question
Solution Approach 1:
The system performs automatic variable selection through the random forest algorithm, which autonomously identifies and selects the most relevant genes from the expression data without requiring manual intervention. The algorithm evaluates all genes and automatically determines the optimal subset for classification, enabling the system to serve itself rather than requiring user configuration for each new research question.
Solution Approach 2:
The patent transforms the classification approach by changing from manual parameter selection to algorithmic parameter selection. The random forest method dynamically adjusts and selects variables (genes) based on the data characteristics, automatically adapting the variable set to each new classification task without requiring manual reconfiguration.
2Adaptability or versatility
If a large number of possible groups are considered in classification, then the classification can cover more biological scenarios, but the accuracy of classification decreases
Solution Approach 1:
The patent performs preliminary action by pre-selecting the most relevant variables (genes) using the random forest algorithm before the actual classification. This preliminary variable selection process identifies and retains only the most discriminative genes, reducing the dimensionality of the data and improving classification accuracy across multiple groups by eliminating noise from less relevant variables.
Solution Approach 2:
The system extracts and isolates the most important variables from the complete gene set using random forest variable selection. By taking out only the most relevant genes and excluding the rest, the system maintains high classification accuracy even when dealing with a large number of possible groups, as the classification is based on a focused subset of discriminative features.
3Measurement precision
If traditional classification methods are used, then the process is simpler, but the accuracy of classification and ability to handle complex samples decreases
Solution Approach 1:
The patent replaces traditional manual classification methods with an automated computational approach using random forest algorithms. This substitution transforms the classification process from a manual, intuition-based method to an automated, data-driven system that objectively selects variables and performs classification, significantly improving accuracy while managing complexity through algorithmic automation.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach improves classification accuracy, automates the process, and allows for the identification of samples not belonging to known groups by using DNA melting curve analysis and random forest classification, enhancing discrimination and confidence in group assignment.
Implementation Method 1
each point corresponding to a quantity proportional or representative of a rate or amount of DNA denaturation of the measurement sample as a function of temperature
Data Source
Figure 1~9
Figure 3
Figure 4~8
AI summary
The present invention relates to a method of classifying a biological measurement sample, comprising: -an acquisition of at least one DNA fusion curve for the biological measurement sample, termed at least one measurement curve, -a determination of a membership of the biological measurement sample in a group determined from among various possible groups, by an analysis of descriptors arising from the at least one measurement curve, characterized in that the descriptors comprise one or more points of the first derivative of each measurement curve and/or comprise one or more points of the second derivative of each measurement curve and/or one or more points of each measurement curve and/or one or more percentiles of each measurement curve. The invention also relates to a device (100) implementing this method.