Machine Learning DNA Mixture Contributor Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for estimating the number of contributors in DNA mixtures are inadequate, often overestimating or underestimating the number of contributors and require extensive processing time, which can lead to inaccurate mixture interpretation in forensic and clinical applications.
Innovation Solution
A machine learning approach using a support vector machine algorithm to probabilistically infer the number of contributors in DNA mixtures based on categorical and quantitative data, such as allele labels, stutter rates, peak heights, and mixture ratios, allowing for rapid and accurate characterization of DNA samples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing methods are used to estimate the number of contributors in DNA mixtures, then processing can be performed, but the accuracy is poor (overestimation or underestimation) and processing time is excessive
Solution Approach 1:
The patent replaces traditional mechanical/statistical processing methods with a machine learning system. The machine learning model is trained on DNA mixture data to automatically predict the number of contributors, substituting manual or conventional computational methods with an intelligent system that learns patterns from training data, thereby improving both accuracy and reducing processing time
Solution Approach 2:
The patent implements preliminary action by training the machine learning model in advance on a comprehensive dataset of DNA mixtures with known contributor numbers. This pre-training phase allows the system to learn optimal prediction patterns before actual analysis, enabling rapid and accurate estimation during forensic analysis without requiring time-consuming iterative processing during the actual case work
2Productivity
If the number of contributors is assumed to be known, then likelihood-based deconvolution can proceed, but the assumption may be erroneous and adversely affect results
Solution Approach 1:
The patent implements feedback by using the machine learning model to predict the number of contributors, which then informs the likelihood-based deconvolution process. The system provides feedback on the predicted contributor number to guide the analysis parameters and expectations, creating a closed-loop system where the prediction results are used to optimize the subsequent interpretation steps
Solution Approach 2:
The machine learning model serves as an intermediary between the raw DNA mixture data and the likelihood-based deconvolution process. It mediates by providing an informed estimate of the number of contributors, which acts as a bridge that connects the raw data to the analytical framework, enabling more accurate and reliable deconvolution without requiring direct assumption of contributor numbers
Data Source
AI summary
A system configured to characterize a number of contributors to a DNA mixture within a sample, the system comprising: a sample preparation module configured to generate initial data about the DNA mixture within the sample; a processor comprising a number of contributors determination module comprising a machine-learning algorithm configured to: (i) receive the generated initial data; (ii) analyze the generated initial data to determine the number of contributors to the DNA mixture within the sample; and an output device configured to receive the determined number of contributors from the processor, and further configured output information about the received determined number of contributors.


