Adaptive Thresholds for Genetic Sequencing Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Genetic sequencing results often contain false positives due to contamination and misidentification of genetic material, making it challenging to accurately classify organisms present in a sample, especially when different organisms share similar genetic sequences or have widespread genetic material.
Innovation Solution
A system that uses adaptive thresholds based on negative control samples to classify genetic sequencing results, employing a neural network to identify diagnostically significant organisms by generating a model from negative control samples and applying it to clinical samples, with features like autoencoders and heatmaps to present significant findings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single threshold is used across all organisms for determining diagnostic relevance, then the classification process is simple, but false positives increase and diagnostic accuracy deteriorates
Solution Approach 1:
The patent implements organism-specific thresholds instead of a universal threshold. Each organism receives a customized threshold value based on its characteristics, contamination likelihood, and diagnostic importance. This local differentiation resolves the contradiction by maintaining simplicity through automation while achieving high precision through tailored classification criteria for each organism type.
Solution Approach 2:
The system dynamically adjusts threshold parameters based on organism identity, sample type, and contextual factors. Rather than using a fixed threshold, the patent modifies the threshold parameter adaptively for different organisms, enabling accurate differentiation between true positives and false positives while maintaining an efficient automated classification process.
2Measurement precision
If organism-specific adaptive thresholds are used, then diagnostic accuracy improves, but system complexity increases
Solution Approach 1:
The system automatically generates and applies organism-specific thresholds without requiring manual configuration for each organism. The classification algorithm self-adjusts based on organism characteristics and historical data, resolving the complexity issue by automating the threshold determination process while maintaining high diagnostic accuracy through personalized classification criteria.
Solution Approach 2:
The patent pre-calculates and stores threshold values for multiple organisms before actual sample analysis. This preliminary preparation of organism-specific parameters eliminates the need for complex real-time calculations during diagnosis, achieving high precision through pre-configured adaptive thresholds while keeping the operational system relatively simple.
3Object-generated harmful factors
If low level results are categorically ignored, then false positives from contamination are reduced, but true positives at low concentrations are missed
Solution Approach 1:
The patent applies different evaluation criteria to different organisms based on their diagnostic importance and typical contamination levels. Clinically significant organisms receive lower detection thresholds with stricter validation, while less important organisms have higher thresholds. This local differentiation resolves the contradiction by enabling low-level detection for critical pathogens while filtering false positives from less relevant organisms.
Data Source
AI summary
In accordance with some embodiments, systems, methods, and media for classifying genetic sequencing results based on pathogen-specific adaptive thresholds are provided. In some embodiments, a system comprises a processor programmed to: receive negative control results, each comprising values indicative of a number of reads detected in the respective negative control sample for an organism; generate a model based on the negative control results; receive a clinical sample result for a clinical sample, comprising values indicative of a number of reads detected in the clinical sample for an organism of a plurality of organisms; identify, utilizing the model, any values in the clinical sample that are likely to be diagnostically significant; generate a report based on the clinical sample result and organisms associated with a value likely to be diagnostically significant; and cause the report to be presented to a user.


