Genotype Estimation Device Using Clustering Confidence Thresholds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing genotyping technologies, such as DNA microarray technology, face challenges in accurately determining genotypes for specimens with fluorescence intensities that deviate from the observed group, leading to unreliable genotype assignments due to difficulties in clustering techniques.

Innovation Solution

A genotype estimation device and method that utilize a k-nearest neighbor algorithm, imputation algorithm, and thresholding method to estimate genotypes by acquiring clustering strengths, selecting reference data, and determining the confidence of clustering, allowing for accurate genotype determination even for specimens with uncertain or missing data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If clustering techniques are used to determine genotypes, then high throughput genotype determination is achieved, but accuracy deteriorates for specimens with fluorescence intensities deviating from the observed group

Engineering Contradiction:
Improvethroughput of genotype determinationVSAvoidaccuracy of genotype assignment
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces a confidence score as an intermediary metric to evaluate the reliability of clustering results. This confidence score is calculated based on the distance of fluorescence intensity from cluster centers and the compactness of clusters. By using this intermediary indicator, the system can identify uncertain specimens and apply alternative genotyping methods, thereby resolving the contradiction between high throughput and accurate genotype assignment for deviating specimens.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements a dynamic genotyping approach where the determination method is adjusted based on the confidence score. For high-confidence specimens, standard clustering is used; for low-confidence specimens, alternative methods such as manual review or re-assay are triggered. This dynamic adaptation allows the system to maintain high throughput for most specimens while ensuring accuracy for uncertain cases.

Inventive Principle:
Principle #15Dynamics

2Reliability

If a threshold is specified for clustering strength to avoid unreliable genotype assignment, then reliability is improved, but productivity deteriorates due to excluded specimens

Engineering Contradiction:
Improvereliability of genotype assignmentVSAvoidthroughput of genotype determination
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements a feedback mechanism where confidence scores are continuously calculated and used to guide subsequent genotyping decisions. Specimens with low confidence scores trigger re-evaluation or alternative methods, and the results feed back into the overall genotyping process. This feedback loop ensures that reliability is maintained while minimizing the loss of productivity by only excluding truly uncertain specimens.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent dynamically adjusts the clustering threshold based on the distribution of confidence scores and the specific characteristics of each specimen. Rather than using a fixed threshold that excludes too many specimens, the system adapts the threshold parameter to balance reliability and productivity, allowing more specimens to be confidently assigned while maintaining high accuracy.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If existing clustering techniques are used, then simplicity is maintained, but measurement precision deteriorates for specimens with uncertain fluorescence intensities

Engineering Contradiction:
Improvesimplicity of clustering processVSAvoidaccuracy of genotype determination
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the genotyping process into distinct stages: initial clustering, confidence scoring, and conditional re-evaluation. This segmentation allows the system to maintain simple clustering for the majority of specimens while applying more sophisticated analysis only where needed. The segmentation resolves the contradiction by isolating the complexity to specific cases rather than applying it universally.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses reference clusters and confidence score calculations as templates that can be applied to multiple specimens. By creating a reusable framework for evaluating clustering confidence and determining when re-evaluation is needed, the system maintains simplicity through standardized procedures while improving precision through systematic application of these copied evaluation patterns.

Inventive Principle:
Principle #26Copying

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The solution enables accurate genotype estimation for specimens with uncertain or missing data by employing algorithms that assess clustering confidence and utilize reference data, improving the reliability of genotype determination and handling ambiguous fluorescence intensities.

Implementation Method 1

unknown base sequence of a specimen is hybridized with a known base sequence near a certain SNP used as a probe to measure the fluorescence intensity

Methodology Applied
Scientific EffectHybridization:

Data Source

PatentUS11355219B2Genotype estimation device, method, and program
Publication Date: 2022.06.07 KK TOSHIBA
  • US11355219B2 patent drawing
  • US11355219B2 patent drawing
  • US11355219B2 patent drawing

AI summary

According to one embodiment, a genotype estimation device includes: an acquirer configured to acquire a clustering strength of genotype data of a plurality of specimens including an unknown specimen whose genotype is not known and known specimens whose genotypes are known; and an estimator configured to estimate the genotype of the unknown specimen on the basis of the genotype data in response to the clustering strength being larger than a first threshold, and output an estimation result.