Polyploid Genotyping via Bayesian Cluster Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genotyping methods for polyploid organisms face challenges in accurately assigning genotypes due to continuous signal scores from microarray data, which are difficult to convert into discrete classes, especially for tetraploid species like those using Illumina GoldenGateTM and Infinium arrays.
Innovation Solution
A cluster model is developed by identifying active genotypes and fitting functions to signal magnitudes from microarray data, using a training set to adjust cluster positions and assign genotypes based on distance from observed data, incorporating unsupervised and supervised machine learning processes, including maximum likelihood estimation and Bayesian methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If continuous signal scores from microarray data are used for genotyping, then measurement information is preserved, but genotype assignment accuracy deteriorates due to difficulty in converting continuous scores to discrete classes
Solution Approach 1:
The patent transforms the continuous signal space into discrete genotype classes by defining cluster regions with specific boundary parameters. The continuous fluorescence intensity signals are converted to discrete genotype calls by determining which cluster region each signal pair falls into, effectively changing the parameter state from continuous to discrete while preserving the underlying information through probabilistic assignment.
Solution Approach 2:
The patent introduces cluster models as intermediary structures that mediate between continuous microarray signals and discrete genotype assignments. These cluster models serve as a bridge, containing multiple discrete genotype classes with associated parameters that map continuous signal values to discrete genotype categories through defined cluster regions and assignment rules.
2Ease of operation
If cluster models with multiple discrete genotype classes are defined, then genotype assignment becomes possible, but difficulty in detecting and measuring increases due to continuous signal variation
Solution Approach 1:
The patent performs preliminary action by pre-defining cluster models with multiple discrete genotype classes, cluster regions, and assignment rules before actual genotype assignment. The cluster models are constructed in advance with specified parameters including mean signal values, standard deviations, and boundary definitions, enabling straightforward genotype assignment by simply determining which pre-defined cluster region each sample signal falls into.
3Ease of manufacture
If diploid clustering methods are applied to polyploid organisms, then existing algorithms can be used, but genotyping accuracy deteriorates due to limitations in handling polyploid complexity
Solution Approach 1:
The patent creates a universal cluster model framework that can handle both diploid and polyploid organisms through multi-functionality. The model defines genotype classes based on allele dosage counts that can accommodate different ploidy levels, allowing the same algorithmic approach to work across different organism types by simply adjusting the number of expected genotype classes and their corresponding signal parameters.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided are methods, systems, and computer products for genotyping polyploid organisms, as well as diploid organisms. The provided methods use an allele-intensity model to generate cluster definitions. The allele-intensity model relates allele counts of different genotypes to signal intensities generated by the genotyping platform. The model also includes a capability to update cluster positions obtained from a maximum likelihood model using a Bayesian method.