Genotype Inference Using Machine Learning and Parental Haplotypes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing genotyping technologies face challenges in accurately inferring the complete genome of offspring from ultra-low coverage sequencing data, particularly in preimplantation genetic testing, due to limitations in detecting rare or structural variants and handling minimal biological samples.

Innovation Solution

The use of machine learning models, including statistical models, in conjunction with known parental genetic information, to generate comprehensive genetic information for offspring by determining a probability distribution based on offspring genotype data and parental haplotypes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If ultra-low coverage sequencing is used to analyze offspring genetic material, then the amount of genetic material required is minimized, but the measurement precision and reliability of genotype inference deteriorates

Engineering Contradiction:
Improveamount of genetic materialVSAvoidgenotype inference accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent introduces machine learning models as an intermediary computational layer that processes ultra-low coverage sequencing data combined with parental genotype information. The model acts as a mediator that infers complete offspring genotypes by learning from the relationship between limited observed data and expected genetic inheritance patterns, thereby recovering measurement precision despite minimal input material

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the problem by changing from direct genotype calling to probabilistic inference. Instead of attempting to directly determine genotypes from limited reads, the system uses machine learning to model the probability distribution of possible genotypes given the observed data and parental information, fundamentally changing the approach from deterministic to probabilistic parameter estimation

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If whole genome sequencing is used to achieve comprehensive genomic coverage, then the completeness of genetic information is improved, but the cost and data processing requirements increase

Engineering Contradiction:
Improvecompleteness of genetic informationVSAvoiddata processing requirements
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by incorporating parental genotype data before analyzing offspring data. The machine learning model is pre-trained or pre-configured with parental genetic information, allowing it to make informed inferences about offspring genotypes from minimal sequencing data, thereby achieving comprehensive genetic information without requiring complete sequencing of the offspring genome

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses parental genome data as a template or copy to infer offspring genotypes. Instead of sequencing the entire offspring genome, the system copies relevant genetic information from parental data and uses machine learning to determine which parental alleles were inherited at each position, significantly reducing sequencing requirements while maintaining information completeness

Inventive Principle:
Principle #26Copying

3Productivity

If array genotyping is used to efficiently detect specific variants, then the productivity and efficiency are improved, but the ability to detect rare or structural variants deteriorates

Engineering Contradiction:
Improvegenotyping efficiencyVSAvoidvariant detection capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal system that can handle multiple types of genetic data (array genotyping, sequencing data, parental information) and infer all types of genetic variants (common and rare SNPs, structural variants). The machine learning model is designed to process diverse input formats and output comprehensive genotype information, making the system adaptable to different input types while maintaining high efficiency

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent combines multiple data sources (array genotyping data, sequencing data, parental genotypes) into a composite input for the machine learning model. This composite approach leverages the strengths of each data type - the efficiency of array genotyping for common variants and the versatility of sequencing for rare and structural variants - to achieve both high productivity and broad variant detection capability

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS20250034645A1Systems and methods for inferring genotypes of biological samples
Publication Date: 2025.01.30 HERASIGHT INC
  • US20250034645A1 patent drawing
  • US20250034645A1 patent drawing
  • US20250034645A1 patent drawing

AI summary

Disclosed herein are methods, systems, and devices for inferring genetic information, or genotypes, of a subject based on low-coverage, or otherwise incomplete, genotype data from the individual and known genetic information of the subject's parents. In particular, disclosed herein are methods for generating a predicted, comprehensive genome of an offspring regardless of the quality of or gaps in coverage in the individual's data and/or the parental data.