HLA Genotyping Using Read Equivalence Grouping and Score Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sequencing systems face challenges in accurately determining human leukocyte antigen (HLA) alleles due to high variation and homology, resulting in imperfect accuracy and the need for multiple assays, which increases computational burden and time.
Innovation Solution
The HLA-aware sequencing system employs alignment-score-based filtering and read-support-equivalence grouping of nucleotide reads, using an expectation-maximization algorithm to determine HLA alleles at one or two-field resolution, improving accuracy and reducing the need for additional assays.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing sequencing systems use conventional alignment methods to determine HLA alleles, then the process can be completed with standard computational resources, but the genotyping accuracy is limited to approximately 88-94% due to high HLA-allele variation and homology
Solution Approach 1:
The patent segments the HLA genotyping process into multiple specialized stages: initial alignment with a reference genome, extraction of HLA-corresponding reads, realignment with HLA-allele-reference sequences, and filtering based on alignment scores. This segmentation allows each stage to be optimized independently, improving overall accuracy while managing computational complexity through progressive refinement rather than attempting to solve the entire problem in one step.
Solution Approach 2:
The patent performs preliminary actions by first aligning reads with a reference genome to identify HLA-corresponding reads before proceeding to the more computationally intensive realignment with HLA-allele-reference sequences. This preliminary filtering step reduces the dataset size and focuses computational resources on the most relevant reads, enabling higher accuracy without proportionally increasing overall computational burden.
2Reliability
If existing sequencing systems perform multiple assays to compensate for limited accuracy, then confidence in HLA genotyping increases, but the computational burden and time required increase significantly
Solution Approach 1:
The patent changes key parameters of the alignment process by using HLA-allele-reference sequences with known high-variation regions instead of a generic reference genome. This parameter change in the reference sequence quality allows a single assay to achieve the reliability that previously required multiple assays, reducing time loss while maintaining or improving confidence in results.
Solution Approach 2:
The patent substitutes the mechanical approach of running multiple separate sequencing assays with a refined computational alignment process using specialized HLA-allele-reference sequences. This replacement achieves similar or superior reliability through improved data analysis methodology rather than repeated experimental runs, significantly reducing the time required.
3Measurement precision
If existing sequencing systems determine HLA alleles at two-field resolution, then the genotyping process is simpler and faster, but the accuracy is insufficient for clinical applications where single-nucleobase differences matter
Solution Approach 1:
The patent implements a dynamic resolution approach where the system can operate at different levels of detail depending on needs. The realignment process with HLA-allele-reference sequences enables full-resolution genotyping when required, while the filtering based on alignment scores allows the system to efficiently handle datasets. This dynamic capability provides high accuracy for critical cases without sacrificing overall productivity.
Solution Approach 2:
The patent performs preliminary filtering of HLA-corresponding reads using alignment score thresholds before proceeding to detailed realignment and genotyping. This preliminary action separates clearly identifiable reads from ambiguous ones, allowing the system to process high-confidence cases quickly at full resolution while managing computational resources efficiently, thus maintaining both accuracy and productivity.
Data Source
AI summary
This disclosure describes methods, non-transitory-computer readable media, and systems that can accurately genotype one or more human leukocyte antigen (HLA) alleles from a genomic sample by using alignment-score-based filtering and read-support-equivalence grouping of reads for genotype inference. To genotype HLA alleles, the disclosed systems extract a genomic sample's reads corresponding to an HLA genomic region and align the extracted reads with HLA-allele-reference sequences. The disclosed systems further select a subset of read alignments for the extracted reads based on alignment scores for alignments between the extracted reads and the HLA-allele-reference sequences. Based on the selected subset of read alignments, the disclosed systems group individual reads into HLA equivalence classes and determine candidate HLA alleles for the genome sample at one or more HLA loci. From among the candidate HLA alleles, the disclosed systems determine genotype calls that a genomic sample includes particular HLA alleles at one or more HLA loci.


