DASH Model for HLA Loss-of-Heterozygosity Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional techniques are inadequate in accurately detecting loss of heterozygosity in human leukocyte antigen (HLA) alleles, which is a significant cause of immune checkpoint blockade resistance in cancer treatment, due to challenges such as poor alignment of sequence reads, complex sequence variations, and unreliable copy number variant algorithms, especially in samples with low tumor purity.
Innovation Solution
A machine-learning model, termed DASH, is trained using allele-specific, subject-specific, and whole-exome features to detect loss of heterozygosity in HLA alleles, utilizing gradient boosting algorithms to process sequence data and generate accurate predictions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional genome-wide copy number interrogation is used to detect HLA loss of heterozygosity, then the detection can be performed, but the accuracy is poor due to polymorphic nature of mutated genes causing poor alignment and complex sequence variations
Solution Approach 1:
The patent introduces an intermediary alignment process that uses a panel of HLA-specific reference sequences as mediators between the sequencing reads and the final detection. Instead of directly aligning to a single reference genome, the system uses multiple HLA-specific references that account for polymorphic variations, thereby improving alignment accuracy and subsequent loss of heterozygosity detection
Solution Approach 2:
The patent creates a composite reference system by combining multiple HLA-specific reference sequences into an alignment panel. This composite reference structure accommodates the diversity of HLA alleles and enables more accurate alignment of sequencing reads from samples with various HLA genotypes, resolving the alignment difficulties caused by polymorphic nature
2Reliability
If conventional copy number variant algorithms are used after alignment to HLA allele-specific reference sequences, then the analysis can be performed, but the sensitivity is poor for biological samples with low tumor purity and subclonal deletions
Solution Approach 1:
The patent performs preliminary normalization of read counts using control samples before applying copy number variant algorithms. This preliminary action establishes a baseline that accounts for technical variations and sample-specific factors, enabling the subsequent analysis to more sensitively detect subtle changes associated with low tumor purity and subclonal deletions
Solution Approach 2:
The system implements feedback by comparing detected copy number changes against control samples and using this information to refine the detection of HLA loss of heterozygosity. The control samples provide feedback on normal variation, allowing the algorithm to distinguish true biological signals from technical noise, thereby improving sensitivity in challenging samples
3Ease of operation
If conventional techniques rely on deletions of flanking regions as a proxy for HLA loss of heterozygosity, then the detection can be performed, but the specificity is reduced because it cannot identify the specific HLA allele that has been deleted
Solution Approach 1:
The patent segments the analysis into two distinct components: first, alignment to HLA-specific reference sequences that preserve allele identity information, and second, copy number analysis that detects loss of heterozygosity. This segmentation allows the system to maintain allele-specific information throughout the analysis pipeline rather than relying on proxy markers from flanking regions
Solution Approach 2:
The patent adds a new dimension to the analysis by incorporating allele-specific copy number information. Instead of relying solely on flanking region deletions (one-dimensional approach), the system analyzes copy number changes within the HLA allele sequences themselves, providing both detection of loss of heterozygosity and identification of the specific affected allele
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method of detecting loss of heterozygosity in HLA alleles is provided. The method can include accessing a trained machine-learning model, which was trained using a training data set that included at least a training data set that includes an adjusted B allele frequency that represents a ratio between a first B allele frequency of heterozygous alleles in the tumor sample that correspond to the genomic region and a second B allele frequency of heterozygous alleles in the genomic region and associated with one or more control samples. The method can also include using the machine-learning model to generate a result corresponding to a probability of whether a loss of heterozygosity exists in an HLA allele identified in the biological sample of the particular subject by processing the sequence data using the machine-learning model.