LOH Region Detection Model Using Low-Depth Genome Sequencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting Loss of Heterozygosity (LOH) regions are time-consuming and labor-intensive, requiring lengthy analysis steps and high-throughput technologies like Chromosome Microarray (CMA) that are costly and inefficient.
Innovation Solution
A method involving low-depth genome sequencing data derived from high-depth data, utilizing a self-attention allocation algorithm to train a LOH region detection model, determining information values and feature matrices to iteratively improve detection efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high-depth genome sequencing data and complex high-throughput methods like CMA are used to detect LOH regions, then detection accuracy is improved, but detection efficiency deteriorates due to lengthy analysis steps and high time cost
Solution Approach 1:
The patent changes the sequencing depth parameter from high-depth to low-depth genome sequencing, fundamentally altering the input data characteristics. This parameter change enables the use of simplified analysis algorithms while maintaining acceptable detection accuracy, thereby resolving the contradiction between accuracy and efficiency
Solution Approach 2:
The patent replaces complex mechanical/high-throughput detection systems (CMA technology with multiple analysis steps) with a computational model-based approach. The LOH region detection model substitutes the need for complex wet-lab procedures and manual analysis, achieving both accuracy and efficiency improvements
2Measurement precision
If high-depth genome sequencing data is used for LOH region detection, then detection accuracy is improved, but time cost increases due to lengthy analysis steps
Solution Approach 1:
The patent performs preliminary action by pre-training the LOH region detection model using high-depth sequencing data to learn accurate detection patterns. Once trained, the model can quickly process low-depth sequencing data without requiring lengthy analysis steps, thus reducing time loss while maintaining accuracy
Solution Approach 2:
The patent creates a computational model that copies and encodes the detection knowledge from high-depth sequencing analysis. This model copy can then rapidly process new data without repeating the lengthy analysis steps of traditional methods, significantly reducing analysis time
3Measurement precision
If complex high-throughput methods are used for LOH region detection, then detection accuracy is improved, but labor cost increases
Solution Approach 1:
The patent implements self-service by enabling the LOH region detection model to automatically process sequencing data and identify LOH regions without requiring manual consultation of literature or databases. The model serves itself by making autonomous detection decisions, eliminating labor-intensive steps while maintaining accuracy
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
Embodiments of the present disclosure provide a method for training a loss of heterozygosity region detection model, an apparatus, a device and a medium, relating to the technical field of computers. The method includes: acquiring first data; determining information values of each region of N regions in a chromosome corresponding to the first data, wherein the information values include mean value information and variance information of the above regions; the N regions including heterozygous regions and the LOH regions, where N is a positive integer; determining feature information of the N regions, and determining an input feature matrix based on the feature information of the N regions and the information values of the N regions, wherein the feature information includes the information of the LOH regions corresponding to the N regions and an association relationship between the information values of each region; and performing iterative training on an initial model based on the input feature matrix to obtain a LOH region detection model, wherein the initial model is determined based on a self-attention allocation algorithm.