Method for detecting neutral heterozygosity deletion variation and calculating variation proportion and application
By calculating the BAF and RD values of SNP loci using whole-genome sequencing technology, the CN-LOH interval can be determined. This solves the problem of insufficient sensitivity in the detection of neutral heterozygous loss variants in hematologic malignancies in existing technologies, and achieves efficient detection of CN-LOH variant ratios, supplementing the detection results of traditional cytogenetic methods.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-03
AI Technical Summary
Existing chromosome karyotype analysis and high-throughput sequencing technologies suffer from insufficient sensitivity, high cost, and limited applicability in detecting loss of heterozygosity variants and their proportions in hematologic malignancies, particularly in FFPE samples and lymphoma detection.
By employing whole-genome sequencing technology, sequencing coverage information of sample DNA is obtained, BAF and RD values of SNP sites are calculated, maximum likelihood values are calculated using beta distribution, chromosomal intervals are determined to be CN-LOH, and the proportion of variation is calculated, thus achieving accurate detection of neutral heterozygous deletion variants.
It enables accurate detection of CN-LOH variants of 10Mb or more and 10% or more in hematologic malignancies, providing more comprehensive diagnostic and treatment information, reducing the false positive rate, and is applicable to various high-throughput sequencing data.
Smart Images

Figure CN121789760A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of detecting neutral heterozygous deletion variants and their proportions in hematological malignancies, specifically relating to a method and application for detecting neutral heterozygous deletion variants and calculating their proportions. Background Technology
[0002] Hematologic malignancies are malignant clonal proliferative diseases caused by the impaired differentiation and development of hematopoietic stem cells. These include acute and chronic leukemia, lymphoma, and multiple myeloma, among others. Their pathogenesis and clinical characteristics are highly heterogeneous, and prognoses vary considerably. Therefore, clinical diagnosis and treatment require the integration of a range of clinical data, primarily relying on integrated MICM diagnostic methods. MICM integrated diagnosis refers to a comprehensive approach to diagnosing hematologic malignancies using morphology and cytochemistry (M), immunology (I), cytogenetics (C), and molecular biology (M). Because each method has its own characteristics, integrated diagnosis can provide a more comprehensive and accurate diagnosis, classification, and prognosis of hematologic malignancies, as well as monitor treatment efficacy, thereby better guiding treatment.
[0003] Cytogenetics (C), a key evidence-providing tool in integrated diagnostic MICM (Mixed Microbiome Therapy), reveals chromosomal abnormalities by culturing cells, examining them under a microscope, and analyzing the number, morphology, and structure of chromosomes. However, this technique requires viable cells that are actively dividing to produce analyzable karyotypes. During culture, cells may exhibit dominant or absent growth, resulting in unrecoverable karyotypes. Furthermore, some tumor cells, particularly indolent lymphomas, plasma cell tumors, and myelodysplastic syndromes (MDS), are less prone to division than normal cells, often producing karyotypes that resemble those of normal cells. Therefore, seemingly normal chromosomal reports can be misleading. Additionally, some minute chromosomal abnormalities are difficult for technicians to identify through karyotype analysis alone, leading to misdiagnosis, missed diagnosis, or complete unrecognization. This necessitates extensive knowledge and experience from karyotype analysts.
[0004] Given these limitations of cellular karyotype analysis, an increasing number of laboratories are widely using molecular karyotyping (chromosome microarray CMA or high-throughput sequencing NGS) techniques to screen for chromosomal variations to compensate for the shortcomings of cellular karyotype analysis. This is because it does not require cell culture, only a sufficient amount of genomic DNA, and can screen for copy number variations (CNVs) across the entire chromosome and genome, especially microdeletions or microduplications. It can also clearly identify the genomic coordinates of specific variant regions and affected genes, with higher resolution. Most importantly, it can simultaneously detect copy-neutral loss of heterozygosity (CN-LOH) variations that cannot be detected by cellular karyotype analysis.
[0005] CN-LOH indicates an allele imbalance without copy number change, meaning the mutated region is homozygous and the heterozygous genotype has been lost. It is a known tumor escape mechanism that occurs during mitosis, causing the chromosome or a portion of a chromosome to lose a copy from its parent chromosome. The remaining copy then serves as a template for replication, leading to loss of heterozygosity in somatic cells. This results in homozygosity of the pathogenic mutation, causing clonal evolution of tumor cells. CN-LOH is considered a common mechanism for cancer cells to achieve homozygous mutations in oncogenes or tumor suppressor genes. Common CN-LOH mutations in hematologic malignancies include CN-LOH1p, CN-LOH4q, CN-LOH7q, CN-LOH9p, CN-LOH11q, CN-LOH4q, and CN-LOH17p, often accompanied by mutations in genes such as MPL, TET2, EZH2, JAK2, CBL, NF1, and TP53. Currently, the most common clinical techniques for CN-LOH detection are chromosome microarray (CMA) and whole-exome sequencing (WES). Chromosomal microarrays (CMA) were first used for the detection of CN-LOH, analyzing whole-genome allele typing through SNP probe hybridization signals. This technique relies on probe design and distribution, resulting in incomplete probe coverage of parts of the genome, making it difficult to definitively identify abnormalities. It is also unsuitable for FFPE samples, limiting its application in detecting hematologic malignancies, particularly lymphomas. WES, on the other hand, is more commonly used clinically due to its low cost, high sequence coverage, and strong sensitivity and specificity. However, WES's targeted capture principle means it only captures whole-exon regions, and capturing GC-rich regions is relatively difficult, leading to errors in capture efficiency. Therefore, its sensitivity and accuracy in detecting CN-LOH are not as good as CMA, and it cannot report the proportion of CN-LOH variants. This proportion typically indicates the presence of acquired clones in hematologic malignancies, and its level can serve as a monitoring indicator of disease progression and evidence of clonal evolution.
[0006] Therefore, there is an urgent need to provide a molecular karyotype analysis technology based on whole-genome sequencing (WGS) to detect neutral heterozygous deletion variants and their proportions, so as to provide more comprehensive information for the integrated diagnosis and treatment of hematological malignancies. Summary of the Invention
[0007] In view of this, the present invention provides a method that can effectively improve the detection rate of low-frequency variants and reduce the false positive rate, and is applicable to neutral heterozygous deletion variants and variant ratios in hematologic malignancies in various high-throughput sequencing data.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: A method for detecting neutral heterozygous deletion variants and calculating the proportion of variants based on whole-genome sequencing technology includes the following steps: S1. Obtain the raw sequencing data of the sample DNA, filter it, and align it to the reference genome to obtain the sequencing coverage information of each site in the genome; S2. Calculate all SNP sites obtained in the whole genome based on the alignment results, extract SNP site information, filter SNP sites based on the population SNP database to obtain effective SNP sites, and calculate their BAF value based on the Reads alignment information of the effective SNP sites. S3. Calculate the average sequencing depth RD of the whole genome based on the sequencing coverage information obtained from the alignment. all The value represents the expected depth of normal diploid sequencing; S4. Divide each chromosome into uniform, non-overlapping intervals and calculate the RD value of each interval based on the sequencing coverage information of the sites within the interval. S5. Compare the RD value with RD all The values are compared to obtain the ratio rd, which represents the degree of deviation between the interval sequencing depth and the overall sample sequencing depth; S6. Based on BAF values of [0-0.4], [0.4-0.6], and (0.6-1], classify the valid SNP sites into the corresponding three band sets, and calculate their maximum likelihood values by traversing the Reads alignment information of the valid SNP sites in each band set, where l0, l 0.5 l1 and l2 represent the maximum likelihood values of the three strip sets, respectively. S7, based on rd, l0, l 0.5 The l1 value is used to determine whether the interval is CN-LOH, as follows: a. When 0.9 ≤ rd ≤ 1.1, l 0.5 When the value is NaN, if the sum of l0 and l1 is 1, then the interval is CN-LOH; b. When 0.9 ≤ rd ≤ 1.1, l 0.5 When it is not NaN, l0 / l0.5 ≥0.1 and l1 / l 0.5 If ≥0.1, then the interval is CN-LOH; c. If conditions a and b are not met, the interval is not CN-LOH; S8. Starting from the chromosome start position, traverse the interval backward, compare the type and likelihood value of the interval with the adjacent intervals, and merge them to obtain the maximized CN-LOH segment. Then, calculate the variation ratio p of the region based on the maximum likelihood value l0 or l1 in the CN-LOH segment.
[0009] Furthermore, the samples in step S1 include bone marrow and peripheral blood, and the reference genome is the hg38 human reference genome; The specific steps for obtaining sequencing coverage information are as follows: The adapters and low-quality bases in the raw sequencing data were removed. Then, the effective data were aligned to the hg38 human reference genome using alignment software. The generated results were sorted and statistically analyzed to obtain the sorted BAM file. The alignment rate and sequencing depth information were obtained. Furthermore, the filtered SNP sites meet the following conditions: The SNP corresponds to a coverage depth ≥7, a quality value ≥20, and a population frequency ≥0.2; The specific details for obtaining valid SNP sites are as follows: The sorted BAM files were analyzed using variant detection software to obtain SNP site information, which included location, coverage depth, and quality value data.
[0010] Furthermore, the BAF value is calculated based on the ratio of the number of reads supporting alt and ref at the SNP site to the number of reads supporting alt and ref, as shown in the following formula:
[0011] Furthermore, in step S3, RD all The value is obtained by calculating the ratio of the total number of bases in all alignments to the size of the reference genome, as shown in the following formula:
[0012] In the formula, the total number of bases and the size of the reference genome are both in bp.
[0013] Furthermore, in step S4, the interval length is divided into 1kb~1Mb. The RD value of each interval is obtained by calculating the ratio of the sum of the coverage depth values of all sites to the interval length, as shown in the following formula:
[0014] In the formula, n The interval length is...RD i This represents the site coverage depth value.
[0015] Furthermore, the maximum likelihood value is obtained in step S6 as follows: The beta distribution is used to calculate the maximum likelihood of the strip set, ensuring that the likelihood distribution is between 0 and 1. The beta distribution is calculated and normalized after each valid SNP locus is counted. After all valid SNP loci in the strip set have been calculated, the value corresponding to the maximum likelihood in L is extracted as the maximum likelihood value l0 or ln for that strip set. 0.5 Or l1, the calculation formula is as follows: [ , , , ..., 1- ] L=
[0016] In the formula, k is the number of references supported by the site reads, and m is the number of alts supported by the site reads. Let Li-1 be the likelihood distribution of the (i-1)th site.
[0017] In some specific embodiments, preferably, in case b of step S7, l0 / l 0.5 and l1 / l 0.5 The threshold range is 0.1-0.3, including but not limited to 0.1, 0.15, 0.2, 0.25, and 0.3.
[0018] Furthermore, the CN-LOH fragment maximized in step S8 is as follows: The adjacent CN-LOH intervals are merged, and the maximum likelihood value of the merged interval is the maximum likelihood value corresponding to the product of the likelihood matrices of each cell. The variation ratio p ranges from 1 to 100%. When p ≥ 90%, the variation is determined to be homozygous; when p < 90%, it is determined to be heterozygous. The calculation formula is as follows:
[0019] The above method is used to detect the loss of heterozygosity variants and their proportions in hematologic malignancies.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: The detection method provided by this invention does not require complex operations such as cell culture. It only requires whole-genome sequencing of whole-genome DNA from various hematologic malignancy samples (including bone marrow / peripheral blood, fresh tissue, and paraffin-embedded tissue) to establish an analytical workflow for detecting neutral heterozygous deletion variants (CN-LOH) and their proportions. This workflow can accurately detect CN-LOH variants of 10 Mb or more and those exceeding 10%, representing the first time that whole-genome sequencing technology has been used to detect this specific variant in hematologic malignancies. This achievement can effectively supplement traditional cytogenetic testing results and provide more comprehensive information for the integrated diagnosis and treatment of hematologic malignancies. Attached Figure Description
[0021] Figure 1 This is a flowchart of the technical process of Embodiment 1 of the present invention.
[0022] Figure 2 This is the sequencing coverage information obtained by comparing the original data of AML patients in Example 1 of the present invention with the hg38 human reference genome after filtering.
[0023] Figure 3 In Embodiment 1 of the present invention, based on the original data of AML patients processed according to rd, l0, l 0.5 l1 determines whether the interval is a CN-LOH result.
[0024] Figure 4 This is the result of maximizing the CN-LOH fragment based on the original data of AML patients in Embodiment 1 of the present invention.
[0025] Figure 5 This is the SNP subtyping map (BAF map) that maximizes CN-LOH based on the original data of AML patients in Embodiment 1 of the present invention.
[0026] Figure 6 In Embodiment 2 of the present invention, based on the processed raw data of MDS patients, according to rd, l0, l 0.5 l1 determines whether the interval is a CN-LOH result.
[0027] Figure 7 This refers to the maximized CN-LOH fragment result obtained from the original data of MDS patients in Embodiment 2 of the present invention.
[0028] Figure 8 This is the SNP subtyping map (BAF map) of the maximized CN-LOH obtained from the original data of MDS patients in Embodiment 2 of the present invention. Detailed Implementation
[0029] The present invention will be further described in detail below with reference to specific embodiments, so that those skilled in the art can more clearly understand the present invention. Unless otherwise specified, the technical means used in the following embodiments are all conventional means well known to those skilled in the art, and all reagents and consumables are commercially available products.
[0030] Example 1 This embodiment uses a patient clinically diagnosed with acute myeloid leukemia (AML) as the research subject. This subject had previously been found to have CN-LOH22q using chromosome microarray technology (CMA). Based on this case, this embodiment provides a method for detecting CN-LOH22q variants and their proportions using whole-genome sequencing (WGS) technology (see the technical flowchart). Figure 1 ), as detailed below: 1. Raw DNA data from the patient's bone marrow sample was obtained using whole-genome sequencing. The raw data was filtered to remove adapters and low-quality bases. BWA or Bowtie alignment software was used to align the valid data to the hg38 human reference genome. The generated results were sorted and statistically analyzed to obtain the sorted BAM file. Alignment rate, sequencing depth, and other information were statistically analyzed. The results are as follows: Figure 2 As shown.
[0031] 2. Use GATK or VarDict variant detection software to analyze the sorted BAM files to obtain SNP site information, including location, coverage depth, quality value, and other data.
[0032] 3. Filter SNP sites based on the population SNP database, with the following criteria: coverage depth ≥7, quality value ≥20, and population frequency ≥0.2; obtain the effective SNP sites, and calculate the B allele frequency (BAF value) corresponding to the effective SNP sites.
[0033] The BAF value is calculated based on the ratio of the number of reads supporting alt and ref at the SNP site to the number of reads supporting alt and ref, as shown in the following formula:
[0034] 4. Divide the whole genome into 50Kb intervals, calculate the RD value of each 50Kb interval, and compare it with the RD value of the entire sample. all The values are compared to obtain the rd value.
[0035] Among them, RD all The value is obtained by calculating the ratio of the total number of bases in all alignments to the size of the reference genome, as shown in the following formula:
[0036] In the formula, the total number of bases and the reference genome size are both in bp.
[0037] 5. Collect the set of SNP sites with BAF values in three bands within the 50Kb interval: [0-0.4], [0.4-0.6], and (0.6-1]. Calculate the maximum likelihood values l0 and l1 for each band. 0.5 、l1.
[0038] The maximum likelihood value is obtained as follows: The beta distribution is used to calculate the maximum likelihood of the strip set, ensuring that the likelihood distribution is between 0 and 1. The beta distribution is calculated and normalized after each valid SNP locus is counted. After all valid SNP loci in the strip set have been calculated, the value corresponding to the maximum likelihood in L is extracted as the maximum likelihood value l0 or ln for that strip set. 0.5 Or l1, the calculation formula is as follows: [ , , , ..., 1- ] L=
[0039] In the formula, k is the number of references supported by the site reads, and m is the number of alts supported by the site reads. Let Li-1 be the likelihood distribution of the (i-1)th site.
[0040] 6. Based on rd, l0, l 0.5 l1 determines whether the interval is CN-LOH. The criteria for this determination are as follows: a. When 0.9 ≤ rd ≤ 1.1, l 0.5 When the value is NaN, if the sum of l0 and l1 is 1, then the interval is CN-LOH; b. When 0.9 ≤ rd ≤ 1.1, l 0.5 For non-NaN, 0.1 ≥ l0 / l 0.5 ≥0.3 and 0.1≥l1 / l 0.5 If ≥0.3, then the interval is CN-LOH; c. If conditions a and b are not met, the interval is not CN-LOH.
[0041] The specific results are as follows: It was found that the long arm of chromosome 22 has a ratio of 0.9 ≤ rd ≤ 1.1. 0.5 If the value is NaN and the sum of l0 and l1 is 1, then there is a CN-LOH. There are a total of 37 segments (e.g., Figure 3(As shown). A sliding window was used to merge adjacent regions. It was determined that q11.21q13.33 on the long arm of chromosome 22 could be merged into a large CN-LOH fragment, with a merged segment size of 32.20 Mb, as shown below. Figure 4 And from the BAF distribution of 22q (e.g. Figure 5 As shown in the figure, 22q11.21q13.33 is a CN-LOH variant, and it can be determined that it is probably a homozygous variant.
[0042] 7. Finally, based on the maximum likelihood value of l0 or l1 in the interval CN-LOH22q11.21q13.33, the variation ratio of this region is calculated as p = 1 - 2 × l0 or p = 2 × l1 - 1, and the variation ratio of CN-LOH22q11.21q13.33 is approximately 100%.
[0043] Example 2 Furthermore, to further verify the accuracy of the detection method of this application, this embodiment also uses patients clinically diagnosed with myelodysplastic tumors / syndromes (MDS) as research subjects. These research subjects were previously detected with CN-LOH22q using chromosome microarray technology (CMA). The specific detection method is basically the same as in Example 1, except that the source of the original data is changed to patients with myelodysplastic tumors / syndromes (MDS), while everything else remains the same.
[0044] The specific test results are as follows: According to rd, l0, l 0.5 The l1 function was used to determine if the interval was a CN-LOH. Similarly, it was found that on the long arm of chromosome 22, 0.9 ≤ rd ≤ 1.1. 0.5 If the value is NaN and the sum of l0 and l1 is 1, then a CN-LOH is determined to exist. There are a total of 23 segments that are determined to be CN-LOH (e.g., Figure 6 (As shown). A sliding window was used to merge adjacent regions. It was determined that q11.21q13.33 on the long arm of chromosome 22 could be merged into a large CN-LOH fragment, with a merged segment size of 32.72 Mb. Figure 7 As shown. And from the BAF distribution of 22q (as shown) Figure 8 As shown in the figure, 22q11.1q13.33 is the entire long arm CN-LOH, and it can be determined that it is probably a chimeric CN-LOH, that is, the segment contains both normal segments and variant segments.
[0045] Furthermore, the variation ratio of the region is calculated based on the maximum likelihood value of l0 or l1 in the CN-LOH22q11.1q13.33 interval, and the variation ratio of CN-LOH22q11.1q13.33 is approximately 25.20%.
[0046] Through the aforementioned series of investigations, the accuracy and effectiveness of the proposed detection method have been demonstrated. It utilizes whole-genome DNA sequencing on samples such as bone marrow / peripheral blood, fresh tissue, and paraffin-embedded tissue, without requiring cell culture, and establishes an analytical workflow for detecting neutral heterozygous deletion variants (CN-LOH) and their proportions. This workflow can accurately detect CN-LOH variants of 10 Mb or more and those exceeding 10%, representing the first time that whole-genome sequencing (WGS) technology has been used to detect this specific variant in hematological malignancies. It can effectively supplement traditional cytogenetic testing results and provide more comprehensive information for the integrated diagnosis and treatment of hematological malignancies.
[0047] Unless otherwise specified, all raw materials used in this invention are existing substances that can be purchased directly from the market.
[0048] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for detecting neutral heterozygous deletion variants and calculating the proportion of variants, characterized in that, Includes the following steps: S1. Obtain the raw sequencing data of the sample DNA, filter it, and align it to the reference genome to obtain the sequencing coverage information of each site in the genome; S2. Calculate all SNP sites obtained in the whole genome based on the alignment results, extract SNP site information, filter SNP sites based on the population SNP database to obtain effective SNP sites, and calculate their BAF value based on the Reads alignment information of the effective SNP sites. S3. Calculate the average sequencing depth RD of the whole genome based on the sequencing coverage information obtained from the alignment. all The value represents the expected depth of normal diploid sequencing; S4. Divide each chromosome into uniform, non-overlapping intervals and calculate the RD value of each interval based on the sequencing coverage information of the sites within the interval. S5. Compare the RD value with RD all The values are compared to obtain the ratio rd, which represents the degree of deviation between the interval sequencing depth and the overall sample sequencing depth; S6. Based on BAF values of [0-0.4], [0.4-0.6], and (0.6-1], classify the valid SNP sites into the corresponding three band sets, and calculate their maximum likelihood values by traversing the Reads alignment information of the valid SNP sites in each band set, where l0, l 0.5 l1 and l2 represent the maximum likelihood values of the three strip sets, respectively. S7, based on rd, l0, l 0.5 The l1 value is used to determine whether the interval is CN-LOH, as follows: a. When 0.9 ≤ rd ≤ 1.1, l 0.5 When the value is NaN, if the sum of l0 and l1 is 1, then the interval is CN-LOH; b. When 0.9 ≤ rd ≤ 1.1, l 0.5 When it is not NaN, l0 / l 0.5 ≥0.1 and l1 / l 0.5 If ≥0.1, then the interval is CN-LOH; c. If conditions a and b are not met, the interval is not CN-LOH; S8. Starting from the chromosome start position, traverse the interval backward, compare the type and likelihood value of the interval with the adjacent intervals, and merge them to obtain the maximized CN-LOH segment. Then, calculate the variation ratio p of the region based on the maximum likelihood value l0 or l1 in the CN-LOH segment.
2. The method according to claim 1, characterized in that, The samples mentioned in step S1 include bone marrow and peripheral blood, and the reference genome is the hg38 human reference genome; The specific steps for obtaining sequencing coverage information are as follows: The adapters and low-quality bases in the raw sequencing data were removed. Then, the valid data were aligned to the hg38 human reference genome using alignment software. The generated results were sorted and statistically analyzed to obtain the sorted BAM file. The alignment rate and sequencing depth information were obtained.
3. The method according to claim 1, characterized in that, The filtered SNP sites meet the following conditions: The SNP corresponds to a coverage depth ≥7, a quality value ≥20, and a population frequency ≥0.2; The specific details for obtaining valid SNP sites are as follows: The sorted BAM files were analyzed using variant detection software to obtain SNP site information, which included location, coverage depth, and quality value data.
4. The method according to claim 3, characterized in that, The BAF value is calculated based on the number of reads supporting alt and ref at the SNP site and the ratio of Nalt to Nref, as shown in the following formula: 。 5. The method according to claim 1, characterized in that, In step S3, RD all The value is obtained by calculating the ratio of the total number of bases in all alignments to the size of the reference genome, as shown in the following formula: ; In the formula, the total number of bases and the size of the reference genome are both in bp.
6. The method according to claim 1, characterized in that, In step S4, the interval length is divided into 1kb~1Mb. The RD value of each interval is obtained by calculating the ratio of the sum of the coverage depth values of all sites to the interval length, as shown in the following formula: ; In the formula, n The interval length is... RD i This represents the site coverage depth value.
7. The method according to claim 1, characterized in that, The maximum likelihood value is obtained in step S6 as follows: The beta distribution is used to calculate the maximum likelihood of the strip set, ensuring that the likelihood distribution is between 0 and 1. The beta distribution is calculated and normalized after each valid SNP locus is counted. After all valid SNP loci in the strip set have been calculated, the value corresponding to the maximum likelihood in L is extracted as the maximum likelihood value l0 or ln for that strip set. 0.5 Or l1, the calculation formula is as follows: [ , , , ..., 1- ]; L= ; In the formula, k is the number of references supported by the site reads, and m is the number of alts supported by the site reads. Let Li-1 be the likelihood distribution of the (i-1)th site.
8. The method according to claim 1, characterized in that, In step S7, case b, l0 / l 0.5 and l1 / l 0.5 The threshold range is 0.1 to 0.
3.
9. The method according to claim 1, characterized in that, The maximized CN-LOH fragment obtained in step S8 is as follows: The adjacent CN-LOH intervals are merged, and the maximum likelihood value of the merged interval is the maximum likelihood value corresponding to the product of the likelihood matrices of each cell. The variation ratio p ranges from 1 to 100%. When p ≥ 90%, the variation is determined to be homozygous; when p < 90%, it is determined to be heterozygous. The calculation formula is as follows: 。 10. The application of the method according to any one of claims 1-9 in detecting the loss of heterozygosity variant and the proportion of variants in hematologic malignancies.