A method for detecting TP53 gene heterozygous deletion and copy number variation based on targeted sequencing

By combining targeted sequencing with customized capture and an improved Hidden Markov Model, we have achieved simultaneous and accurate detection of TP53 gene heterozygosity loss and copy number variation, solving the problems of high detection cost and long cycle in existing technologies, and providing efficient and accurate prognostic assessment and treatment decision-making basis.

CN121281628BActive Publication Date: 2026-05-01SHANGHAI TISSUEBANK GENE TECH CO LTD +3
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI TISSUEBANK GENE TECH CO LTD
Filing Date
2025-12-09
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies are insufficient for accurately detecting loss of heterozygosity and copy number variations of the TP53 gene at low cost and within a short timeframe, which affects prognostic assessment and treatment decisions for patients with myeloid tumors such as AML/MDS.

Method used

A customized capture panel was used for targeted capture library construction. Combined with high-throughput sequencing and an improved hidden Markov model, allele frequencies were analyzed by nuclear density estimation to achieve simultaneous and accurate detection of pathogenic mutations, copy number variations, and LOH/CN-LOH in the TP53 gene.

Benefits of technology

It achieves accurate determination of TP53 gene multiple hit status, shortens the detection cycle by 50%, reduces costs by 60%-70%, and has a 100% concordance rate with whole exome sequencing, making it suitable for routine clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121281628B_ABST
    Figure CN121281628B_ABST
Patent Text Reader

Abstract

The application is suitable for the field of bioinformatics and molecular diagnosis technology, and provides a TP53 gene heterozygous deletion and copy number variation detection method based on targeted sequencing. Through core technical innovations such as customized probe design, dynamic reference set correction, improved HMM model and KDE bimodal judgment, the application realizes the synchronous and accurate detection of TP53 gene LOH / CNV and mutation. Compared with traditional detection technology, the application effectively solves the problems of CN-LOH missed detection and insufficient integrated analysis of mutations and structural variations, and the coincidence rate with whole exon sequencing reaches 100%. The method has the outstanding advantages of high sensitivity, high specificity, rapid efficiency and controllable cost, can accurately determine the TP53 multiple hit state, and provides reliable molecular diagnosis basis for risk stratification, prognosis evaluation and individualized treatment decision of AML / MDS and other myeloid tumor patients, and is more suitable for clinical routine popularization and application.
Need to check novelty before this filing date? Find Prior Art

Description

A method for detecting loss of heterozygosity and copy number variation in the TP53 gene based on targeted sequencing Technical Field

[0001] This application belongs to the field of bioinformatics and molecular diagnostics technology, and in particular relates to a method for detecting TP53 gene loss of heterozygosity (LOH), copy number neutral loss of heterozygosity (CN-LOH) and copy number variation (CNV) based on targeted high-throughput sequencing, as well as the corresponding customized capture panel and judgment system. Background Technology

[0002] The TP53 gene, located on the short arm of human chromosome 17 (17p13.1), encodes the p53 protein, a classic tumor suppressor gene that plays a crucial role in cell cycle regulation, DNA damage repair, and apoptosis. Under normal conditions, p53 protein levels are low, but it is activated when cells are subjected to stress, promoting cell damage repair or apoptosis to maintain genomic stability. Loss of function of the TP53 gene is closely associated with the development of various malignant tumors, particularly myeloid leukemia (AML) and myelodysplastic syndromes (MDS), where the mutation rate is approximately 10%-15% and 10%, respectively. In treatment-related myeloid tumors, the mutation rate can reach as high as 35%, often accompanied by complex karyotypes and extremely poor prognosis. In myeloid tumors, pathogenic variations of TP53 are primarily manifested as missense mutations, nonsense mutations, frameshift insertions / deletions, and splicing site variations. According to the "double-hit" theory, biallelic loss of function of TP53 can be achieved through multiple mechanisms, including deletion of chromosome 17p (LOH), a new pathogenic mutation in the second allele, or copy number neutral heterozygous deletion (CN-LOH). CN-LOH is more common in TP53 mutant AML / MDS, leading to the paired presence of mutant alleles in the genome, significantly amplifying their pathogenic effect. Clinical studies show that compared to monoallelic mutations, patients with biallelic loss (multiple-hit) TP53 mutations have a very poor prognosis and a significantly shortened median survival.

[0003] Existing methods for detecting TP53 gene status mainly include conventional cytogenetic and fluorescence in situ hybridization (FISH), targeted NGS detection based on multi-gene panels, whole exome sequencing (WES) or whole genome sequencing (WGS), and SNP microarrays. However, these methods all have significant drawbacks, such as the lack of CN-LOH detection, inconsistent judgment criteria, limitations in detection cost and cycle time, and insufficient integration of mutations and structural variations. Specifically, conventional detection procedures are difficult to reliably resolve CN-LOH, and are prone to missed detection in patients with normal karyotypes or without obvious chromosomal deletions, leading to misjudgment of allele status. The method of inferring CN-LOH based on variant allele frequency (VAF) is not accurate enough under different sequencing depths and tumor purity, and is prone to classification errors. WES / WGS is costly and has a long analysis cycle (10-12 days), making it unsuitable for routine clinical testing. Most methods lack a combined analysis process of mutation information, copy number variations, and CN-LOH, making it impossible to accurately determine multiple hit status, affecting prognostic assessment and treatment decisions.

[0004] Therefore, there is an urgent need for a novel detection method that can accurately detect TP53 mutations, copy number variations, and CN-LOH simultaneously at a lower cost and in a shorter timeframe, in order to improve the accuracy of stratified management and prognostic assessment for patients with myeloid tumors such as AML / MDS. Summary of the Invention

[0005] The purpose of this application is to provide a method for detecting TP53 gene heterozygosity loss and copy number variation based on targeted sequencing, aiming to achieve simultaneous and accurate detection of TP53 pathogenic mutations, copy number variations, and LOH / CN-LOH, shorten the detection cycle, reduce detection costs, and provide reliable multiple hit status assessment results for clinical use.

[0006] The embodiments of this application are implemented as follows: a method for detecting TP53 gene heterozygosity loss and copy number variation based on targeted sequencing includes the following steps:

[0007] S1: Sample preparation and genomic DNA extraction: Genomic DNA is extracted from bone marrow or peripheral blood samples.

[0008] S2: Targeted capture and library construction sequencing: A customized capture panel is used for targeted capture library construction. The customized capture panel covers all exons of TP53 and specific flanking regions (e.g., 100kb to 1Mb), and multiple preset reference regions are set throughout the genome. After library construction, a high-throughput sequencing platform is used for sequencing to obtain sequencing data of the target region.

[0009] S3: Data preprocessing and variant detection: The sequencing data obtained in step S2 is subjected to quality control filtering. The cleaned sequences are aligned to the human reference genome (such as GRCh37 / hg19, GRCh38 or its updated version), SNV / INDEL detection is performed, and pathogenicity annotation is performed in conjunction with clinical databases to screen out pathogenic mutations.

[0010] S4: Coverage Depth Standardization and Copy Number Variation Analysis: A reference sample set with similar sequencing depth distribution to the sample to be tested is dynamically selected from historical samples. The coverage depth of the sample to be tested and the reference sample set are summarized according to the capture probes. After normalization and bias correction by LOESS regression, the copy ratio of each capture probe region is calculated by combining the preset reference region data and GC content. The copy number status is divided into deletion, neutral or increase by introducing an improved hidden Markov model (HMM) with physical distance constraints.

[0011] S5: LOH / CN-LOH determination: High polymorphic SNP sites are screened in a specific range (e.g., 100kb to 1Mb) flanking TP53. Allele frequencies <95% are defined as effective heterozygous sites. After analyzing the allele frequency distribution by nuclear density estimation, the LOH type is determined in combination with the copy number status in step S4.

[0012] S6: Comprehensive determination of multiple hit status: Integrate the pathogenic mutation results of step S3 with the LOH / CN-LOH determination results of step S5. If there is ≥1 pathogenic mutation accompanied by LOH / CN-LOH, it is determined to be a double hit; if there are ≥2 pathogenic mutations, or a single pathogenic mutation accompanied by copy number increase, it is determined to be a multiple hit.

[0013] This application also provides a TP53 gene multiple hit status determination system based on the above method, including:

[0014] Sequencing data input module: Receives the raw sequencing files generated by targeted sequencing;

[0015] Preprocessing module: performs quality control filtering on the raw sequencing files, aligns the cleaned sequences to the GRCh37 / hg19 reference genome, performs single nucleotide variant and insertion / deletion variant detection, and performs pathogenicity annotation in conjunction with clinical databases to screen for pathogenic mutations;

[0016] Copy number status analysis module: Dynamically screens reference sample sets with similar sequencing depth distribution to the sample to be tested, summarizes the coverage depth of the sample to be tested and the reference sample set according to the capture probe, combines the preset reference region data and GC content for normalization and bias correction through LOESS regression, calculates the copy ratio of each capture probe region, and determines the copy number status as missing, neutral or increased by an improved hidden Markov model that introduces physical distance constraints.

[0017] LOH determination module: Screening for highly polymorphic SNP sites in the 500kb flanking region of TP53, defining sites with allele frequencies <95% and simultaneously detecting two alleles as effective heterozygous sites, analyzing the allele frequency distribution of the effective heterozygous sites through nuclear density estimation, and combining the determination results of the copy number status analysis module to complete the determination of loss of heterozygosity and copy number neutral loss of heterozygosity;

[0018] Results output module: Generates a comprehensive report, which includes information on TP53 pathogenic mutation sites, copy number variation status, loss of heterozygosity and copy number neutral loss of heterozygosity types, and the determination of double hit or multiple hit.

[0019] This application utilizes core technologies such as customized capture panel design, dynamic reference set selection, improved HMM segmentation, and KDE bimodal analysis to achieve simultaneous and accurate detection of TP53 pathogenic mutations (SNV / INDEL), copy number variations (CNV), and LOH / CN-LOH. It effectively solves the problems of missed detection of CN-LOH and insufficient integration of mutations and structural variations in existing technologies. The concordance rate with whole-exome sequencing (WES) is 100%, accurately distinguishing between classic LOH, CN-LOH, and LOH-free states. Compared to WES, this application reduces sequencing data volume by approximately 90% (only 0.5-1.0 Gb), shortens the detection cycle by approximately 50% (5-6 days), and reduces the cost per test by 60%-70%, making it more suitable for routine clinical application. Furthermore, it can accurately determine TP53 multiple hit status, providing strong evidence for risk stratification, prognostic assessment, and individualized treatment decisions for patients with myeloid tumors such as AML / MDS. It combines comprehensiveness, accuracy, high efficiency, economy, and clear clinical application value. Attached Figure Description

[0020] Figure 1 is a flowchart of the detection of TP53 gene heterozygosity loss and copy number variation based on targeted sequencing provided in the embodiments of this application;

[0021] Figure 2 shows the TP53 LOH detection results of sample 2024NS2974 provided in the embodiments of this application;

[0022] Figure 3 shows the TP53 LOH detection results of sample 2024NS3992 provided in the embodiments of this application;

[0023] Figure 4 shows the TP53 LOH detection results of sample 2025NS0936 provided in the embodiments of this application;

[0024] Figure 5 shows the TP53 LOH detection results of the 2025NS1197 sample provided in the embodiments of this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0026] To overcome the shortcomings of existing technologies, this application provides a multidimensional data analysis method based on targeted capture-high-throughput sequencing (TargetedNGS), aiming to simultaneously analyze pathogenic mutations (SNV / INDEL), copy number variations (CNV), and loss of heterozygosity events (LOH and CN-LOH) of the TP53 gene. This method solves the technical challenge of conventional panel detection in identifying copy number neutral heterozygous loss of heterozygosity (CN-LOH) through customized probe design, dynamic background modeling, and multi-algorithm fusion judgment. It achieves accurate stratification of TP53 "single-hit / multiple-hit" states, and the detection results are highly consistent with whole-exome sequencing (WES), which can be effectively used for clinical prognostic assessment of myeloid tumors such as AML / MDS. As shown in Figure 1, the core workflow of this method includes: sample and library construction → targeted capture sequencing → quality control and alignment → variant detection → coverage depth standardization and CNV identification → SNP allele ratio distribution (KDE) → LOH / CN-LOH determination → multi-hit comprehensive judgment.

[0027] Specifically, this application provides a method for detecting TP53 gene heterozygosity loss and copy number variation based on targeted sequencing, including the following steps:

[0028] S1: Sample preparation and genomic DNA extraction: Extract genomic DNA from bone marrow or peripheral blood samples to be tested, ensuring a concentration ≥10 ng / μL and an OD260 / 280 ratio of 1.8-2.0;

[0029] S2: Targeted Capture and Library Construction Sequencing: A customized capture panel was used for targeted capture library construction. This panel covered all exons of TP53 and 500 kb flanking regions upstream and downstream, and multiple pre-defined reference regions were set throughout the genome. After library construction, paired-end 150 bp sequencing was performed using the Illumina NovaSeq platform. The average sequencing depth of the target region was ≥1500×, and the Q30 base ratio was ≥85%.

[0030] S3: Data preprocessing and variant detection: The sequencing data obtained in step S2 is filtered for quality control, the cleaned sequences are aligned to the GRCh37 / hg19 reference genome, SNV / INDEL detection is performed, and pathogenicity annotation is performed in conjunction with clinical databases to screen out pathogenic mutations;

[0031] S4: Coverage Depth Standardization and Copy Number Variation Analysis: Reference sample sets with similar sequencing depth distribution to the sample to be tested are dynamically selected from historical samples. The coverage depth of the sample to be tested and the reference sample sets are summarized according to the capture probes. After normalization and bias correction by LOESS regression, the copy ratio of each capture probe region is calculated by combining the preset reference region data and GC content. The copy number status is divided into missing, neutral or increased by introducing an improved hidden Markov model with physical distance constraints.

[0032] S5: LOH / CN-LOH determination: High polymorphic SNP sites are screened in the 500 kb region flanking TP53. Allele frequencies <95% are defined as effective heterozygous sites. After analyzing the allele frequency distribution by nuclear density estimation, the LOH type is determined in combination with the copy number status in step S4.

[0033] S6: Comprehensive determination of multiple hit status: Integrate the pathogenic mutation results of step S3 with the LOH / CN-LOH determination results of step S5. If there is ≥1 pathogenic mutation accompanied by LOH / CN-LOH, it is determined to be a double hit; if there are ≥2 pathogenic mutations, or a single pathogenic mutation accompanied by copy number increase, it is determined to be a multiple hit.

[0034] The screening criteria for the preset reference region in step S2 are as follows: the copy number variation rate is <0.1% in the 1,000 genome sequencing data, the GC content is distributed between 40% and 60%, there are no known highly repetitive sequences in the region, and the preset reference region is distributed on autosomes other than chromosome 17.

[0035] In step S2, the probe length of the customized capture panel is 120 bp, the hybridization capture conditions are a hybridization temperature of 65℃ and a hybridization time of 16 hours, and the washing is performed with a 2×SSC+0.1% SDS washing buffer preheated to 65℃.

[0036] In step S4, the screening criteria for the reference sample set is a Pearson correlation coefficient ≥ 0.95, and the number of samples screened is controlled between 5 and 20. If there are fewer than 5 samples that meet the criteria, the 5 samples with the highest correlation coefficients are selected.

[0037] In step S4, the improved Hidden Markov Model introduces a physical distance constraint. The state transition probability P is related to the physical distance d between adjacent exons or probes, where d is in kilobases. The formula for calculating the state transition probability is P = 1 / (1 + d / 1000). The copy number state is divided into deletion, neutrality, or increase. The deletion corresponds to a log2 copy ratio < -0.3, the neutrality corresponds to -0.3 ≤ log2 copy ratio ≤ 0.3, and the increase corresponds to a log2 copy ratio > 0.3.

[0038] Step S4 also includes quality control and correction steps: marking or removing low coverage segments and abnormal fluctuation segments; when the sample to be tested is judged to be missing in a certain segment while the reference sample generally shows an increase in the same segment, the result verification mechanism is triggered.

[0039] Specifically, step S5 can be as follows: Screening for highly polymorphic SNP sites within a specific region (e.g., a 500 kb region) flanking TP53, where the minimum allele frequency of the highly polymorphic SNP sites is 40%-60%, and defining sites with an allele frequency <95% as effective heterozygous sites; analyzing the allele frequency distribution of effective heterozygous sites using nuclear density estimation; if a unimodal distribution with a peak value between 0.45 and 0.55, it is determined to be without LOH; if a bimodal distribution with a center distance ≥0.2 and a height difference ≤2 times the standard deviation of the overall distribution of allele frequencies of effective heterozygous sites, it is determined to have an LOH signal; combining the copy number status from step S4 to determine the LOH type: when the copy number is in a deletion state, it is determined to be a classic LOH; when the copy number is in a neutral state, it is determined to be a CN-LOH.

[0040] In step S5, the screening of effective heterozygous sites must also meet the following requirements: both alleles must be detected simultaneously, and the site must be used to exclude interference from homozygous sites and homozygous mutations.

[0041] The customized capture panel, used to implement the above detection method, includes the following core components: 1. Target exon region: covering the full-length exons (Exon 1 to Exon 11) and intron regions of the TP53 gene; 2. Flanking LOH analysis region: covering the genomic region from 500 kb upstream of the transcription start site to 500 kb downstream of the transcription termination site of the TP53 gene, containing ≥50 highly polymorphic SNP sites with a minimum allele frequency (MAF) of 40%-60%, ensuring high-density probe coverage; 3. Preset reference region: screening 30-50 copy number-stable regions across the entire genome (excluding chromosome 17), with the screening criteria being: copy number variation rate <0.1% and GC content distribution between 40%-60% in the 1000-person genome sequencing data. There are no known highly repetitive sequences between or within the regions; example coordinates of reference regions include chr1:9777016-9777166, chr3:37048481-37048554, chr20:31016127-31016225, etc., each region is 10-20kb in length, covering non-hotspot regions of autosomes.

[0042] This application also provides a TP53 gene multiple hit status determination system, which is used to implement the above detection method and includes the following functional modules:

[0043] Sequencing data input module: Receives raw Fastq format sequencing files generated by targeted sequencing, supporting sequencing data formats from the Illumina NovaSeq platform;

[0044] Preprocessing module: Integrates Fastp, BWA-MEM, Picard, GATK and ANNOVAR tools to automatically perform quality control filtering of sequencing data, reference genome alignment, repetitive sequence removal, SNV / INDEL detection and pathogenicity annotation, and screen out pathogenic mutations;

[0045] Copy number status analysis module: Dynamically screens reference sample sets with similar sequencing depth distribution to the sample to be tested, summarizes the coverage depth of the sample to be tested and the reference sample set according to the capture probe, combines the preset reference region data and GC content for normalization and bias correction through LOESS regression, calculates the copy ratio of each capture probe region, and determines the copy number status as missing, neutral or increased by an improved hidden Markov model that introduces physical distance constraints.

[0046] LOH determination module: Screening for highly polymorphic SNP sites in the 500kb flanking region of TP53, defining sites with allele frequencies <95% and simultaneously detecting two alleles as effective heterozygous sites, analyzing the allele frequency distribution of the effective heterozygous sites through nuclear density estimation, and combining the determination results of the copy number status analysis module to complete the determination of loss of heterozygosity and copy number neutral loss of heterozygosity;

[0047] Results output module: Generates a comprehensive report, which includes information on TP53 pathogenic mutation sites (exon position, nucleotide variation, amino acid variation, VAF value), copy number variation status, loss of heterozygosity and copy number neutral loss of heterozygosity type, as well as the determination of double hit or multiple hit, providing a reference for risk stratification and prognostic assessment of myeloid tumors such as AML / MDS in clinical practice.

[0048] The following detailed description, in conjunction with specific embodiments, elaborates on the proposed method for detecting TP53 gene loss of heterozygosity (LOH) and copy number variation (CNV) based on targeted sequencing. It should be noted that the experimental procedures and parameter settings described in this section are applicable to all embodiments and can ensure complete reproduction by those skilled in the art. This experiment used samples from multiple patients with acute myeloid leukemia (AML) / myelodysplastic syndrome (MDS) as research subjects; the relevant detection steps can be extended to other tumor or clinical sample types.

[0049] I. Sample processing and core detection steps:

[0050] 1. Sample source and DNA extraction:

[0051] Bone marrow or peripheral blood samples were collected from subjects such as AML / MDS patients, and genomic DNA was extracted using a commercial DNA extraction kit (such as the QIAamp DNA Blood Mini Kit). The following quality control standards were strictly adhered to: DNA concentration ≥10 ng / µL, total amount ≥500 ng, and OD260 / 280 ratio between 1.8 and 2.0, ensuring DNA purity and integrity to meet the requirements of subsequent experiments. If the purity of tumor cells in the bone marrow sample is low, it is recommended to simultaneously perform flow cytometry to determine the proportion of primitive cells, providing a reference for subsequent result interpretation.

[0052] 2. Customized targeted panel design:

[0053] To simultaneously achieve the dual objectives of TP53 gene mutation detection and LOH analysis, this method designs a customized capture panel comprising three core components, as follows:

[0054] (1) Target coding region: Full coverage of the full-length exons (Exon 1 to Exon 11) and key intron regions of the TP53 gene to ensure no blind spots in mutation detection.

[0055] (2) Flanking LOH analysis region: Focusing on the 17p13.1 site where the TP53 gene is located, covering the genomic region from 500 kb upstream of its transcription start site to 500 kb downstream of its transcription termination site (based on the GRCh37 / hg19 reference genome, coordinate range approximately chr17:7,000,000-8,100,000). Based on the East Asian population database of the 1000 Genomes Project, high polymorphic SNP sites with a minimum allele frequency (MAF) of 0.4-0.6 were screened in this region, and high-density probe coverage was used to ensure that at least 20 effective heterozygous SNP sites could be stably captured in each sample (see Table 1 for details).

[0056] (3) Internal reference regions: At least 30 copy number-stable regions were selected from the entire genome, excluding chromosome 17, as deep normalized internal references. Reference region selection must meet three core criteria: copy number variation rate <0.1% in the 1000 Genomes Data, GC content between 40% and 60%, and no known highly repetitive sequences within the region to exclude background interference. Typical reference region coordinates are: chr1:9777016-9777166, chr3:37048481-37048554, chr20:31016127-31016225.

[0057] Table 1. Probe position information for TP53 LOH panel:

[0058] ;

[0059] 3. Targeted capture and high-throughput sequencing experiments:

[0060] The experimental procedure is strictly standardized, and the specific steps are as follows:

[0061] (1) Library construction: The extracted high-quality genomic DNA was randomly fragmented into 150-250 bp fragments, and standardized sequencing libraries were constructed by end repair, adding A tails to the 3' end and ligating adapters.

[0062] (2) Hybridization capture: Biotin-labeled custom probes (probe length approximately 120 bp) were used for liquid-phase hybridization with the library. Core reaction conditions: hybridization temperature 65℃, isothermal hybridization for 16 hours; during the washing stage, rigorous washing was performed using 2×SSC + 0.1% SDS buffer preheated to 65℃ to efficiently remove non-specific binding fragments.

[0063] (3) High-throughput sequencing: After elution and recovery of the target fragment, the fragment was amplified and enriched by PCR, and then sequenced at 150 bp (PE150) pairs on the Illumina NovaSeq6000 platform. Sequencing quality control standards: average sequencing depth of TP53 exons and flanking regions ≥1500×, Q30 base ratio ≥85%, to ensure high accuracy and high coverage of sequencing data.

[0064] 4. Data preprocessing and basic variation detection:

[0065] Sequencing data were processed using a standardized bioinformatics workflow: adapter sequence removal and low-quality base trimming were performed on the raw sequencing data (Fastq format) using Fastp software; the clean data were aligned to the human reference genome GRCh37 / hg19 using the BWA-MEM algorithm; single nucleotide variants (SNVs) and insertions / deletions (INDELs) were detected using the GATK HaplotypeCaller workflow, and pathogenicity annotations were performed using authoritative databases such as ClinVar and COSMIC to clarify the clinical significance of the variants.

[0066] 5. Coverage depth standardization and accurate CNV identification:

[0067] To accurately determine copy number variations in TP53 and its flanking regions, a multi-step standardized analysis process was established, the core of which is as follows:

[0068] (1) Construction of dynamic reference set: Establish a candidate database containing no less than 50 healthy control samples (all samples were verified to be free of 17p abnormalities by whole exome sequencing or whole genome sequencing). For each sample to be tested, calculate the Pearson correlation coefficient between its whole genome coverage depth distribution and each sample in the candidate database, and select samples with a correlation coefficient ≥0.95 (5-20 cases) to form a "dynamic reference set"; if there are fewer than 5 samples that meet the criteria, select the 5 cases with the highest correlation coefficient as the reference set to effectively eliminate the batch effect of the experiment.

[0069] (2) Depth standardization and Log2 Ratio calculation: Calculate the standardized coverage depth ratio of the test sample and the dynamic reference set sample in each probe region, and take the logarithm to obtain log2 (copy ratio); at the same time, based on the GC content of the probe region, the log2 value is smoothed and corrected by LOESS regression to eliminate the influence of GC content deviation on the results.

[0070] (3) Improved copy number segmentation of the HMM model: The Hidden Markov Model (HMM) is used to decode the continuous log2 values ​​of the TP53 region and its flanks. The copy number states are defined as follows: Deletion log2 < -0.3, Neutral -0.3 ≤ log2 ≤ 0.3, and Duplication log2 > 0.3. To reduce the misjudgment of fragmentation in long intron regions, a physical distance constraint is introduced to optimize the transition probability: the state transition probability P is related to the physical distance d (in kb) between adjacent exons / probes, and the calculation formula is P = 1 / (1 + d / 1000), which improves the segmentation accuracy.

[0071] 6. LOH / CN-LOH determination and overall conclusion (KDE bimodal analysis model):

[0072] Kernel density estimation (KDE) analysis based on SNP allele frequency (VAF) combined with CNV detection results achieves accurate differentiation between LOH and copy number neutral LOH (CN-LOH). The specific process is as follows:

[0073] (1) Screening of effective heterozygous sites: VAF values ​​of all SNP sites were extracted in the 500 kb flanking region upstream and downstream of TP53. Homozygous sites with VAF ≥ 95% or VAF ≤ 5% were removed, and effective heterozygous sites were retained for subsequent analysis.

[0074] (2) KDE analysis and peak determination: Gaussian kernel density estimation is performed on the VAF distribution of effective heterozygous sites to generate VAF distribution curves. The distribution peak is identified by the local maximum detection algorithm. The determination criteria are as follows:

[0075] No LOH (negative): The curve shows a single-peak distribution, with the peak value located in the range of 0.45-0.55 (close to 0.5), and the AIC value of the single-peak fitting is ≥10 lower than that of the double-peak fitting;

[0076] LOH signal present (positive): The curve shows a clear bimodal distribution, with the two peaks located on both sides of 0.5 (such as near 0.2 and 0.8), and the distance between the peak centers is ≥0.2, and the difference in peak height is ≤2 times the standard deviation (suggesting two segregated allele populations).

[0077] (3) Comprehensive judgment logic:

[0078] Classical LOH: KDE bimodal distribution + CNV analysis showed that the TP53 region was "missing";

[0079] CN-LOH: KDE bimodal distribution + CNV analysis showed that the TP53 region was in a "neutral" state.

[0080] 7. Layered status for multiple strikes:

[0081] Based on the above mutation detection, CNV, and LOH analysis results, the samples were divided into two categories:

[0082] Single Hit: Only one pathogenic TP53 mutation was detected, with no LOH / CN-LOH and no TP53 copy number deletion.

[0083] A multi-hit is defined as one of the following: (a) ≥2 pathogenic TP53 mutations are detected (must be confirmed to be located in different alleles, or VAF > 50%); (b) 1 pathogenic TP53 mutation is detected, accompanied by 17p deletion (LOH); (c) 1 pathogenic TP53 mutation is detected, accompanied by CN-LOH (mutant VAF is usually significantly higher than 50%, often close to homozygous).

[0084] II. Verification Results and Performance Advantages:

[0085] 1. Clinical sample validation results:

[0086] In 14 validation samples, this method detected 7 positive samples (including classic LOH and CN-LOH) and 7 negative samples, which was completely consistent with the results of whole exome sequencing (WES) (see Table 2 for details). The SNP-VAF distribution in the TP53 flanking regions of all positive samples showed a typical bimodal pattern, while the negative samples showed a unimodal or randomly discrete distribution, validating the high accuracy of the method. Typical sample analysis results are as follows:

[0087] (1) Classic LOH positive sample (2024NS2974), Figure 2 shows the TP53 LOH detection results of sample 2024NS2974:

[0088] KDE analysis: The flanking SNPs of TP53 showed a bimodal distribution, with the peak centers located at 0.19 and 0.79, suggesting an allele imbalance;

[0089] CNV analysis: The Log2 Ratio of the TP53 region was significantly low, indicating a decrease in coverage depth, which was determined to be a copy number loss.

[0090] Mutation detection: Exon 4 showed a c.184G>T nonsense mutation, with a VAF of only 3.34%;

[0091] Conclusion: The results conform to the classic LOH (17p deletion) characteristics, confirming that this method can still accurately identify chromosomal structural variations in the context of low tumor burden and subclonal mutations.

[0092] (2) Classic LOH positive sample (2024NS3992), Figure 3 shows the TP53 LOH detection results of sample 2024NS3992:

[0093] KDE analysis: The TP53 flanking SNPs show a bimodal distribution, with the peak centers located at 0.14 and 0.80, indicating a LOH signal;

[0094] CNV analysis: The TP53 region was determined to be a copy number missing region;

[0095] Mutation detection: Exon 7 was found to have a c.700T>A (p.Tyr234Asn) missense mutation, with a VAF as high as 65.08%;

[0096] Conclusion: It conforms to the classic LOH characteristics, and the high VAF indicates homozygosity of the mutant allele, belonging to the TP53 "multiple hit", which is completely consistent with the WES conclusion.

[0097] (3) CN-LOH positive sample (2025NS0936), Figure 4 shows the TP53 LOH detection results of sample 2025NS0936:

[0098] KDE analysis: The TP53 flanking SNPs show a bimodal distribution, with the peak centers located at 0.29 and 0.70, indicating a LOH signal;

[0099] CNV analysis: Copy number in the TP53 region is neutral;

[0100] Mutation detection: Exon 5 was found to have the c.376-1G>A mutation, with a VAF of 45.74%;

[0101] Conclusion: The tumor was classified as CN-LOH. Although the VAF did not reach high frequency due to the influence of tumor purity, the haplotype imbalance characteristics were clear, which was consistent with the WES conclusion.

[0102] (4) Negative sample (2025NS1197), Figure 5 shows the TP53 LOH detection results of sample 2025NS1197:

[0103] Key results: The Log2 Ratio in the TP53 region remained stable around 0 (normal copy number), and the SNP-VAF showed a single-peak distribution with the peak center located at 49.553% (close to 0.5), indicating a balanced allele ratio.

[0104] Conclusion: No LOH / CN-LOH, completely consistent with WES verification results.

[0105] Table 2 Summary of TP53 LOH / CN-LOH detection results:

[0106] ;

[0107] 2. Advantages in testing efficiency and cost:

[0108] Compared to conventional WES technology, this method, relying on the regional focusing characteristics of the TP53 targeting panel, significantly reduces the amount of sequencing data and the computational power consumption for analysis, demonstrating outstanding advantages in detection time-to-analysis (TAT) and cost control (see Table 3 for details): the amount of sequencing data is reduced by 90%, greatly reducing data storage and analysis time; the detection time is shortened by about 50%, meeting the needs of rapid clinical diagnosis; and the cost per test is reduced by 60%-70%, significantly improving the clinical accessibility of the technology.

[0109] Table 3 Comparison of the performance of the method in this embodiment with whole exome sequencing (WES):

[0110] ;

[0111] In summary, this method achieves simultaneous and accurate detection of TP53 gene LOH / CNV and mutations through core technological innovations such as customized probe design, dynamic reference set calibration, improved HMM model, and KDE bimodal determination. It has the advantages of high sensitivity, high specificity, rapid efficiency, and controllable cost, providing reliable molecular diagnostic evidence for the clinical diagnosis and treatment of diseases such as AML / MDS.

[0112] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for detecting TP53 gene heterozygosity loss and copy number variation based on targeted sequencing, characterized in that, Includes the following steps: S1: Sample preparation and genomic DNA extraction: Genomic DNA is extracted from bone marrow or peripheral blood samples. S2: Targeted capture and library construction sequencing: A customized capture panel is used for targeted capture library construction. The customized capture panel covers all exons of TP53 and 500 kb flanking regions upstream and downstream, and sets multiple preset reference regions throughout the whole genome. After library construction, high-throughput sequencing was performed using a high-throughput sequencing platform to obtain sequencing data for the target region. S3: Data Preprocessing and Variant Detection: The sequencing data obtained in step S2 is quality-controlled and filtered. The cleaned sequences are aligned to the human reference genome, and SNV / INDEL detection is performed. Pathogenicity annotation is performed in conjunction with clinical databases to screen for pathogenic mutations. S4: Coverage Depth Normalization and Copy Number Variation Analysis: A reference sample set with a Pearson correlation coefficient ≥0.95 with the sequencing depth distribution of the sample to be tested is dynamically screened from historical samples. The number of samples screened is controlled between 5 and 20. If there are fewer than 5 samples that meet the criteria, the 5 samples with the highest correlation coefficients are selected. The coverage depth of the sample to be tested and the reference sample set is summarized according to the capture probes. After normalization and bias correction by LOESS regression based on the preset reference region data and GC content, the copy ratio of each capture probe region is calculated. The copy number status is classified as deletion, neutral, or increase by introducing an improved Hidden Markov model with physical distance constraints. S5: LOH / CN-LOH Determination: Highly polymorphic SNPs are screened in the flanking regions of 500kb upstream and downstream of TP53. Loci are defined as effective heterozygous loci with allele frequencies <95%. After analyzing the allele frequency distribution by nuclear density estimation, the LOH type is determined in combination with the copy number status in step S4. S6: Comprehensive determination of multiple hit status: Integrate the pathogenic mutation results of step S3 with the LOH / CN-LOH determination results of step S5. If there is ≥1 pathogenic mutation accompanied by LOH / CN-LOH, it is determined to be a double hit; if there are ≥2 pathogenic mutations, or a single pathogenic mutation accompanied by copy number increase, it is determined to be a multiple hit.

2. The method for detecting TP53 gene heterozygosity loss and copy number variation based on targeted sequencing according to claim 1, characterized in that, The screening criteria for the preset reference region in step S2 are: copy number variation rate <0.1% in the 1,000 genome sequencing data, GC content between 40% and 60%, no known highly repetitive sequences in the region, and the preset reference region is distributed on autosomes other than chromosome 17.

3. The method for detecting TP53 gene heterozygosity loss and copy number variation based on targeted sequencing according to claim 1, characterized in that, The probe length of the customized capture panel described in step S2 is 120 bp. The hybridization capture conditions are a hybridization temperature of 65°C and a hybridization time of 16 hours. During washing, a rigorous washing process is performed using 2×SSC + 0.1% SDS washing buffer preheated to 65°C.

4. The method for detecting TP53 gene heterozygosity loss and copy number variation based on targeted sequencing according to claim 1, characterized in that, In step S4, the improved Hidden Markov Model introduces physical distance constraints. The state transition probability P is related to the physical distance d between adjacent exons or probes, where d is in kilobases. The formula for calculating the state transition probability is P = 1 / (1 + d / 1000). The copy number state is divided into deletion, neutrality, or increase. The deletion corresponds to a log2 copy ratio < -0.3, the neutrality corresponds to -0.3 ≤ log2 copy ratio ≤ 0.3, and the increase corresponds to a log2 copy ratio > 0.

3.

5. The method for detecting TP53 gene heterozygosity loss and copy number variation based on targeted sequencing according to claim 1, characterized in that, Step S4 also includes quality control and correction steps: marking or removing low coverage segments and abnormal fluctuation segments; when the test sample is judged to be missing in a certain segment while the reference sample generally shows an increase in the same segment, the result verification mechanism is triggered.

6. The method for detecting TP53 gene heterozygosity loss and copy number variation based on targeted sequencing according to claim 1, characterized in that, Step S5 specifically involves: screening for highly polymorphic SNP sites within 500kb flanking regions upstream and downstream of TP53. The minimum allele frequency of these highly polymorphic SNP sites is 40%-60%, and sites with an allele frequency <95% are defined as effective heterozygous sites. The allele frequency distribution of effective heterozygous sites is analyzed using nuclear density estimation. If the distribution is unimodal with a peak value between 0.45 and 0.55, it is determined to be without LOH. If the distribution is bimodal with a center distance ≥0.2 and a height difference ≤2 times the standard deviation of the overall allele frequency distribution of effective heterozygous sites, it is determined to have an LOH signal. The LOH type is determined in conjunction with the copy number status from step S4: a copy number deletion indicates a classic LOH; a copy number neutral indicates a CN-LOH.

7. The method for detecting TP53 gene heterozygosity loss and copy number variation based on targeted sequencing according to claim 1, characterized in that, The screening of effective heterozygous sites in step S5 must also meet the following requirements: both alleles are detected simultaneously, and the site is used to exclude interference from homozygous sites and homozygous mutations.

8. The method for detecting TP53 gene heterozygosity loss and copy number variation based on targeted sequencing according to claim 1, characterized in that, The customized capture panel includes: a region covering all 11 exons of the TP53 gene; a flanking region covering 500 kb upstream and 500 kb downstream of the TP53 gene, containing ≥50 highly polymorphic SNP sites with a minimum allele frequency of 40%-60%; and 30-50 pre-defined reference regions throughout the genome, each 10-20 kb in length, covering non-hotspot regions on autosomes.

9. A TP53 gene multiple hit status determination system based on the method of any one of claims 1-8, characterized in that, include: Sequencing data input module: Receives the raw sequencing files generated by targeted sequencing; Preprocessing module: performs quality control filtering on the raw sequencing files, aligns the cleaned sequences to the GRCh37 / hg19 reference genome, performs single nucleotide variant and insertion / deletion variant detection, and performs pathogenicity annotation in conjunction with clinical databases to screen for pathogenic mutations; Copy number status analysis module: Dynamically screens a reference sample set with a Pearson correlation coefficient ≥0.95 between the sequencing depth distribution of the test samples and the target samples. The number of samples screened is controlled between 5 and 20. If fewer than 5 samples meet the criteria, the 5 samples with the highest correlation coefficients are selected. The coverage depth of the test samples and the reference sample set is summarized according to the capture probes. Combined with the preset reference region data and GC content, normalization and bias correction are performed using LOESS regression. After calculating the copy ratio of each capture probe region, the copy number status (deletion, neutral, or increase) is determined by an improved Hidden Markov model that incorporates physical distance constraints. LOH determination module: Screens for highly polymorphic SNP sites in the 500kb flanking region of TP53, with allele frequencies <95%. The site where two alleles are detected simultaneously is designated as a valid heterozygous site. The allele frequency distribution of the valid heterozygous site is analyzed by nuclear density estimation. The determination of loss of heterozygosity and copy number neutral loss of heterozygosity is completed by combining the determination results of the copy number status analysis module. The result output module generates a comprehensive report, which includes TP53 pathogenic mutation site information, copy number variation status, loss of heterozygosity and copy number neutral loss of heterozygosity type, and the determination result of double hit or multiple hit.

Citation Information

Patent Citations

  • Gene homozygosis and heterozygosis deletion detection method and system based on NGS platform

    CN115631788A

  • Kit, library construction method thereof and target variation detection and analysis method

    CN119552946A