Method for accurately analyzing copy number variation of whole genome sequencing data
By employing multi-level comparison quality assessment and deep standardization methods, the problems of comparison interference and resolution in CNV detection were solved, enabling accurate identification and confidence marking of copy number variations, thereby improving the accuracy and efficiency of detection.
Patent Information
- Application Number
- CN202511567008.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-01-09
AI Technical Summary
Existing technologies for CNV detection suffer from severe alignment interference, limited high resolution, and inaccurate boundary recognition, which affect the accuracy and efficiency of detection.
A multi-level comparison quality assessment, depth standardization, and structural breakpoint identification method is adopted. By dividing 1kb intervals, constructing target interval files, calculating standardization depth, identifying difference multiples and Z-scores, and combining breakpoint comparison information, the boundary of copy number variation is determined.
It improves the sensitivity and specificity of CNV detection, reduces the false positive and false negative rates, and achieves precise localization and confidence marking of copy number variations.
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of bioinformatics analysis and genome data processing, and in particular to a method for accurately analyzing copy number variations in whole-genome sequencing data. Background Technology
[0002] With the rapid development of next-generation sequencing (NGS) technology and its continuous cost reduction, NGS has become an indispensable tool in medical fields such as genetic disease screening, personalized cancer treatment, drug guidance, and prenatal screening. In particular, whole-genome sequencing (WGS), due to its ability to comprehensively cover and detect multiple variant types at once, has been widely used in clinical practice.
[0003] Copy number variations (CNVs) refer to large duplications or deletions of DNA sequences in the genome. These structural variations are widespread in the human genome and have a significant impact on gene expression, chromosome stability, and individual phenotypes. Numerous studies have shown that CNVs are closely related to various genetic diseases, neurodevelopmental disorders, mental illnesses, and malignant tumors. Therefore, developing an accurate and efficient CNV detection method is of great significance for the early detection and molecular subtyping of diseases.
[0004] Currently, CNV analysis based on NGS data mainly employs sequencing depth comparison, which involves calculating the sequencing coverage depth of the target sample across different genomic regions and comparing it with the average depth of a reference sample to identify potential copy number variations. However, existing methods have the following technical problems: 1. Serious alignment interference problem: There are a large number of highly homologous regions in the human genome, such as homologous genes and pseudogenes, which leads to non-unique sequencing reads and thus interferes with the true copy number signal.
[0005] 2. Limited resolution: Traditional methods often use a fixed window to divide intervals. If the window is too large, small CNVs will be missed, and if the window is too small, false positive results will be easily introduced, affecting the efficiency and accuracy of the analysis.
[0006] 3. Inaccurate boundary identification: CNVs identified by mean depth changes often cannot accurately determine the start and end points of the variation, especially for breakpoints that are crucial for clinical interpretation, where existing methods have poor identification capabilities.
[0007] Therefore, there is an urgent need for a novel CNV analysis strategy that combines multi-level comparative quality assessment, deep standardization, and structural breakpoint identification to improve the sensitivity, specificity, and breakpoint localization accuracy of detection, reduce false positive and false negative rates, and meet the needs of clinical diagnosis and scientific research. Summary of the Invention
[0008] To address the shortcomings of existing technologies, the present invention aims to provide a method for accurately analyzing copy number variations in whole-genome sequencing data.
[0009] A method for accurately analyzing copy number variations in whole-genome sequencing data includes the following steps: S1: Divide the reference genome into multiple fixed-length 1kb intervals and generate the first interval file; S2: Extract all exon positions based on gene annotation information and generate a second interval file; S3: Merge the first interval file and the second interval file and sort them according to chromosome position to obtain the target interval file; S4: Perform quality control, alignment, sorting, and duplicate labeling on the sequencing data of the sample to be tested to obtain the alignment file; S5: Under different alignment quality thresholds, calculate the average sequencing depth of the alignment file in each interval of the target interval file, and standardize the average sequencing depth based on the whole genome average depth to obtain a standardized depth file; S6: Based on the standardized depth file of multiple samples, calculate the mean and standard deviation of each target interval to construct a baseline depth file; S7: Compare the standardized depth file of the sample to be analyzed with the baseline depth file, and calculate the difference factor and Z-score for each target interval; S8: Based on the preset difference fold threshold and Z-score threshold, candidate intervals for copy number variation are identified; S9: Extract the alignment information of candidate intervals from the alignment file, count the number of broken alignments and non-broken alignments, and if there is a break point where the proportion of broken alignments exceeds a set threshold, then the break point is used as the boundary position of copy number variation and marked as high confidence. If it does not exceed the set threshold, then the boundary of the candidate interval is used as the variation boundary and marked as low confidence.
[0010] Furthermore, the alignment quality thresholds in step S5 are Q0 and Q20; where Q0 represents alignment results without quality filtering, and Q20 represents the filtering standard of "average base sequencing quality not lower than 20".
[0011] Furthermore, in step S8, the threshold for the difference factor is less than 0.75 or greater than 1.25, and the Z-score threshold is greater than 3.
[0012] Furthermore, in step S9, the break point is the location where the compared read segments show pair separation or break alignment.
[0013] Furthermore, the tools used to process the alignment files in step S4 include, in sequence, the fastp tool for quality control, the bwa mem tool for sequence alignment, samtools for sorting, and sambamba for duplicate marking.
[0014] Furthermore, the standardization process involves dividing the average sequencing depth of each target region by the average sequencing depth of the corresponding whole genome to obtain the relative depth.
[0015] Compared with the prior art, the beneficial effects of the present invention are mainly reflected in the following aspects: (1) Improve detection accuracy: This invention introduces different alignment quality thresholds to perform deep extraction and standardization of sequencing data, which significantly reduces signal interference caused by non-unique alignment of homologous genes and pseudogene regions, and improves the reliability of CNV detection in complex regions.
[0016] (2) Enhanced variant recognition resolution: By combining equal-length intervals and exon sites to construct a target interval system, the system takes into account both the breadth of the genome and the precision of functional regions, effectively identifying copy number variations down to the exon level, and avoiding the problems of missed detection and false detection caused by window settings in the traditional window method.
[0017] (3) Achieve precise location of breakpoints: Based on the candidate variant region, this invention further extracts abnormal comparison information (such as split reads and split pairs), and judges the start and end boundaries of the variant by statistically analyzing the number of break support comparisons, which significantly improves the accuracy of CNV boundary identification and facilitates the interpretation of clinical pathogenicity.
[0018] (4) Constructing a baseline model to improve specificity: By constructing a baseline data model using the mean and standard deviation of the standardized depth of multiple samples, CNV identification becomes more statistically robust. Combined with the Z-score and fold difference dual indicators, the specificity of variant screening is improved and false positives are reduced.
[0019] (5) Automatic assignment of confidence level to identification results: The present invention automatically marks the mutation results as high confidence or low confidence based on the support strength of the breakpoint, which helps downstream clinical or research personnel to perform stratified processing and accurate judgment of mutation results. Detailed Implementation
[0020] The present invention will now be described in detail with reference to the embodiments.
[0021] Example 1 This embodiment aims to illustrate how to perform copy number variation (CNV) analysis on high-throughput whole-genome sequencing data of a clinical sample based on the method of the present invention, and obtain high-confidence CNV results.
[0022] I. Preparation of Raw Data We selected WGS data from a clinical sample (patient Zhang, female, 7 years old at the time of consultation, clinical phenotype: onset at age 6, severe hearing loss in the left ear, profound hearing loss in the right ear, slurred speech, and bilateral strabismus. A pathogenic heterozygous deletion of chr6:1477648-2176231 was detected using this method). The sequencing platform was MGISEQ-2000, with a read length of 150 bp and a sequencing depth of approximately 30×. Simultaneously, 20 sex-matched normal population samples were selected as a reference cohort. All samples were from the same platform and underwent quality control screening.
[0023] II. Data Processing and Comparison 1. Use fastp (v0.20.1) to perform quality control on the raw FASTQ data, remove low-quality bases and adapter contaminants, and obtain the Clean.fastq file; 2. Use bwa mem (v0.7.17) to align Clean.fastq to the human reference genome (GRCh37) and generate the alignment result Sample.sam; 3. Sort Sample.sam using samtools sort, then use sambamba markdup (v0.8.2) to remove PCR duplicates, obtaining the final alignment file Sample.bam; 4. Perform the same steps as above on the reference samples to generate comparison files Reference1.bam to Reference20.bam respectively.
[0024] III. Constructing the target interval file 1. Divide the human genome (GRCh37) into 1kb units and generate 1kb.bed files; 2. Extract all exon regions from the GTF format gene annotation file and generate exome.bed; 3. Merge 1kb.bed and exome.bed and sort them by chromosome and starting coordinates to generate the target interval file target.bed.
[0025] IV. Calculate the standardization depth and establish a reference baseline 1. At alignment quality thresholds Q0 and Q20, respectively, calculate the average sequencing depth of Sample.bam in each target.bed interval; 2. The depth of each region is standardized using the genome-wide average depth of Sample.bam, generating a standardized depth file Sample.norm.bed; 3. Perform the same processing on the reference samples to generate Reference1.norm.bed to Reference20.norm.bed; 4. Calculate the relative mean depth and standard deviation of each interval in the 20 reference samples, and construct the reference depth baseline file reference.bed.
[0026] V. Identification and Statistics of Difference Intervals 1. Compare Sample.norm.bed with reference.bed; 2. Calculate the relative depth difference factor (FC) and Z-score for each interval; 3. Select the intervals with FC < 0.75 or > 1.25 and Z-score > 3 as candidate copy number variation regions.
[0027] VI. Breakpoint Identification and Confidence Assessment 1. For each candidate variant region, call samtools view in Sample.bam to extract the alignment information of ±10 kb upstream and downstream of that region; 2. Use split reads and split pair information to determine if there are any breakpoints, and count the number of supported alignments and the number of non-breakpoint alignments. 3. If the support ratio of the breakpoint exceeds 20%, the variant region is marked as High Confidence (HC), and the variant boundary is updated to the breakpoint location; 4. If the percentage does not exceed 20%, the original candidate region boundary is retained and marked as Low Confidence (LC).
[0028] VII. Output of Analysis Results The final output file includes: Sample.norm.bed: Standardization depth of each interval of the sample; Candidate_CNV.bed: List of candidate CNV regions; CNV_HC.bed and CNV_LC.bed: High-confidence and low-confidence CNV result files, containing mutation start and end sites, Z-score, fold change, breakpoint information, etc.
[0029] VIII. Case Results In this embodiment, a total of 22 candidate CNV segments were detected, including 12 high-confidence CNV regions (9 of which were duplicates and 3 were deletions), 1 deletion that fell in a clinically relevant gene region, and 10 low-confidence regions for subsequent manual review and verification.
[0030] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A method for accurately analyzing copy number variations in whole-genome sequencing data, characterized in that, Includes the following steps: S1: Divide the reference genome into multiple fixed-length 1kb intervals and generate the first interval file; S2: Extract all exon positions based on gene annotation information and generate a second interval file; S3: Merge the first interval file and the second interval file and sort them according to chromosome position to obtain the target interval file; S4: Perform quality control, alignment, sorting, and duplicate labeling on the sequencing data of the sample to be tested to obtain the alignment file; S5: Under different alignment quality thresholds, calculate the average sequencing depth of the alignment file in each interval of the target interval file, and standardize the average sequencing depth based on the whole genome average depth to obtain a standardized depth file; S6: Based on the standardized depth file of multiple samples, calculate the mean and standard deviation of each target interval to construct a baseline depth file; S7: Compare the standardized depth file of the sample to be analyzed with the baseline depth file, and calculate the difference factor and Z-score for each target interval; S8: Based on the preset difference fold threshold and Z-score threshold, candidate intervals for copy number variation are identified; S9: Extract the alignment information of candidate intervals from the alignment file, count the number of broken alignments and non-broken alignments, and if there is a break point where the proportion of broken alignments exceeds a set threshold, then the break point is used as the boundary position of copy number variation and marked as high confidence. If it does not exceed the set threshold, then the boundary of the candidate interval is used as the variation boundary and marked as low confidence.
2. The method for accurately analyzing copy number variations in whole-genome sequencing data according to claim 1, characterized in that, The alignment quality thresholds in step S5 are Q0 and Q20; where Q0 represents alignment results without quality filtering, and Q20 represents the filtering standard of "average base sequencing quality not lower than 20".
3. The method for accurately analyzing copy number variations in whole-genome sequencing data according to claim 1, characterized in that, In step S8, the threshold for the difference factor is less than 0.75 or greater than 1.25, and the Z-score threshold is greater than 3.
4. The method for accurately analyzing copy number variations in whole-genome sequencing data according to claim 1, characterized in that, In step S9, the break point is the location where the compared read segments show either pairing separation or break alignment.
5. The method for accurately analyzing copy number variations in whole-genome sequencing data according to claim 1, characterized in that, The tools used to process the alignment files in step S4 include, in order: fastp for quality control, bwa mem for sequence alignment, samtools for sorting, and sambamba for duplicate marking.
6. The method for accurately analyzing copy number variations in whole-genome sequencing data according to claim 1, characterized in that, The standardization process involves dividing the average sequencing depth of each target region by the average sequencing depth of the corresponding whole genome to obtain the relative depth.
Citation Information
Patent Citations
Method for detecting mitochondrial copy number variation
CN117727364A
Variant detection using improved sequence data alignments
WO2025217057A1