Low-depth whole genome sequencing-based copy number variation detection method, apparatus and device, and storage medium
Through the pre-processing of the dynamic negative reference library and reference genome data, combined with dynamic Z test and multi-dimensional indicators, the copy number variation detection method for low-deep whole genome sequencing is optimized, which solves the problem of low detection accuracy and reliability, and achieves efficient and reliable CNV detection.
Patent Information
- Application Number
- CN202510441873.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
The existing low-deep whole genome sequencing technology has problems with poor detection accuracy and low reliability in copy number variation detection, mainly due to the lack of scientific data processing methods and the complexity and complexity of genetic data, especially in the establishment of negative reference libraries, individual differences were not fully considered.
The sample gene data was preprocessed by using dynamic negative reference library and reference genome. Through sequencing quality control, dynamic Z test and improved copy number variation detection, including data quality control, data preprocessing and CNV detection links, window division strategy and GC content correction were optimized, and CNV filtering and classification were combined with local weighted regression method and multi-dimensional indicators.
It significantly improves the accuracy and sensitivity of detection, reduces systematic deviations, enhances the reliability and automation of detection, and improves the convenience and efficiency of clinical applications through the online CNV analysis platform.
Smart Images

Figure CN120340607A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of medical data processing, and particularly relates to a copy number variation detection method, device, equipment, and storage medium based on low-depth whole-genome sequencing. Background Art
[0002] Birth defects are major public health problems worldwide, imposing a heavy burden on families and society. Research shows that chromosomal copy number variations (CNVs) play an important role in the occurrence of many birth defects. Although traditional detection methods such as fluorescence in situ hybridization and microarray technology are relatively mature, they are difficult to fully meet clinical and research needs due to high costs, complex operations, and limited resolution. Low-pass whole genome sequencing (LP-WGS) is an efficient high-throughput sequencing method characterized by its ability to rapidly sequence the entire genome at a relatively low cost. This technology achieves cost-effective genome analysis by reducing the sequencing coverage, i.e., the average number of sequencing reads per genomic locus, hence the name "low-depth sequencing". This technical strategy significantly reduces sequencing costs (sequencing coverage as low as 0.1 - 1X) while ensuring the acquisition of genomic information, providing a feasible technical solution for large-scale genomic research. An innovative tool based on LP-WGS, copy number variation sequencing (CNV-seq), has significantly improved detection capabilities with its advantages such as high resolution, high efficiency, low cost, and flexibility, especially being outstanding in the analysis of low-copy number regions.
[0003] However, in the actual clinical application of CNV-seq, existing methods still face some problems. Due to the lack of scientific data processing means and the complexity and enormity of genetic data, current CNV detection methods still have problems of poor detection accuracy and low reliability.
[0004] Therefore, there is an urgent need for a scientific copy number variation detection method based on low-depth whole-genome sequencing to solve the problems of the existing technology. Summary of the Invention
[0005] The main objective of this application is to provide a copy number variation detection method, device, equipment, and storage medium based on low-depth whole-genome sequencing, aiming to solve the technical problems of poor detection accuracy and low reliability existing in current CNV detection methods.
[0006] To achieve the above object, the present application proposes a copy number variation detection method based on low-depth whole-genome sequencing, which includes: performing sequencing quality control on sample gene data to obtain filtered gene data;
[0007] Performing data preprocessing on the filtered gene data through a dynamic negative reference library and a reference genome to obtain processed gene data and dynamic Z-test results; the dynamic negative reference library is determined according to the GC content of the comparative negative sample set;
[0008] Performing improved copy number variation detection based on the processed gene data and the dynamic Z-test results to obtain copy number variation detection results.
[0009] In one embodiment, the step of performing data preprocessing on the filtered gene data through a dynamic negative reference library and a reference genome to obtain processed gene data and dynamic Z-test results includes:
[0010] Performing sequence alignment on the filtered gene data based on the reference genome to obtain normalized gene data; the reference genome is segmented according to a set window size, and multiple overlapping windows are defined by means of a sliding window, and the step size of the sliding window is smaller than the set window size;
[0011] Performing genome window division and counting on the normalized gene data according to the divided windows of the reference genome to obtain window gene data; the number of reads in each window of the window gene data is greater than the minimum ratio of the coverage of the corresponding region in the reference genome, and the N-base ratio corresponding to each window in the reference genome is less than the N-base ratio threshold;
[0012] Performing data correction on the window gene data through the locally weighted regression method to obtain processed gene data;
[0013] Performing a Z-test on the dynamic negative reference library and the processed gene data to obtain dynamic Z-test results.
[0014] In one embodiment, before performing a Z-test on the dynamic negative reference library and the processed gene data to obtain dynamic Z-test results, it further includes:
[0015] Setting a shared threshold range according to the GC content of the sample gene data and a preset floating range;
[0016] Setting a gradually expanding unit based on the GC content of the sample gene data and a preset screening step size;
[0017] Generating a test GC content corresponding to the sample gene data according to the shared threshold range and the gradually expanding unit;
[0018] Evaluate the coefficient of variation of the control negative sample set through the GC content of the said test, and generate a dynamic negative reference library according to the evaluation result.
[0019] In one embodiment, the step of performing improved copy number variation detection based on the processed gene data and the dynamic Z-test result to obtain a copy number variation detection result includes:
[0020] Perform copy number status marking based on the processed gene data and the dynamic Z-test result to obtain marked gene data;
[0021] Classify the copy number variation types of the marked gene data through the marked Z-test result corresponding to the marked gene data and the variation statistical count to obtain classified gene data;
[0022] Perform variation detection on the chromosome based on the classified gene data to generate a copy number variation detection result.
[0023] In one embodiment, the step of performing copy number status marking based on the processed gene data and the dynamic Z-test result to obtain marked gene data includes:
[0024] Determine the target repeat region through the dynamic Z-test result and the continuous peak of the processed gene data to obtain the target repeat region;
[0025] Determine the target deletion region through the dynamic Z-test result and the continuous valley of the processed gene data to obtain the target deletion region;
[0026] Perform status marking on the processed gene data based on the target repeat region and the target deletion region to obtain marked gene data.
[0027] In one embodiment, the step of classifying the copy number variation types of the marked gene data through the marked Z-test result corresponding to the marked gene data and the variation statistical count to obtain classified gene data includes:
[0028] Standardize the marked Z-test result corresponding to the marked gene data through the variation statistical count corresponding to the marked gene data to obtain a standard Z-test result;
[0029] Determine the variation statistical ratio according to the positive and negative directions of the standard Z-test result and the variation statistical count;
[0030] Classify the copy number variation types based on the standard Z-test result and the variation statistical ratio to obtain classified gene data.
[0031] In one embodiment, the step of performing variant detection on a chromosome based on the classified gene data to generate a copy number variant detection result includes:
[0032] Determining the variant type of the chromosome based on the variant window data ratio corresponding to the classified gene data;
[0033] Calculating the copy number variant chimeric ratio by modifying the ratio calculation formula and the copy number corresponding to the classified gene data;
[0034] Generating a copy number variant detection result according to a preset database, the classified gene data, the variant type of the chromosome, and the copy number variant chimeric ratio.
[0035] In addition, to achieve the above object, the present application also proposes a copy number variant detection device based on low-depth whole-genome sequencing. The copy number variant detection device based on low-depth whole-genome sequencing includes:
[0036] A data quality control module for performing sequencing quality control on sample gene data to obtain filtered gene data;
[0037] A data preprocessing module for preprocessing the filtered gene data through a dynamic negative reference library and a reference genome to obtain processed gene data and a dynamic Z-test result;
[0038] A copy number variant detection module for performing improved copy number variant detection based on the processed gene data and the dynamic Z-test result to obtain a copy number variant detection result.
[0039] In addition, to achieve the above object, the present application also proposes a copy number variant detection device based on low-depth whole-genome sequencing. The device includes: a memory, a processor, and a copy number variant detection program based on low-depth whole-genome sequencing stored on the memory and executable on the processor. The copy number variant detection program based on low-depth whole-genome sequencing is configured to implement the steps of the copy number variant detection method based on low-depth whole-genome sequencing as described above.
[0040] In addition, to achieve the above object, the present application also provides a storage medium. The storage medium is a computer-readable storage medium, and a program for implementing the copy number variant detection method based on low-depth whole-genome sequencing is stored on the computer-readable storage medium. The program for implementing the copy number variant detection method based on low-depth whole-genome sequencing is executed by a processor to implement the steps of the copy number variant detection method based on low-depth whole-genome sequencing as described above.
[0041] The present application provides a method, apparatus, device and storage medium for detecting copy number variations based on low-depth whole-genome sequencing. The method includes: performing sequencing quality control on sample gene data to obtain filtered gene data; performing data preprocessing on the filtered gene data through a dynamic negative reference library and a reference genome to obtain processed gene data and dynamic Z-test results; determining the dynamic negative reference library according to the GC content of the comparison negative sample set; performing improved copy number variation detection based on the processed gene data and the dynamic Z-test results to obtain copy number variation detection results. The present application first performs sequencing quality control to ensure data quality. Then, data preprocessing is performed based on a dynamically constructed dynamic negative reference library. Finally, CNV detection is performed based on the optimized processed gene data and dynamic Z-test results, thereby effectively reducing systematic biases and improving detection reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0044] Figure 1 It is a schematic diagram of the first process of the first embodiment of the method for detecting copy number variations based on low-depth whole-genome sequencing of the present application;
[0045] Figure 2 It is a schematic diagram of the second process of the first embodiment of the method for detecting copy number variations based on low-depth whole-genome sequencing of the present application;
[0046] Figure 3 It is a schematic diagram of the process of the second embodiment of the method for detecting copy number variations based on low-depth whole-genome sequencing of the present application;
[0047] Figure 4 It is a schematic diagram of the brief process of the method for detecting copy number variations based on low-depth whole-genome sequencing of the present application;
[0048] Figure 5 It is a visualization diagram of the sample whole-chromosome CNV results of the method for detecting copy number variations based on low-depth whole-genome sequencing of the present application;
[0049] Figure 6 It is a visualization diagram of the sample single-chromosome CNV results of the method for detecting copy number variations based on low-depth whole-genome sequencing of the present application;
[0050] Figure 7 This is a visualization diagram of the detection results of whole-chromosome CNV for the copy number variation detection method based on low-depth whole-genome sequencing in this application;
[0051] Figure 8 This is a visualization diagram of the detection results of single-chromosome CNV for the copy number variation detection method based on low-depth whole-genome sequencing in this application;
[0052] Figure 9 This is a CNV annotation result diagram and a schematic diagram of the chimeric ratio calculation tool for the copy number variation detection method based on low-depth whole-genome sequencing in this application;
[0053] Figure 10 This is a schematic diagram of the module structure of the copy number variation detection device based on low-depth whole-genome sequencing in the embodiment of this application;
[0054] Figure 11 This is a schematic diagram of the device structure of the hardware operating environment involved in the copy number variation detection method based on low-depth whole-genome sequencing in the embodiment of this application.
[0055] The realization of the purpose, functional features, and advantages of this application will be further described with reference to the accompanying drawings in combination with the embodiments. Detailed Embodiments
[0056] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.
[0057] In order to better understand the technical solutions of this application, the following will be described in detail in combination with the accompanying drawings of the specification and specific embodiments.
[0058] The main solution of this application is: performing sequencing quality control on sample gene data to obtain filtered gene data; performing data preprocessing on the filtered gene data through a dynamic negative reference library and a reference genome to obtain processed gene data and dynamic Z-test results; the dynamic negative reference library is determined according to the GC content of the comparative negative sample set; performing improved copy number variation detection based on the processed gene data and dynamic Z-test results to obtain copy number variation detection results.
[0059] Currently, the bioinformatics detection process of CNV-seq includes multiple links such as data quality control, data preprocessing, and CNV detection. Due to the lack of scientific data processing means, as well as the complexity and enormity of gene data, and in the data preprocessing stage of the existing process, the establishment of the negative reference library does not consider individual differences, resulting in the detection accuracy being affected. The existing CNV detection methods have problems of poor detection accuracy and low reliability.
[0060] To address the above problems, the present application first performs quality control on the sequencing quality of sample gene data to ensure analysis reliability and guarantee data quality. Then, data preprocessing is carried out based on a dynamically constructed negative reference library, which is constructed in real time based on the GC content of the comparison samples. Finally, the present application performs CNV detection based on the optimized processed gene data and dynamic Z-test results. Therefore, the CNV detection results of the present application can effectively reduce systematic biases and improve detection reliability.
[0061] It should be noted that the execution entity of this embodiment can be a copy number variation detection system based on low-depth whole-genome sequencing, or a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or a copy number variation detection device based on low-depth whole-genome sequencing that can implement the above functions. This embodiment does not make specific limitations in this regard. Hereinafter, taking a copy number variation detection device based on low-depth whole-genome sequencing (abbreviated as detection device) as the execution entity as an example, this embodiment and the following embodiments will be described.
[0062] Based on this, the embodiments of the present application provide a copy number variation detection method based on low-depth whole-genome sequencing, referring to Figure 1 , Figure 1 which is the first process schematic diagram of the first embodiment of the copy number variation detection method based on low-depth whole-genome sequencing of the present application.
[0063] In this embodiment, the copy number variation detection method based on low-depth whole-genome sequencing includes steps S10 to S30:
[0064] Step S10, perform quality control on the sequencing quality of the sample gene data to obtain filtered gene data;
[0065] It should be understood that the bioinformatics detection process of the copy number variation detection method (CNV-seq) based on low-depth whole-genome sequencing provided in this embodiment includes multiple links such as data quality control, data preprocessing, and CNV detection. Among them, the data quality control link can comprehensively evaluate the quality of the detection samples in advance to ensure a high-quality data basis for subsequent CNV detection.
[0066] In this embodiment, the process of the above sequencing quality control can be as follows: First, use the high-throughput sequencing data quality assessment tool FastQC to comprehensively count the quality scores, GC content distribution, adapter contamination, repetitive sequences and other indicators of the sample gene data, so as to comprehensively understand the quality status of the original sequencing data. Subsequently, the high-throughput sequencing data quality control tool Trimmomatic can be used to filter the sample gene data, including trimming adapter sequences, removing low-quality bases, and filtering short fragments, to generate higher-quality sequence data, that is, the above-mentioned filtered gene data, laying a reliable foundation for subsequent analysis.
[0067] It is easy to understand that the above specific filtering rules may include: 1) Allow a small number of base mismatches when aligning adapters, and use an appropriate score as the shear threshold for the moving window; 2) Set reasonable maximum and minimum ranges for the alignment length of the adapter sequence; 3) If an adapter sequence is detected, retain the shorter sequence; 4) Remove the bases with lower quality at the beginning and end; 5) Set the sliding window length to several bases, and when the average quality value in the window is lower than a certain standard, shear the window and the subsequent sequence. 6) In this embodiment, only sequences that reach a certain length requirement after shearing may be retained, and sequences shorter than this length will be discarded. It can be understood that in this embodiment, a first quality inspection report corresponding to the sample gene data can be generated and sent to the user in the sequencing quality control link. If the user confirms that the preliminary quality verification of the sample gene data passes based on the first quality inspection report, subsequent data preprocessing and CNV detection can be carried out.
[0068] Step S20, perform data preprocessing on the filtered gene data through a dynamic negative reference library and a reference genome to obtain processed gene data and dynamic Z-test results; the dynamic negative reference library is determined according to the GC content of the comparative negative sample set;
[0069] It can be understood that after confirming that the filtered gene data with controllable quality is obtained, this embodiment can perform data preprocessing on the filtered gene data based on the dynamic negative reference library and the reference genome, providing a good data detection basis for subsequent CNV detection. In a feasible implementation manner, referring to Figure 2 , Figure 2 is the second process schematic diagram of the first embodiment of the copy number variation detection method based on low-depth whole-genome sequencing in this application. In this embodiment, step S20 may include steps A1 to A4:
[0070] Step A1, perform sequence alignment on the filtered gene data based on the reference genome to obtain standardized gene data; the reference genome is segmented according to the set window size, and multiple overlapping windows are defined by means of a sliding window, and the step size of the sliding window is smaller than the set window size;
[0071] It should be noted that in this embodiment, the data preprocessing process may include the sequence alignment, genome window partitioning and counting, data correction, and statistical test steps referred to in the above steps A1 to A4. Exemplarily, in this embodiment, the filtered gene data may first be aligned to the human hg19 reference genome using the BWA software, and the alignment result may be converted into the BAM format. Subsequently, the BAM file is further processed using GATK (Genome Analysis Toolkit) to identify and mark the repetitive sequences generated during the PCR amplification process, that is, multiple reads generated by multiple amplifications of the same DNA fragment, to obtain the normalized gene data. By this method, this embodiment can effectively avoid the bias that the repetitive sequences may cause to the downstream CNV analysis. In addition, after the sequence alignment, this embodiment may also perform alignment quality quality control on the normalized gene data. The alignment quality quality control is the statistics and analysis of the key quality indicators in the normalized gene data, including the number of aligned sequences (Mapped reads), the number of unaligned sequences (Unmapped reads), the alignment rate (Mapped rate(%)), the average sequencing depth of the whole genome (Mean depth of whole genome), the number of uniquely aligned sequences (Unique mapping reads), the duplication rate (Duplication Rate(%)), and the GC content (GC content) and other indicators. Therefore, this embodiment can comprehensively evaluate the quality of the sample gene data by combining the results of the sequencing quality control and the alignment quality quality control, ensuring a high-quality data basis for the subsequent CNV detection.
[0072] Step A2, partitioning and counting the normalized gene data according to the partitioning windows of the reference genome to obtain window gene data; the number of reads in each window of the window gene data is greater than the minimum ratio of the coverage of the corresponding region in the reference genome, and the N base ratio corresponding to each window in the reference genome is less than the N base ratio threshold;
[0073] Furthermore, in this embodiment, genomic window partitioning and counting can be performed on the standardized gene data. First, during the genomic window partitioning and counting process, the reference genome can be segmented according to the set window size, and multiple overlapping windows can be defined by means of a sliding window. Information such as the starting position of each window, GC content (i.e., the proportion of guanine and cytosine in the DNA molecule), and the number of covered bases will be counted. At the same time, the step size of the sliding window is set to be smaller than the set window size, aiming to ensure overlap between windows, increase data coverage, improve analysis accuracy, and avoid missing data, so as to capture the local features of the genome more precisely. Then, within each partitioned window of the reference genome, characteristic indicators of the sample to be detected, such as GC content and the number of reads, can be counted. In this link, strict quality control will be carried out by this module. Among them, when partitioning the genomic window, it is necessary to avoid partitioning into special regions of the chromosome, such as highly repetitive regions, regions close to the centromere and telomere, and homologous regions of the Y chromosome.
[0074] In addition, when performing genomic window counting, it is required that the number of reads in the window is greater than the minimum proportion of the coverage of this region in the reference genome to ensure sufficient data. For example, assume that the average coverage of a certain region in the reference genome is 100×. If the set minimum proportion is 50%, then the number of reads in the current window needs to be at least 100×50% = 50 reads. If there are only 30 reads in this window, it will be filtered out. Thus, it is possible to exclude regions with too low coverage and avoid problems such as false positives or analysis biases (such as SNP detection errors) caused by insufficient data in the follow-up. At the same time, the proportion of N bases (unknown bases, that is, regions where assembly or sequencing fails) in each window of the reference genome must be lower than the set maximum threshold to avoid excessive missing data affecting the analysis result. For example, if the set threshold is 10%, regions where the proportion of N bases in the window exceeds 10% will be excluded. Assume that the length of a certain window is 1000bp, and 150bp of it is N bases, then the proportion of N bases is 15%, and the set maximum threshold is 10%, so the data of this window does not meet the requirements and needs to be filtered. Thus, it is possible to exclude regions with poor quality of the reference genome itself (such as assembly gaps or repetitive regions) and avoid analyzing unreliable genomic positions.
[0075] It is easy to understand that by meeting the above window partitioning conditions, this embodiment can ensure that the analyzed window region has sufficient sequencing depth and avoid accidental errors; at the same time, ensure high quality of the reference genome (less N bases) and avoid interference of the defects of the reference sequence itself on the results.
[0076] Step A3, perform data correction on the window gene data by the locally weighted regression method to obtain processed gene data;
[0077] It is understandable that in the data correction step, the local weighted regression method (LOWESS) can be used in this embodiment to smooth the filtered window gene data. Specifically, in this embodiment, the GC content and the read count of the window gene data can be weighted and smoothed to obtain more stable smoothed data. Then, by comparing the original data with the smoothed data, the deviation of each data point is calculated, and further the read count of each window is corrected. Finally, the processed gene data after correction can more accurately reflect the true trend of the data and remove unnecessary fluctuations.
[0078] Step A4: Perform a Z-test on the dynamic negative reference library and the processed gene data to obtain a dynamic Z-test result.
[0079] It should be understood that in the prior art, when performing data preprocessing, the establishment of the negative reference library did not fully consider the influence of individual GC content differences, resulting in the impact on detection accuracy. Therefore, to improve the detection accuracy, in this embodiment, a matching negative reference library, that is, the above-mentioned dynamic negative reference library, can be dynamically constructed by comparing the GC content of the negative sample set (that is, the gene data set of negative samples (or healthy populations) that do not contain CNV). Therefore, in a feasible implementation manner, in this embodiment, before step A4, steps S00 to S04 are further included:
[0080] Step S00: Set a shared threshold range according to the GC content of the sample gene data and a preset floating range;
[0081] Step S01: Set a gradually expanding unit based on the GC content of the sample gene data and a preset screening step size;
[0082] Step S02: Generate a test GC content corresponding to the sample gene data according to the shared threshold range and the gradually expanding unit;
[0083] Step S03: Evaluate the coefficient of variation of the comparison negative sample set through the test GC content, and generate a dynamic negative reference library according to the evaluation result.
[0084] It should be noted that in CNV-seq, the negative reference library can play a key role in improving the accuracy of CNV detection. In this process, the GC content is a key factor. Selecting samples with similar GC content helps to reduce GC bias and improve the reliability of the analysis results. Therefore, in this embodiment, in this step, by analyzing the GC content percentage of all uniquely mapped reads (that is, the short reads (reads) generated by sequencing that can be uniquely and clearly mapped to specific positions of the reference genome) in the comparison negative sample set, the comparison samples most suitable for the sample gene data can be screened to construct the above-mentioned dynamic negative reference library.
[0085] It is easy to understand that, first, in this embodiment, the GC content of the sample gene data can be used as a benchmark, combined with a preset floating range, such as ±0.02, as the shared threshold range; second, with the GC content of the sample gene data as the center, combined with a preset screening step size, such as ±0.001, as the unit value for gradual expansion; then, by gradually increasing the step size unit value of the GC content until the preset threshold (±0.02) is reached, several test GC contents are generated, and the coefficient of variation (CV) of the selected reference samples, that is, the negative samples without CNV, is calculated after each expansion to evaluate the quality of the negative reference library; finally, this embodiment can select the reference sample set with the lowest CV value and compare it with the negative sample set as the final negative reference library to ensure its high matching and stability with the test sample in terms of GC content. Compared with the existing method of constructing a static reference library, this embodiment fully considers the differences in GC content among different individuals, can effectively improve the optimization degree of the negative reference library, and provides more reliable data support for subsequent analysis.
[0086] Exemplarily, for example, this embodiment can perform CNV-seq analysis on 1 test sample (GC content = 0.415) and screen for matching reference samples from 200 negative reference libraries to construct a dynamic negative reference library. The specific process is as follows:
[0087] First, with the GC content of 0.415 of the test sample as a benchmark, the preset floating range is set to ±0.02 (i.e., 0.395 - 0.435); second, with the GC content of 0.415 of the test sample as the center, the preset screening step size is set to ±0.001 as the unit value for gradual expansion; then, by gradually increasing the step size unit value of the GC content until the preset threshold (±0.02) is reached, and the coefficient of variation (CV) of the selected reference samples is calculated after each expansion to evaluate the quality of the negative reference library.
[0088] 1) In the first round of expansion, screen for GC ∈ [0.414, 0.416] (±0.001), and 15 samples are obtained, with CV = 0.18;
[0089] 2) In the second round of expansion, expand to [0.413, 0.417] (±0.002), and 28 samples are obtained, with CV = 0.15;
[0090] 3) Gradually expand to ±0.01 (GC ∈ [0.405, 0.425]), and 50 samples are obtained, with CV dropping to 0.09;
[0091] 4) Continue to expand to ±0.02, and CV may rise back to 0.12.
[0092] Finally, select the reference sample set with the lowest CV value as the final negative reference library to ensure its high matching and stability with the test samples in terms of GC content. Here, finally, 50 samples with GC ∈ [0.405, 0.425] (CV = 0.09) are selected as the negative reference library. The average GC value of this group of samples is 0.416 (highly matched with the test sample 0.415). It should be understood that after determining the most suitable dynamic negative reference library, this embodiment can implement the processed test samples through the Z-test, that is, the comparison between the above-mentioned processed gene data and the dynamic negative reference library. Exemplarily, this embodiment can standardize through the mean and standard deviation of the window data points and the corresponding negative reference library data points, and the formula is as follows:
[0093]
[0094] Where X is the value of the window data point in the processed gene data, μ is the mean of the dynamic negative reference library data points, and σ is the standard deviation of the dynamic negative reference library data points.
[0095] Through this method, this embodiment can effectively compare the processed gene data with the dynamic negative reference library, determine whether some data points in the sample significantly deviate from the reference data, and then identify potential abnormalities, which is applicable to the downstream CNV analysis.
[0096] Step S30, perform improved copy number variation detection according to the processed gene data and the dynamic Z-test result to obtain a copy number variation detection result.
[0097] It should be understood that this embodiment can subsequently perform multi-dimensional copy number variation detection through the processed gene data and the dynamic Z-test result. The output copy number variation detection result includes, but is not limited to, the name of each chromosome, Z-score, total number of windows, proportion of variant windows, standardized Z-score, preliminary classification result, and copy number status to comprehensively describe the chromosomal variation situation. In addition, this embodiment can perform CNV annotation by combining information from multiple databases. For example, annotate the chromosomal band position through cytoBand to locate the physical position of the CNV; combine the OMIM database to annotate the disease or phenotype information of related genes; query the DECIPHER database to identify whether there are similar structural variations and their phenotypic associations in clinical samples; use the DGV database to determine whether the CNV belongs to normal variations in the population; at the same time, combine the ClinGen database to provide an expert-reviewed assessment of the clinical importance of genes and copy number variations to further support the clinical interpretation of CNV. Combining these information provides location, function, and clinical relevance information for the interpretation of CNV, facilitating subsequent analysis and explanation.
[0098] It should be understood that CNV visualization can also be performed in this embodiment. The detection device can take CNV-related data as input, including the copy number ratio file of the sample, the database annotation file, the blacklist region file, and the line data required for plotting, etc., and generate two types of visualization charts: one is the whole-genome copy number distribution map, which is used to display the copy number distribution of the sample on different chromosomes; the other is the copy number distribution map of a single chromosome, which precisely visualizes the CNV variation events on each chromosome. These charts intuitively present the copy number variation characteristics of each chromosome in the genome, providing strong support for researchers to identify potential copy number abnormal regions, evaluate data quality, and explore the associations of potential genetic variations. It should be noted that in order to further improve the reliability and clinical application value of CNV detection results, this embodiment can also combine an online CNV analysis platform to visually display the above copy number variation detection results. This analysis platform can integrate CNV annotation, visualization analysis, etc. into an online CNV-seq analysis platform, facilitating access at any time and interactive data interpretation, and significantly improving the interpretation efficiency and convenience of CNV.
[0099] In summary, the detection device in this embodiment can integrate an online CNV analysis platform to support whole-genome and single-chromosome level analysis; and integrate authoritative databases such as OMIM, DECIPHER, DGV, and ClinGen to provide information on the physical location, functional impact, and clinical relevance of CNV, helping to facilitate the interpretation of variations. At the same time, the visualization module intuitively displays the distribution characteristics of copy number variations, facilitating a quick assessment of their potential impacts.
[0100] In this embodiment, a scientific CNV-seq detection method is provided. This embodiment fully considers the influence of GC content on the detection results, and in real time matches the optimal reference library according to the GC content of the sample to be tested, so as to reduce systematic bias by dynamically constructing a negative reference library and improve the detection reliability; and optimizes the window division strategy, divides the window by means of a sliding window. At the same time, in order to ensure that there is overlap between windows, increase data coverage, improve the analysis accuracy and avoid missing data, the step size of the sliding window needs to be less than the window size, so that the local characteristics of the genome can be captured more precisely. When dividing the genomic window, it is required that the number of reads in the window is greater than the minimum ratio of the coverage of this region in the reference genome to ensure sufficient data, and the N base ratio of this window in the reference genome must be lower than the set maximum threshold to avoid excessive missing data affecting the analysis results and ensure proper processing of abnormal windows; at the same time, a reasonable gc content correction is performed by the locally weighted regression method during window statistical counting, so that the corrected data can more accurately reflect the true trend of the data, remove unnecessary fluctuations and reduce bias. In addition, this embodiment implements quality control throughout the entire process from sequencing data, alignment data to CNV results through a strict quality control system to ensure the reliability of the analysis. By monitoring key indicators such as alignment rate, GC content, and duplication rate, the data quality is guaranteed. And an efficient automated analysis process is set up. This method adopts modular and parallel analysis to improve the computing efficiency. Combined with the optimization of cloud computing resources, the analysis cycle is shortened, which is suitable for large-scale CNV research. Therefore, compared with the prior art, the CNV-seq detection method proposed in this embodiment can greatly improve the detection accuracy, sensitivity and automation degree, and at the same time integrate an online CNV analysis platform, greatly enhancing the convenience and efficiency of clinical applications.
[0101] This embodiment provides a method for detecting copy number variations based on low-depth whole-genome sequencing. The method includes: performing sequencing quality control on sample gene data to obtain filtered gene data; performing sequence alignment on the filtered gene data based on a reference genome to obtain normalized gene data; segmenting the reference genome according to a set window size, and defining multiple overlapping windows by means of a sliding window, where the step size of the sliding window is smaller than the set window size; performing genome window partitioning and counting on the normalized gene data according to the divided windows of the reference genome to obtain window gene data; the read count of each window in the window gene data is greater than the minimum ratio of the coverage of the corresponding region in the reference genome, and the N base ratio corresponding to each window in the reference genome is less than the N base ratio threshold; performing data correction on the window gene data by means of locally weighted regression to obtain processed gene data; setting a shared threshold range according to the GC content of the sample gene data and a preset floating range; setting a gradually expanding unit based on the GC content of the sample gene data and a preset screening step size; generating a test GC content corresponding to the sample gene data according to the shared threshold range and the gradually expanding unit; evaluating the coefficient of variation of a negative control sample set by means of the test GC content, and generating a dynamic negative reference library according to the evaluation result. Performing a Z-test on the dynamic negative reference library and the processed gene data to obtain a dynamic Z-test result; the dynamic negative reference library is determined according to the GC content of the negative control sample set; performing improved copy number variation detection according to the processed gene data and the dynamic Z-test result to obtain a copy number variation detection result. Compared with the prior art, the CNV-seq detection method proposed in this embodiment can greatly improve the detection accuracy, sensitivity and automation degree, and at the same time integrate an online CNV analysis platform, greatly enhancing the convenience and efficiency of clinical applications.
[0102] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter.
[0103] On the basis of the first embodiment, please refer to Figure 3 , Figure 3 which is a schematic flowchart of the second embodiment of the method for detecting copy number variations based on low-depth whole-genome sequencing of the present application. In this embodiment, step S30 includes steps S31 to S33:
[0104] Step S31, performing copy number status marking according to the processed gene data and the dynamic Z-test result to obtain marked gene data;
[0105] It should be understood that in the existing CNV-seq during the CNV detection stage, the filtering conditions are not optimized enough, resulting in a large number of false positive results, increasing the difficulty of interpretation. At the same time, there is a lack of a scientific calculation method for the CNV chimeric ratio, and the detection limit is ambiguous, affecting the accuracy of the detection results. To solve the above problems, in this embodiment, the filtering conditions can be scientifically set during the improvement of the copy number variation detection process to reduce the false positive rate. At the same time, the calculation method of the chimeric ratio is improved, and the sensitivity and lower limit of the detection are strictly controlled, thereby improving the detection accuracy and clinical application value of CNV-seq. It can be understood that the above copy number status marking can refer to identifying and marking the duplication and deletion regions in the processed gene data. This embodiment can combine the dynamic Z-test results and the gender information in the processed gene data, read the data of each chromosome line by line and perform analysis. At the same time, during this process, basic quality control is required to ensure the integrity and correct format of the data through a scientific CNV filtering strategy, so as to ensure the accuracy of subsequent analysis.
[0106] Therefore, in a feasible implementation manner, in this embodiment, step S31 may include steps B1 to B3:
[0107] Step B1, determining the target duplication region through the dynamic Z-test result and the continuous peak value of the processed gene data, and obtaining the target duplication region;
[0108] Step B2, determining the target deletion region through the dynamic Z-test result and the continuous valley value of the processed gene data, and obtaining the target deletion region;
[0109] Step B3, performing status marking on the processed gene data based on the target duplication region and the target deletion region, and obtaining the marked gene data.
[0110] It should be noted that this embodiment can use the Z-score value combined with the continuous peak value of the sequencing depth and the continuous valley value of the sequencing depth for comprehensive judgment to ensure the accuracy and reliability of subsequent CNV analysis, scientifically set the filtering conditions, and reduce the false positive rate. Exemplarily, for the determination of the duplication region, this embodiment can use the Z-score value, that is, the dynamic Z-test result combined with the continuous peak count method for comprehensive judgment. The specific rules are as follows:
[0111] 1) When the Z-score is greater than or equal to 0.3 and less than 1.8, it is regarded as the starting point of duplication, and it is required that the number of continuous peaks is greater than 1 and the Z-score value is greater than 1.5 before it can be marked as a valid duplication region.
[0112] 2) If the Z-score value is less than 0.3 and the continuous peaks reach more than 3 times, it is marked as the end of the duplication region.
[0113] Among them, the number of consecutive peaks can be the number of window intervals corresponding to consecutive Z-scores ≥ 1.5 in the processed gene data.
[0114] Furthermore, the final determination of the repetitive region also needs to meet the conditions that the length is greater than 100 kb (kilobase) or the consecutive peaks exceed 1 time, and the peak count is greater than 2. For the determination of the deletion region, in this embodiment, a method combining the Z-score value and the consecutive valley count can be used for comprehensive judgment. The specific rules can be as follows:
[0115] 1) When the Z-score is greater than or equal to -1.8 and less than -0.3, it is regarded as the starting point of deletion. Depending on the consecutive valley count, only when the Z-score is less than or equal to -1.5 and the length of the deletion region is greater than 100 kb or the number of consecutive valleys is greater than 1, it is regarded as a valid deletion region.
[0116] 2) If the Z-score value exceeds -0.3, it is marked as the end of the deletion region.
[0117] In summary, this embodiment innovatively designs an algorithm based on multi-dimensional indicators to strictly control the quality of genomic repetitive or deletion regions, thereby ensuring the accuracy and reliability of subsequent CNV analysis.
[0118] Specifically, this embodiment can make a determination by comprehensively considering the following three key indicators: 1) the number of consecutive peaks; 2) the lengths of the duplication and deletion regions; 3) the Z-score value. Only when these three conditions simultaneously meet the above-set threshold conditions can a certain region in the processed gene data be marked as a repetitive or deletion region and included in the subsequent CNV detection and analysis.
[0119] Step S32: Classify the marked gene data according to the marked Z-test result and the mutation statistical count corresponding to the marked gene data to obtain classified gene data;
[0120] It should be understood that after the copy number status is marked, the present embodiment can perform variant statistical counting, that is, recalculate the length, Z-score value, peak value (count_peak), trough value (count_trough), peak top (count_peak_top), and trough top (count_trough_top) count and other characteristics of the repeated or deleted regions corresponding to the marker gene data. Then, by comprehensively analyzing the Z-score value, copy number status (duplication or deletion), and the count characteristics such as count_peak, count_trough, count_peak_top, and count_trough_top of the variant region, the status of the chromosome is determined, and the CNV is accurately classified in multiple dimensions. In a feasible implementation manner, in the present embodiment, step S32 may include steps C1 to C3:
[0121] Step C1, standardize the marker Z-test result corresponding to the marker gene data through the variant statistical counting corresponding to the marker gene data to obtain a standard Z-test result;
[0122] It should be understood that the present embodiment first needs to calculate the standardized Z-score, that is, the standard Z-test result, based on the above-mentioned variant statistical counting. The corresponding formula is as follows:
[0123]
[0124] Among them, Z-score', count_peak, and count_trough respectively represent the Z-score value, peak count, and trough count of the repeated or deleted region recalculated in the previous step.
[0125] Step C2, determine the variant statistical ratio according to the positive and negative directions of the standard Z-test result and the variant statistical counting;
[0126] It is easy to understand that in this embodiment, the peak ratio or valley ratio (percent1) and the ratio of the peak or valley (percent2), that is, the above-mentioned variation statistical ratio, can be calculated according to the positive and negative directions of the standardized Z-score. Exemplarily, it is assumed that the total number of abnormal windows confirmed in the previous step = the number of windows with standardized Z-score ≥ 0.5 + the number of windows with standardized Z-score ≤ -0.5. Then percent1 is: when standardized Z-score <= 0, percent1 = the ratio of the number of windows with standardized Z-score ≤ -0.5 to the total number of abnormal windows; when standardized Z-score > 0, percent1 = the ratio of the number of windows with standardized Z-score ≥ 0.5 to the total number of abnormal windows. percent2 is: when standardized Z-score <= 0, percent2 = the ratio of the number of windows with standardized Z-score ≤ -1 to the total number of abnormal windows; when standardized Z-score > 0, percent2 = the ratio of the number of windows with standardized Z-score ≥ 1 to the total number of abnormal windows.
[0127] Step C3, classify the copy number variation type based on the standard Z-test result and the variation statistical ratio to obtain classified gene data.
[0128] It can be understood that the preliminary classification rule of chromosomes in this embodiment can be determined based on the above-mentioned standardized Z-score (i.e., the standard Z-test result) and the variation statistical ratio (i.e., percent1 and percent2). Exemplarily, when the standardized Z-score reaches a certain positive value and both ratio parameters exceed the specified threshold, for example, when standardized Z-score ≥ 3 and percent1 >= 0.7 and percent2 >= 0.5, it can be determined as "significant amplification"; if only some conditions are met, for example, when standardized Z-score ≥ 3 and percent1 >= 0.7, it is determined as "mild amplification". Similarly, when the standardized Z-score reaches a certain negative value and combined with the ratio parameters and both ratio parameters exceed the specified threshold, it can be determined as "significant deletion"; if only some conditions are met, it is determined as "mild deletion". In other cases, it is marked as "not significantly variant". Therefore, the preliminary classification results included in the classified gene data in this embodiment can be: significant amplification, mild amplification, significant deletion, and not significantly variant. This embodiment can use the standardized Z-score combined with multi-dimensional statistical counting indicators such as continuous peaks of sequencing depth and continuous valleys of sequencing depth for CNV classification to improve the accuracy of CNV determination.
[0129] Step S33, perform variation detection on the chromosome based on the classified gene data to generate a copy number variation detection result.
[0130] In a feasible embodiment, in this embodiment, step S33 may include steps D1 to D3:
[0131] Step D1, determining the mutation type of the chromosome based on the proportion of mutation window data corresponding to the classified gene data;
[0132] Step D2, determining the copy number variation chimeric ratio by correcting the proportion calculation formula and the copy number corresponding to the classified gene data;
[0133] Step D3, generating a copy number variation detection result according to a preset database, the classified gene data, the mutation type of the chromosome, and the copy number variation chimeric ratio.
[0134] It should be noted that the above mutation window data may be the proportion of the number of duplication or deletion windows in the CNV region calculated based on the classified gene data. This embodiment can further clarify the copy number variation type of the chromosome based on it. Exemplarily, when the proportion exceeds a certain critical value, the chromosome can be marked as "aneuploid"; otherwise, it can be marked as "euploid". In addition, this embodiment also designs a scientific method for calculating the CNV chimeric ratio to control the detection lower limit and enhance its value in clinical applications. According to industry standards, the lower limit of the chimeric ratio routinely reported by CNV-seq is 20%. In actual applications, the CNV-seq technology has high detection reliability for samples with a chimeric ratio reaching or exceeding 30%. Exemplarily, assuming that the copy number of the genomic target segment is CN, the theoretical calculation formula derivation of the copy number variation chimeric ratio (P) can be as follows:
[0135] 1) For a chimera with a copy number of 3, 3*P + 2*(1 - P) = CN, so P = CN - 2;
[0136] 2) For a chimera with a copy number of 1, 1*P + 2*(1 - P) = CN, so P = 2 - CN.
[0137] However, due to the influence of objective factors such as wet laboratory operations, the actual calculation of the chimeric ratio needs to be corrected based on the theoretical formula (see Table 1 for details):
[0138] Table 1: Derivation of the chimeric ratio calculation formula
[0139]
[0140]
[0141] It should be understood that the proposed reference log2(ratio) values in Table 1 have been adjusted once in the actual test. Therefore, in this embodiment, the corrected calculation formula determined according to Table 1 can be: 1) For the chimerism with a copy number of 3, P = (CN - 2) / 0.93; 2) For the chimerism with a copy number of 1, P = (2 - CN) / 0.89.
[0142] Therefore, in this embodiment, the detection device can adopt an optimized corrected ratio calculation formula to improve the detection ability of low-frequency chimeric variations. The corrected formula combines the theoretical log2(ratio) adjustment, thereby reducing errors and improving the stability of the detection limit. This method shows high reliability when detecting samples with a chimeric ratio ≥ 20%. The detection device in this embodiment also has an automatic chimeric ratio calculation function to evaluate the degree of cell chimerism and provide strong support for clinical diagnosis and scientific research analysis. Improve the chimeric ratio calculation method, strictly control the sensitivity and lower limit of detection, and then improve the detection accuracy and clinical application value of CNV-seq.
[0143] In summary, the copy number variation detection results finally generated in this embodiment can include the name of each chromosome, Z-score, total number of windows, proportion of variant windows, standardized Z-score, preliminary classification result, copy number status (i.e., the specific copy number of CNV, such as 0, 1, 2, etc.), and copy number variation chimeric ratio, so as to comprehensively describe the chromosomal variation situation. Therefore, this embodiment can further optimize the filtering strategy and chimeric calculation on the basis of optimizing the reference library, window division, and GC correction, greatly improving the detection accuracy, sensitivity, and automation degree. At the same time, an online CNV analysis platform is integrated, greatly enhancing the convenience and efficiency of clinical applications.
[0144] This embodiment discloses determining a target repeated region by using the dynamic Z-test result and the continuous peak value of the processed gene data; determining a target deletion region by using the dynamic Z-test result and the continuous valley value of the processed gene data; performing status marking on the processed gene data based on the target repeated region and the target deletion region to obtain marked gene data. Standardizing the marked Z-test result corresponding to the marked gene data by using the variant statistical count corresponding to the marked gene data to obtain a standard Z-test result; determining the variant statistical proportion according to the positive and negative directions of the standard Z-test result and the variant statistical count; classifying the copy number variation types based on the standard Z-test result and the variant statistical proportion to obtain classified gene data. Determining the chromosomal variation type based on the proportion of variant window data corresponding to the classified gene data; determining the copy number variation chimeric ratio by using the corrected ratio calculation formula and the copy number corresponding to the classified gene data; generating a copy number variation detection result according to a preset database, the classified gene data, the chromosomal variation type, and the copy number variation chimeric ratio.
[0145] This embodiment proposes a scientific CNV filtering and classification strategy. This method uses the standardized Z-score combined with multi-dimensional indicators such as continuous peaks of sequencing depth and continuous valleys of sequencing depth for CNV filtering and classification, improving the accuracy of CNV determination. At the same time, this embodiment improves the calculation method of the chimeric ratio, adopts an optimized ratio calculation formula, and improves the detection ability of low-frequency chimeric variations. Combining with the theoretical log2(ratio) adjustment, the error is reduced and the stability of the detection limit is improved. This method shows high reliability when detecting samples with a chimeric ratio ≥ 20%.
[0146] Exemplarily, to help understand the technical concept or technical principle of the copy number variation detection method based on low-depth whole-genome sequencing after combining this embodiment with the above-mentioned Embodiment 1 and Embodiment 2, please refer to Figure 4 , Figure 4 which is a schematic diagram of the brief process of the copy number variation detection method based on low-depth whole-genome sequencing of this application, specifically as follows:
[0147] The entire analysis process of the copy number variation detection method based on low-depth whole-genome sequencing proposed in this application can be composed of three main modules, including a data quality control module, a data preprocessing module, and a CNV detection module. The specific detection process can be seen in Figure 4 .
[0148] Figure 4 In, RaW FASTQ is the FASTQ data of the untreated sample, that is, the sample gene data, CleanFASTQ is the filtered gene data after processing, RaW BAM is the BAM format alignment result after aligning the filtered gene data and the human hg19 reference genome using the BWA software, and Clean BAM is the final alignment result obtained by identifying and removing the repetitive sequences generated during the PCR amplification process through GATK, that is, the standardized gene data. In addition, Figure 4 in, Raw Counts is the window gene data obtained after genome window partitioning and counting, Normalized Counts is the processed gene data after data correction, Zscore Counts is the dynamic Z-test result obtained after performing the Z-test, Annotated CNVs is the refined annotated copy number variation detection result based on the authoritative database, providing physical location, functional impact, and clinical relevance information of CNV, and Visualized CNVs is the visualized copy number variation detection result. On the basis of Figure 4 , this embodiment can utilize simulated CNV data and clinical CNV data to discuss in detail the rationality of the present invention in aspects such as CNV filtering criteria, CNV detection resolution, and detection limit of CNV detection, further demonstrating its reliability.
[0149] Exemplarily, first, the present application can perform verification data collection:
[0150] 1) Construct simulated CNV data:
[0151] First, avoid the centromere and telomere regions of chromosomes, and randomly select regions in the genome at the scale of hundreds of KB to MB. Using the clinically positive sample with the least influence of CNV variation as the base sample, use the bedtools intersect tool to accurately extract the target region from the BAM file of the base sample. Subsequently, use samtools to perform downsampling on the target region to simulate heterozygous CNV. Then, perform deletion or overlay operations on the target region in the base sample through bedtools to construct simulated heterozygous or homozygous CNV (including deletion and duplication). Regarding clinical data, the present application can collect 48 samples confirmed to be CNV positive by chromosomal microarray analysis (CMA). These samples cover common trisomies, monosomies, chimeras, typical syndromes, complex variations, and limited CNV types. As the clinical gold standard for detecting CNV, CMA has the characteristics of high resolution and high sensitivity, and can accurately locate and quantify copy number changes, especially showing significant advantages in detecting microdeletions, microduplications, and complex chromosomal abnormalities. Therefore, the positive samples pre-constructed in the present application based on CMA detection provide high-quality reference data for subsequent research. Then, select the chimerism ratio data, and select 9 samples whose chimerism ratios have been confirmed by CMA. And use the detection device to calculate the chimerism ratios of these samples.
[0152] 2) Set CNV filtering criteria:
[0153] Under the filtering criteria set in the CNV detection module, the detection device analyzed 48 clinical samples confirmed to be CNV positive by CMA. Among these 48 samples, the CNV detection results of 44 samples were consistent with the CMA results, and the remaining 4 samples were complex chimeric samples (see Table 2 for details). In addition, through genetic counseling experts combined with this detection device, CNV visualization charts can be generated, as shown in Figure 5 and Figure 6 , Figure 5 is the visualization diagram of the sample whole-chromosome CNV results of the copy number variation detection method based on low-depth whole-genome sequencing of the present application, Figure 6 is the visualization diagram of the sample single-chromosome CNV results of the copy number variation detection method based on low-depth whole-genome sequencing of the present application. Figure 6In it, BlackList indicates genomic regions on the genome that are unreliable or have high noise, usually located in regions with high GC content, repetitive sequences, telomeres, centromeres, etc. CopyNum is the detected CNV copy number. Chrn (n = 1 - 22, x, y) represents the chromosome number where the CNV variation is located, and chr1 refers to chromosome 1. And Figure 6 In it, different colored blocks on the abscissa can be used to show the specific positions of CNV variations on the chromosome. Using cytogenetic nomenclature, "p" represents the short arm, "q" represents the long arm, and the numbers indicate the band regions (such as p11.2). At the same time, after artificial analysis through the CNV functional annotation results, all these complex chimeric samples showed obvious signal abnormalities. This result further verified the detection performance of the CNV detection method based on low-depth whole-genome sequencing proposed in this application in clinical positive samples and supported the rationality of setting the CNV filtering criteria.
[0154] Table 2: Complex chimeric CNV samples
[0155]
[0156] 3) CNV detection resolution:
[0157] Through the detection results of known CNVs and simulated CNVs in 48 positive samples, the CNV detection resolution of the CNV detection method based on low-depth whole-genome sequencing proposed in this application was evaluated. A total of 16 simulated CNVs were designed (see Table 3 for details), among which 10 CNVs had a length less than 1 MB, 6 CNVs had a length below 150 KB, and the shortest length was 80 KB.
[0158] In addition, a total of 88 CNVs were included in the clinical positive samples, among which 8 had a length less than 1 MB (see Table 4 for details), and the shortest length was 101 KB. After using this detection device to detect CNVs in positive samples and simulated samples, all simulated CNVs passed the filtering criteria set by the CNV detection module, and the shortest detection length was 80 KB. At the same time, the shortest length of the successfully detected CNVs in the clinical positive samples was 101 KB. These results further verified the rationality of setting the resolution of this detection device to 100 KB.
[0159] Table 3: Simulated CNV information
[0160]
[0161] Table 4: Real data CNV information
[0162]
[0163] 4) Detection limit of CNV detection:
[0164] Furthermore, the present application screened out 9 samples containing partial chromosomal regions or whole chromosomal chimeras from 48 positive samples, including CNVseq008, CNVseq009, CNVseq024, CNVseq035, CNVseq036, CNVseq037, CNVseq038, CNVseq039, and CNVseq041. The chimerism ratios of these samples (ranging from 20% to 60%) have been determined by CMA detection, and a total of 13 chimeric CNVs (9 DUPs and 4 DELs) have been identified. The chimerism ratios of these samples were calculated using the detection device, and the results showed that the calculated results of 6 out of 9 samples were basically consistent with the chimerism ratios detected by CMA (the difference was within 5 percentage points); the chimerism ratios of another 2 samples were slightly higher than the reference standard, and that of 1 sample was slightly lower than the reference standard (see Table 5 for details). These results further demonstrated the high reliability of the detection device in detecting chimeric CNVs with a chimerism ratio greater than 20%, and clarified the detection limit of CNV detection.
[0165] Table 5: Detection results of chimerism ratios of chimeric CNVs
[0166]
[0167] 5) Construct a CNV analysis platform:
[0168] The CNV analysis platform proposed in the present application is an interactive online platform that integrates CNV visualization, CNV annotation, and chimerism ratio calculation functions, significantly improving the efficiency and convenience of CNV variant interpretation. The platform supports the visualization of CNV results for whole chromosomes and single chromosomes, and provides convenient switching and zooming interaction functions, enabling analysts to quickly browse the CNV distribution across the whole genome and also focus on the CNV details of a single chromosome (see Figure 7 and Figure 8 ), Figure 7 which is a visualization diagram of the whole chromosome CNV results detected by the copy number variation detection method based on low-depth whole-genome sequencing of the present application, Figure 8 and Figure 7 is a visualization diagram of the single chromosome CNV results detected by the copy number variation detection method based on low-depth whole-genome sequencing of the present application. Figure 8The visualization results not only clearly show the distribution characteristics of copy number variations but also mark the genomic blacklist regions, which helps to evaluate potential impacts. At the same time, the CNV analysis platform integrates authoritative databases such as OMIM, DECIPHER, DGV, and ClinGen, provides information on the genomic location, functional impact, and clinical relevance of CNVs, and supports one-click jumping to the official websites of the corresponding databases, avoiding the cumbersome manual search and improving the interpretation efficiency. The platform also has an automatic calculation function for chimeric ratio, providing important support for CNV analysis. In addition, the platform provides complete original data, including the total number of windows involved in CNV calculation, Z-score values of multiple window regions, copy numbers, the number of uniquely mapped reads aligned to the CNV region, and CNV classification information (see details in Figure 9 ), Figure 9 which is the CNV annotation result graph and the schematic diagram of the chimeric ratio calculation tool for the copy number variation detection method based on low-depth whole-genome sequencing of this application. Figure 9 The data shown provides a solid foundation for the identification, classification, reliability assessment, and subsequent analysis of CNVs. As an online platform, users can access it anytime and anywhere without complex installation procedures, greatly improving flexibility and usability, and providing an integrated, efficient, and accurate solution for CNV analysis.
[0169] In summary, the present invention provides a method for detecting copy number variations based on low-depth whole-genome sequencing. The whole method consists of three main modules: data quality control, data preprocessing, and CNV detection. This CNV-seq detection technology demonstrates significant advantages in the following aspects: 1) Precise bioinformatics methods: In the data preprocessing module, by dynamically constructing a negative reference library, the influence of individual GC content differences on the detection results is fully considered to improve the reliability of the detection. At the same time, the window division strategy is optimized to ensure more accurate processing of abnormal windows, and reasonable GC content correction is performed on the window statistical counts to minimize sequencing bias to the greatest extent. 2) In the CNV detection module, the filtering criteria are scientifically set to effectively reduce the false positive rate; the calculation method of the chimeric ratio is improved to strictly control the detection sensitivity and the lowest detection limit to enhance the detection ability for low-frequency chimeric variations. The whole method is fully verified by clinical real samples and simulated data, ensuring its accuracy and stability in practical applications. 3) A comprehensive and strict quality control mechanism runs through the entire process: From the acquisition of the original sequencing data, through key analysis links such as sequence alignment and CNV detection, a strict quality control system is implemented. This high-standard quality control not only significantly improves the reliability and consistency of the analysis results but also provides a solid and credible data basis for subsequent research. 4) An efficient automated analysis process: The automated, modular, and parallel analysis process designed based on workflow language significantly improves the data processing efficiency. This advanced analysis architecture not only greatly shortens the analysis cycle but also accelerates the research progress by optimizing the utilization of computing resources, providing strong support for the rapid analysis of large-scale data.
[0170] It should be noted that the above examples are only for understanding the present application and do not constitute a limitation to the method for detecting copy number variations based on low-depth whole-genome sequencing of the present application. More forms of simple transformations based on this technical concept are within the protection scope of the present application.
[0171] The present application also provides a device for detecting copy number variations based on low-depth whole-genome sequencing. Please refer to Figure 10 , Figure 10 which is the schematic diagram of the module structure of the device for detecting copy number variations based on low-depth whole-genome sequencing in the embodiment of the present application. In this embodiment, the device for detecting copy number variations based on low-depth whole-genome sequencing includes:
[0172] A data quality control module T1, which is used to perform sequencing quality control on the sample gene data to obtain filtered gene data;
[0173] A data preprocessing module T2, which is used to perform data preprocessing on the filtered gene data through a dynamic negative reference library and a reference genome to obtain processed gene data and dynamic Z-test results;
[0174] The copy number variation detection module T3 is used to perform improved copy number variation detection based on the processed gene data and the dynamic Z - test result to obtain a copy number variation detection result.
[0175] In a feasible implementation manner, in this embodiment, the data pre - processing module T2 is further used to perform sequence alignment on the filtered gene data based on a reference genome to obtain standardized gene data; the reference genome is segmented according to a set window size, and a plurality of overlapping windows are defined by a sliding window manner, and the step size of the sliding window is smaller than the set window size; genomic window division and counting are performed on the standardized gene data according to the divided windows of the reference genome to obtain window gene data; the read count of each window in the window gene data is greater than the minimum ratio of the coverage of the corresponding region in the reference genome, and the N - base ratio of each window corresponding to the reference genome is less than the N - base ratio threshold; data correction is performed on the window gene data by the locally weighted regression method to obtain processed gene data; a Z - test is performed on the dynamic negative reference library and the processed gene data to obtain a dynamic Z - test result.
[0176] In a feasible implementation manner, in this embodiment, the data pre - processing module T2 is further used to set a shared threshold range according to the GC content of the sample gene data and a preset floating range; set a gradually expanding unit based on the GC content of the sample gene data and a preset screening step size; generate a test GC content corresponding to the sample gene data according to the shared threshold range and the gradually expanding unit; perform a coefficient of variation evaluation on the comparative negative sample set through the test GC content, and generate a dynamic negative reference library according to the evaluation result.
[0177] In a feasible implementation manner, in this embodiment, the copy number variation detection module T3 is further used to perform copy number status marking on the processed gene data and the dynamic Z - test result to obtain marked gene data; perform copy number variation type classification on the marked gene data through the marked Z - test result corresponding to the marked gene data and the variation statistical count to obtain classified gene data; perform variation detection on the chromosome based on the classified gene data to generate a copy number variation detection result.
[0178] In a feasible implementation manner, in this embodiment, the copy number variation detection module T3 is further used to determine a target repeat region through the continuous peaks of the dynamic Z - test result and the processed gene data; determine a target deletion region through the continuous valleys of the dynamic Z - test result and the processed gene data; perform status marking on the processed gene data based on the target repeat region and the target deletion region to obtain marked gene data.
[0179] In a feasible embodiment, in this embodiment, the copy number variation detection module T3 is further configured to standardize the labeled Z-test result corresponding to the labeled gene data through the variation statistical count corresponding to the labeled gene data to obtain a standard Z-test result; determine a variation statistical ratio according to the positive and negative directions of the standard Z-test result and the variation statistical count; classify the copy number variation types based on the standard Z-test result and the variation statistical ratio to obtain classified gene data.
[0180] In a feasible embodiment, in this embodiment, the copy number variation detection module T3 is further configured to determine the variation type of the chromosome based on the variation window data ratio corresponding to the classified gene data; determine the copy number variation chimeric ratio through a corrected ratio calculation formula and the copy number corresponding to the classified gene data; generate a copy number variation detection result according to a preset database, the classified gene data, the variation type of the chromosome, and the copy number variation chimeric ratio.
[0181] The copy number variation detection device based on low-depth whole-genome sequencing provided by the present application adopts the copy number variation detection method based on low-depth whole-genome sequencing in the above embodiment, and can solve the technical problems of copy number variation detection based on low-depth whole-genome sequencing. Compared with the prior art, the beneficial effects of the copy number variation detection device based on low-depth whole-genome sequencing provided by the present application are the same as those of the copy number variation detection method based on low-depth whole-genome sequencing provided by the above embodiment, and other technical features in the copy number variation detection device based on low-depth whole-genome sequencing are the same as the features disclosed in the method of the above embodiment, and will not be elaborated here.
[0182] The present application provides a copy number variation detection device based on low-depth whole-genome sequencing. The copy number variation detection device based on low-depth whole-genome sequencing includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the copy number variation detection method based on low-depth whole-genome sequencing in the first embodiment above.
[0183] Next, refer to Figure 11, which shows a schematic structural diagram of a copy number variation detection device suitable for implementing the embodiments of the present application based on low-depth whole-genome sequencing. The copy number variation detection device based on low-depth whole-genome sequencing in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 11 The shown copy number variation detection device based on low-depth whole-genome sequencing is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0184] As Figure 11 shown, the copy number variation detection device based on low-depth whole-genome sequencing may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the random access memory 1004, various programs and data required for the operation of the copy number variation detection device based on low-depth whole-genome sequencing are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the copy number variation detection device based on low-depth whole-genome sequencing to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a copy number variation detection device with various systems based on low-depth whole-genome sequencing, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.
[0185] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a copy number variation detection program product based on low-depth whole-genome sequencing, which includes a copy number variation detection program based on low-depth whole-genome sequencing carried on a computer-readable medium. The copy number variation detection program based on low-depth whole-genome sequencing contains program code for performing the methods shown in the flowcharts. In such an embodiment, the copy number variation detection program based on low-depth whole-genome sequencing can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the copy number variation detection program based on low-depth whole-genome sequencing is executed by a processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0186] The copy number variation detection device based on low-depth whole-genome sequencing provided by the present application adopts the copy number variation detection method based on low-depth whole-genome sequencing in the above embodiments, and can solve the technical problems of copy number variation detection based on low-depth whole-genome sequencing. Compared with the prior art, the beneficial effects of the copy number variation detection device based on low-depth whole-genome sequencing provided by the present application are the same as those of the copy number variation detection method based on low-depth whole-genome sequencing provided in the above embodiments, and other technical features in the copy number variation detection device based on low-depth whole-genome sequencing are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0187] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0188] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0189] The present application provides a storage medium having computer-readable program instructions stored thereon (i.e., a copy number variation detection program based on low-depth whole-genome sequencing), and the computer-readable program instructions are used to execute the copy number variation detection method based on low-depth whole-genome sequencing in the above embodiments.
[0190] The storage medium provided by the present application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of the storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0191] The above storage medium can be included in a copy number variation detection device based on low-depth whole-genome sequencing; it can also exist independently and not be assembled into a copy number variation detection device based on low-depth whole-genome sequencing.
[0192] The above storage medium carries one or more programs. When the one or more programs are executed by a copy number variation detection device based on low-depth whole-genome sequencing, the copy number variation detection device based on low-depth whole-genome sequencing is enabled to perform copy number variation detection based on low-depth whole-genome sequencing.
[0193] The copy number variation detection program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0194] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and copy number variation detection program products based on low-depth whole-genome sequencing according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0195] The modules involved in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0196] The readable storage medium provided by this application is a storage medium, which stores computer-readable program instructions for executing the above copy number variation detection method based on low-depth whole-genome sequencing (i.e., the copy number variation detection program based on low-depth whole-genome sequencing), and can solve the technical problem of how to improve the detection reliability of copy number variation detection based on low-depth whole-genome sequencing. Compared with the prior art, the beneficial effects of the storage medium provided by this application are the same as those of the copy number variation detection method based on low-depth whole-genome sequencing provided in the above embodiments, and will not be elaborated here.
[0197] The above are only partial embodiments of this application, and do not limit the patent scope of this application. Any equivalent structural transformation made under the technical concept of this application by using the content of the specification and drawings of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application.
Claims
1. A method for detecting copy number variations based on low-depth whole-genome sequencing, characterized in that, The method includes: Performing sequencing quality control on sample gene data to obtain filtered gene data; Performing data preprocessing on the filtered gene data through a dynamic negative reference library and a reference genome to obtain processed gene data and dynamic Z-test results; the dynamic negative reference library is determined according to the GC content of a comparative negative sample set; Performing improved copy number variation detection based on the processed gene data and the dynamic Z-test results to obtain copy number variation detection results.
2. The copy number variation detection method based on low-depth whole genome sequencing according to claim 1, wherein The step of performing data preprocessing on the filtered gene data through a dynamic negative reference library and a reference genome to obtain processed gene data and dynamic Z-test results includes: Performing sequence alignment on the filtered gene data based on the reference genome to obtain normalized gene data; the reference genome is segmented according to a set window size, and a plurality of overlapping windows are defined by means of a sliding window, and the step size of the sliding window is smaller than the set window size; Performing genome window partitioning and counting on the normalized gene data according to the partitioned windows of the reference genome to obtain window gene data; the number of reads of each window in the window gene data is greater than the minimum ratio of the coverage of the corresponding region in the reference genome, and the proportion of N bases corresponding to each window in the reference genome is less than the N base proportion threshold; Performing data correction on the window gene data by means of local weighted regression to obtain processed gene data; Performing a Z-test on the dynamic negative reference library and the processed gene data to obtain dynamic Z-test results.
3. The copy number variation detection method based on low-depth whole-genome sequencing according to claim 1, characterized in that, Before performing the Z-test on the dynamic negative reference library and the processed gene data to obtain dynamic Z-test results, it further includes: Setting a shared threshold range according to the GC content of the sample gene data and a preset floating range; Setting a gradually expanding unit based on the GC content of the sample gene data and a preset screening step size; Generating a test GC content corresponding to the sample gene data according to the shared threshold range and the gradually expanding unit; Evaluating the coefficient of variation of the comparative negative sample set through the test GC content, and generating a dynamic negative reference library according to the evaluation result.
4. The copy number variation detection method based on low-depth whole-genome sequencing according to claim 1, wherein The step of performing improved copy number variation detection based on the processed gene data and the dynamic Z-test results to obtain copy number variation detection results includes: Performing copy number status marking based on the processed gene data and the dynamic Z-test results to obtain marked gene data; Performing copy number variation type classification on the marked gene data through the marked Z-test results corresponding to the marked gene data and variation statistical counting to obtain classified gene data; Performing variation detection on chromosomes based on the classified gene data to generate copy number variation detection results.
5. The copy number variation detection method based on low-depth whole-genome sequencing according to claim 4, wherein The step of performing copy number status marking based on the processed gene data and the dynamic Z-test results to obtain marked gene data includes: Determining a target repeat region through the continuous peaks of the dynamic Z-test results and the processed gene data; Determining a target deletion region through the continuous valleys of the dynamic Z-test results and the processed gene data; Perform status marking on the processed gene data based on the target repeated region and the target deletion region to obtain marked gene data.
6. The copy number variation detection method based on low-depth whole-genome sequencing according to claim 4, wherein The step of classifying the copy number variation types of the marked gene data through the marked Z-test results corresponding to the marked gene data and the variant statistical count to obtain classified gene data includes: Standardize the marked Z-test results corresponding to the marked gene data through the variant statistical count corresponding to the marked gene data to obtain standard Z-test results; Determine the variant statistical proportion according to the positive and negative directions of the standard Z-test results and the variant statistical count; Perform copy number variation type classification based on the standard Z-test results and the variant statistical proportion to obtain classified gene data.
7. The copy number variation detection method based on low-depth whole-genome sequencing according to claim 4, wherein, The step of performing variant detection on chromosomes based on the classified gene data to generate copy number variation detection results includes: Determine the variant type of the chromosome based on the variant window data proportion corresponding to the classified gene data; Determine the copy number variation chimeric proportion through the corrected proportion calculation formula and the copy number corresponding to the classified gene data; Generate copy number variation detection results according to a preset database, the classified gene data, the variant type of the chromosome, and the copy number variation chimeric proportion.
8. A copy number variation detection device based on low-depth whole-genome sequencing, characterized in that, The copy number variation detection device based on low-depth whole-genome sequencing includes: A data quality control module for performing sequencing quality control on sample gene data to obtain filtered gene data; A data preprocessing module for performing data preprocessing on the filtered gene data through a dynamic negative reference library and a reference genome to obtain processed gene data and dynamic Z-test results; A copy number variation detection module for performing improved copy number variation detection according to the processed gene data and the dynamic Z-test results to obtain copy number variation detection results.
9. A copy number variation detection device based on low-depth whole-genome sequencing, characterized in that, The copy number variation detection device based on low-depth whole-genome sequencing includes: a memory, a processor, and a copy number variation detection program based on low-depth whole-genome sequencing stored on the memory and executable on the processor. The copy number variation detection program based on low-depth whole-genome sequencing is configured to implement the steps of the copy number variation detection method based on low-depth whole-genome sequencing according to any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a storage medium, and a copy number variation detection program based on low-depth whole-genome sequencing is stored on the storage medium. When the copy number variation detection program based on low-depth whole-genome sequencing is executed by a processor, it implements the steps of the copy number variation detection method based on low-depth whole-genome sequencing according to any one of claims 1 to 7.
Citation Information
Cited By
Integrated genome analysis method based on low-depth sequencing
CN121999865A