A method and apparatus for detecting gene copy number variations using targeted sequencing
By establishing a sample library, implementing quality control, dividing gene regions, and performing in-depth comparisons, combined with negative sample selection, the problems of high cost and low accuracy in targeted sequencing detection have been solved, achieving efficient and low-cost detection of gene copy number variations.
Patent Information
- Application Number
- CN202211046725.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-30
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-08-30
AI Technical Summary
Existing targeted sequencing methods for detecting gene copy number variations require the collection of normal samples from the same individual, which is costly and unavailable in some cases. Detection methods with multiple background samples present challenges in sample selection and representativeness, while detection without control samples cannot accurately correct sample characteristics.
By acquiring the samples to be tested and associated negative samples to establish a sample library, quality control and comparison are carried out, gene intervals are divided, depth indicators are statistically analyzed, gene sets are screened in combination with a preset scheme, the best negative samples are selected for in-depth comparison and correction, and the gene-level copy number ratio is calculated, thereby reducing costs and improving detection accuracy.
It effectively removes potential copy number variations, reduces detection costs, and improves detection results. It also selects the nearest normal sample as a control to improve detection accuracy and reduce false positive results.
Smart Images

Figure CN115497557B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene detection technology, and in particular to a method and apparatus for detecting gene copy number variations using targeted sequencing. Background Technology
[0002] Traditional methods for detecting copy number variations include fluorescence in situ hybridization (FISH), multiplex ligation-dependent probe amplification (MLPA), digital PCR (ddPCR), and chromosome microarray (CMA). FISH is based on the hybridization of sequence-specific fluorescently labeled probes, performing microscopic detection on a given fluorescence signal that indicates the presence or absence of a specific target DNA sequence. ddPCR, by diluting template DNA into thousands of nanoscale droplets, allows for absolute quantification of the target copy number without the need for standard assays. In addition to traditional methods, next-generation sequencing (NGS)-based detection methods are also widely used. NGS sequencing for copy number variation detection includes targeted sequencing (Target-NGS), whole-exome sequencing (WES), and whole-genome sequencing (WGS). NGS-based copy number variation detection relies on differences in sequencing depth to identify genes or genomic regions with copy number variations; however, different sequencing methods cover different genomic regions, thus requiring different identification algorithms. Targeted sequencing covers smaller genomic regions and typically identifies copy number variations (CNVs) by using differences in relative sequencing depth across these regions. Whole-exome sequencing and whole-genome sequencing, on the other hand, cover a larger genomic area. Besides direct identification based on depth differences, they can also incorporate neural networks and wavelet transforms for CNV signal identification. Targeted sequencing-based copy number detection allows for targeted design for specific genes, offering high specificity and lower cost compared to WES and WGS, and is currently widely used in related fields. Targeted sequencing-based copy number detection methods can be further categorized into paired-sample-based CNV detection, multi-background-pool CNV detection, and control-free CNV detection. Paired-sample-based CNV detection uses normal tissue or blood cells from the same individual as controls to detect CNVs in tumor tissue. Multi-background-pool detection involves selecting multiple normal samples to construct a background pool. Control-free detection does not rely on control samples and directly identifies CNVs based on the depth differences within the samples themselves.
[0003] Among copy number variation (CNV) detection methods based on targeted sequencing, paired-sample-based detection is theoretically the optimal approach. However, it requires collecting normal samples from the same individual, which is sometimes unavailable, and sequencing paired samples significantly increases the overall cost. Multiple-background-sample detection methods face several challenges in constructing the sample pool, including: lack of consideration for copy number variation in the target sample when selecting background samples; difficulty in defining the number of background samples; and the potential for the sample pool to remain unchanged over time, potentially failing to represent the characteristics of new samples. Unpaired-sample detection is low-cost, but it relies solely on population genomic characteristics (such as GC content and repetitive sequence distribution) for depth correction, failing to address the characteristics of the sample itself. Summary of the Invention
[0004] To address the problems existing in the prior art, embodiments of the present invention provide a method and apparatus for detecting gene copy number variations using targeted sequencing.
[0005] This invention provides a method for detecting gene copy number variations using targeted sequencing, comprising:
[0006] Obtain the sample to be tested, and obtain the associated negative sample based on the sample to be tested; establish a sample library based on the sample to be tested and the negative sample.
[0007] The sequencing data of the samples in the sample library are subjected to quality control, and the quality-controlled sequencing data of the samples are compared with the reference genome to obtain the position information of each sequence fragment in the sequencing data of the samples.
[0008] The preset targeted sequencing regions of the reference genome are obtained, and the targeted sequencing regions are divided into continuous intervals to obtain each gene interval. The read depth is statistically analyzed by combining the position information of each sequence fragment in the sample sequencing data to obtain the depth index of each gene interval.
[0009] Based on the depth index of each gene region and combined with the preset copy number variation gene screening scheme, the gene set obtained by the copy number variation gene screening is obtained.
[0010] Based on the gene set obtained from the initial screening of the copy number variation genes, the gene intervals corresponding to the gene set in the test sample and the negative sample are removed, and the maximum and minimum values of the depth index are normalized for the removed test sample and negative sample. The distance between the test sample and the negative sample is calculated, and the best negative sample is selected based on the distance.
[0011] Based on the best negative sample, the relative depth ratio of each gene region in the sample to be tested is calculated, and the adjusted relative copy number ratio of each gene region is calculated according to the relative depth ratio. The gene-level copy number ratio is calculated according to the adjusted relative copy number ratio. Based on the gene-level copy number ratio, the copy number variation genes in the sample to be tested are calculated.
[0012] In one embodiment, the method further includes:
[0013] The gene intervals of the sample to be tested are sorted according to the depth index, and the corresponding target genes are selected one by one according to the sorting for gene screening. The gene screening includes: calculating the standard deviation of the depth index of the intervals of the remaining genes excluding the target genes, and comparing the interval depth of the target gene with the standard deviation.
[0014] Based on the comparison results of the initial gene screening, the gene set obtained from the initial screening of the copy number variation gene is determined.
[0015] In one embodiment, the method further includes:
[0016] The GC ratio of each gene region is calculated based on the sequence information of the reference genome;
[0017] The GC ratio range is divided into windows, and the screening gene range in which the depth index of the sample to be tested occupies the top 5% in each window is calculated.
[0018] Each interval is compared with the screening gene interval one by one. When the target gene satisfies that ≥60% of the gene intervals belong to the screening gene interval, the target gene belongs to the gene set obtained from the initial screening of copy number variation genes.
[0019] In one embodiment, the method further includes:
[0020] Remove adapter sequences, low-quality sequences at both ends, and sequences containing multiple consecutive N bases or with a length below a preset threshold from the sample sequencing data.
[0021] In one embodiment, the reference genome includes:
[0022] GRCh37, GRCh38.
[0023] This invention provides an apparatus for detecting gene copy number variations using targeted sequencing, comprising:
[0024] The acquisition module is used to acquire the sample to be tested, acquire the associated negative sample based on the sample to be tested, and establish a sample library based on the sample to be tested and the negative sample.
[0025] The quality control module is used to perform quality control on the sequencing data of the samples in the sample library, and to compare the quality-controlled sequencing data of the samples with the reference genome to obtain the position information of each sequence fragment in the sequencing data of the samples.
[0026] The interval partitioning module is used to obtain the preset targeted sequencing region of the reference genome, and to continuously partition the targeted sequencing region into intervals to obtain each gene interval. The module also combines the position information of each sequence fragment in the sample sequencing data to perform read depth statistics to obtain the depth index of each gene interval.
[0027] The initial screening module is used to obtain the gene set obtained by the initial screening of the copy number variation gene based on the depth index of each gene region and in combination with the preset copy number variation gene initial screening scheme.
[0028] The selection module is used to remove gene intervals corresponding to the gene set obtained from the initial screening of the copy number variant genes in the test sample and the negative sample, and to perform depth index normalization on the removed test sample and negative sample, calculate the distance between the test sample and the negative sample, and select the best negative sample based on the distance.
[0029] The calculation module is used to calculate the relative depth ratio of each gene region in the sample to be tested based on the best negative sample, and to calculate the adjusted relative copy number ratio of each gene region based on the relative depth ratio, and to calculate the gene-level copy number ratio based on the adjusted relative copy number ratio, and to calculate the copy number variation gene in the sample to be tested based on the gene-level copy number ratio.
[0030] In one embodiment, the device further includes:
[0031] The sorting module is used to sort the gene intervals of the sample to be tested according to the depth index, and select the corresponding target genes one by one according to the sorting for gene screening. The gene screening includes: calculating the standard deviation of the depth index of the intervals of the remaining genes excluding the target genes, and comparing the interval depth of the target gene with the standard deviation.
[0032] The determination module is used to determine the gene set obtained from the initial screening of copy number variant genes based on the comparison results of the initial gene screening.
[0033] In one embodiment, the device further includes:
[0034] The second calculation module is used to calculate the GC ratio of each gene region based on the sequence information of the reference genome.
[0035] The partitioning module is used to divide the GC ratio interval into windows and calculate the screening gene intervals in which the depth index of the sample to be tested occupies the top 5% in each window.
[0036] The comparison module is used to compare each interval with the screening gene interval one by one. When the target gene satisfies that ≥60% of the gene intervals belong to the screening gene interval, the target gene belongs to the gene set obtained from the initial screening of copy number variation genes.
[0037] This invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the method described above for detecting gene copy number variations using targeted sequencing.
[0038] This invention provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method for detecting gene copy number variations using targeted sequencing.
[0039] This invention provides a method and apparatus for detecting gene copy number variations using targeted sequencing. The method involves acquiring a sample to be tested and obtaining associated negative samples based on the sample to be tested. A sample library is established based on the sample to be tested and the negative samples. Quality control is performed on the sequencing data of the samples in the sample library, and the quality-controlled sequencing data is compared with a reference genome to obtain the position information of each sequence fragment in the sequencing data. A preset targeted sequencing region of the reference genome is obtained, and the targeted sequencing region is divided into continuous intervals to obtain each gene interval. Read depth statistics are performed based on the position information of each sequence fragment in the sequencing data to obtain the depth index of each gene interval. Based on the depth index of each gene interval, and combined with the preset... A preliminary screening scheme for copy number variant (CNV) genes was developed, resulting in a gene set obtained from the initial screening. Based on this gene set, gene regions corresponding to the CNV gene set in the test sample and negative sample were removed. The depth index of the removed test and negative samples was then normalized to its maximum value. The distance between the test and negative samples was calculated, and the optimal negative sample was selected based on this distance. The relative depth ratio of each gene region in the test sample was calculated based on the optimal negative sample. The adjusted relative copy number ratio of each gene region was then calculated based on the relative depth ratio. The gene-level copy number ratio was then calculated based on the adjusted relative copy number ratio. Based on the gene-level copy number ratio, the copy number variant genes in the test sample were identified. This method removes potential copy number variants from the test sample and selects the nearest normal sample from the background sample pool (negative samples) as a control for copy number variant detection in the test sample, reducing detection costs and improving detection efficiency. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a flowchart of a method for detecting gene copy number variations using targeted sequencing, as described in an embodiment of the present invention.
[0042] Figure 2 This is a structural diagram of a device for detecting gene copy number variations using targeted sequencing, as described in an embodiment of the present invention.
[0043] Figure 3 This is a schematic diagram of the electronic device structure in an embodiment of the present invention. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] Figure 1 This is a flowchart illustrating a method for detecting gene copy number variations using targeted sequencing, as provided in an embodiment of the present invention. Figure 1 As shown, this embodiment of the invention provides a method for detecting gene copy number variations using targeted sequencing, comprising:
[0046] Step S101: Obtain the sample to be tested, and obtain the associated negative sample based on the sample to be tested, and establish a sample library based on the sample to be tested and the negative sample.
[0047] Specifically, the process involves acquiring the sample to be tested and obtaining associated negative samples. For example, when the sample to be tested is tumor tissue cells, the associated negative samples could be copy number variation-negative tissues from the same individual or different individuals, blood cell samples, negative corporate reference materials, etc. Furthermore, to ensure the diversity of the negative sample library, multiple (more than 10) negative samples from different times, batches, and material types need to be selected, and then a corresponding sample library is established based on the sample to be tested and the negative samples.
[0048] Step S102: Perform quality control on the sequencing data of the samples in the sample library, and compare the quality-controlled sequencing data with the reference genome to obtain the position information of each sequence fragment in the sequencing data.
[0049] Specifically, quality control is performed on the sequencing data of samples in the sample library, including the sequencing data of samples to be tested and the sequencing data of negative samples. Quality control may include: removing adapter sequences, low-quality sequences at both ends, and sequences containing multiple consecutive N bases or with a length below a threshold from the sequencing data. Then, sequence alignment software is used to align the quality-controlled data to the reference genome to obtain the alignment results at each position, thus obtaining the position information of each sequence fragment in the sample sequencing data on the reference genome. The human reference genome may be GRCh37 or GRCh38, and the sequence alignment software may be bwa or bowtie2.
[0050] Step S103: Obtain the preset targeted sequencing region of the reference genome, divide the targeted sequencing region into continuous intervals to obtain each gene interval, and combine the position information of each sequence fragment in the sample sequencing data to perform read depth statistics to obtain the depth index of each gene interval.
[0051] Specifically, a pre-defined targeted sequencing region is obtained from the reference genome, and the targeted sequencing region is divided into continuous intervals, the size of which can be 100-1000 bp, to obtain each gene interval. The read depth is then statistically analyzed by combining the position information of each sequence fragment in the sample sequencing data. The average depth or median depth of each interval can then be taken as the depth index of that interval.
[0052] Step S104: Based on the depth index of each gene region and combined with the preset copy number variation gene screening scheme, the gene set obtained by the copy number variation gene screening is obtained.
[0053] Specifically, based on the depth index of each gene region and combined with the pre-set copy number variation gene screening scheme, a gene set obtained from the initial screening of copy number variation genes is obtained. This allows for the preliminary identification of some genes with copy number variations, especially those with significant copy number variations. By removing the identified "abnormal" gene regions in subsequent steps, the final selected best negative control sample has a higher similarity to the sample to be tested.
[0054] In addition, the preset primary screening protocol for copy number variation genes can be:
[0055] The genomic regions of the samples to be tested are sorted from highest to lowest read depth. Target genes are selected one by one according to the sorted region order. The average (μ) and standard deviation (σ) of the remaining genomic regions excluding the target genes are calculated, and the relative relationship between the region depth of the gene and the remaining genomic regions is determined. A target gene is considered to have copy number variation if more than 50% of its regions have a read depth ≥ μ + 2σ, or more than 20% of its regions have a read depth ≥ μ + 3σ. For genes identified as having copy number variation, all their corresponding regions are excluded from the remaining regions. This process is repeated for all genes to complete the initial screening and obtain a set of all genes with copy number variation and their corresponding genomic regions. The determination formula is as follows:
[0056]
[0057] Where RegionDepth represents the depth of the gene region, and TotalRegion represents the number of regions contained in the gene.
[0058] In addition, the preset copy number variation gene screening scheme can also be:
[0059] The GC ratio (GC_ratio) of each region is calculated based on the reference genome sequence information. The reference gene can be GRCh37 or GRCh38, and must be consistent with the reference genome used for genome alignment. Considering that the sequencing depth of regions with excessively high / low GC content will be significantly affected by GC content, regions with GC_ratio < 0.3 or GC_ratio > 0.8 are removed during the initial screening of copy number variant genes. The remaining GC_ratio range (target range) [0.3-0.8] is divided into windows with a window length of 0.05. The gene regions in which the sequencing depth of the sample to be tested occupies the top 5% in each GC window are calculated and denoted as GC_top5 (screening gene regions). The results of all GC_ratio windows are combined for judgment: when a gene satisfies ≥60% of its regions belonging to GC_top5, the gene is considered to have a copy number variant. This scheme is used to complete the initial screening of all genes one by one and obtain the set of all genes with copy number variants.
[0060] Step S105: Based on the gene set obtained from the initial screening of copy number variation genes, remove the gene intervals corresponding to the gene set in the test sample and negative sample, and perform depth index normalization on the removed test sample and negative sample, calculate the distance between the test sample and the negative sample, and select the best negative sample based on the distance.
[0061] Specifically, based on the gene set obtained from the initial screening of copy number variation genes, the corresponding gene regions in the test samples and negative samples are removed. Then, the maximum and minimum values of the depth index (AveDepth) are normalized in the region depth of the test samples and negative samples to eliminate the influence of sequencing depth differences between samples. This normalization method can be replaced with other types of normalization methods. Based on AveDepth, the Euclidean distance between negative samples in the negative sample library and the test samples is calculated one by one. The sample in the negative sample library with the smallest Euclidean distance (Dist) to the test sample is determined as the near-by-control sample, i.e., the optimal negative sample. The formula for calculating the Euclidean distance (Dist) is as follows:
[0062]
[0063] Where n represents the total number of remaining intervals of the sample, t represents the sample to be tested, and n represents the negative sample being compared.
[0064] Step S106: Calculate the relative depth ratio of each gene region in the sample to be tested based on the best negative sample, and calculate the adjusted relative copy number ratio of each gene region based on the relative depth ratio, and calculate the gene-level copy number ratio based on the adjusted relative copy number ratio, and calculate the copy number variation gene in the sample to be tested based on the gene-level copy number ratio.
[0065] Specifically, using the determined close-related control samples (best negative samples) and the GC ratio (GC_ratio) of each interval, the sequencing depth of the sample to be tested is corrected and transformed based on the Log2 function to obtain the relative depth ratio (RD) of each interval of the sample to be tested. The difference between the relative depth ratio of each interval in the sample to be tested and the median relative depth ratio (MedianLog2) of all intervals is calculated and denoted as the adjusted relative copy number ratio (AdjustRatio). The calculation formula is as follows, where RD All Represents the relative depth ratio across all intervals:
[0066] AdjustRatio = RD - median(RD) All )
[0067] The gene-level copy number ratio of a specified gene in the sample to be tested is calculated based on AdjustRatio, denoted as ResRatio. The calculation formula is as follows, where n represents the number of intervals contained in the gene:
[0068]
[0069] Based on the gene-level copy number ratio ResRatio, the copy number is calculated using the following formula, where ResRatio is the gene-level copy number ratio:
[0070]
[0071] This allows us to obtain the copy number variation gene CN in the sample to be tested.
[0072] This invention provides a method for detecting gene copy number variations using targeted sequencing. The method involves acquiring a sample to be tested and obtaining associated negative samples based on the sample to be tested. A sample library is then established based on the sample to be tested and the negative samples. Quality control is performed on the sequencing data of the samples in the sample library, and the quality-controlled sequencing data is compared with a reference genome to obtain the position information of each sequence fragment in the sequencing data. A preset targeted sequencing region of the reference genome is obtained, and the targeted sequencing region is divided into continuous intervals to obtain each gene interval. Read depth statistics are then performed based on the position information of each sequence fragment in the sequencing data to obtain the depth index of each gene interval. Based on the depth index of each gene interval, and combined with the preset... The copy number variation (CNV) gene screening protocol obtains a gene set from the initial screening. Based on this gene set, gene regions corresponding to the CNV gene set are removed from both the test sample and the negative sample. The depth index of the removed test and negative samples is then normalized to its maximum value. The distance between the test and negative samples is calculated, and the optimal negative sample is selected based on this distance. The relative depth ratio of each gene region in the test sample is calculated based on the optimal negative sample. The adjusted relative copy number ratio of each gene region is then calculated based on the relative depth ratio. The gene-level copy number ratio is then calculated based on the adjusted relative copy number ratio. Based on the gene-level copy number ratio, the copy number variation genes in the test sample are identified. This approach removes potential copy number variations from the test sample and selects the nearest normal sample from the background sample pool (negative samples) as a control for copy number variation detection in the test sample, reducing detection costs and improving detection efficiency.
[0073] In another embodiment of the present invention, panel sequencing was performed on 13 reference samples. The methods used for copy number variation detection included: detection based on a negative sample pool, random selection of a negative sample as a control, detection without a control, and the detection method of the present invention. All detection methods were consistent except for the selection of the control sample. The results of the four detection methods for the 13 reference samples were compared and analyzed. These 13 reference samples were sequenced from a panel containing 11 genes. The copy number variation information of these 13 samples is shown in the table below:
[0074]
[0075] Copy number variation (CNV) detection was performed on 13 samples using four methods: the method of this invention, a detection method based on a negative sample pool (containing 10 negative samples), a detection method using randomly selected negative samples as a control, and a detection method without a control. The results showed that all four methods effectively detected the CNV genes present in the 13 samples. Specifically, the method of this invention had no false positives; the detection method based on the negative sample pool had 4 false positives (3.1%); the method using randomly selected negative samples as a control had 4 false positives (3.1%); and the detection method without a control had 1 false positive (0.8%).
[0076] The above results indicate that, in this embodiment, the detection method based on negative sample pool detection can achieve a good detection effect.
[0077] Figure 2 An apparatus for detecting gene copy number variations using targeted sequencing, provided in an embodiment of the present invention, includes: an acquisition module S201, a quality control module S202, a region partitioning module S203, a preliminary screening module S204, a selection module S205, and a calculation module S206, wherein:
[0078] The acquisition module S201 is used to acquire the sample to be tested, acquire the associated negative sample based on the sample to be tested, and establish a sample library based on the sample to be tested and the negative sample.
[0079] The quality control module S202 is used to perform quality control on the sample sequencing data in the sample library, and compare the quality-controlled sample sequencing data with the reference genome to obtain the position information of each sequence fragment in the sample sequencing data.
[0080] The interval partitioning module S203 is used to obtain the preset targeted sequencing region of the reference genome, and to continuously partition the targeted sequencing region into intervals to obtain each gene interval. The module also combines the position information of each sequence fragment in the sample sequencing data to perform read depth statistics to obtain the depth index of each gene interval.
[0081] The initial screening module S204 is used to obtain the gene set obtained by the initial screening of the copy number variation gene based on the depth index of each gene region and in combination with the preset copy number variation gene initial screening scheme.
[0082] The selection module S205 is used to remove the gene intervals corresponding to the gene set in the sample to be tested and the negative sample based on the gene set obtained from the initial screening of the copy number variation gene, and to perform depth index normalization on the removed sample to be tested and the negative sample, calculate the distance between the sample to be tested and the negative sample, and select the best negative sample based on the distance.
[0083] The calculation module S206 is used to calculate the relative depth ratio of each gene region in the sample to be tested based on the best negative sample, and to calculate the adjusted relative copy number ratio of each gene region based on the relative depth ratio, and to calculate the gene-level copy number ratio based on the adjusted relative copy number ratio, and to calculate the copy number variation gene in the sample to be tested based on the gene-level copy number ratio.
[0084] In one embodiment, the apparatus may further include:
[0085] The sorting module is used to sort the gene intervals of the sample to be tested according to the depth index, and select the corresponding target genes one by one according to the sorting for gene screening. The gene screening includes: calculating the standard deviation of the depth index of the intervals of the remaining genes excluding the target genes, and comparing the interval depth of the target genes with the standard deviation.
[0086] The determination module is used to determine the gene set obtained from the initial screening of copy number variant genes based on the comparison results of the initial gene screening.
[0087] In one embodiment, the apparatus may further include:
[0088] The second calculation module is used to calculate the GC ratio of each gene region based on the sequence information of the reference genome.
[0089] The partitioning module is used to divide the GC ratio interval into windows and calculate the screening gene interval in which the depth index of the sample to be tested occupies the top 5% in each window.
[0090] The comparison module is used to compare each interval with the screening gene interval one by one. When the target gene satisfies that ≥60% of the gene intervals belong to the screening gene interval, the target gene belongs to the gene set obtained from the initial screening of copy number variation genes.
[0091] Specific limitations regarding the device for detecting gene copy number variations using targeted sequencing can be found in the above description of the methods for detecting gene copy number variations using targeted sequencing, and will not be repeated here. Each module in the aforementioned device for detecting gene copy number variations using targeted sequencing can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0092] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include: a processor 301, a memory 302, a communication interface 303, and a communication bus 304. The processor 301, memory 302, and communication interface 303 communicate with each other via the communication bus 304. The processor 301 can call logical instructions in the memory 302 to execute the following methods: acquiring a sample to be tested, acquiring associated negative samples based on the sample to be tested, and establishing a sample library based on the sample to be tested and the negative samples; performing quality control on the sequencing data of the samples in the sample library, and comparing the quality-controlled sequencing data with a reference genome to obtain the position information of each sequence fragment in the sequencing data; acquiring a preset targeted sequencing region of the reference genome, dividing the targeted sequencing region into continuous intervals to obtain each gene interval, and performing read depth statistics based on the position information of each sequence fragment in the sequencing data to obtain the depth index of each gene interval; and based on the depth index of each gene interval, combining the preset... The copy number variation (CNV) gene screening scheme is used to obtain a gene set obtained from the initial screening. Based on the gene set obtained from the initial screening, the gene intervals corresponding to the gene set in the test sample and negative sample are removed. The depth index of the removed test sample and negative sample is normalized to the maximum and minimum values. The distance between the test sample and the negative sample is calculated, and the best negative sample is selected based on the distance. Based on the best negative sample, the relative depth ratio of each gene interval in the test sample is calculated, and the adjusted relative copy number ratio of each gene interval is calculated based on the relative depth ratio. The gene-level copy number ratio is calculated based on the adjusted relative copy number ratio. Based on the gene-level copy number ratio, the copy number variation genes in the test sample are calculated.
[0093] Furthermore, the logical instructions in the aforementioned memory 302 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0094] On the other hand, embodiments of the present invention also provide a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, this computer program implements the transmission methods provided in the above embodiments, including, for example,: acquiring a sample to be tested, acquiring associated negative samples based on the sample to be tested, and establishing a sample library based on the sample to be tested and the negative samples; performing quality control on the sequencing data of the samples in the sample library, and comparing the quality-controlled sequencing data with a reference genome to obtain the position information of each sequence fragment in the sequencing data; acquiring a preset targeted sequencing region of the reference genome, dividing the targeted sequencing region into continuous intervals to obtain each gene interval, and performing read depth statistics based on the position information of each sequence fragment in the sequencing data to obtain the position information of each gene interval. Depth metrics; Based on the depth metrics of each gene interval, combined with the preset copy number variation gene screening scheme, the gene set obtained from the copy number variation gene screening is obtained; Based on the gene set obtained from the copy number variation gene screening, the gene intervals corresponding to the gene set in the test sample and negative sample are removed, and the maximum and minimum values of the depth metrics of the removed test sample and negative sample are normalized. The distance between the test sample and the negative sample is calculated, and the best negative sample is selected based on the distance; Based on the best negative sample, the relative depth ratio of each gene interval in the test sample is calculated, and the adjusted relative copy number ratio of each gene interval is calculated based on the relative depth ratio. The gene-level copy number ratio is calculated based on the adjusted relative copy number ratio, and the copy number variation gene in the test sample is calculated based on the gene-level copy number ratio.
[0095] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0096] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting gene copy number variations using targeted sequencing, characterized in that, include: Obtain the sample to be tested, and obtain the associated negative sample based on the sample to be tested; establish a sample library based on the sample to be tested and the negative sample. The sequencing data of the samples in the sample library are subjected to quality control, and the quality-controlled sequencing data of the samples are compared with the reference genome to obtain the position information of each sequence fragment in the sequencing data of the samples. The preset targeted sequencing regions of the reference genome are obtained, and the targeted sequencing regions are divided into continuous intervals to obtain each gene interval. The read depth is statistically analyzed by combining the position information of each sequence fragment in the sample sequencing data to obtain the depth index of each gene interval. Based on the depth index of each gene region and combined with the preset copy number variation gene screening scheme, the gene set obtained by the copy number variation gene screening is obtained. The gene intervals of the sample to be tested are sorted according to the depth index, and the corresponding target genes are selected one by one according to the sorting for gene screening. The gene screening includes: calculating the standard deviation of the depth index of the intervals of the remaining genes excluding the target genes, and comparing the interval depth of the target gene with the standard deviation. Based on the comparison results of the initial gene screening, the gene set obtained from the initial screening of the copy number variation gene is determined; Based on the gene set obtained from the initial screening of the copy number variation genes, the gene intervals corresponding to the gene set in the test sample and the negative sample are removed, and the maximum and minimum values of the depth index are normalized for the removed test sample and negative sample. The distance between the test sample and the negative sample is calculated, and the best negative sample is selected based on the distance. Based on the best negative sample, the relative depth ratio of each gene region in the sample to be tested is calculated, and the adjusted relative copy number ratio of each gene region is calculated according to the relative depth ratio. The gene-level copy number ratio is calculated according to the adjusted relative copy number ratio. Based on the gene-level copy number ratio, the copy number variation genes in the sample to be tested are calculated.
2. The method for detecting gene copy number variations using targeted sequencing according to claim 1, characterized in that, The gene set obtained from the initial screening of copy number variation genes is obtained by combining the depth index of each gene region with a preset copy number variation gene screening scheme, including: The GC ratio of each gene region is calculated based on the sequence information of the reference genome; The GC ratio range is divided into windows, and the screening gene range in which the depth index of the sample to be tested occupies the top 5% in each window is calculated. Each interval is compared with the screening gene interval one by one. When the target gene satisfies that ≥60% of the gene intervals belong to the screening gene interval, the target gene belongs to the gene set obtained from the initial screening of copy number variation genes.
3. The method for detecting gene copy number variations using targeted sequencing according to claim 1, characterized in that, The quality control of the sequencing data of samples in the sample library includes: Remove adapter sequences, low-quality sequences at both ends, and sequences containing multiple consecutive N bases or with a length below a preset threshold from the sample sequencing data.
4. The method for detecting gene copy number variations using targeted sequencing according to claim 1, characterized in that, The reference genome includes: GRCh37, GRCh38.
5. A device for detecting gene copy number variations using targeted sequencing, characterized in that, The device includes: The acquisition module is used to acquire the sample to be tested, acquire the associated negative sample based on the sample to be tested, and establish a sample library based on the sample to be tested and the negative sample. The quality control module is used to perform quality control on the sequencing data of the samples in the sample library, and to compare the quality-controlled sequencing data of the samples with the reference genome to obtain the position information of each sequence fragment in the sequencing data of the samples. The interval partitioning module is used to obtain the preset targeted sequencing region of the reference genome, and to continuously partition the targeted sequencing region into intervals to obtain each gene interval. The module also combines the position information of each sequence fragment in the sample sequencing data to perform read depth statistics to obtain the depth index of each gene interval. The initial screening module is used to obtain the gene set obtained by the initial screening of the copy number variation gene based on the depth index of each gene region and in combination with the preset copy number variation gene initial screening scheme. The sorting module is used to sort the gene intervals of the sample to be tested according to the depth index, and select the corresponding target genes one by one according to the sorting for gene screening. The gene screening includes: calculating the standard deviation of the depth index of the intervals of the remaining genes excluding the target genes, and comparing the interval depth of the target gene with the standard deviation. The determination module is used to determine the gene set obtained from the initial screening of the copy number variation gene based on the comparison results of the initial screening. The selection module is used to remove gene intervals corresponding to the gene set obtained from the initial screening of the copy number variant genes in the test sample and the negative sample, and to perform depth index normalization on the removed test sample and negative sample, calculate the distance between the test sample and the negative sample, and select the best negative sample based on the distance. The calculation module is used to calculate the relative depth ratio of each gene region in the sample to be tested based on the best negative sample, and to calculate the adjusted relative copy number ratio of each gene region based on the relative depth ratio, and to calculate the gene-level copy number ratio based on the adjusted relative copy number ratio, and to calculate the copy number variation gene in the sample to be tested based on the gene-level copy number ratio.
6. The apparatus for detecting gene copy number variations by targeted sequencing according to claim 5, characterized in that, The device further includes: The second calculation module is used to calculate the GC ratio of each gene region based on the sequence information of the reference genome. The partitioning module is used to divide the GC ratio range into windows and calculate the screening gene range in which the depth index of the sample to be tested occupies the top 5% in each window. The comparison module is used to compare each interval with the screening gene interval one by one. When the target gene satisfies that ≥60% of the gene intervals belong to the screening gene interval, the target gene belongs to the gene set obtained from the initial screening of copy number variation genes.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for detecting gene copy number variations by targeted sequencing as described in any one of claims 1 to 4.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method for detecting gene copy number variations for targeted sequencing as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Copy number variation detection method and application thereof
CN113674803A
Method and device for detecting chromosomal variations
WO2018161245A1