Processing Method, Device and Equipment for Genome Copy Number Variation Analysis Data

The method enhances the accuracy and readability of CNV detection by filtering and classifying ultra-low-depth whole genome sequencing data, addressing the complexity of result interpretation in CNV analysis.

CN119864076BActive Publication Date: 2025-07-15广州凯普医学检验所有限公司 +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510354669.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-15
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

It is difficult to interpret the results of the CNV analysis of the whole genome in low-deep and ultra-low-deep genome, and it is difficult for the prior art to accurately present the genome copy number variation.

Method used

By obtaining the copy number of genomic sequence sample data, combining adjacent partition intervals with similar copy numbers, forming a combined interval, and filtering and classification according to the interval size and copy number, excluding interference information, and drawing and displaying using R language.

Benefits of technology

It improves the accuracy and readability of genomic CNV detection results, reduces controversy over the interpretation of results, and provides intuitive data presentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119864076B_ABST
    Figure CN119864076B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of bioinformatics technology, and discloses a method, device and equipment for processing genomic copy number variation analysis data. The method includes: obtaining the copy numbers of each partition interval of genomic sequence sample data; merging adjacent partition intervals with similar copy numbers to form combined intervals; screening the combined intervals according to the interval size to obtain target combined intervals greater than a preset threshold; classifying the partition intervals within the target combined intervals respectively based on the copy numbers, and counting the interval numbers of each classification; for each target combined interval, screening the partition intervals in the target combined interval based on the corresponding multiple interval numbers to obtain target partition intervals; and performing plotting and display based on the data in the target partition intervals. The present invention provides data screening logics and conditions that can accurately present the results of whole-genome CNV, exclude interference information, reduce controversial sites, and improve the accuracy and readability of the interpretation of detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics technology, and particularly relates to a method, apparatus and device for processing genomic copy number variation analysis data. Background Art

[0002] Genomic copy number variation (CNV) is a genetic test used to identify parts of an individual's genomic DNA sequence that are increased or decreased relative to a reference sequence. These changes may include duplications, deletions or insertions of genes, and they may be related to certain genetic diseases or cancers. Currently, the main methods for detecting CNV include array-based comparative genomic hybridization (aCGH), single nucleotide polymorphism array (SNParray), multiplex ligation-dependent probe amplication (MLPA), etc. Among them, the MLPA method is the gold standard for detecting chromosomal copy number variation. With the completion of the Human Genome Project, DNA sequencing technology has entered a stage of rapid development. Genomic copy number variation detection (CNV) based on next-generation sequencing (NGS) technology and bioinformatics analysis provides a new means for prenatal diagnosis and genetic disease screening.

[0003] CNV is an important part of human gene diversity and plays an important role in diagnosing patients with unclear clinical symptoms and identifying established fetal chromosomal disease syndromes. High-depth paired-end whole-genome sequencing (>30×) has high sensitivity and resolution for CNV detection, but it is costly and has a relatively long test cycle. At the same time, it also has relatively high requirements for the hardware conditions of the computer and the analysts. Recent studies have found that low-depth whole-genome sequencing technology (1× - 5×) or even ultra-low-depth whole-genome sequencing technology (0.1×) can detect aneuploidy variations of 23 pairs of chromosomes, and microdeletions / microduplications greater than 50kb / 100kb / 500kb, which can be used to investigate the genetic causes of various diseases such as recurrent spontaneous abortion, congenital malformations, and developmental retardation. Compared with high-depth sequencing, low-depth whole-genome sequencing has lower costs, can reduce the detection cost, and can quickly obtain genomic information due to the short detection time. However, the interpretation of the results of low-depth whole-genome CNV analysis is difficult. Summary of the Invention

[0004] In view of this, the present invention provides a method, an apparatus and a device for processing genomic copy number variation analysis data, so as to solve the problem of difficult interpretation of whole-genome CNV analysis results with low depth or even ultra-low depth.

[0005] In a first aspect, the present invention provides a method for processing genomic copy number variation analysis data based on ultra-low depth whole-genome high-throughput sequencing, and the method includes:

[0006] Obtain the copy number of each partition interval of the genomic sequence sample data;

[0007] Merge adjacent partition intervals with similar copy numbers to form combined intervals;

[0008] Screen the combined intervals according to the interval size to obtain target combined intervals greater than a preset threshold;

[0009] Based on the copy number, classify the partition intervals within the target combined intervals respectively, and count the number of intervals in each classification;

[0010] For each of the target combined intervals, based on the multiple interval numbers corresponding to the target combined interval, screen the partition intervals in the target combined interval to obtain target partition intervals;

[0011] Perform plotting and display based on the data in the target partition intervals.

[0012] In an optional implementation manner, before the step of merging adjacent partition intervals with similar copy numbers to form combined intervals, the method further includes:

[0013] For a first partition interval corresponding to the Y chromosome in the partition intervals, if the copy number is less than a first preset threshold, delete the first partition interval.

[0014] In an optional implementation manner, the value range of the first preset threshold is 0.4 - 0.6.

[0015] In an optional implementation manner, the step of, for each of the target combined intervals, based on the multiple interval numbers of the target combined interval, screening the partition intervals in the target combined interval to obtain target partition intervals includes:

[0016] Based on the multiple interval numbers of the target combined interval, determine the corresponding partition interval screening rule;

[0017] According to the determined partition interval screening rule, screen the partition intervals in the target combined interval.

[0018] In an alternative embodiment, the partition interval screening rule for screening the partition intervals in the target combination interval is: screening out the partition intervals in the target combination interval with copy numbers within a preset range, where the preset range is a set value or is determined based on the copy numbers of the target combination interval.

[0019] In an alternative embodiment, based on the copy numbers, classifying the partition intervals in the target combination interval respectively and counting the interval quantities of each classification includes:

[0020] Dividing the partition intervals with copy numbers less than a second preset threshold into a first category, and obtaining a first interval quantity by statistics;

[0021] Dividing the partition intervals with copy numbers greater than or equal to the second preset threshold and less than or equal to a third preset threshold into a second category, and obtaining a second interval quantity by statistics;

[0022] Dividing the partition intervals with copy numbers greater than the third preset threshold into a third category, and obtaining a third interval quantity by statistics;

[0023] Wherein, the second preset threshold is less than the third preset threshold.

[0024] In an alternative embodiment, the value range of the second preset threshold is 1.6 - 1.8, and the value range of the third preset threshold is 2.2 - 2.4.

[0025] In an alternative embodiment, determining the corresponding partition interval screening rule based on the multiple interval quantities of the target combination interval includes:

[0026] If the sum of the first interval quantity and the third interval quantity is less than or equal to the second interval quantity, determining the corresponding partition interval screening rule as: screening out the partition intervals in the target combination interval with copy numbers within a first preset range;

[0027] If the sum of the second interval quantity and the third interval quantity is less than or equal to the first interval quantity, determining the corresponding partition interval screening rule as: screening out the partition intervals in the target combination interval with copy numbers within a second preset range;

[0028] If the sum of the first interval quantity and the second interval quantity is less than or equal to the third interval quantity, determining the corresponding partition interval screening rule as: screening out the partition intervals in the target combination interval with copy numbers within a second preset range;

[0029] Wherein, the first preset range is a copy number range set based on the second preset threshold and the third preset threshold; the second preset range is a copy number range set based on the copy number of the target combination interval.

[0030] In an alternative embodiment, the first preset range is 1.9 - 2.1; and / or, the starting value of the second preset range is the copy number of the target combination interval minus 0.1, and the ending value of the second preset range is the copy number of the target combination interval plus 0.1.

[0031] In an alternative embodiment, obtaining the copy number of each partition interval of the genomic sequence sample data includes:

[0032] Obtaining the copy number ratio of each partition interval of the genomic sequence sample data;

[0033] If the genomic sequence sample data includes Y chromosome data, then for the second partition interval corresponding to the Y chromosome and the second partition interval corresponding to the X chromosome, determine the copy number as ; for other partition intervals other than the second partition interval, determine the copy number as ;

[0034] If the genomic sequence sample data does not include Y chromosome data, then the copy number of each partition interval is determined as ;

[0035] Wherein, ratio is the copy number ratio.

[0036] In a second aspect, the present invention provides a processing device for genomic copy number variation analysis data based on ultra-low depth whole-genome high-throughput sequencing, and the device includes:

[0037] An interval copy number acquisition module, configured to acquire the copy number of each partition interval of the genomic sequence sample data;

[0038] A first screening module, configured to, for the first partition interval corresponding to the Y chromosome, if the copy number is less than the first preset threshold, delete the first partition interval;

[0039] A merging module, configured to merge adjacent partition intervals with similar copy numbers to form a combined interval;

[0040] A second screening module, configured to screen the combined intervals according to the interval size to obtain target combined intervals greater than a preset threshold;

[0041] A classification and statistics module, configured to classify the partition intervals within the target combined intervals respectively based on the copy number, and count the number of intervals in each classification;

[0042] A third screening module, configured to screen the partition intervals in each of the target combination intervals based on the multiple interval quantities corresponding to the target combination intervals, so as to obtain target partition intervals;

[0043] A drawing display module, configured to perform drawing display based on the data in the target partition intervals.

[0044] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the processing method for genomic copy number variation analysis data according to the first aspect or any corresponding embodiment thereof.

[0045] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the processing method for genomic copy number variation analysis data according to the first aspect or any corresponding embodiment thereof.

[0046] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the processing method for genomic copy number variation analysis data according to the first aspect or any corresponding embodiment thereof.

[0047] Aiming at the problem that the CNV result graph of low-depth, even ultra-low-depth (0.1×) whole genome mainly shows a discrete dot graph, which brings challenges to identifying the actual situation of CNV, this embodiment proposes a processing method, device and equipment for genomic copy number variation analysis data based on ultra-low-depth whole genome high-throughput sequencing, provides a data screening logic and conditions that can accurately present the whole genome CNV results, excludes interference information, reduces controversial sites, is beneficial to the intuitive presentation of results, reduces the controversy in result interpretation, is convenient for correct interpretation, and improves the accuracy and readability of detection results.

[0048] In addition, in the embodiment of the present invention, by deleting the first partition interval corresponding to the Y chromosome with a small copy number, the error interference caused by factors such as reagents and sequencing errors is excluded, and the accuracy and readability of the detection results are further improved. Description of the Drawings

[0049] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the related art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the related art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0050] Figure 1 It is a schematic flowchart of a method for processing genomic copy number variation analysis data based on ultra-low-depth whole-genome high-throughput sequencing according to an embodiment of the present invention;

[0051] Figure 2 It is a schematic flowchart of another method for processing genomic copy number variation analysis data based on ultra-low-depth whole-genome high-throughput sequencing according to an embodiment of the present invention;

[0052] Figure 3 It is a schematic diagram of the male gene CNV analysis result according to an embodiment of the present invention;

[0053] Figure 4 It is a schematic diagram of a combination interval and the corresponding copy number ratio value according to an embodiment of the present invention;

[0054] Figure 5 It is a schematic diagram of a screened combination interval and the corresponding copy number ratio value according to an embodiment of the present invention;

[0055] Figure 6 It is a schematic diagram of a screened high-confidence partition interval and the corresponding copy number ratio value according to an embodiment of the present invention;

[0056] Figure 7 It is a schematic diagram of optimized interval copy number visualization according to an embodiment of the present invention;

[0057] Figure 8 It is a schematic diagram of interval copy number visualization before optimization according to an embodiment of the present invention;

[0058] Figure 9 It is a schematic diagram of the female gene CNV analysis result according to an embodiment of the present invention;

[0059] Figure 10 It is a schematic diagram of another combination interval and the corresponding copy number ratio value according to an embodiment of the present invention;

[0060] Figure 11 It is a schematic diagram of another screened combination interval and the corresponding copy number ratio value according to an embodiment of the present invention;

[0061] Figure 12 It is a schematic diagram of another screened high-confidence partition interval and the corresponding copy number ratio value according to an embodiment of the present invention;

[0062] Figure 13 It is a schematic diagram of another optimized interval copy number visualization according to an embodiment of the present invention;

[0063] Figure 14Another schematic diagram of interval copy number visualization before optimization according to an embodiment of the present invention;

[0064] Figure 15 It is a structural block diagram of a processing device for genomic copy number variation analysis data based on ultra-low-depth whole-genome high-throughput sequencing according to an embodiment of the present invention;

[0065] Figure 16 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed implementation manners

[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0067] According to an embodiment of the present invention, an embodiment of a method for processing genomic copy number variation analysis data based on ultra-low-depth whole-genome high-throughput sequencing is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of executable computer instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0068] In this embodiment, a method for processing genomic copy number variation analysis data based on ultra-low-depth whole-genome high-throughput sequencing is provided, which can be used in various computer devices, including terminals, servers, etc. Figure 1 It is a flowchart of a method for processing genomic copy number variation analysis data based on ultra-low-depth whole-genome high-throughput sequencing according to an embodiment of the present invention. As Figure 1 shown, the process includes the following steps:

[0069] Step S101, obtain the copy number of each partition interval (bin interval) of the genomic sequence sample data.

[0070] Specifically, the genomic sequence sample data, such as human genomic sequence sample data, is data obtained by whole-genome sequencing, for example, data obtained by ultra-low-depth (0.1×) whole-genome high-throughput sequencing.

[0071] Taking the human genome sequence sample data as an example, the sample can be first subjected to whole-genome sequencing, and then the obtained whole-genome sequencing data can be quality controlled and filtered to obtain high-quality genomic sequence sample data. Then, the high-quality genomic sequence sample data is aligned, sorted, and deduplicated with the human reference genome to obtain the genomic sequence data of the person in the sample, that is, the human genome sequence sample data. Finally, the human genome sequence sample data is compared with the detection baseline to obtain chromosome information (X chromosome, Y chromosome, or chromosomes 1-22), the starting position of the partition interval, and the copy number ratio (ratio) value, etc.

[0072] Among them, the above detection baseline can be determined by the following method: First, the human genome data in the 1000 Genomes Project is aligned and deduplicated with the reference genome to obtain a high-quality 1000 Genomes reference sequence; then, the high-quality 1000 Genomes reference sequence obtained in the previous step is divided into multiple equal-length consecutive intervals according to a preset size (for example, 100 kb) as the detection baseline. Since the subsequent CNV analysis is to compare the human genome sequence sample data with the detection baseline, the partitioning method of the human genome sequence sample data is consistent with the corresponding partitioning method of the detection baseline.

[0073] In the embodiments of the present invention, the 1000 Genomes data is used to construct the alignment baseline, which prevents the influence of individual differences on the results, improves the reliability of statistical analysis and detection sensitivity, and helps to more comprehensively understand the genomic variation. In addition, the baseline constructed using a large amount of data is more likely to represent the genomic characteristics of the entire population or a specific group. This makes the analysis results more universal, helps to more widely and comprehensively understand the genomic variation, and can be applied to a wider range of population or disease studies.

[0074] In some optional specific embodiments, the above step S101, that is, obtaining the copy number of each partition interval of the genomic sequence sample data, includes:

[0075] Obtaining the copy number ratio of each partition interval of the genomic sequence sample data;

[0076] If the genomic sequence sample data includes Y chromosome data, for the second partition interval corresponding to the Y chromosome and the second partition interval corresponding to the X chromosome, the copy number is determined as ; for other partition intervals other than the second partition interval, the copy number is determined as ;

[0077] If the genomic sequence sample data does not include Y chromosome data, the copy number of each partition interval is determined as ;

[0078] Among them, ratio is the copy number ratio.

[0079] In the embodiments of the present invention, the calculation methods of the copy numbers in different partition intervals may be different, which are expressed by the formula as:

[0080] ;

[0081] .

[0082] Step S102: Combine adjacent partition intervals with similar copy numbers to form combined intervals. Specifically, relevant technologies can be used to complete the determination of copy number similarity and interval combination. The copy numbers are considered similar, for example, when the difference in copy numbers is less than a set threshold, that is, the copy numbers are determined to be similar.

[0083] Specifically, after forming the combined intervals, obtain the chromosome information (X chromosome, Y chromosome, or chromosome numbers 1-22), the starting positions of the combined intervals, and the copy number ratio (ratio) values of each combined interval, etc.

[0084] Step S103: Screen the combined intervals according to the interval size to obtain target combined intervals greater than a preset threshold.

[0085] Among them, the preset threshold can be, for example, 4.5 Mb or 4.0 Mb.

[0086] Step S104: Classify the partition intervals within the target combined intervals respectively based on the copy numbers, and count the number of intervals in each classification.

[0087] Step S105: For each target combined interval, screen the partition intervals in the target combined interval based on the multiple interval numbers corresponding to the target combined interval to obtain target partition intervals.

[0088] Step S106: Perform plotting and display based on the data in the target partition intervals. For example, use the R language for plotting and then display.

[0089] Aiming at the problem that the CNV result graph of low-depth, even ultra-low-depth (0.1×) whole-genome mainly shows a discrete dot graph, which brings challenges to identifying the actual situation of CNV, this embodiment proposes a processing method for genomic copy number variation analysis data based on ultra-low-depth whole-genome high-throughput sequencing, provides a data screening logic and conditions that can accurately present the whole-genome CNV results, excludes interference information, reduces controversial sites, is conducive to the intuitive presentation of the results, reduces the controversy in result interpretation, facilitates correct interpretation, and improves the accuracy and readability of the detection results.

[0090] In this embodiment, a method for processing genomic copy number variation analysis data based on ultra-low-depth whole-genome high-throughput sequencing is provided, which can be used in various computer devices. Figure 2 It is a flowchart of the method for processing genomic copy number variation analysis data based on ultra-low-depth whole-genome high-throughput sequencing according to an embodiment of the present invention. As Figure 2 shown, this process includes the following steps:

[0091] Step S201, obtain the copy number of each partition interval of the genomic sequence sample data. For details, please refer to Figure 1 step S101 of the embodiment shown here, which will not be elaborated here.

[0092] Step S202, for the first partition interval corresponding to the Y chromosome in the partition interval, if the copy number is less than the first preset threshold, then delete the first partition interval. The first partition interval is the partition interval corresponding to the Y chromosome among the multiple partition intervals corresponding to the genomic sequence sample data. In the embodiment of the present invention, before performing interval division, the X chromosome sequence, Y chromosome sequence, and other chromosome sequences are first distinguished, and then interval division is performed on the X chromosome sequence, Y chromosome sequence, and other chromosome sequences respectively.

[0093] Of course, if there is no Y chromosome in the genomic sequence sample data, such as a female genomic sequence sample data, then this screening step does not need to be executed.

[0094] The value range of the first preset threshold is 0.4 - 0.6, and it can be 0.5 for example. The calculation method of the copy number of the first partition interval corresponding to the Y chromosome is: , where ratio is the copy number ratio value.

[0095] In the embodiment of the present invention, by deleting the first partition interval corresponding to the Y chromosome with a smaller copy number, the error interference caused by factors such as reagents and sequencing errors is excluded, and the accuracy and readability of the detection results are further improved.

[0096] Step S203, merge adjacent partition intervals with similar copy numbers to form combined intervals. For details, please refer to Figure 1 step S102 of the embodiment shown here, which will not be elaborated here.

[0097] Step S204, screen the combined intervals according to the interval size to obtain target combined intervals greater than the preset threshold. For details, please refer to Figure 1 step S103 of the embodiment shown here, which will not be elaborated here.

[0098] Step S205, based on the copy number, classify the partition intervals within the target combined intervals respectively, and count the number of intervals in each classification.

[0099] Specifically, in the above step S205, that is, based on the copy number, the partition intervals within the target combination interval are classified respectively, and the number of intervals in each classification is counted, including:

[0100] Step S2051: The partition intervals with a copy number less than the second preset threshold are classified into the first category, and the number of the first intervals is counted.

[0101] Step S2052: The partition intervals with a copy number greater than or equal to the second preset threshold and less than or equal to the third preset threshold are classified into the second category, and the number of the second intervals is counted.

[0102] Step S2053: The partition intervals with a copy number greater than the third preset threshold are classified into the third category, and the number of the third intervals is counted.

[0103] Wherein, the second preset threshold is less than the third preset threshold.

[0104] Specifically, the value range of the second preset threshold is 1.6 - 1.8, and the value range of the third preset threshold is 2.2 - 2.4. For example, the second preset threshold can be 1.7, and the third preset threshold can be 2.3.

[0105] Step S206: For each of the target combination intervals, based on the multiple numbers of intervals corresponding to the target combination interval, the partition intervals in the target combination interval are screened to obtain the target partition intervals.

[0106] In some optional specific implementation manners, step S206, that is, for each of the target combination intervals, based on the multiple numbers of intervals of the target combination interval, the partition intervals in the target combination interval are screened to obtain the target partition intervals, including:

[0107] Step S2061: Based on the multiple numbers of intervals of the target combination interval, the corresponding partition interval screening rule is determined.

[0108] In some optional specific implementation manners, the partition interval screening rule for screening the partition intervals in the target combination interval is: the partition intervals with a copy number within a preset range in the target combination interval are screened out, where the preset range is a set value or is determined based on the copy number of the target combination interval.

[0109] In some optional specific implementation manners, the above step S2061, that is, based on the multiple numbers of intervals of the target combination interval, determining the corresponding partition interval screening rule, includes:

[0110] If the sum of the number of the first intervals and the number of the third intervals is less than or equal to the number of the second intervals, determine the corresponding partition interval screening rule as: screening out the partition intervals in the target combined interval whose copy numbers are within the first preset range;

[0111] If the sum of the number of the second intervals and the number of the third intervals is less than or equal to the number of the first intervals, determine the corresponding partition interval screening rule as: screening out the partition intervals in the target combined interval whose copy numbers are within the second preset range;

[0112] If the sum of the number of the first intervals and the number of the second intervals is less than or equal to the number of the third intervals, determine the corresponding partition interval screening rule as: screening out the partition intervals in the target combined interval whose copy numbers are within the second preset range;

[0113] Wherein, the first preset range is a copy number range set based on the second preset threshold and the third preset threshold; the second preset range is a copy number range set based on the copy number of the target combined interval. The starting value of the second preset range can be the copy number of the target combined interval minus a preset value, and the ending value can be the copy number of the target combined interval plus a preset value. The value range of this preset value can be 0.05 - 0.2. For example, the starting value of the first preset range is greater than the second preset threshold, and the ending value is less than the third preset threshold. For example, the second preset threshold is 1.7 and the third preset threshold is 2.3, then the first preset range is 1.9 - 2.1. The second preset range can be a range centered on the copy number of the target combined interval. For example, the starting value of the second preset range is the copy number of the target combined interval minus 0.1, and the ending value is the copy number of the target combined interval plus 0.1.

[0114] Step S2062, screen the partition intervals in the target combined interval according to the determined partition interval screening rule.

[0115] Step S207, perform plotting and display based on the data in the target partition interval. For example, use the R language to plot the copy numbers in the target partition interval and then display.

[0116] In a specific embodiment, for a male gene sample, by comparing with a detection baseline, 30,894 partition intervals and corresponding copy number ratio (ratio) values are obtained ( Figure 3 ). On this basis, adjacent partition intervals with similar copy numbers are merged to obtain 73 combined intervals (also called segment intervals) representing the copy number level of the fragments and corresponding copy number ratio values ( Figure 4). Subsequently, the copy number ratio values were converted into corresponding copy number values, filtering out empty partition intervals without sequencing reads coverage, and at the same time filtering out partition intervals on the Y chromosome with copy number values less than 0.5 to exclude contamination. Filtering out combined intervals less than 4 Mb, 61 combined intervals of more than 4 Mb and corresponding copy number values were obtained ( Figure 5 ). For each combined interval, copy number statistics and screening were performed on the partition intervals within it, and 6071 high-confidence partition intervals and corresponding copy number values were obtained ( Figure 6 ). Visualizing the copy number values of the partition intervals, it can be seen that the optimized copy number map ( Figure 7 , Figure 7 where the horizontal axis chr1-chr22, as well as chrX and chrY in the figure represent 23 pairs of chromosomes) is more intuitive than the unfiltered copy number map ( Figure 8 , where the horizontal axis chr1-chr22, as well as chrX and chrY in the figure represent 23 pairs of chromosomes) and is easier to judge.

[0117] In another specific embodiment, for a female gene sample, 30321 partition intervals and corresponding copy number ratios were obtained by comparing with the detection baseline ( Figure 9 ). On this basis, adjacent partition intervals with similar copy numbers were merged, and 88 combined intervals representing the copy number level of the fragments and corresponding copy number ratio values were obtained ( Figure 10 ). Subsequently, the copy number ratio values were converted into corresponding copy number values, filtering out empty partition intervals without sequencing reads coverage. Filtering out combined intervals less than 4 Mb, 74 combined intervals of more than 4 Mb and corresponding copy number values were obtained ( Figure 11 ). For each combined interval, copy number statistics and screening were performed on the partition intervals within it, and 6113 high-confidence partition intervals and corresponding copy number values were obtained ( Figure 12 ). Visualizing the copy number values of the partition intervals, it can also be seen that the optimized copy number map ( Figure 13 , where the horizontal axis chr1-chr22, as well as chrX and chrY in the figure represent 23 pairs of chromosomes) is more intuitive than the unfiltered copy number map ( Figure 14 , where the horizontal axis chr1-chr22, as well as chrX and chrY in the figure represent 23 pairs of chromosomes) and is easier to judge.

[0118] In this embodiment, a processing device for genomic copy number variation analysis data is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0119] This embodiment provides a processing device for genomic copy number variation analysis data based on ultra-low depth whole-genome high-throughput sequencing, as Figure 15 shown, including:

[0120] An interval copy number acquisition module 1501 for acquiring the copy number of each partition interval of the genomic sequence sample data;

[0121] A first screening module 1502 for deleting the first partition interval corresponding to the Y chromosome in the partition intervals if the copy number is less than a first preset threshold;

[0122] A merging module 1503 for merging adjacent partition intervals with similar copy numbers to form a combined interval;

[0123] A second screening module 1504 for screening the combined intervals according to the interval size to obtain target combined intervals greater than a preset threshold;

[0124] A classification and statistics module 1505 for classifying the partition intervals within the target combined intervals respectively based on the copy number and counting the interval numbers of each classification;

[0125] A third screening module 1506 for screening the partition intervals in each target combined interval based on the multiple interval numbers corresponding to the target combined interval to obtain target partition intervals;

[0126] A plotting and display module 1507 for plotting and displaying based on the data in the target partition intervals.

[0127] In some alternative embodiments, the third screening module 1506 includes:

[0128] A screening rule determination unit for determining the corresponding partition interval screening rule based on the multiple interval numbers of the target combined interval;

[0129] A screening unit for screening the partition intervals in the target combined interval according to the determined partition interval screening rule.

[0130] In some alternative embodiments, the partition interval screening rule for screening the partition intervals in the target combined interval is: screening out the partition intervals with copy numbers within a preset range in the target combined interval, where the preset range is a set value or is determined based on the copy number of the target combined interval.

[0131] In some alternative embodiments, the classification and statistics module 1505 includes:

[0132] The first classification and statistics unit is used to divide the partition intervals with copy numbers less than the second preset threshold into the first category, and statistically obtain the number of the first intervals;

[0133] The second classification and statistics unit is used to divide the partition intervals with copy numbers greater than or equal to the second preset threshold and less than or equal to the third preset threshold into the second category, and statistically obtain the number of the second intervals;

[0134] The third classification and statistics unit is used to divide the partition intervals with copy numbers greater than the third preset threshold into the third category, and statistically obtain the number of the third intervals;

[0135] Wherein, the second preset threshold is less than the third preset threshold.

[0136] In some optional embodiments, the screening rule determination unit is specifically configured to, if the sum of the number of the first intervals and the number of the third intervals is less than or equal to the number of the second intervals, determine the corresponding partition interval screening rule as: screening out the partition intervals with copy numbers within the first preset range in the target combined interval; if the sum of the number of the second intervals and the number of the third intervals is less than or equal to the number of the first intervals, determine the corresponding partition interval screening rule as: screening out the partition intervals with copy numbers within the second preset range in the target combined interval; if the sum of the number of the first intervals and the number of the second intervals is less than or equal to the number of the third intervals, determine the corresponding partition interval screening rule as: screening out the partition intervals with copy numbers within the second preset range in the target combined interval;

[0137] Wherein, the first preset range is a copy number range set based on the second preset threshold and the third preset threshold; the second preset range is a copy number range set based on the copy numbers of the target combined interval.

[0138] In some optional embodiments, the interval copy number acquisition module 1501 includes:

[0139] The copy number ratio acquisition unit is used to acquire the copy number ratios of the respective partition intervals of the genomic sequence sample data;

[0140] The first interval copy number acquisition unit is used to, if the genomic sequence sample data includes Y chromosome data, for the second partition interval corresponding to the Y chromosome and the second partition interval corresponding to the X chromosome, determine the copy number as ; for other partition intervals other than the second partition interval, determine the copy number as ;

[0141] The second interval copy number acquisition unit is configured to, if the genomic sequence sample data does not include Y chromosome data, determine the copy number of each of the partition intervals as ;

[0142] where ratio is the copy number ratio.

[0143] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.

[0144] The processing device for genomic copy number variation analysis data in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0145] An embodiment of the present invention further provides a computer device having the above Figure 15 shown processing device for genomic copy number variation analysis data.

[0146] Please refer to Figure 16 , Figure 16 which is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. As Figure 16 shown, the computer device includes: one or more processors 10, a memory 20, and an interface for connecting each component, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 16 One processor 10 is taken as an example in

[0147] The processor 10 can be a central processor, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above hardware chip can be an application specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.

[0148] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.

[0149] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0150] The memory 20 may include a volatile memory, for example, a random access memory; the memory may also include a non-volatile memory, for example, a flash memory, a hard disk, or a solid-state drive; the memory 20 may also include a combination of the above types of memories.

[0151] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means. Figure 16 Taking connection through a bus as an example.

[0152] The input device 30 may receive input digital or character information, and generate key signal inputs related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (for example, an LED), and a tactile feedback device (for example, a vibration motor), etc. The above display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.

[0153] The computer device further includes a communication interface for the computer device to communicate with other devices or communication networks.

[0154] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0155] A part of the present invention can be applied as a computer program product, for example, computer program instructions, which when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should be able to understand that the forms of existence of computer program instructions in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.

[0156] Although the embodiments of the present invention are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for processing genomic copy number variation analysis data based on ultra-low depth whole-genome high-throughput sequencing, characterized in that, The method includes: Obtaining the copy number of each partition interval of the genomic sequence sample data; the partition interval is a bin interval; For the first partition interval corresponding to the Y chromosome in the partition interval, if the copy number is less than the first preset threshold, then delete the first partition interval; Merging adjacent partition intervals with similar copy numbers to form a combined interval; Screening the combined intervals according to the interval size to obtain target combined intervals greater than the preset threshold; Based on the copy number, classifying the partition intervals within the target combined intervals respectively, and counting the interval numbers of each classification; For each of the target combined intervals, based on the multiple interval numbers corresponding to the target combined interval, screening the partition intervals in the target combined interval to obtain target partition intervals; Performing drawing display based on the data in the target partition intervals.

2. The method according to claim 1, wherein The value range of the first preset threshold is 0.4 - 0.

6.

3. The method according to claim 1, characterized in that, The step of, for each of the target combined intervals, based on the multiple interval numbers of the target combined interval, screening the partition intervals in the target combined interval to obtain target partition intervals includes: Based on the multiple interval numbers of the target combined interval, determining the corresponding partition interval screening rule; According to the determined partition interval screening rule, screening the partition intervals in the target combined interval.

4. The method according to claim 3, wherein The partition interval screening rule for screening the partition intervals in the target combined interval is: screening out the partition intervals with copy numbers within a preset range in the target combined interval, where the preset range is a set value or is determined based on the copy number of the target combined interval.

5. The method according to claim 4, wherein The step of, based on the copy number, classifying the partition intervals within the target combined intervals respectively, and counting the interval numbers of each classification includes: Dividing the partition intervals with copy numbers less than the second preset threshold into the first category, and statistically obtaining the first interval number; Dividing the partition intervals with copy numbers greater than or equal to the second preset threshold and less than or equal to the third preset threshold into the second category, and statistically obtaining the second interval number; Dividing the partition intervals with copy numbers greater than the third preset threshold into the third category, and statistically obtaining the third interval number; Among them, the value range of the second preset threshold is 1.6 - 1.8, and the value range of the third preset threshold is 2.2 - 2.

4.

6. The method according to claim 5, wherein The step of, based on the multiple interval numbers of the target combined interval, determining the corresponding partition interval screening rule includes: If the sum of the first interval number and the third interval number is less than or equal to the second interval number, then determine the corresponding partition interval screening rule as: screening out the partition intervals with copy numbers within the first preset range in the target combined interval; If the sum of the second interval number and the third interval number is less than or equal to the first interval number, then determine the corresponding partition interval screening rule as: screening out the partition intervals with copy numbers within the second preset range in the target combined interval; If the sum of the number of the first intervals and the number of the second intervals is less than or equal to the number of the third intervals, determine the corresponding partition interval screening rule as: screening out the partition intervals in the target combined interval whose copy numbers are within the second preset range; Wherein, the first preset range is a copy number range set based on the second preset threshold and the third preset threshold; the second preset range is a copy number range set based on the copy number of the target combined interval.

7. The method according to claim 6, characterized in that, The first preset range is 1.9 - 2.1; and / or, the starting value of the second preset range is the copy number of the target combined interval minus 0.1, and the ending value of the second preset range is the copy number of the target combined interval plus 0.

1.

8. The method according to any one of claims 1 to 3, characterized in that, The obtaining of the copy numbers of the respective partition intervals of the genomic sequence sample data includes: Obtaining the copy number ratio of each partition interval of the genomic sequence sample data; If the genomic sequence sample data includes Y chromosome data, for the second partition interval corresponding to the Y chromosome and the second partition interval corresponding to the X chromosome, determine that the copy number is ; for other partition intervals outside the second partition interval, determine that the copy number is ; If the genomic sequence sample data does not include Y chromosome data, the copy number of each of the partition intervals is determined as ; Wherein, ratio is the copy number ratio.

9. A processing device for genomic copy number variation analysis data based on ultra-low depth whole-genome high-throughput sequencing, characterized in that, The device includes: An interval copy number obtaining module, configured to obtain the copy numbers of the respective partition intervals of the genomic sequence sample data; the partition interval is a bin interval; A first screening module, configured to, for a first partition interval corresponding to the Y chromosome in the partition intervals, if the copy number is less than a first preset threshold, delete the first partition interval; A merging module, configured to merge adjacent partition intervals with similar copy numbers to form a combined interval; A second screening module, configured to screen the combined intervals according to the interval size to obtain target combined intervals greater than a preset threshold; A classification and statistics module, configured to classify the partition intervals within the target combined interval respectively based on the copy numbers, and count the number of intervals in each classification; A third screening module, configured to, for each of the target combined intervals, screen the partition intervals in the target combined interval based on the multiple numbers of intervals corresponding to the target combined interval to obtain target partition intervals; A plotting and display module, configured to perform plotting and display based on the data in the target partition intervals.

10. A computer device, characterized in that, Including: A memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the processing method of genomic copy number variation analysis data based on ultra-low-depth whole-genome high-throughput sequencing according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Genome copy number variation analysis method and device and storage medium

    CN117995269A