Method, device and equipment for generating simulated CNV-seq data and computer readable storage medium

By analyzing the reads distribution characteristics of CNV-seq data of normal population samples, the whole genome sequencing data is processed, and simulated CNV-seq data is generated, which solves the problem of existing methods ignoring batch effects, and generates simulated data closer to the characteristics of the real sample to meet the development and testing needs of CNV detection tools.

CN119964637APending Publication Date: 2025-05-09AEGICARE (SHENZHEN) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411844474.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing simulated CNV-seq data generation method ignores the batch effect in whole genome sequencing, resulting in the generated data not enough to reflect the distribution characteristics of the real samples and cannot meet the development and testing requirements of CNV detection tools.

Method used

By obtaining the CNV-seq data of normal population samples, dividing the chromosome into multiple segments, counting the number of reads in each segment, and obtaining the reads distribution information. Then, based on the whole genome sequencing data, reads are randomly selected in each segment according to the reads distribution information, simulated CNV-seq data, and copy number variation events can be simulated to meet specific needs.

Benefits of technology

The generated simulated CNV-seq data is closer to the real sample characteristics, and can overcome the sequencing preferences of whole genome sequencing data while maintaining the preferences of CNV-seq, and meet the development and testing needs of CNV detection tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964637A_ABST
    Figure CN119964637A_ABST
Patent Text Reader

Abstract

The invention discloses a method, device and equipment for generating simulated CNV-seq data and a computer readable storage medium, and the method comprises the following steps: obtaining CNV-seq data of at least one normal population sample, dividing a chromosome into a plurality of sections according to a preset length, counting the reads number of each section, and calculating the read number of each section; obtaining reads distribution information of each chromosome segment of the CNV-seq data; the method comprises the following steps: acquiring whole genome sequencing data, dividing a chromosome position into a plurality of sections according to a preset length, dividing all reads of the whole genome sequencing data into corresponding sections, and randomly extracting the reads divided into the sections in each section according to read distribution information of each chromosome section of CNV-seq data to obtain a corresponding number of reads; and obtaining the simulated CNV-seq data. According to the method and the device, the simulated CNV-seq data closer to real sample characteristics can be generated, so that the development and test requirements of copy number variation detection tools or software can be better met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of copy number variation analysis, and in particular to a method, apparatus, device and computer-readable storage medium for generating simulated CNV-seq data. Background Art

[0002] Copy number variation (CNV) refers to the increase or decrease in the number of copies of large fragments of DNA (usually larger than 1kb) in the genome. CNVs are widely present in the human genome and are closely related to a variety of genetic diseases, cancers and other complex diseases. Genome copy number variation sequencing (CNV-seq) based on next generation sequencing (NGS) technology provides a new means for the detection of copy number variations. CNV-seq is a high-throughput sequencing technology specifically used to detect CNVs. CNV-seq uses NGS technology to perform low-depth whole-genome sequencing of sample DNA, and identifies regions with abnormal copy numbers by comparing the sequencing depth differences between the sample to be tested and the reference sample. Compared with other technologies, CNV-seq technology has the following advantages: 1) Wide detection range: covering whole chromosome aneuploidy, large deletion / duplication and whole genome CNVs; 2) High throughput: can better alleviate the contradiction of severe shortage of current genetic testing service supply; 3) Simple operation: simple experimental process, high degree of automation of data analysis, clear quality control standards, short reporting cycle, which can significantly save manpower and reduce the risk of human error; 4) Detection of low proportion of chimeras: CMA technology cannot accurately analyze chimeras <30%, while CNV-seq technology can detect lower proportion of chimeras, and under ideal conditions, it can detect as low as 5% of chromosome aneuploidy chimeras; 5) Detection of low DNA sample amount: Studies have shown that CNV-seq technology can accurately detect DNA samples as low as 10-50ng, which is more clinically applicable.

[0003] In the process of developing and optimizing CNV detection tools or software, a large amount of high-quality CNV-seq data is needed for testing, especially for samples of rare diseases, which is crucial to evaluating the performance of the algorithm and improving the reliability of the tool. However, the cost and difficulty of obtaining samples in reality are high, especially for samples with clear phenotypes or etiologies. It is even more difficult to obtain CNV-seq data resources for these samples. In order to solve this problem, generating simulated CNV-seq data based on existing whole genome sequencing (WGS) data can not only significantly reduce the cost of sample collection, but also improve the efficiency and breadth of tool testing, and help the in-depth development of research. At present, the tools used to simulate sequencing data mainly include Seqtk, Wigsam, VarBen, and software developed by Beijing Kexun Biotechnology Co., Ltd. (patent number: CN108229101A). Among them, Seqtk simulates CNVseq sequencing by directly sampling to reduce the sequencing depth; Wigsam and VarBen simulate variant data by changing reads; Beijing Kexun's tool uses normal distribution to determine the copy number to generate simulated data. However, these methods usually ignore the batch effects in whole genome sequencing, so the generated data are not sufficient to reflect the distribution characteristics of sequencing data within the same batch, and are insufficient in the accuracy of the true sample distribution.

[0004] Therefore, there is an urgent need for a CNV-seq simulation data that can generate more realistic sample characteristics, so as to better meet the development and testing needs of CNV detection tools. Summary of the invention

[0005] Based on the above problems, the present application provides a method, apparatus, device and computer-readable storage medium for generating simulated CNV-seq data.

[0006] This application adopts the following technical solutions:

[0007] The first aspect of the present application discloses a method for generating simulated CNV-seq data, comprising: obtaining CNV-seq data of at least one normal population sample, dividing the chromosome into multiple segments according to a predetermined length, counting the number of reads in each segment, and obtaining reads distribution information of each chromosome segment of the CNV-seq data; obtaining whole genome sequencing data, dividing the chromosome position into multiple segments according to the predetermined length, and dividing all reads of the whole genome sequencing data into corresponding segments, and then according to the reads distribution information of each chromosome segment of the CNV-seq data, in each segment, randomly extracting a corresponding number of reads from the reads divided into the segment; and obtaining the simulated CNV-seq data. It should be noted that in the present application, the sequencing preference of the CNV-seq data at different chromosome positions is obtained by analyzing the reads distribution characteristics of the CNV-seq sequencing data of the real sample; then, the whole genome sequencing data is processed, and reads sampling is performed according to the reads distribution characteristics of the CNV-seq sequencing data. Randomness is introduced in the sampling process to generate simulated CNV-seq data with similar or identical reads distribution characteristics to the CNV-seq sequencing data. This can overcome the sequencing preference problem of the whole genome sequencing data while maintaining the sequencing preference of CNV-seq, and generate simulated CNV-seq data that is closer to the characteristics of the real samples, thereby better meeting the development and testing requirements of copy number variation detection tools or software.

[0008] In one implementation of the present application, after extracting a corresponding number of reads in each segment, it also includes: simulating copy number variation events to obtain the simulated CNV-seq data containing specific copy number variation. It should be noted that by simulating copy number variation events, simulated CNV-seq data containing specific copy number variation in a specific region can be obtained, and these simulated CNV-seq data containing specific copy number variation are used for the development or testing of copy number variation analysis tools or software, which can help evaluate the detection performance of the tool or software for the specific copy number variation.

[0009] In one implementation of the present application, the simulation of copy number variation events includes: obtaining information about the copy number variation event, the information about the copy number variation event includes the variation type, the variation ratio and the region where the variation is located, and counting the number of reads in the region where the variation is located, denoted as n; determining the copy number variation type and adjusting the number of reads, including: if the variation type is deletion, randomly deleting the reads in the region where the variation is located until the number of reads is n*(1-the variation ratio); if the variation type is repetition, continuing to randomly extract reads in the region where the variation is located until the number of reads is n*(1+the variation ratio).

[0010] In one implementation of the present application, the acquired CNV-seq data or the whole genome sequencing data is a BAM file.

[0011] In one implementation of the present application, the sequencing depth of the obtained CNV-seq data is 1X to 3X. It should be noted that the sequencing depth of the conventional CNV-seq data is 1X to 3X, and the sequencing depth of the CNV-seq data used in the present application is also 1X to 3X, thereby enabling the sequencing data to meet the analysis requirements of copy number variation.

[0012] In one implementation of the present application, the sequencing depth of the whole genome sequencing data obtained is ≥10X. It should be noted that the sequencing depth of the whole genome sequencing data needs to be higher than the sequencing depth of the CNV-seq data. In the present application, the sequencing depth of the genome sequencing data is ≥10X, which can provide high-quality sequencing data and facilitate more redundant reads for us to extract in the future.

[0013] In one implementation of the present application, the sequencing depth of the whole genome sequencing data obtained is 10X to 50X. It should be noted that if the sequencing depth of the whole genome sequencing data is too low, the number of reads in some segments may be insufficient for extraction, and if the sequencing depth of the whole genome sequencing data is too high, the file may be large, which is not conducive to processing.

[0014] In one implementation of the present application, the predetermined length is 1 kb to 20 kb. It should be noted that when dividing multiple segments, an appropriate segmentation unit is required. If the segmentation unit is too small, the processing is more complicated. If the segmentation unit is too large, the preference analysis may be too rough.

[0015] In one implementation of the present application, the predetermined length is 5 kb to 15 kb.

[0016] In one implementation of the present application, the number of CNV-seq data of the normal population sample is not less than 10, and the calculation of the number of reads in each segment is to calculate the average number of reads in each segment of multiple samples. It should be noted that if the number of real CNV-seq samples used is too low, there may be individual differences in the preference analysis of CNV-seq data.

[0017] The second aspect of the present application discloses a device for generating simulated CNV-seq data, including: a CNV-seq data reads distribution analysis module, used to obtain CNV-seq data of at least one normal population sample, divide the chromosome into multiple segments according to a predetermined length, calculate the number of reads in each segment, and obtain the reads distribution information of each chromosome segment of the CNV-seq data; a whole genome sequencing data processing module, used to obtain whole genome sequencing data, divide the chromosome position into multiple segments according to the predetermined length, and divide all reads of the whole genome sequencing data into corresponding segments, and then according to the reads distribution information of each chromosome segment of the CNV-seq data, in each segment, randomly extract a corresponding number of reads from the reads divided into the segment to obtain the simulated CNV-seq data.

[0018] In one implementation of the present application, a copy number variation event simulation module is also included, which is used to simulate copy number variation events to obtain the simulated CNV-seq data containing specific copy number variation.

[0019] In one implementation of the present application, the copy number variation event simulation module includes: obtaining information about the copy number variation event, the information about the copy number variation event includes the variation type, the variation ratio and the region where the variation is located, and counting the number of reads in the region where the variation is located, denoted as n; and for determining the copy number variation type and adjusting the number of reads, including: if the variation type is deletion, randomly deleting the reads in the region where the variation is located until the number of reads is n*(1-the variation ratio); if the variation type is repetition, continuing to randomly extract reads in the region where the variation is located until the number of reads is n*(1+the variation ratio). For example, the information of the simulated copy number variation event is "Chr4:331568-2010962, DEL, 0.5", where "Chr4:331568-2010962" represents the base region from 331568 to 2010962 of chromosome 4, which is the variation region of the copy number variation event; "DEL" represents deletion, which is the variation type of the copy number variation event; "0.5" represents the deletion ratio of 50%, which is the variation ratio (the normal copy number should be 2, and the deletion variation ratio is 0.5, which means that the copy number of this region is 2*(1-0.5)=1).

[0020] The third aspect of the present application discloses a device for generating simulated CNV-seq data, comprising a memory and a processor, wherein the memory is used to store a program, and the processor executes the program stored in the memory to implement the method for generating simulated CNV-seq data as described in the first aspect of the present application.

[0021] The fourth aspect of the present application discloses a computer-readable storage medium, in which a program is stored. The program can be executed by a processor to implement the method for generating simulated CNV-seq data as described in the first aspect of the present application.

[0022] The beneficial effects of this application are:

[0023] In the present application, the reads distribution characteristics of CNV-seq sequencing data of real samples are analyzed to obtain the sequencing preferences of CNV-seq data at different chromosome positions; then, the whole genome sequencing data is processed, and reads sampling is performed according to the reads distribution characteristics of the CNV-seq sequencing data. Randomness is introduced in the sampling process to generate simulated CNV-seq data with similar or identical reads distribution characteristics to the CNV-seq sequencing data. This can overcome the sequencing preference problem of whole genome sequencing data while maintaining the sequencing preference of CNV-seq, and generate simulated CNV-seq data that is closer to the characteristics of real samples, thereby better meeting the development and testing requirements of copy number variation detection tools or software. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a flowchart of the method for generating simulated CNV-seq data involved in the present application.

[0025] Figure 2 It is a reads distribution diagram of CNV-seq data involved in a specific embodiment of the present application.

[0026] Figure 3 Schematic diagram of a device for generating simulated CNV-seq data involved in the present application.

[0027] Figure 4 is a schematic diagram of an apparatus for generating simulated CNV-seq data involved in the present application. DETAILED DESCRIPTION

[0028] The present invention is further described in detail below by specific embodiments in conjunction with the accompanying drawings. In the following embodiments, many detailed descriptions are intended to enable the present application to be better understood. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other materials and methods. In some cases, some operations related to the present application are not shown or described in the specification, in order to avoid the core part of the present application being overwhelmed by too much description, and for those skilled in the art, it is not necessary to describe these related operations in detail, and the related operations can be fully understood according to the description in the specification and the general technical knowledge in the art.

[0029] In addition, the features, operations or characteristics described in the specification can be combined in any appropriate manner to form various implementations. At the same time, the steps or actions in the method description can also be interchanged or adjusted in a manner that is obvious to those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for the purpose of clearly describing a certain embodiment and are not meant to be a required sequence, unless otherwise specified that a certain sequence must be followed.

[0030] In the process of developing and optimizing CNV (Copy Number Variation) detection tools, a large amount of high-quality CNV-seq (copy number variation sequencing) data is needed as a reference, especially for samples of rare diseases, which is crucial to evaluating the performance of the algorithm and improving the reliability of the tool. For example, to test the detection ability of a new software for a certain type of CNV, it is usually necessary to obtain biological samples containing this type of CNV and obtain its CNV-seq data. However, in reality, it is difficult to collect biological samples containing all types of this type of CNV. Therefore, in reality, high-quality CNV-seq data resources are very limited, especially for samples with clear phenotypes or etiologies, and the cost of obtaining and organizing them is high.

[0031] In order to solve this problem, generating simulated CNV-seq data based on existing WGS (Whole Genome Sequencing) data can not only significantly reduce the cost of sample collection, but also improve the efficiency and breadth of tool testing, and help the in-depth development of research. At present, the tools used to simulate sequencing data mainly include Seqtk, Wigsam, VarBen and software developed by Beijing Kexun Biotechnology Co., Ltd. (patent number CN108229101A). Among them, Seqtk simulates CNVseq sequencing by directly sampling and reducing the sequencing depth; Wigsam and VarBen simulate variant data by changing reads; Beijing Kexun's tools use normal distribution to determine the copy number to generate simulated data. However, these methods usually ignore the batch effect in whole genome sequencing, so the generated data is not enough to reflect the distribution characteristics of sequencing data in the same batch, and there are deficiencies in the accuracy of the real sample distribution.

[0032] In view of this, the present application generates CNV-seq copy number variation simulation data that is closer to the characteristics of real samples by modeling the distribution characteristics of real CNV-seq sequencing data, combined with simulated copy number variability events, so as to better meet the needs of CNV detection tool development and testing. The advantages of this application include: 1) Improving the availability of test data: The present invention can use existing WGS samples to generate simulated CNVseq data with specified characteristics, solving the problem of scarce test data. 2) Retaining the sequencing preference of real samples: By sampling based on the sequencing preference of real CNV-seq samples, the generated data is more in line with the characteristics of real CNV-seq samples, making the test results more valuable for reference. 3) Customizable copy number variation simulation: Users can freely define the variation position, type and variation ratio to generate simulated samples with diversity and complexity, which are suitable for various testing scenarios.

[0033] The present application provides a method, apparatus, device and computer-readable storage medium for generating simulated CNV-seq data.

[0034] As described above, the present application relates to a method for generating simulated CNV-seq data.

[0035] Figure 1 It is a flowchart of the method for generating simulated CNV-seq data involved in the present application.

[0036] like Figure 1As shown, in a specific embodiment, the method for generating simulated CNV-seq data may include: step S11, obtaining CNV-seq data of at least one normal population sample. The CNV-seq data of the normal population sample refers to the sequencing data obtained by performing CNV-seq on the samples of the phenotypically normal population. The gender of the population can be divided into male (46, XY) and female (46, XX). For the copy number variation of the sex chromosomes, it is necessary to distinguish between male samples and female samples. For the copy number variation of the autosomes, it is not necessary to distinguish between male samples and female samples. The number of samples can be multiple, for example, the number of samples is not less than 10, so that the data bias caused by individual differences can be avoided by statistically calculating the mean of the number of reads of multiple samples. In a specific embodiment, the sequencing depth of the obtained CNV-seq data can be no less than 1X. Preferably, the sequencing depth of the CNV-seq data is 1X to 3X. In a specific embodiment, the file format of the acquired CNV-seq data is in BAM format. BAM (Binary Alignment / Map) file is a standard file format in high-throughput sequencing data analysis, which is used to store aligned sequencing reads information. BAM files are widely used in genomics, transcriptomics, epigenetics and other fields, especially in alignment, variation detection, expression quantification and other analyses. Through some common tools, BAM files can be easily processed and analyzed to extract target information.

[0037] In a specific embodiment, the method for generating simulated CNV-seq data may include: step S12, for the acquired CNV-seq data, the chromosome is divided into a plurality of segments according to a predetermined length. In a specific embodiment, the predetermined length may be 1kb to 20kb. Preferably, the predetermined length may be 5kb to 15kb. For example, the chromosome may be divided into a plurality of segments according to a length of 10kb, the first base to the 10k base of chromosome 1 being the first segment, the 10k+1 base to the 20k base of chromosome 1 being the second segment, and so on, all chromosomes are segmented.

[0038] In a specific embodiment, the method for generating simulated CNV-seq data may include: step S13, counting the number of reads in each segment to obtain the reads distribution information of each chromosome segment of the CNV-seq data. In a specific embodiment, for the case of multiple samples, the average number of reads in each segment of each sample is counted. It should be noted that reads refer to the sequences generated by the sequencer, and a read refers to a short sequence output by the sequencer. For paired-end sequencing, that is, sequencing the two ends of the same DNA fragment separately, two related reads will be obtained, namely R1 reads and R2 reads. For single-end sequencing, only one end of the DNA fragment is sequenced, and each fragment generates a separate read. In a specific embodiment, for a pair of reads generated by paired-end sequencing, it can be counted as 1 read when counting the number of reads.

[0039] In a specific embodiment, the method for generating simulated CNV-seq data may include: step S21, obtaining whole genome sequencing data. Whole genome sequencing (WGS) data refers to data containing all DNA sequence information of an individual obtained by high-throughput sequencing of the entire genome. WGS data is a large-scale and complex biological information resource that covers all types of variations in the genome, including single nucleotide polymorphisms (SNPs), insertions / deletions (InDels), copy number variations (CNVs), structural variations (SVs), etc. In a specific embodiment, the sequencing depth of the acquired whole genome sequencing data may be no less than 10X. Preferably, the sequencing depth of the whole genome sequencing data is 10X to 50X. In a specific embodiment, the file format of the acquired whole genome sequencing data is in BAM format. In a specific embodiment, the file format of the acquired whole genome sequencing data may also be in FASTQ format.

[0040] In a specific embodiment, the method for generating simulated CNV-seq data may include: step S22, similar to the segmentation of CNV-seq data, the whole genome sequencing data also divides the chromosome into multiple segments, and divides all reads of the whole genome sequencing data into corresponding segments. For example, as in step S12, the chromosome can be divided into multiple segments according to the length of 10kb, the first segment is the first base to the 10k base of chromosome 1, the second segment is the 10k+1 base to the 20k base of chromosome 1, and so on, all chromosomes are segmented.

[0041] In a specific embodiment, the method for generating simulated CNV-seq data may include: step S23, according to the reads distribution information of each chromosome segment of the CNV-seq data, randomly extracting a corresponding number of reads from the reads divided into the segment in each segment to obtain a preliminary simulation file. That is, according to the reads distribution information of each chromosome segment of the CNV-seq data obtained in step 13, then, in each segment, a corresponding number of reads are randomly extracted from the reads divided into the segment (for example, the number of reads of the CNV-seq data in the segment from the 1st base to the 10kth base of chromosome 1 is a, and the number of reads in the segment from the 10k+1th base to the 20kth base of chromosome 1 is b; then, for the whole genome sequencing data, in the segment from the 1st base to the 20kth base of chromosome 1 In the segment from base 1 to base 10k, a reads or a* coefficient c (c>1, so that reads can be extracted in a certain proportion) are randomly extracted from the reads from base 1 to base 10k of chromosome 1, and in the segment from base 10k+1 to base 20k of chromosome 1, b reads or b* coefficient c reads are randomly extracted from the reads from base 10k+1 to base 20k of chromosome 1. It should be noted that by introducing randomness, data bias can be reduced. When extracting reads, a pair of reads generated by double-end sequencing needs to be extracted together. The preliminary simulation file obtained in step 23 has the same distribution of reads in each chromosome segment as the distribution of reads in each chromosome segment of the CNV-seq data. This can overcome the problem of sequencing preference differences in whole genome sequencing data and generate simulated CNV-seq data that is closer to the characteristics of real samples.

[0042] In a specific embodiment, the method for generating simulated CNV-seq data may include: step S31, simulating copy number variation events. It should be noted that if it is desired to generate simulated CNV-seq data of normal human samples, it is not necessary to perform the step of simulating copy number variation events; if it is desired to generate simulated CNV-seq data with specific copy number variation, the step of simulating copy number variation events may be performed. By simulating copy number variation events, simulated CNV-seq data containing specific copy number variation in a specific region can be obtained. These simulated CNV-seq data containing specific copy number variation are used for the development or testing of copy number variation analysis tools or software, which can help evaluate the detection performance of the tool or software for the specific copy number variation.

[0043] In a specific embodiment, the method for generating simulated CNV-seq data may include: step S32, obtaining information of copy number variation events, and counting the number of reads in the region where the variation is located, recorded as n. Wherein, the information of the copy number variation event includes the variation type, variation ratio and variation region. For example, the information of the simulated copy number variation event is "Chr4:331568-2010962, DEL, 0.5", where "Chr4:331568-2010962" represents the base region from 331568 to 2010962 of chromosome 4, which is the variation location of the copy number variation event; "DEL" represents deletion, which is the variation type of the copy number variation event; "0.5" represents that the deletion ratio is 50%, which is the variation ratio (the normal copy number should be 2, and the deletion variation ratio is 0.5, which means that the copy number of this region is 2*(1-0.5)=1). For example, the information of the simulated copy number variation event is "Chr4:331568-2010962, DUP, 0.5", which means that there is a 50% duplication (duplicate) in the base region from 331568 to 2010962 of chromosome 4 (the normal copy number should be 2, and the duplication variation ratio is 0.5, which means that the copy number of this region is 2*(1+0.5)=3).

[0044] In a specific embodiment, the method for generating simulated CNV-seq data may include: step S33, determining the variation type.

[0045] In a specific embodiment, the method for generating simulated CNV-seq data may include: step S341, if the variation type is deletion (DEL), the reads in the region where the variation is located are randomly deleted until the number of reads is n*(1-variation ratio). For example, if the information of the simulated copy number variation event is "Chr4:331568-2010962, DEL, 0.5", the reads in "Chr4:331568-2010962" are randomly deleted until the number of reads is n*(1-0.5)=0.5n.

[0046] In a specific embodiment, the method for generating simulated CNV-seq data may include: step S342, if the variation type is duplication (DUP), then continue to randomly extract reads in the region where the variation is located until the number of reads is n*(1+variation ratio). For example, if the information of the simulated copy number variation event is "Chr4:331568-2010962, DUP, 0.5", then in the region "Chr4:331568-2010962", continue to randomly extract reads from the reads in this region of the whole genome sequencing data until the number of reads is n*(1+0.5)=1.5n.

[0047] In a specific embodiment, the generated simulated CNV-seq data is a BAM file.

[0048] In a specific embodiment, the overall steps of the method for generating simulated CNV-seq data can be as follows:

[0049] 1) Analyze 10 BAM files of CNVseq sequencing data with a sequencing depth of 1-3X. First, bin all chromosomes, and set the length of each bin to 10kb. By traversing each BAM file, determine the bin region to which each read belongs. Then, count the number of reads in each bin, and calculate the total number of reads in all bins in all chromosomes. Finally, average the results of the 10 samples to obtain the average result of the read distribution. Figure 2 is a reads distribution diagram of CNV-seq data involved in a specific embodiment of the present application, wherein: Figure 2 Only the read distribution of some regions of chromosome 1 is schematically shown. Figure 2 In order to visualize the area with 0 reads, the number of reads is adjusted to 1 when the number of reads is 0. Otherwise, the value after log is negative infinity, which affects the visualization effect.

[0050] 2) Bin the BAM file of whole genome sequencing (sequencing depth is 30X, hereinafter referred to as raw.bam) according to the same bin length, divide the reads into different bins according to the chromosome position, and perform 1:1 sampling of the reads in each bin according to the average result of the reads distribution obtained in step 1) to generate a preliminary sampling result, temporarily named tmp.bam. (Sampling is based on the number of reads in the corresponding interval of CNV-seq data. For example, if the number of reads in a certain interval of CNV-seq data is 100, 100 reads are randomly selected in the corresponding interval of the whole genome sequencing data).

[0051] 3) Simulate CNV mutation events, and input mutation parameters as "Chr4:331568-2010962, DEL, 0.5" (indicating that there is a 50% mosaic deletion at position 331568-2010962 of chromosome 4, which will lead to Wolf-Hirschhorn syndrome (WHS). WHS is a rare genetic disease mainly caused by partial deletion of the short arm of chromosome 4. The phenotype of WHS includes severe intellectual disability, developmental delay, unique facial features and multiple congenital abnormalities). First, count the reads in the Chr4:331568-2010962 region of tmp.bam in step 2), record the number of reads in this region as n, and record the final target number of reads in this region as new_n (new_n = (1-0.5)*n). Randomly delete the number of reads in this region of tmp.bam in step 2) to new_n to generate the final simulated BAM file.

[0052] The present application also relates to a device for generating simulated CNV-seq data.

[0053] Figure 3 8 is a schematic diagram of an apparatus 800 for generating simulated CNV-seq data according to the present application. Figure 3 As shown, the device 800 for generating simulated CNV-seq data may include a CNV-seq data reads distribution analysis module 801 and a whole genome sequencing data processing module 802. The device 800 for generating simulated CNV-seq data may also include a copy number variation event simulation module 803. The functions of each module may be as follows:

[0054] In a specific embodiment, the CNV-seq data reads distribution analysis module 801 can be used to obtain CNV-seq data of at least one normal population sample, divide the chromosome into multiple segments according to a predetermined length, calculate the number of reads in each segment, and obtain the reads distribution information of each chromosome segment of the CNV-seq data. The specific functions can refer to the corresponding content in the aforementioned method for generating simulated CNV-seq data, which will not be repeated here.

[0055] In a specific embodiment, the whole genome sequencing data processing module 802 can be used to obtain whole genome sequencing data, divide the chromosome position into multiple segments according to a predetermined length, and divide all reads of the whole genome sequencing data into corresponding segments, and then randomly extract a corresponding number of reads from the reads divided into the segment in each segment according to the reads distribution information of each chromosome segment of the CNV-seq data. The specific functions can refer to the corresponding contents in the aforementioned method for generating simulated CNV-seq data, which will not be repeated here.

[0056] In a specific embodiment, the copy number variation event simulation module 803 can be used to obtain information about copy number variation events, and count the number of reads in the variation region of the copy number variation event, recorded as n; and used to determine the copy number variation type and adjust the number of reads, including: if the variation type is deletion, the reads in the variation region are randomly deleted until the number of reads is n*(1-variation ratio); if the variation type is repetition, reads are continuously randomly extracted in the variation region until the number of reads is n*(1+variation ratio). The specific functions can refer to the corresponding contents in the aforementioned method for generating simulated CNV-seq data, which will not be described in detail here.

[0057] The present application also relates to a device for generating simulated CNV-seq data.

[0058] Figure 4 is a schematic diagram of an apparatus 100 for generating simulated CNV-seq data involved in the present application. Figure 4 As shown, the device 100 for generating simulated CNV-seq data may include a processor 10, a memory 20, and a computer program 21 (also referred to as a computer-readable storage medium 21) stored in the memory 20. The computer program 21 may be run on the processor 10, and when the processor 10 executes the computer program 21, the method for generating simulated CNV-seq data is implemented, for example, Figure 1 Alternatively, when the processor 60 executes the computer program 62, the functions of each module in the above-mentioned device 800 for generating simulated CNV-seq data are implemented, for example Figure 3 The functions of module 801, module 802 and module 803 are shown.

[0059] The device 100 for generating simulated CNV-seq data may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The device 100 for generating simulated CNV-seq data may include, but is not limited to, a processor 10 and a memory 20. Those skilled in the art will appreciate that Figure 4The device 100 for generating simulated CNV-seq data is merely an example and does not constitute a limitation on the device 100 for generating simulated CNV-seq data. For example, the device 100 for generating simulated CNV-seq data may include more or fewer components than shown in the figure, or may combine certain components, or may have different components. For example, the device 100 for generating simulated CNV-seq data may also include input and output devices, network access devices, buses, etc.

[0060] The processor 10 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0061] The memory 20 may be an internal storage unit of the device 100 for generating simulated CNV-seq data, such as a hard disk or a memory. The memory 20 may also be an external storage device of the device 100 for generating simulated CNV-seq data, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the device 100 for generating simulated CNV-seq data. Further, the memory 20 may include both an internal storage unit of the device 100 for generating simulated CNV-seq data and an external storage device.

[0062] The memory 20 may be used to store the computer program 21 and other programs and data required by the apparatus 100 for generating simulated CNV-seq data.

[0063] The memory 20 may also be used to temporarily store data that has been output or is to be output.

[0064] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0065] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0066] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0067] In the embodiments provided in the present application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0068] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0069] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0070] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0071] The above contents are further detailed descriptions of the present application in combination with specific implementation methods, and it cannot be determined that the specific implementation of the present application is limited to these descriptions. For ordinary technicians in the technical field to which the present application belongs, several simple deductions or substitutions can be made without departing from the concept of the present application.

Claims

1. A method for generating simulated CNV-seq data, characterized in that: include: Obtaining CNV-seq data of at least one normal population sample, dividing the chromosome into multiple segments according to a predetermined length, counting the number of reads in each segment, and obtaining reads distribution information of each chromosome segment of the CNV-seq data; Acquire whole genome sequencing data, divide the chromosome position into multiple segments according to the predetermined length, and divide all reads of the whole genome sequencing data into corresponding segments, and then randomly extract a corresponding number of reads from the reads divided into the segment in each segment according to the reads distribution information of each chromosome segment of the CNV-seq data; The simulated CNV-seq data is obtained.

2. The method according to claim 1, characterized in that After extracting a corresponding number of reads in each segment, the method further includes: simulating copy number variation events to obtain the simulated CNV-seq data containing specific copy number variation.

3. The method according to claim 2, characterized in that The simulated copy number variation event includes: Obtaining information about the copy number variation event, the information about the copy number variation event including the variation type, variation ratio, and variation region, and counting the number of reads in the variation region, recorded as n; Determining the copy number variation type and adjusting the number of reads includes: if the variation type is deletion, randomly deleting the reads in the region where the variation is located until the number of reads is n*(1-the variation ratio); if the variation type is duplication, continuing to randomly extract reads in the region where the variation is located until the number of reads is n*(1+the variation ratio).

4. The method according to claim 1, characterized in that The obtained CNV-seq data or the whole genome sequencing data is a BAM file; Preferably, the sequencing depth of the CNV-seq data obtained is 1X to 3X; Preferably, the sequencing depth of the whole genome sequencing data obtained is ≥10X; Preferably, the sequencing depth of the acquired whole genome sequencing data is 10X to 50X.

5. The method according to claim 1, characterized in that: The predetermined length is 1 kb to 20 kb; Preferably, the predetermined length is 5 kb to 15 kb.

6. The method according to claim 1, characterized in that The number of CNV-seq data of the normal population sample is not less than 10, and the calculation of the number of reads in each segment is to calculate the average value of the number of reads in each segment of multiple samples.

7. A device for generating simulated CNV-seq data, characterized in that include: A CNV-seq data reads distribution analysis module, used to obtain CNV-seq data of at least one normal population sample, divide the chromosome into multiple segments according to a predetermined length, calculate the number of reads in each segment, and obtain the reads distribution information of each chromosome segment of the CNV-seq data; The whole genome sequencing data processing module is used to obtain the whole genome sequencing data, divide the chromosome position into multiple segments according to the predetermined length, and divide all the reads of the whole genome sequencing data into corresponding segments, and then according to the reads distribution information of each chromosome segment of the CNV-seq data, in each segment, randomly extract a corresponding number of reads from the reads divided into the segment to obtain the simulated CNV-seq data.

8. The device according to claim 7, characterized in that Also included is a copy number variation event simulation module, which is used to simulate copy number variation events to obtain the simulated CNV-seq data containing specific copy number variation; Preferably, the copy number variation event simulation module includes: obtaining information about the copy number variation event, the information about the copy number variation event includes the variation type, the variation ratio and the region where the variation is located, and counting the number of reads in the region where the variation is located, denoted as n; and for determining the copy number variation type and adjusting the number of reads, including: if the variation type is deletion, randomly deleting the reads in the region where the variation is located until the number of reads is n*(1-the variation ratio); if the variation type is repetition, continuing to randomly extract reads in the region where the variation is located until the number of reads is n*(1+the variation ratio).

9. A device for generating simulated CNV-seq data, characterized in that The method comprises a memory and a processor, wherein the memory is used to store a program, and the processor implements the method for generating simulated CNV-seq data according to any one of claims 1 to 6 by executing the program stored in the memory.

10. A computer-readable storage medium, characterized in that: The storage medium stores a program, which can be executed by a processor to implement the method for generating simulated CNV-seq data as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Targeted sequencing data simulation method and device based on NGS

    CN108229101A