A method and system for determining copy number variations using single nucleotide polymorphisms
By calculating the frequency of single nucleotide polymorphisms to assist in the determination of gene copy number variations, the problem of large errors in existing technologies is solved, and more accurate gene copy number variation detection is achieved, which is suitable for high-throughput sequencing platforms.
Patent Information
- Application Number
- CN202211451852.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-11-21
AI Technical Summary
Existing probabilistic NGS methods suffer from large errors, low efficiency, and subjectivity when detecting copy number variations, and cannot accurately determine the state of gene copy number variations.
By calculating the frequency of single nucleotide polymorphisms (SNPs) within a gene range, and combining the sequencing data of reference and test samples, stable SNP loci groups are screened. The frequency of SNPs is used to help determine gene copy number variations, and a high-throughput sequencing platform is used for data processing.
It improves the accuracy of gene copy number variation determination, reduces errors caused by insufficient sequencing data quality and depth, and provides more accurate copy number variation results.
Smart Images

Figure CN115762630B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a method for determining copy number variation and a data processing system, in particular to a method for determining copy number variation using single nucleotide polymorphism and a data processing system for implementing the method. BACKGROUND
[0002] With the development of individualized medicine and the concept of "precision medicine", tumor drug treatment has developed rapidly, and more gene mutations related to drug treatment efficacy prediction have been found and confirmed in clinical research. Traditional gene mutation detection methods such as Sanger sequencing, pyrosequencing and real-time fluorescent PCR can only detect single gene or part of exon mutations of single gene. Simultaneous detection of multiple genes using the above traditional gene mutation detection methods requires large sample size, longer detection time and more work. High-throughput sequencing (High-throughput sequencing, HTS) or next-generation sequencing technology (Next-generation sequencing technology, NGS) can simultaneously sequence millions or even billions of DNA fragments, allowing simultaneous detection of up to hundreds of tumor-related genes, whole exons and whole genomes at a relatively low cost, without increasing the sample size. Due to its advantages in throughput, cost and efficiency, NGS has shown great application prospects in solid tumor somatic gene mutations.
[0003] Copy number variation plays an important role in human diseases and biology. For example, germ cell copy number variation has a huge impact on development, such as trisomy 21, 18 and 13, which leads to Down syndrome, Edwards syndrome and Patau syndrome, respectively. On the other hand, somatic copy number variation is often found in cancer patients, and copy number variation is a major driver of tumor development and drug resistance. The Pan-Cancer Genome Analysis report found that copy number amplification of MYC and copy number deletion of PTEN and TP53 were found in many tumor types. In acute myeloid leukemia (AML), deletion involving chromosome 5 and 7 often occurs in patients with cytogenetic risk. In multiple myeloma, deletion of chromosome 17 is associated with more aggressive disease, and acquisition of chromosome 17 deletion during disease progression leads to a worse prognosis. In addition, TP53 chromosome deletion or chromosome 1 amplification leads to abnormality of myeloma pathogenesis-related genes (such as CKS1B, MCL1), which is associated with poor prognosis. Therefore, the description of cancer-related copy number variation events is of great significance for determining patient subgroups and insights into prognosis and potential treatment strategies.
[0004] Currently, NGS methods for detecting CNV mainly use 1) read depth: according to the read depth of a sliding window to indicate the copy number amplification and deletion; 2) pair-end method: according to the distance between the two ends of the pair-end and the difference on the reference genome to confirm the copy number variation; 3) sequence assembly method: after assembling short reads, the difference between them and the reference genome is found to confirm the copy number variation. The first method based on read depth is the most widely used method, and the last two methods are mainly used for the detection of other structural variations, such as transversion. The core technology of the read depth detection method is mainly based on a probability statistical model. The detection method based on probability statistics has a hypothesis premise: the read depth and the number of copy number variations are linearly related, that is, we assume that the sequencing process is uniform, and the read depth of a specific window on the chromosome is subject to a certain distribution, such as Poisson distribution, Gaussian distribution, etc. If the read depth of the sliding window increases or decreases, it means that the copy number amplification or deletion occurs. However, the cumulative error in the sequencing process makes the read depth and the number of copy number variations not linearly related, so the method is based on the wrong assumption, and the error of the result is large.
[0005] In addition, according to the cell marker gene, it can also be identified manually, which is low in efficiency and has a lot of subjectivity.
[0006] In order to solve the limitations of the above methods, a method is developed to assist in determining the state of gene copy number variation based on the frequency of many single nucleotide polymorphisms in the gene range. SUMMARY
[0007] To solve the above problems, the application provides a method for assisting in judging copy number variation by using single nucleotide polymorphism, which comprises the following steps:
[0008] Step 1: calculating the gene CNV score according to the third data of the detection sample and the reference sample, screening the genes that may have copy number variation according to the CNV score, and selecting the second genome;
[0009] Step 2: performing snp analysis according to the third data of the reference sample to screen out a first site group with stable mutation frequency from the sites of the second genome; performing snp analysis according to the third data of the detection sample to obtain the snp base information of the detection sample at the first site group as the fifth data;
[0010] Step 3: judging the gene copy number variation of the detection sample according to the fifth data;
[0011] The third data is obtained by removing low-quality sequences, aligning to a reference genome, de-duplication and quality control from the original sequencing data.
[0012] The low-quality sequences refer to sequences with average base quality and read length lower than the set value, and an exemplary low-quality sequence removal method is given in Embodiment One of the present application.
[0013] The present application can be used for identifying the copy number variation state of a first genome from the sequencing results of a second-generation sequencing platform, also known as a high-throughput sequencing platform (such as illumina or MGI sequencing data).
[0014] The third data is obtained by the following method:
[0015] Sequencing is performed on a detection sample or a reference sample to obtain original sequencing data including first genome information, and the first genome is one or more genes selected according to the analysis purpose;
[0016] The clean data obtained by removing low-quality sequences from the original sequencing data is used as first data, and the first data is aligned to a reference genome, so that the sequencing sequences in the first data are positioned on the relevant genes to obtain alignment result data as second data.
[0017] De-duplication and quality control are performed on the second data to obtain third data.
[0018] The method for aligning to a reference genome is: using the bwa software, using the mem algorithm, aligning the first data obtained by removing low-quality sequences to the human reference genome hg19, so that the sequencing sequences are positioned on the relevant genes.
[0019] The reference sample is selected as a human white blood cell sample, and to ensure the diversity of the sample and the statistical significance of the results, blood samples of multiple people can be selected, and in the present application, blood samples of 20 people are selected to construct the CNV and snp baseline.
[0020] As a preferred scheme, step 1 specifically includes the following steps:
[0021] Step 1.1 obtains the corrected normalized sequencing depth of the detection sample and the reference sample;
[0022] When multiple reference samples are used, the corrected normalized sequencing depth of all reference samples is averaged to calculate the average reference sequencing depth;
[0023] Step 1.2 calculates the CNV copy number score of the detection sample; the calculation method is: for each region, the CNV copy number score of the region = 2 x corrected normalized sequencing depth / average reference sequencing depth;
[0024] Step 1.3, screening samples with 0≤CNV copy number score≤1 or 4≤CNV copy number score<6 as genes possibly having copy number variation.
[0025] As a preferred solution, step 2 specifically comprises the following steps:
[0026] Step 2.1, obtaining base information of snp sites of the second genome from the third data of the reference sample, including reads supporting base mutation and reads of reference base; removing sites with sequencing depth lower than a set value to obtain fourth data of the reference sample;
[0027] Using the fourth data of the reference sample to perform the following calculation and screening:
[0028] snp mutation frequency = number of reads supporting base mutation / (number of reads of reference base + number of reads supporting base mutation) ;
[0029] The screening standard is that if the mutation frequency of the snp site is between 0.4 and 0.6, it is determined to be stable and is classified into the first site group.
[0030] Step 2.2, obtaining base information of snp sites of the first site group from the third data of the detection sample, including reads supporting base mutation and reads of reference base; removing sites with sequencing depth lower than a set value to obtain fifth data of the reference sample.
[0031] Step 3.1, using the fifth data of the detection sample to perform the following calculation:
[0032] snp mutation frequency = number of reads supporting base mutation / (number of reads of reference base + number of reads supporting base mutation) ;
[0033] Step 3.2, if the distribution frequency of the snp is <0.3, it indicates that the copy number may have deletion, and if the distribution frequency of the snp after gene standardization correction is >0.7, it indicates that the copy number may have amplification.
[0034] The application further provides a system for judging copy number variation using single nucleotide polymorphism, comprising:
[0035] a memory storing executable instructions, and; and
[0036] one or more processors in communication with the memory to execute the executable instructions to accomplish the following operations:
[0037] Step 1: according to the third data of the detection sample and the reference sample, calculate the CNV score of the gene, screen the gene with possible copy number variation according to the CNV score, and select into the second genome;
[0038] Step 2: according to the third data of the reference sample, perform snp analysis, and screen out a first site group with stable mutation frequency from the sites of the second genome; according to the third data of the detection sample, perform snp analysis, and obtain the snp mutation frequency data of the detection sample in the first site group as the fifth data;
[0039] Step 3: according to the fifth data, judge the copy number variation of the gene in the second genome of the detection sample;
[0040] The third data is obtained after removing low-quality sequences, aligning to the reference genome, de-duplication and quality control of the original sequencing data.
[0041] As preferred, the one or more processors communicate with the memory to execute executable instructions to accomplish the following operations:
[0042] Step 1.1: obtaining the corrected normalized sequencing depth of the detection sample and the reference sample;
[0043] When multiple reference samples are used, the corrected normalized sequencing depths of all reference samples are averaged to calculate the average reference sequencing depth;
[0044] Step 1.2: calculating the CNV copy number score of the detection sample;
[0045] The calculation method is: for each region, the CNV copy number score of the region = 2 x corrected normalized sequencing depth / average reference sequencing depth;
[0046] Step 1.3: screening the sample with 0≤CNV copy number score≤1 or 4≤CNV copy number score<6 as the gene with possible copy number variation;
[0047] Step 2.1: obtaining the base information of the snp site of the second genome from the third data of the reference sample, including reads supporting base mutation and reads of reference base; removing sites with sequencing depth lower than a set value to obtain the fourth data of the reference sample;
[0048] The fourth data of the reference sample is used to calculate and screen as follows:
[0049] snp mutation frequency = number of reads supporting base mutation / (number of reads of reference base + number of reads supporting base mutation);
[0050] The standard for screening is that the mutation frequency of the snp site is between 0.4 and 0.6, which is determined as stable and belongs to the first site group;
[0051] Step 2.2 obtains the base information of the snp site of the first site group from the third data of the detection sample, including reads supporting base mutation and reads of reference base; removes the sites with a sequencing depth lower than a set value, and obtains the fifth data of the reference sample;
[0052] Step 3.1 calculates as follows by using the fifth data of the detection sample:
[0053] The snp mutation frequency = the number of reads supporting base mutation / (the number of reads of reference base + the number of reads supporting base mutation) ;
[0054] Step 3.2, if the distribution frequency of the snp is less than 0.3, it indicates that the copy number may be deleted, and if the distribution frequency of the snp after gene standardization correction is greater than 0.7, it indicates that the copy number may be amplified.
[0055] The application has the advantages that:
[0056] The snp distribution frequency is used to assist in judging the gene copy number variation, and compared with the traditional simple sequencing depth judgment, the result is more accurate.
[0057] In the preferred scheme, errors caused by low sequencing data quality and insufficient depth can be excluded. BRIEF DESCRIPTION OF DRAWINGS
[0058] Other features, objects and advantages of the application will become more apparent from the following detailed description of non-limiting embodiments made with reference to the accompanying drawings:
[0059] Figure 1 Figure 1 is a gene CNV score distribution graph of Example 1;
[0060] Figure 2 Figure 2 is a snp mutation frequency baseline graph of the reference sample (white blood cells) of Example 1;
[0061] Figure 3 Figure 3 is a snp mutation frequency distribution graph of the first site group of the detection sample of Example 1;
[0062] Figure 4 Figure 4 is a technical solution implementation flowchart of the application;
[0063] Figure 5 Figure 5 is a schematic diagram of the data processing system shown in Example 2 of the application. DETAILED DESCRIPTION
[0064] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0065] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0066] Example 1
[0067] In this embodiment, tumor-related genes were selected as the first genome.
[0068] A tumor sample submitted by a hospital in Guangdong was selected as the test sample; 20 white blood cell samples extracted from blood samples from 20 healthy adults were selected as the reference samples.
[0069] Sequencing and preprocessing of sequencing data can be accomplished using conventional methods and software. This embodiment provides an exemplary operation as follows:
[0070] The Illumina next-generation sequencing platform was used to sequence the test samples and reference samples to obtain raw sequencing data.
[0071] Preprocessing of raw sequencing data of the test samples: a. Removal of low-quality sequences from the raw data: Due to limitations of the sequencer, the base quality of sequencing data often fluctuates and is low during the sequencing start and end stages. Even with short inserts, sequencing may still occur, meaning the sequence may contain sequencing adapters. Therefore, to ensure the accuracy of the sample analysis results, it is necessary to remove low-quality sequences or sequencing adapters introduced during sequencing. The removal criteria are as follows: remove Illumina adapters used in sequencing; remove bases with an average base quality below 15 using a 4bp sliding window; and remove reads shorter than 26 bytes to obtain clean data, which will be referred to as the first data for ease of description.
[0072] b. Alignment with the reference genome: Using the bwa algorithm, the clean data sequence after removing low-quality sequences is aligned with the human reference genome, thereby locating the sequencing sequence to the relevant gene and obtaining the second data.
[0073] c. Deduplication of alignment results: Because PCR amplification is involved in the experiment, in order to avoid the influence of amplification data, the MarkDuplicates module of Picard software is used to deduplicatize the aligned data (i.e., the second data).
[0074] d. Quality control of sample sequencing: The sequencing quality and alignment quality of the samples are evaluated based on the alignment results to determine whether the sample quality meets the requirements. Samples that do not meet the requirements cannot ensure the accuracy of the analysis results. Quality control is performed on the deduplicated data, with the following standards: target rate > 60%, Q30 > 85%, alignment rate > 95%, and average depth after deduplication > 1000x. The second set of data undergoes deduplication and quality control to obtain the third set of data.
[0075] The raw sequencing data of each reference sample were processed in the same way to obtain the third data for each reference sample.
[0076] Step 1
[0077] Calculation and screening of gene CNV scores:
[0078] 1.1 Obtaining the corrected standardized sequencing depth
[0079] The multicov function of the bedtools software was used to calculate the sequencing depth of each gene in the first genome region of the sample as the third data; the total sequencing depth of the entire sample in the first genome region was calculated; the sequencing depth of each gene and the calculated total sequencing depth were normalized; and the Loess method was used to correct the normalized sequencing depth results based on GC content.
[0080] For the reference samples, the data were processed in the same way. After obtaining the processing results for each reference sample, the average value was taken to obtain the average reference sequencing depth.
[0081] 1.2 Calculate the CNV copy number score of the detected sample
[0082] The calculation method is as follows: For each region, the CNV copy number score of that region = 2 × calibrated normalized sequencing depth / average reference sequencing depth;
[0083] 1.3 Screening for genes that may have copy number variations
[0084] The screening criteria are as follows: if the gene CNV copy number score is ≥4 and <6, amplification may occur; if the gene CNV score is ≤1, deletion may occur; genes in both cases are selected into the second genome.
[0085] See the instruction manual appendix Figure 1 The CNV copy number scores of the EGFR, MET, and ERBB2 genes marked in the figure meet the screening criteria, indicating that they are likely to have copy number variations.
[0086] Step 2
[0087] The processing of SNP data specifically includes:
[0088] 2.1 Screening for sites with stable SNP mutation frequencies
[0089] By analyzing and summarizing the mutation frequencies of SNP sites in 20 white blood cell samples, sites with stable mutation frequencies in the second genome were screened out.
[0090] Specifically, using the `mpileup` function of the `bcftools` software, the base information of the SNP sites of the second genome was obtained from the third data of each reference sample. This information was then used as input and read using R language to obtain all relevant reads in the second genome of the sample, including reads supporting base mutations and reads supporting reference bases. For each SNP site, if the sum of the number of reads supporting base mutations and the number of reads supporting reference bases was less than 200, it indicated insufficient sequencing depth, which might affect the accuracy of the analysis results, and these reads needed to be removed. The data obtained after the above processing was called the fourth data of the reference sample.
[0091] The following calculations and screenings were performed using the fourth data from the reference sample:
[0092] SNP mutation frequency = number of reads supporting base mutation / (number of reads for reference base + number of reads supporting base mutation)
[0093] The screening criteria were that if the mutation frequency of the SNP site was between 0.4 and 0.6, it was considered stable and classified into the first site group.
[0094] To visually display the results, a baseline of the white blood cell SNP mutation frequency can be established; see the instruction manual appendix. Figure 2 .
[0095] SNP sites with unstable mutation frequencies in leukocytes generally also exhibit instability in other cells, making them unsuitable for analysis using the method of this invention. Removing data from these sites can avoid erroneous results caused by abnormalities in some SNP sites.
[0096] 2.2 Obtain the fifth data of the test sample
[0097] The base information of the SNP sites of the first site group is obtained from the third data of the test sample, including reads that support base mutations and reads of reference bases; sites with sequencing depths lower than a set value are removed, and the resulting data is called the fifth data of the reference sample.
[0098] Specifically, using the mpileup function of the bcftools software, the base information of the SNP site of the first locus group is obtained from the third data of the test sample. For each SNP site, if the number of reads supporting the base mutation plus the number of reads of the reference base is less than 200, it indicates that the sequencing depth is insufficient and may affect the accuracy of the analysis results, and these reads need to be removed. The data obtained after the above processing is called the fifth data of the test sample.
[0099] Step 3: Determining copy number mutation.
[0100] 3.1 Calculation of SNP mutation frequency in the detection sample
[0101] The following calculations were performed using the fifth data point from the test sample:
[0102] SNP mutation frequency = number of reads supporting base mutation / (number of reads for reference base + number of reads supporting base mutation);
[0103] 3.2 Copy Number Variation Judgment
[0104] Based on the SNP mutation frequency data calculated in step 3.1, genes with SNP mutation frequencies < 0.3 are selected as the fourth genome that may have copy number deletions, and genes with SNP mutation frequencies > 0.7 are selected as the fifth genome that may have copy number amplifications.
[0105] For a more intuitive display, the calculated SNP mutation frequency data can be used to create graphs, as shown in the instruction manual. Figure 3 As shown, from Figure 3 The differences in the sites where copy number variations may occur among different genes are clearly visible. In the absence of CNV deletions and mutations, the distribution of SNP mutation frequencies is generally between 0.4 and 0.6, close to the baseline SNP frequency in leukocytes. CDKN2A and CDKN2B genes may have copy number deletions, while ERBB2 and MET genes may have copy number amplifications.
[0106] As an example of the advantages of the technical solution of the present invention, according to Figure 1 Judging from the results, the EGFR gene, which is very likely to have copy number variation, is likely to be affected. Figure 3 Based on this, the possibility of copy number mutation can be largely ruled out. This proves that the SNP-assisted judgment used in this invention has higher accuracy than existing technologies.
[0107] It should be noted that the significance of drawing is to present the data more intuitively. Even without drawing, the technical solution of this invention can be completed through data processing.
[0108] Example 2
[0109] This application also provides a system for determining copy number variations using single nucleotide polymorphisms, which can be implemented via mobile terminals, personal computers (PCs), tablets, servers, etc. See below for reference. Figure 5 It shows a schematic diagram of a system suitable for implementing the embodiments of this application, which uses single nucleotide polymorphism to determine copy number variations.
[0110] like Figure 5 As shown, the computer system 300 includes one or more processors, a communication unit, etc. The one or more processors include, for example, one or more central processing units (CPUs) 301, and / or one or more graphics processing units (GPUs) 313, etc. The processors can perform various appropriate actions and processes according to executable instructions stored in read-only memory (ROM) 302 or executable instructions loaded from storage unit 308 into random access memory (RAM) 303. The communication unit 312 may include, but is not limited to, a network interface card (NIC), which may include, but is not limited to, an InfiniBand (IB) NIC.
[0111] The processor can communicate with read-only memory 302 and / or random access memory 303 to execute executable instructions, and is connected to communication unit 312 via bus 304 and communicates with other target devices via communication unit 312, thereby completing the operation corresponding to any of the methods provided in the embodiments of this application, for example:
[0112] Step 1.2 Calculate the CNV copy number score of the detected sample;
[0113] The calculation method is as follows: For each region, the CNV copy number score of that region = 2 × calibrated normalized sequencing depth / average reference sequencing depth;
[0114] Step 1.3 Select samples with 0 ≤ CNV copy number score ≤ 1 or 4 ≤ CNV copy number score < 6 as genes that may have copy number variations;
[0115] Step 2.1 Obtain the base information of the SNP site of the second genome from the third data of the reference sample, including reads supporting base mutations and reads of reference bases; remove sites with sequencing depths lower than a set value, and the resulting data is called the fourth data of the reference sample;
[0116] The sequencing depth setting can be adjusted by the operator based on factors such as the number of amplifications during sequencing, with the aim of reducing interference and improving analysis accuracy. In this embodiment, steps 2.1 and 2.2 are both set to 200.
[0117] The following calculations and screenings were performed using the fourth data from the reference sample:
[0118] SNP mutation frequency = number of reads supporting base mutation / (number of reads for reference base + number of reads supporting base mutation);
[0119] The screening criteria were that if the mutation frequency of the SNP site was between 0.4 and 0.6, it was considered stable and classified into the first site group.
[0120] Step 2.2 Obtain the base information of the SNP sites of the first site group from the third data of the test sample, including reads that support base mutations and reads of reference bases; remove sites with sequencing depths lower than a set value, and the resulting data is called the fifth data of the reference sample;
[0121] Step 3.1 Using the fifth data point from the test sample, perform the following calculations:
[0122] SNP mutation frequency = number of reads supporting base mutation / (number of reads for reference base + number of reads supporting base mutation);
[0123] Step 3.2: If the distribution frequency of SNP is <0.3, it indicates that copy number deletion may occur; if the distribution frequency of SNP after gene standardization correction is >0.7, it indicates that copy number amplification may occur.
[0124] Furthermore, RAM 303 can also store various programs and data required for device operation. CPU 301, ROM 302, and RAM 303 are interconnected via bus 304. With RAM 303 present, ROM 302 is an optional module. RAM 303 stores executable instructions, or executable instructions are written to ROM 302 during runtime. These executable instructions cause processor 301 to perform operations corresponding to the aforementioned communication method. Input / output interface (I / O interface) 305 is also connected to bus 304. Communication unit 312 can be integrated or configured with multiple sub-modules (e.g., multiple IB network cards) linked on the bus.
[0125] The following components are connected to I / O interface 305: input section 306 including keyboard, mouse, etc.; output section 307 including cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; storage section 308 including hard disk, etc.; and communication section 309 including network interface card, such as LAN card, modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed.
[0126] It needs to be explained, such as Figure 5The architecture shown is only one optional implementation. In practice, the above can be modified according to actual needs. Figure 5 The number and type of components can be selected, deleted, added, or replaced; different functional components can also be implemented by separate or integrated configurations. For example, the GPU and CPU can be separated or the GPU can be integrated into the CPU, the communication unit 312 can be separated or integrated into the CPU or GPU, and so on. All these alternative implementations fall within the protection scope of this invention.
[0127] Specifically, according to this application, the reference process Figure 5 The described process can be implemented as a computer program product. For example, this application provides a computer program product including computer-readable instructions that, when executed by a processor, perform the following operations: In such an embodiment, the computer program product can be downloaded and installed from a network via a communication unit 309, and / or read and installed from a removable medium 311. When the computer program product is executed by a central processing unit (CPU) 301, it performs the functions defined in the method of this application.
[0128] The technical solutions of this application may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The order of steps used to illustrate the method is provided only for the purpose of more clearly illustrating the technical solutions. Unless specifically limited, the method steps of this application are not limited to the order specifically described above. Furthermore, in some embodiments, this application may also be implemented as a storage medium for storing computer program products.
Claims
1. A method for assisting in calling copy number variations using single nucleotide polymorphisms, the method comprising: obtaining a plurality of single nucleotide polymorphisms; obtaining a plurality of copy number variations; and determining a relationship between the plurality of single nucleotide polymorphisms and the plurality of copy number variations. The method comprises the following steps: Step 1: calculating gene CNV scores according to the third data of the detection sample and the reference sample, screening genes with possible copy number variations according to the CNV scores, and selecting into a second genome; Step 2: performing snp analysis according to the third data of the reference sample, and screening a first site group with stable mutation frequency from the sites of the second genome; performing snp analysis according to the third data of the detection sample to obtain snp base information of the detection sample at the first site group as fifth data; The method for screening the first site group with stable mutation frequency comprises the following steps: Step 2.1: obtaining base information of snp sites of the second genome from the third data of the reference sample, including reads supporting base mutations and reads of reference bases; removing sites with sequencing depth lower than a set value to obtain fourth data of the reference sample; The fourth data of the reference sample is used to perform the following calculations and screening: snp mutation frequency = number of reads supporting base mutations / (number of reads of reference bases + number of reads supporting base mutations); The screening standard is that if the snp site mutation frequency is between 0.4 and 0.6, it is determined to be stable and belongs to the first site group; Step 3: judging the gene copy number variation of the detection sample according to the fifth data, comprising the following steps: Step 3.1: using the fifth data of the detection sample to perform the following calculations: snp mutation frequency = number of reads supporting base mutations / (number of reads of reference bases + number of reads supporting base mutations); Step 3.2: if the distribution frequency of snp is less than 0.3, it indicates that the copy number may be deleted, and if the distribution frequency of snp after standardization correction is greater than 0.7, it indicates that the copy number may be amplified; The third data is obtained after removing low-quality sequences, aligning to a reference genome, de-duplicating and quality control on raw sequencing data, wherein the low-quality sequences refer to sequences with average base quality and read length lower than a set value.
2. The method of claim 1, wherein, The third data is obtained by the following method: sequencing the detection sample or the reference sample to obtain raw sequencing data including information of a first genome, wherein the first genome is one or more genes selected according to analysis purposes; removing low-quality sequences from the raw sequencing data to obtain clean data as first data, and aligning the first data to a reference genome to position the sequencing sequences in the first data to related genes to obtain alignment result data as second data; de-duplicating and quality controlling the second data to obtain the third data.
3. The method of claim 1, wherein, The reference sample is selected from human white blood cell samples.
4. The method of claim 1, wherein, The calculation of the gene CNV score comprises the following steps: Step 1.1: obtaining corrected standardization sequencing depth of the detection sample and the reference sample; when multiple reference samples are used, the corrected standardization sequencing depth of all reference samples is averaged to calculate an average reference sequencing depth; Step 1.2: calculating CNV copy number scores of the detection sample; The calculation method is: for each region, the CNV copy number score of the region = 2 x corrected normalized sequencing depth / average reference sequencing depth.
5. The method of claim 4, wherein, The method for screening the gene possibly having copy number variation comprises the following steps: Step 1.3, screening the sample with 0≤CNV copy number score≤1 or 4≤CNV copy number score<6 as the gene possibly having copy number variation.
6. The method of claim 1, wherein, The method for acquiring the fifth data comprises the following steps: Step 2.2, acquiring the base information of the snp sites in the first site group from the third data of the detection sample, including reads supporting base mutation and reads of reference base; removing the sites with sequencing depth lower than a set value to obtain the fifth data of the detection sample.
7. A data processing system comprising: a memory storing executable instructions; and, one or more processors in communication with the memory to execute the executable instructions to accomplish the following operations: Step 1: calculating the CNV score of the gene according to the third data of the detection sample and the reference sample, screening the gene possibly having copy number variation according to the CNV score, and selecting the second gene group; Step 2: performing snp analysis according to the third data of the reference sample, and screening the first site group with stable mutation frequency from the sites in the second gene group; performing snp analysis according to the third data of the detection sample to obtain the snp mutation frequency data of the detection sample in the first site group as the fifth data; The method for screening the first site group with stable mutation frequency comprises the following steps: Step 2.1, acquiring the base information of the snp sites in the second gene group from the third data of the reference sample, including reads supporting base mutation and reads of reference base; removing the sites with sequencing depth lower than a set value to obtain the fourth data of the reference sample; The following calculation and screening are performed by using the fourth data of the reference sample: snp mutation frequency = number of reads supporting base mutation / (number of reads of reference base + number of reads supporting base mutation); The screening standard is that the snp site with mutation frequency between 0.4 and 0.6 is determined to be stable and belongs to the first site group; Step 3: judging the copy number variation of the gene in the second gene group of the detection sample according to the fifth data, comprising the following steps: Step 3.1, performing the following calculation by using the fifth data of the detection sample: snp mutation frequency = number of reads supporting base mutation / (number of reads of reference base + number of reads supporting base mutation); Step 3.2, if the distribution frequency of the snp is less than 0.3, it indicates that the copy number may be deleted, and if the distribution frequency of the snp is greater than 0.7, it indicates that the copy number may be amplified. The third data is obtained by removing low-quality sequences, aligning to a reference genome, removing duplicates and quality control from raw sequencing data, and the low-quality sequences refer to sequences with average base quality and read length lower than a set value.
8. The data processing system of claim 7, wherein, The one or more processors are in communication with the memory to execute the executable instructions to accomplish the following operations: Step 1.1 Obtain the corrected normalized sequencing depth of the detection sample and the reference sample; When multiple reference samples are used, the corrected normalized sequencing depths of all reference samples are averaged to calculate the average reference sequencing depth; Step 1.2 Calculate the CNV copy number score of the detection sample; The calculation method is: for each region, the CNV copy number score of the region = 2 x corrected normalized sequencing depth / average reference sequencing depth; Step 1.3 Screen the samples with 0≤CNV copy number score≤1 or 4≤CNV copy number score<6 as genes that may have copy number variations; Step 2.1 Obtain the base information of the snp sites of the second genome from the third data of the reference sample, including reads supporting base mutations and reads of reference bases; remove sites with sequencing depth lower than a set value to obtain the fourth data of the reference sample; Using the fourth data of the reference sample, the following calculations and screenings are performed: snp mutation frequency = number of reads supporting base mutations / (number of reads of reference bases + number of reads supporting base mutations); The screening standard is that if the snp site mutation frequency is between 0.4 and 0.6, it is determined to be stable and is classified into the first site group; Step 2.2 Obtain the base information of the snp sites of the first site group from the third data of the detection sample, including reads supporting base mutations and reads of reference bases; remove sites with sequencing depth lower than a set value to obtain the fifth data of the detection sample; Step 3.1 Use the fifth data of the detection sample to perform the following calculations: snp mutation frequency = number of reads supporting base mutations / (number of reads of reference bases + number of reads supporting base mutations); Step 3.2 If the distribution frequency of snp is <0.3, it indicates that the copy number may have a deletion, and if the distribution frequency of the standardized and corrected snp of the gene is >0.7, it indicates that the copy number may have an amplification.
Citation Information
Patent Citations
Methods for detecting copy-number variations in next-generation sequencing
CN108292327A
Method for detecting haploid copy number variation of tumor unicell genome
CN110029157A