Method for detecting copy number change based on amplicon capture technology and application thereof
Through the sequencing depth stability sorting and control sample selection of internal reference areas, combined with random forest algorithm and subject working characteristic curve analysis, the deviation problems caused by sequencing data and PCR reactions in amplicon capture technology were solved, and accurate and stable analysis of copy number changes in the target area was achieved.
Patent Information
- Application Number
- CN202510433697.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
AI Technical Summary
Existing amplicon capture techniques have differences in sequencing data and PCR reaction bias when analyzing copy number changes, resulting in poor analysis accuracy and stability.
In the amplicon capture technology, the sequencing depth stability sorting of the internal reference region and the selection of control samples, combined with the random forest algorithm and subject working characteristic curve analysis, the deviations between the sequencing data and PCR reactions are eliminated, and the copy number of the target region is accurate and stable.
Accurate and stable analysis of copy number changes in the target area is achieved, the deviation caused by sequencing data and PCR reactions is reduced, and the reliability of the analysis is improved.
Smart Images

Figure CN120350099A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of bioinformatics processing, and particularly to a method for detecting copy number variation based on amplicon capture technology and its application. Background Art
[0002] Amplicon Sequencing is a technology used to amplify specific nucleotide sequences. The specific operations include: first, designing specific primers for the target region, and amplifying the target region into a large number of replicates, namely amplicons, through polymerase chain reaction (PCR reaction); then enriching these amplicons by using the index sequences carried by the primers. High-throughput sequencing is performed on the enriched amplicons to obtain the sequencing data of the amplicons, and the obtained sequencing data can be widely used to study the variations of specific genes or genomic regions. This method can achieve the amplification and sequencing of a large number of samples and multiple target regions in each sample simultaneously, with high processing efficiency and relatively low cost.
[0003] Amplicon Sequencing usually adopts the multiplex PCR method. However, the preference and instability of the multiplex PCR reaction are likely to affect the amplification of the target regions of the samples, resulting in significant differences in the sequencing data between samples in the same batch and between different batches, affecting the statistics and analysis of the sequencing results, and it is difficult to obtain accurate information on the relative copy number of the target regions of the samples, making the analysis accuracy and stability poor. This defect limits the application of amplicon capture sequencing technology in the analysis of copy number variation (CNV). Summary of the Invention
[0004] In view of the problems in the prior art, the present application provides a method for detecting copy number variation based on amplicon capture technology to accurately and stably analyze the copy number variation of the target region.
[0005] The method for detecting copy number variation based on amplicon capture technology includes the following steps:
[0006] S10. Obtaining the amplicon sequencing data of multiple amplification regions in each DNA sample, where the DNA samples include samples to be analyzed and control samples;
[0007] S20. Aligning the amplicon sequencing data with the genomic reference sequence to obtain the number of Reads and the sequencing depth of each amplification region in each DNA sample;
[0008] S30. Extracting the remaining regions except the target region from all DNA samples, and selecting multiple stable remaining regions as internal reference regions according to the stability ranking of the sequencing depths of all DNA samples based on the same data volume in the remaining regions;
[0009] S40. Calculate the ratios of the number of Reads in the target region to the number of Reads in multiple reference regions for each DNA sample respectively, and calculate the average of the multiple ratios to obtain the A value of the target region in each DNA sample.
[0010] S50. Calculate the ratio of the A value of the target region in the sample to be analyzed to the A value of the target region in the control sample, which is the Ratio value of the target region in the sample to be analyzed.
[0011] S60. Determine the copy number variation of the target region in the sample to be analyzed according to the Ratio value of the target region.
[0012] Optionally, extract candidate samples from all DNA samples, and determine the control sample from the candidate samples through the following steps:
[0013] (1) Determine a target region, and obtain the A value of the target region corresponding to the reference region in each candidate sample.
[0014] (2) Calculate the median Median of the target region in the set of A values of the target region in all candidate samples.
[0015] (3) Select the expected number of samples with the lowest degree of dispersion from all candidate samples according to the degree of dispersion of the A value of the target region relative to the median Median as the control sample of the target region.
[0016] Optionally, there are multiple target regions. Perform steps (1) to (3) for each target region to obtain the control sample of each target region; take the union of the control samples of all target regions and count the number of occurrences of the same control sample to obtain the final control sample set.
[0017] Optionally, extracting candidate samples from all DNA samples includes:
[0018] Perform predictive screening for negative samples among all DNA samples, and use the DNA samples obtained by predictive screening as candidate samples; or use all DNA samples as candidate samples as a whole.
[0019] Optionally, use the DNA sample group with known copy of the target region and the A value of the target region in the DNA sample group as the training set to train and obtain a random forest algorithm training model, and use the random forest algorithm training model to perform predictive screening for negative samples among all DNA samples.
[0020] Optionally, the target region has a corresponding theoretical Ratio value in the theoretical sample. Compare the Ratio value of the target region in the sample to be analyzed with the corresponding theoretical Ratio value to determine the copy number variation of the target region in the sample to be analyzed.
[0021] Optionally, the method includes:
[0022] Providing a DNA sample database with known target region copies, performing steps S10 to S50, and obtaining the Ratio value of the target region in all DNA samples in the DNA sample database;
[0023] Performing receiver operating characteristic curve analysis based on the target region copies of all DNA samples in the DNA sample database and the Ratio value of the target region, and obtaining a Ratio threshold for determining whether the copy of the target region is positive;
[0024] Comparing the Ratio value of the target region of the sample to be analyzed with the corresponding Ratio threshold, and obtaining the change in the copy number of the target region in the sample to be analyzed according to the comparison result.
[0025] Optionally, there are multiple control samples. Calculate the ratios of the target region A values of the sample to be analyzed to those of the multiple control samples respectively, and obtain the average value of the multiple ratios to obtain the Ratio value of the target region of the sample to be analyzed.
[0026] Optionally, the amplification primers used for the amplification region include primers designed upstream and downstream of the differential bases between the target region and its homologous region. The method includes:
[0027] Previously shielding the homologous region in the genomic reference, and then aligning the amplicon sequencing data with the genomic reference sequence to obtain the Ratio value of the primer amplification region;
[0028] According to the differential bases between the target region and the homologous region, count the ratio of the number of target regions or homologous regions in the primer amplification region to the number of the primer amplification region to obtain the VAF value;
[0029] Combined with the analysis of the Ratio value of the primer amplification region and the VAF value, determine the change in the copy number of the target region in the sample to be analyzed.
[0030] Optionally, the specific way of combined analysis of the Ratio value and the VAF value is:
[0031] Providing multiple theoretical samples, each theoretical sample having a known Ratio value and VAF value according to the copy numbers of the target region and its homologous region, and all theoretical samples are arranged and distributed in ascending order according to the VAF value;
[0032] All theoretical samples are subjected to the first-round screening according to the VAF value of the sample to be analyzed to obtain candidate theoretical samples, and the VAF value of the candidate theoretical samples is adjacent to the VAF value of the sample to be analyzed compared with the VAF values of the remaining theoretical samples;
[0033] The candidate theoretical samples are subjected to a second round of screening according to the Ratio value of the sample to be analyzed to obtain the final theoretical samples. The Ratio value of the final theoretical samples is closer to the Ratio value of the sample to be analyzed than the Ratio values of the other theoretical samples.
[0034] Output the VAF value and Ratio value of the final theoretical samples as the VAF value and Ratio value of the sample to be analyzed, and judge the copy number variation of the target region according to the output VAF value and Ratio value.
[0035] The present application also provides a kit for analyzing the copy number variation of an amplified region, including a specific amplification reaction solution Equinox Multiplex PCR Master Mix and an enrichment amplification reaction solution Equinox Library AmplificationKit.
[0036] Compared with the prior art, the present application can eliminate the sequencing data differences brought by amplicon sequencing and the biases brought by the PCR reaction, so as to realize accurate and stable analysis of the copy number variation of the target region. This method can be applied to analyze the copy number variation of specific genes or genomic regions. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flowchart of a method for detecting copy number variation based on amplicon capture technology in an embodiment;
[0038] Figure 2 It is a flowchart of determining a control sample from candidate samples in an embodiment;
[0039] Figure 3 It is a flowchart of analyzing the copy number of the SMN gene based on amplicon capture technology in an embodiment;
[0040] Figure 4 It is a distribution diagram of known VAF values and Ratio values of some theoretical samples of the SMN gene. Among them, the gray dots represent the distribution of VAF and Ratio values of the training set samples, the red squares represent the theoretical values of VAF and Ratio values, and the blue X represents the distribution of VAF and Ratio values of the sample to be analyzed;
[0041] Figure 5 It is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The technical solutions of the present application will be further described below in conjunction with specific embodiments, but the present application is not limited thereto.
[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0044] See Figure 1 , this application provides a method for detecting copy number variations based on amplicon capture technology, comprising the following steps:
[0045] S10. Obtain amplicon sequencing data of multiple amplification regions in each DNA sample. The DNA samples include samples to be analyzed and control samples. In amplicon capture technology, the amplification region refers to the range of DNA sequences that need to be specifically amplified determined by specific primer design;
[0046] S20. Align the amplicon sequencing data with the genomic reference sequence to obtain the number of Reads and sequencing depth of each amplification region in each DNA sample;
[0047] S30. Extract the remaining regions other than the target region in all DNA samples, and sort according to the stability of the sequencing depth of all DNA samples based on the same data volume in the remaining regions, and select multiple stable remaining regions as internal reference regions;
[0048] S40. Obtain the ratio of the number of Reads in the target region in each DNA sample to the number of Reads in multiple internal reference regions respectively, and calculate the average value of multiple ratios to obtain the A value of the target region in each DNA sample;
[0049] S50. Calculate the ratio of the A value of the target region in the sample to be analyzed to the A value of the target region in the control sample to obtain the Ratio value of the target region in the sample to be analyzed;
[0050] S60. Judge the copy number variation of the target region in the sample to be analyzed according to the Ratio value of the target region.
[0051] This application first performs in-sample normalization on the target regions of all DNA samples, which can eliminate the difference in sequencing depth caused by inconsistent sequencing data volumes of different DNA samples, and then uses the control sample for inter-sample normalization processing, which can reduce the bias caused by PCR reactions (especially multiplex PCR), and this bias includes the bias caused by PCR amplification efficiency and preference. Obtaining the Ratio value of the target region can be used to accurately indicate the copy number variation of the target region.
[0052] To improve the amplification consistency of different DNA samples, an amplicon mixture is obtained by performing an amplification reaction test on all DNA samples in the same batch using the same polymerase chain reaction system. Among them, the polymerase chain reaction system includes primers and polymerase. For multiple amplification regions, multiple pairs of primers are usually designed in sequence, and each DNA sample undergoes a multiplex PCR reaction.
[0053] The polymerase chain reaction system used in this application includes a specific reaction solution, and the specific reaction solution can use Equinox Multiplex PCR Master Mix (Watchmaker Genomics).
[0054] The polymerase chain reaction system used in this application also includes an enrichment amplification reaction solution, and the enrichment amplification reaction solution can use Equinox Library Amplification Kit (Watchmaker Genomics).
[0055] The amplification reaction in this application includes the steps of: constructing a first amplification reaction system of a DNA sample, specific primers, and a specific amplification reaction solution, performing a first round of amplification reaction, and obtaining a first amplification product mixture after purification; constructing a second amplification reaction system of the first amplification product mixture, sequencing adapter primers, and an enrichment amplification reaction solution, performing a second round of amplification reaction, and obtaining a second amplification product mixture; purifying the second amplification product mixture to obtain an amplicon mixture for sequencing.
[0056] Among them, the final concentrations of each primer in the corresponding amplification reaction system are 0.02 - 0.03 μM respectively, preferably 0.025 μM.
[0057] In one embodiment, the total sequencing data volume of each DNA sample is uniformly corrected so that the total sequencing data volume of each DNA sample is the same. For example, if the application requires a total sequencing data volume of 1G, the total sequencing data volume of each DNA sample obtained by sequencing will fluctuate. The total sequencing data volume of each DNA sample is proportionally converted so that the total sequencing data volume of each DNA sample approaches 1G, that is, uniform correction.
[0058] The so-called stability ranking is: statistically analyze the differences between the remaining sequence regions among all DNA samples to obtain the coefficient of variation CV of each remaining sequence region, and sort all the coefficients of variation CV from smallest to largest. After sorting, extract the expected number of remaining sequence regions with the smallest coefficient of variation CV as the internal reference regions. The expected number can be 3 - 10.
[0059] The control sample is generally selected as a negative sample, that is, a sample with a normal copy in the target region. In one embodiment, the control sample is a known negative sample, such as a known negative sample that can be detected by multiplex ligation-dependent probe amplification (MLPA). Since the position of this control sample in the DNA sample in high-throughput sequencing is known, it is only necessary to extract the amplicon sequencing data obtained from the detection channel corresponding to this control sample after high-throughput sequencing.
[0060] Another embodiment is to determine the control sample from all DNA samples, including the following steps: extracting candidate samples from all DNA samples; selecting the control sample from the candidate samples.
[0061] Specifically, to extract candidate samples from all DNA samples, the following methods can be used: performing predictive screening for negative samples among all DNA samples, and taking the DNA samples obtained from the predictive screening as candidate samples; or taking all DNA samples as a whole as candidate samples.
[0062] The specific steps of predictive screening include: using the DNA sample group with known target region copies and the A value of the target region obtained by the method of the present application as a training set, training to obtain a random forest algorithm training model, and using the random forest algorithm training model to perform predictive screening for negative samples among all DNA samples, and taking the DNA samples obtained from the predictive screening as candidate samples. Among them, the DNA samples with known copy numbers in the amplification region include positive samples with abnormal copies in the target region and negative samples with normal copies in the target region.
[0063] See Figure 2 , the specific steps to determine the control sample from the candidate samples are as follows: (1) determining a target region, and obtaining the A value obtained by the corresponding internal reference region of this target region in each candidate sample; (2) in the set of A values of this target region of all candidate samples, calculating the median Median of this target region; (3) according to the degree of dispersion of the A value of this target region relative to the median Median, selecting the expected number of samples with the lowest degree of dispersion among all candidate samples as the control sample of this target region; the expected number can be 10 to 30.
[0064] The control sample is for each target region. For different target regions, it is necessary to determine the control sample. And one of the purposes of the method of the present application based on high-throughput sequencing analysis means is to realize the analysis of multiple amplification regions. To facilitate the analysis of multiple amplification regions, the method of the present application further includes:
[0065] In one embodiment, there are multiple target regions. For each target region, steps (1) to (3) are executed to obtain the control sample of each target region. Take the union of the control samples of all target regions and count the number of occurrences of the same control sample to obtain the final control sample set.
[0066] When the number of control samples in the final control sample set ≥ 10, extract the top 10 control samples ranked by the number of occurrences during the union operation as the control samples in step S50 and include them in the calculation. When the number of control samples in the final control sample set < 10, extract all control samples in the control sample set as the control samples in step S50 and include them in the calculation.
[0067] In one embodiment, there are multiple control samples. Calculate the ratios of the A values of the target regions of the sample to be analyzed to those of multiple control samples respectively, and obtain the average of the multiple ratios as the Ratio value of the target region in the sample to be analyzed.
[0068] The number of control samples is preferably not less than 5, more preferably 10, which can reduce the analysis workload while ensuring the analysis accuracy.
[0069] For the control samples obtained by predictive screening, each control sample can be further analyzed and verified by the method of the present application. Specifically, in the control sample set, calculate the ratios of the A values of each target region of each control sample to those of all control samples respectively, and obtain the average of the multiple ratios as the Ratio value of the target region in each control sample.
[0070] Copy number variations are divided into copy deletions and copy duplications. The target regions have corresponding theoretical Ratio values for different copy situations. By comparing the actually measured Ratio value of the target region with the corresponding theoretical Ratio value, it can be determined whether the copy number of the target region has changed.
[0071] Taking the human gene as an example: (1) If the copy of the target region is normal, its theoretical Ratio value is 1; (2) If half of the target region is deleted, its theoretical Ratio value is 0.5; (3) If all of the target region is deleted, its theoretical Ratio value is 0. Compare the actually measured Ratio value of the target region with the theoretical Ratio value, thereby indicating the copy number variation of the target region in the sample to be analyzed.
[0072] However, there will be certain deviations in the theoretical Ratio value in actual analysis. To more accurately analyze the copy number variation of the target region, the present application performs ROC analysis on a large amount of data of clinical positive samples and negative samples to obtain a suitable threshold. The specific method includes the following steps:
[0073] Provide a DNA sample database with known copy numbers of target regions, and execute step S10 to step S50 to obtain the Ratio values of the target regions in all DNA samples in the DNA sample database;
[0074] According to the target region copies of all DNA samples in the DNA sample database and the Ratio values of the target regions, a receiver operating characteristic (ROC) curve analysis is performed to obtain a Ratio threshold for determining whether the copy of the target region is positive;
[0075] Compare the Ratio value of the target region of the sample to be analyzed with the corresponding Ratio threshold, and according to the comparison result, obtain the change in the copy number of the target region in the sample to be analyzed.
[0076] For some target regions, there are highly homologous regions in chromosomal DNA. When using primers to amplify the target regions, the homologous regions are often amplified simultaneously, which greatly interferes with the accuracy of the analysis of the copy number of the target regions. To solve this problem, the method of this application includes:
[0077] Pre-screen the homologous regions in the genomic reference sequence in advance, and then align the amplicon sequencing data with the genomic reference sequence, so that the primer amplification regions are aligned at the same position to obtain the Ratio value of the primer amplification region, effectively reducing the difficulty of alignment and avoiding the uncertainty of alignment;
[0078] During the analysis process, according to the differential bases between the target region and the homologous region, count the ratio of the number of target regions or homologous regions in the primer amplification region to the number of primer amplification regions to obtain the VAF value;
[0079] Combine the analysis of the Ratio value and the VAF value of the primer amplification region to judge the change in the copy number of the target region in the sample to be analyzed.
[0080] Among them, the primer is designed upstream and downstream of the differential bases between the target region and its homologous region, and the size of the primer amplification region is 50-300bp, preferably 80-250bp.
[0081] However, the actual situation will deviate from the theoretical situation. To improve the accuracy of the analysis, in one embodiment, the specific method of combining the analysis of the Ratio value and the VAF value is:
[0082] Provide multiple theoretical samples. Each theoretical sample has a known Ratio value and VAF value according to the copy numbers of the target region and its homologous region. All theoretical samples are arranged and distributed in ascending order according to the VAF value;
[0083] All theoretical samples are screened in the first round according to the VAF value of the sample to be analyzed to obtain candidate theoretical samples. The VAF value of the candidate theoretical samples is closer to the VAF value of the sample to be analyzed than the VAF values of the other theoretical samples;
[0084] The candidate theoretical samples are subjected to a second round of screening based on the Ratio value of the sample to be analyzed to obtain the final theoretical samples. The Ratio value of the final theoretical samples is closer to the Ratio value of the sample to be analyzed compared to the Ratio values of the other theoretical samples.
[0085] Output the VAF value and Ratio value of the final theoretical samples as the VAF value and Ratio value of the target region in the sample to be analyzed. Combine the output VAF value and Ratio value to obtain the copy number variation of the target region.
[0086] Some diseases are related to copy number variations of specific genes or genomic regions. Through the method for detecting copy number variations based on amplicon capture technology in this application, evidence for studying the variations of specific genes or genomic regions can be obtained.
[0087] Spinal Muscular Atrophy (SMA) is a group of autosomal recessive genetic diseases characterized by the degenerative lesions of motor neurons in the anterior horn of the spinal cord, mainly affecting the control of muscle movement.
[0088] Studies have found that the survival motor neuron gene (SMN) located in the chromosomal region 5q11.2 - 13.3 is a decisive factor in the onset of SMA. The SMN gene is a gene family, including highly homologous SMN1 gene and SMN2 gene, which only differ by 5 bases, and 2 of these bases are located in exon 7 (c.840C / T) and exon 8 (c.*239G / A). The SMN1 gene is the pathogenic gene of SMA. Most (90.0 - 98.6%) patients have homozygous deletions or mutations of exon 7 and exon 8 (abbreviated as E7 and E8) or only exon 7 of the SMN1 gene; while the SMN2 gene is a regulatory gene for the severity of SMA symptoms, and its copy number is related to the severity of SMA. Usually, the more copies of the SMN2 gene, the relatively more functional SMN protein can be produced in the patient's body, and the condition may be relatively mild.
[0089] As can be seen from the above, obtaining accurate information on the copy numbers of the SMN1 gene and SMN2 gene is a key step in accurately analyzing the SMN genotype of the sample. The existing technology mainly realizes the detection of the copy numbers of the SMN1 gene and SMN2 gene through multiplex ligation-dependent probe amplification technology (MLPA), but there are problems such as complex probe design and high detection costs.
[0090] See Figure 3 , to solve the problems of the existing technology, this application provides a method for analyzing the copy number of the SMN gene based on amplicon capture technology, including the following steps:
[0091] S100. Amplify each DNA sample using a polymerase chain reaction (PCR) system to obtain an amplicon mixture of multiple amplified regions. The polymerase chain reaction system includes primers designed upstream and downstream of the differential bases of the SMN1 gene and the SMN2 gene.
[0092] S200. Use a method for detecting copy number variation based on amplicon capture technology to obtain the Ratio value of the primer amplification region.
[0093] S300. According to the differential bases, obtain the ratio of the number of SMN1 genes or SMN2 genes to the number of primer amplification regions, which is the VAF value. The quantity in this step refers to the number of Reads in next-generation sequencing.
[0094] S400. Combine and analyze the Ratio value and the VAF value to judge the copy number situation of the SMN gene in the sample to be analyzed.
[0095] The method for detecting copy number variation based on amplicon capture technology in this application can eliminate the biases brought by multiplex PCR reactions and sequencing technologies, and obtain information that can be used to accurately indicate the copy number situation of the amplified region, that is, the Ratio value. According to the Ratio value, it can be judged whether the SMN gene is deleted. Further, by combining the VAF value, the influence of homologous regions or genes can be eliminated, so as to accurately analyze the copy number situation of the SMN1 gene and the SMN2 gene.
[0096] In one embodiment, the primers include primer E7 designed upstream and downstream of the differential bases in exon 7, and primer E8 designed upstream and downstream of the differential bases in exon 8, as shown in SEQ ID NO: 1-4 respectively. The two pairs of primers can amplify exons 7 and 8 of the SMN1 gene and the SMN2 gene without difference, containing differential bases.
[0097] Correspondingly, in step 200, exons 7 and 8 of the SMN2 gene are pre-screened in the genomic reference sequence to obtain the Ratio value of the primer E7 amplification region and the Ratio value of the primer E8 amplification region respectively.
[0098] In step S300, according to the differential bases in exon 7, obtain the ratio of the number of SMN2 genes to the number of primer E7 amplification regions, which is the VAF1 value; according to the differential bases in exon 8, obtain the ratio of the number of SMN2 genes to the number of primer E8 amplification regions, which is the VAF2 value.
[0099] In step S400, according to the Ratio value and VAF1 value of the amplification region of primer E7, the copy number of exon 7 in the SMN1 gene and the SMN2 gene is obtained; according to the Ratio value and VAF2 value of the amplification region of primer E8, the copy number of exon 8 in the SMN1 gene and the SMN2 gene is obtained, and by synthesizing the results, the copy number of the SMN gene in the sample to be analyzed is obtained.
[0100] In one embodiment, the specific manner of combining the analysis of the Ratio value and the VAF value may be:
[0101] Provide multiple theoretical samples. Each theoretical sample has known Ratio values and VAF values according to the copy number of the SMN gene, the copy number of the SMN1 gene, and the copy number of the SMN2 gene. See Table 1. Compare the Ratio value of the primer amplification region with the theoretical Ratio value, and compare the VAF value with the theoretical VAF value to obtain the copy number of the SMN1 gene and the copy number of the SMN2 gene.
[0102] However, the actual situation will deviate from the theoretical situation. To improve the accuracy of the analysis, the specific manner of combining the analysis of the Ratio value and the VAF value is:
[0103] See Table 1. Provide multiple theoretical samples. Each theoretical sample has known Ratio values and VAF values according to the copy number of the SMN gene, the copy number of the SMN1 gene, and the copy number of the SMN2 gene. All theoretical samples are arranged and distributed in ascending order according to the VAF value, as Figure 4 shown, a distribution example of the VAF values and Ratio values of some theoretical samples is provided, and the same applies to other theoretical samples;
[0104] All theoretical samples are subjected to the first round of screening according to the VAF value of the sample to be analyzed to obtain candidate theoretical samples. The VAF value of the candidate theoretical samples is closer to the VAF value of the sample to be analyzed than the VAF values of the remaining theoretical samples;
[0105] The candidate theoretical samples are subjected to the second round of screening according to the Ratio value of the sample to be analyzed to obtain the final theoretical samples. The Ratio value of the final theoretical samples is closer to the Ratio value of the sample to be analyzed than the Ratio values of the remaining theoretical samples;
[0106] Output the VAF value and Ratio value of the final theoretical samples as the VAF value and Ratio value of the target region in the sample to be analyzed. Combine the output VAF value and Ratio value to judge the copy number change of the target region.
[0107] Table 1 VAF and Ratio values of theoretical samples
[0108]
[0109]
[0110] The above technical solution will be further described below in combination with specific application steps and specific parameters.
[0111] Application Example 1
[0112] 1. Provide multiple DNA samples. The DNA samples include 5 samples to be analyzed and 10 known negative samples. An amplicon mixture of each DNA sample is obtained by polymerase chain reaction in the same batch. The specific steps are as follows:
[0113] (1) Material preparation
[0114] Extract human peripheral blood genomic DNA according to the operating instructions of nucleic acid extraction or purification reagents. The recommended detection amount range of DNA: 2 ng ≤ DNA concentration ≤ 4 ng, purity: 260 / 280 > 1.8, 260 / 230 > 1.6.
[0115] Primers E7 and E8 as shown in SEQ ID NO: 1-4, and specific primers for 54 other sequence regions as shown in SEQ ID NO: 5-112. The 54 regions are chr1:45965892-45966144, chr1:45973019-45973276, chr1:45973821-45974091, chr6:49399032-49399297, chr6:49399357-49399632, chr6:49403128-49403415, chr6:49404191-49404468, chr6:49407785-49408112, chr6:49409493-49409707, chr6:49412245-49412511, chr6:49415325-49415574, chr6:49416458-49416714, chr6:49421221-49421496, chr7:107301191-107301391, chr7:107302029-107302291, chr7:107303683-107303951, chr7:107312536-107312760, chr7:107315337-107315602, chr7:107329439-107329697, chr7:107330503-107330779, chr7:107336319-107336542, chr7:107338386-107338655, chr7:107340433-107340707, chr7:107341427-107341698, chr7:107344693-107344904, chr7:107350465-107350676, chr7:107352933-107353211, chr7:107355810-107356033, chr7:107356455-107356725, chr12:103234102-103234345, chr12:103236767-103236990, chr12:103237359-103237626, chr12:103240535-103240802, chr12:103245345-103245624, chr12:103246518-103246791, chr12:103248230-103248461, chr12:103248466-103248738,chr12:103260227-103260493, chr12:103271099-103271378, chr12:103288476-103288741, chr12:103306515-103306762, chr12:103310791-103311022, chr13:52507387-52507604, chr13:52509695-52509868, chr13:52513113-52513362, chr13:52515172-52515413, chr13:52516476-52516736, chr13:52523730-52524008, chr13:52524064-52524340, chr13:52524348-52524611, chr13:52531588-52531830, chr13:52534249-52534509, chr13:52535852-52536135, chr13:52542520-52542799.
[0116] Table 2 Primer Sequences
[0117] Primer Specific sequence (5’-3’) Sequence number E7-F1 cagacgtgtgctcttccgatctCCTTAATTTAAGGAATGTGAGCACC SEQ ID NO:1 E7-R1 ctacacgacgctcttccgatctAATGCTTTTTAACATCCATATAAAGCT SEQ ID NO:2 E8-F1 cagacgtgtgctcttccgatctAAAGGAAGTGGAATGGGTAACTCT SEQ ID NO:3 E8-R1 ctacacgacgctcttccgatctAATTTTCTCAACTGCCTCACCAC SEQ ID NO:4
[0118] (2) Construct the first-round amplification reaction system:
[0119] Let the specific amplification reaction solution and the amplicon capture Panel thaw at room temperature, vortex for 3 - 5 seconds to mix evenly and centrifuge briefly, then place on ice for later use. Prepare the reaction amplification system according to the dosage in Table 3 below. Dispense 15 μL of the reaction mixture into each well of the PCR plate. Then add 5 μL of the DNA sample to be tested into each reaction well. After sealing the film, vortex and shake evenly, and centrifuge briefly to remove air bubbles at the bottom of the tube, and collect the liquid droplets on the tube wall and the sealing film to the bottom of the tube.
[0120] Table 3 Component Dosages of the Amplification Reaction System
[0121]
[0122] Among them, the specific amplification reaction solution is Equinox Multiplex PCR Master Mix (Watchmaker Genomics); the final concentration of each primer in the amplification reaction system is 0.025 μM.
[0123] (3) Perform the first-round amplification reaction
[0124] Perform the first-round amplification reaction according to Table 4. After amplification, purify the PCR plate with magnetic beads to obtain the purified first amplification product mixture.
[0125] 4 First-round amplification reaction conditions
[0126]
[0127] (4) Construct the second-round amplification reaction system
[0128] According to Table 5, let the enriched amplification reaction solution and the sequencing adapter primer set thaw on ice, vortex for 3 - 5 seconds to mix evenly and centrifuge briefly, then place on ice for later use; take a new PCR plate, dispense 25 μL of the enriched amplification reaction solution into each well of the PCR plate; add 1 μL of the corresponding forward and reverse Index primers to the reaction wells according to the required Index positions, add 23 μL of the purified first-round library product, seal the PCR plate with a sealing film, vortex for 5 - 10 s, briefly centrifuge to expel the bubbles at the bottom of the tube, and collect the liquid droplets on the tube wall and the sealing film to the bottom of the tube.
[0129] Table 5 Second-round amplification reaction system
[0130]
[0131] Among them, the enriched amplification reaction solution is Equinox Library Amplification Kit (Watchmaker Genomics). The sequencing adapter primer set is shown in SEQ ID NO: 113 - 114, where (nnnnnnnn) is the index sequence of the random base combination, and different samples use different index combinations as identity identifiers.
[0132] (5) Perform the second-round amplification reaction
[0133] Perform the second-round amplification reaction according to Table 6. After amplification, purify with magnetic beads to obtain the purified second amplification product mixture, that is, obtain the amplicon mixture of each DNA sample.
[0134] Table 6 Second-round amplification reaction conditions
[0135]
[0136]
[0137] 2. Sequence the amplicon mixture on the machine to obtain the amplicon sequencing data of each DNA sample.
[0138] 3. Mask exons 7 and 8 of the SMN2 gene in the genomic reference sequence in advance, and then align the amplicon sequencing data with the genomic reference sequence to obtain the number of Reads and the sequencing depth of each amplified region.
[0139] 4. Extract the remaining regions except for the amplified regions of primer E7 and primer E8 from all DNA samples, and count the differences in the sequencing depths of the remaining sequence regions among all DNA samples to obtain the coefficient of variation CV of each remaining sequence region. Then sort all the coefficients of variation CV from smallest to largest, and extract the top 5 remaining sequence regions after sorting as the internal reference regions.
[0140] 5. Calculate the ratio of the number of Reads in the amplified region of primer E7 in each DNA sample to the number of Reads in the 5 internal reference regions respectively, and obtain the average value of the 5 ratios as the A value of the amplified region of primer E7; similarly, obtain the A value of the amplified region of primer E8.
[0141] 6. Calculate the ratio of the A value of the amplified region of primer E7 in each sample to be analyzed to the A value of the amplified region of primer E7 in 10 negative samples respectively, and obtain the average value of the multiple ratios as the Ratio value of the amplified region of primer E7; similarly, obtain the Ratio value of the amplified region of primer E8.
[0142] 7. According to the differential bases of exon 7, obtain the ratio of the number of SMN2 genes to the number of amplified regions of primer E7 as the SMN2-VAF1 value; according to the differential bases of exon 8, obtain the ratio of the number of SMN2 genes to the number of amplified regions of primer E8 as the SMN2-VAF2 value.
[0143] 8. Combine and analyze the Ratio value and VAF1 value of the amplified region of primer E7 to obtain the copy number situation of exon 7 in the SMN1 gene and the SMN2 gene; combine and analyze the Ratio value and VAF2 value of the amplified region of primer E8 to obtain the copy number situation of exon 8 in the SMN1 gene and the SMN2 gene. Based on the comprehensive results, obtain the copy number situation of the SMN gene in the sample to be analyzed and output the SMN genotype.
[0144] In the genotype, the first number represents the copy number of exon 7 of the SMN1 gene, the second number represents the copy number of exon 8 of the SMN1 gene, the third number represents the copy number of exon 7 of the SMN2 gene, and the fourth number represents the copy number of exon 8 of the SMN2 gene.
[0145] Analysis results
[0146] The analysis results of Application Example 1 are shown in Table 7 and are verified by the MLPA method.
[0147] Table 7 Analysis results of Application Example 1
[0148]
[0149] Application Example 2
[0150] 1. Provide 46 DNA samples, and refer to steps (1) to (5) of Application Example 1 to obtain the amplicon mixture of each DNA sample.
[0151] 2. Sequence the amplicon mixture on a sequencer to obtain the amplicon sequencing data of each DNA sample.
[0152] 3. Pre-screen exons 7 and 8 of the SMN2 gene in the genomic reference sequence, and then align the amplicon sequencing data with the genomic reference sequence to obtain the number of Reads and sequencing depth of each amplified region.
[0153] 4. Extract the remaining regions except for the primer E7 amplified region and the primer E8 amplified region in all DNA samples, and count the differences in the sequencing depth of each remaining sequence region among all DNA samples to obtain the coefficient of variation CV of each remaining sequence region. Then sort all the coefficients of variation CV from smallest to largest, and extract the top 5 remaining sequence regions after sorting as the internal reference regions.
[0154] 5. Calculate the ratio of the number of Reads in the primer E7 amplified region of each DNA sample to the number of Reads in the 5 internal reference regions respectively, and calculate the average value of the 5 ratios as the A value of the primer E7 amplified region; similarly, obtain the A value of the primer E8 amplified region.
[0155] 6. Use the random forest algorithm model to perform predictive screening of negative samples in all DNA samples to obtain candidate samples. Execute steps (1) to (3), and select the samples with the lowest degree of dispersion of the expected number in all candidate samples as the control samples for exons 7 and 8. Take the union of the control samples for exons 7 and 8. If the number of control samples in the union result exceeds 10, extract the top 10 control samples ranked by the number of occurrences during the union process to obtain the final control sample set.
[0156] 7. Calculate the ratio of the A value of the primer E7 amplified region of each sample to be analyzed to the A values of the primer E7 amplified region in the 10 control samples respectively, and calculate the average value of the multiple ratios as the Ratio value of the primer E7 amplified region; similarly, obtain the Ratio value of the primer E8 amplified region; in the final control sample set, calculate the ratio of each control sample to the 10 control samples respectively, and calculate the average value of the multiple ratios as the Ratio value of each target region in each control sample.
[0157] 8. Obtain the ratio of the number of SMN2 genes to the number of primer E7 amplification regions based on the differential bases in exon 7, which is the SMN2-VAF1 value; obtain the ratio of the number of SMN2 genes to the number of primer E8 amplification regions based on the differential bases in exon 8, which is the SMN2-VAF2 value.
[0158] 9. Combine and analyze the Ratio value and VAF1 value of the primer E7 amplification region to obtain the copy number of exon 7 in the SMN1 gene and the SMN2 gene; combine and analyze the Ratio value and VAF2 value of the primer E8 amplification region to obtain the copy number of exon 8 in the SMN1 gene and the SMN2 gene. Based on the comprehensive results, obtain the copy number of the SMN gene in the sample to be analyzed and output the SMN genotype.
[0159] Analysis result
[0160] The analysis result of Application Example 2 is shown in Table 8 and is verified by the MLPA method.
[0161] Table 8 Analysis result of Application Example 2
[0162]
[0163]
[0164]
[0165] Application Example 3
[0166] The specific analysis steps refer to Application Example 3, where the control samples used are the 10 negative samples detected in Application Example 2. The analysis result is shown in Table 9 and is verified by the MLPA method.
[0167] Table 9 Analysis result of Application Example 3
[0168]
[0169] See Figure 5 , this application also provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement steps S10 to S60.
[0170] This application also provides a computer program product, including computer instructions. When the computer instructions are executed by the processor, steps S10 to S60 are implemented.
[0171] In this application, the computer program product includes a program code portion for performing the method steps in the embodiments of this application when the computer program product is executed by one or more computing devices. The computer program product can be stored on a computer-readable recording medium. The computer program product can also be provided for downloading via a data network (e.g., through a RAN, via the Internet, and / or through an RBS). Alternatively or additionally, the method can be encoded in a field-programmable gate array (FPGA) and / or an application-specific integrated circuit (ASIC), or the functionality can be provided for downloading by means of a hardware description language.
[0172] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0173] The above-described embodiments merely represent several implementation manners of this application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
Claims
1. A method for detecting copy number variations based on amplicon capture technology, characterized in that, Including the following steps: S10. Obtain the amplicon sequencing data of multiple amplification regions in each DNA sample, where the DNA samples include samples to be analyzed and control samples; S20. Align the amplicon sequencing data with the genomic reference sequence to obtain the number of Reads and the sequencing depth of each amplification region in each DNA sample; S30. Extract the remaining regions other than the target region from all DNA samples, and select multiple stable remaining regions as internal reference regions according to the stability ranking of the sequencing depths of all DNA samples based on the same data volume in the remaining regions; S40. Calculate the ratios of the number of Reads in the target region to the number of Reads in multiple internal reference regions in each DNA sample respectively, and calculate the average value of the multiple ratios to obtain the A value of the target region in each DNA sample; S50. Calculate the ratio of the A value of the target region in the sample to be analyzed to the A value of the target region in the control sample, which is the Ratio value of the target region in the sample to be analyzed; S60. Judge the copy number change of the target region in the sample to be analyzed according to the Ratio value of the target region.
2. The method according to claim 1, characterized in that Extract candidate samples from all DNA samples, and determine the control sample from the candidate samples through the following steps: (1) Determine a target region, and obtain the A value of the target region corresponding to the internal reference region in each candidate sample; (2) Calculate the median Median of the target region in the set of A values of the target region in all candidate samples; (3) Select the expected number of samples with the lowest degree of dispersion from all candidate samples according to the degree of dispersion of the A value of the target region relative to the median Median as the control sample of the target region.
3. The method according to claim 2, wherein There are multiple target regions. Perform steps (1) to (3) on each target region to obtain the control sample of each target region; take the union of the control samples of all target regions and count the number of occurrences of the same control sample to obtain the final control sample set.
4. The method according to claim 2, wherein Extract candidate samples from all DNA samples, including: Perform predictive screening of negative samples among all DNA samples, and use the DNA samples obtained by predictive screening as candidate samples; or use all DNA samples as candidate samples as a whole.
5. The method according to claim 4, wherein Use the DNA sample group with known target region copies and the A value of the target region in the DNA sample group as a training set to train and obtain a random forest algorithm training model, and use the random forest algorithm training model to perform predictive screening of negative samples among all DNA samples.
6. The method according to claim 1, wherein The target region has a corresponding theoretical Ratio value in the theoretical sample. Compare the Ratio value of the target region in the sample to be analyzed with the corresponding theoretical Ratio value to judge the copy number change of the target region in the sample to be analyzed.
7. The method according to claim 1, wherein There are multiple control samples. Calculate the ratios of the A value of the target region in the sample to be analyzed to the A value of the target region in multiple control samples respectively, and calculate the average value of the multiple ratios to obtain the Ratio value of the target region of the sample to be analyzed.
8. The method according to claim 1, wherein The amplification primers used for the amplification region include primers designed upstream and downstream of the differential bases in the target region and its homologous region. The method includes: Previously shielding the homologous region in the genomic reference, and then aligning the amplicon sequencing data with the genomic reference sequence to obtain the Ratio value of the primer amplification region; According to the differential bases between the target region and the homologous region, counting the ratio of the number of target regions or homologous regions in the primer amplification region to the number of the primer amplification region to obtain the VAF value; Combining and analyzing the Ratio value of the primer amplification region and the VAF value to judge the copy number change of the target region in the sample to be analyzed.
9. The method according to claim 8, characterized in that, The specific way of combining and analyzing the Ratio value and the VAF value is: Providing a plurality of theoretical samples, each theoretical sample having a known Ratio value and VAF value according to the copy numbers of the target region and its homologous region, and all the theoretical samples are arranged and distributed in ascending order according to the VAF value; All the theoretical samples are subjected to the first round of screening according to the VAF value of the sample to be analyzed to obtain candidate theoretical samples, and the VAF value of the candidate theoretical samples is closer to the VAF value of the sample to be analyzed than the VAF values of the remaining theoretical samples; The candidate theoretical samples are subjected to the second round of screening according to the Ratio value of the sample to be analyzed to obtain the final theoretical samples, and the Ratio value of the final theoretical samples is closer to the Ratio value of the sample to be analyzed than the Ratio values of the remaining theoretical samples; Outputting the VAF value and Ratio value of the final theoretical samples as the VAF value and Ratio value of the sample to be analyzed, and judging the copy number change of the target region according to the output VAF value and Ratio value.
10. A kit for analyzing copy number variations in amplified regions, characterized in that, Including the specific amplification reaction solution Equinox Multiplex PCR Master Mix and the enrichment amplification reaction solution Equinox Library AmplificationKit.
Citation Information
Cited By
CYP21A2 gene mutation detection method and system and storage medium
CN122117026A
A method and system for detecting a CYP21A2 gene mutation and a storage medium
CN122117026B