A method for detecting sequence changes and copy number changes in a target region / homologous region
By designing semi-specific primers and homologous region marker sites, combined with PCR amplification and sequencing technology, the detection problem of sequence changes and copy number changes in target region/homologous region across gene regions by CYP21A2 whole gene sequence + TNXB Exon 32-44 was solved, and efficient and economical variation analysis was achieved.
Patent Information
- Application Number
- CN202310294000.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-03-24
AI Technical Summary
The prior art is difficult to perform high-throughput 2 generation sequencing across gene regions on the CYP21A2 whole gene sequence + exon 32-44, and it is impossible to effectively detect sequence changes and copy number changes in the target region/homologous region. The traditional method is time-consuming and economically beneficial.
The first primer pair and the second primer pair were respectively PCR amplified, and the amplified products were sequenced to analyze the sequence changes and copy number changes of the target region/homologous region. The homologous region carries a homologous region marker site. Through semi-specific primer design, the target region sequence can be amplified when the target region exists and the homologous region sequence is amplified when the target region is lacking.
It realizes the detection of sequence changes and copy number changes in the target area/homologous area in an experimental plan, saving time and reducing economic costs, and avoiding the shortcomings of the traditional method.
Smart Images

Figure CN116121351B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and particularly to a method for detecting sequence variations and copy number variations in a target region / homologous region. Background Art
[0002] For genotype studies of the CYP21A2 full gene sequence + TNXB Exon 32-44 transgene region (the interval of chr6:32006083-32013949 of the Hg19 reference sequence), the biggest problem is that there are highly homologous regions in the genome in this section [i.e., CYP21A1P-TNXA (chr6:31973359-31981050), with a sequence identity of up to ~98%], and random recombination often occurs between the sequences in this interval and their highly homologous or allelic genes, resulting in random sequence exchanges, large fragment duplications or deletions ( Figures 1-3 ). These situations make the standard process of general high-throughput second-generation sequencing (High-throughput NextGeneration Sequencing, HtNGS) inapplicable to the variant analysis of the CYP21A2 full gene sequence + TNXB Exon 32-44 transgene region. Therefore, traditionally, most of the methodologies used in studies on this region have to separately handle different situations such as large fragment rearrangement / copy number variant (Copy number Variant, CNV) and small range variations such as single nucleotide variant (Single Nucleotide Variant, SNV) or small fragment insertion / deletion or exchange (Small insertion / duplication or deletion) through multiplex ligation-dependent probe amplification (Multplex ligation-dependent probeAmplication, MLPA) plus long-range polymerase chain reaction + multiple Sanger sequencing (Long-range PCR+multiple Sanger sequence, LR-PCR / mSanger-seq) ( Figures 1-3)。Moreover, since MLPA only designs probes for common specific sites in the above-mentioned partial genomic regions (the entire CYP21A2 gene sequence + the transgene region of TNXB Exons 35 - 44 and CYP21A1P), if there are sequences outside the probes or small-segment sequence exchanges, it may be impossible to sequence small-segment variations or polymorphisms different from the reference sequence in the entire CYP21A2 gene sequence + the transgene region of TNXB Exons 32 - 44. When using LR-PCR / mSanger-seq alone, in order to ensure highly specific fragments, the primer design for LR-PCR must find regions with completely non-repeating sequences to ensure that in any case, only the fragments of CYP21A2 or TNXB can be amplified. Therefore, the sequencing range will be narrowed (the reverse primer sequence can only be found from Exon 35 of TNXB because TNXA has a sequence lacking ~120 bp in the homologous region). And because Sanger sequencing can only measure a reliable sequence of 800 - 1000 bp in one reaction, multiple sequencing primers need to be designed to sequence the entire interval, which invisibly increases the cost, and the effect is only to find variations in small-segment regions. Moreover, due to polymorphisms that may exist in different intervals, it affects the success rate of sequencing the entire interval. In addition, when copy number variations are encountered, it is completely inapplicable and MLPA needs to be used instead. Therefore, at present, for the entire CYP21A2 gene sequence + the transgene region of TNXB Exons 35 - 44, a time-consuming and extremely uneconomical method of MLPA plus LR-PCR / mSanger-seq is required to cover the variations in the entire CYP21A2 gene sequence + the transgene region of TNXB Exons 35 - 44, but some interval variations may still be missed.
[0003] In recent years, although some literature has adopted the method of long-fragment polymerase chain reaction + next-generation sequencing (LR-PCR / NGS) to detect small-segment sequence variations ( Figure 2 ), since it aims to ensure that the LR-PCR products mainly contain the sequence of CYP21A2, the analysis and consideration of copy number variations are not included in the one-time solution. It still adopts the method of separately detecting copy number variations and small-segment variations, and still cannot solve the problem of small-segment variations plus copy number recombination variations in one experimental protocol. Summary of the Invention
[0004] The present invention discloses a method for detecting sequence changes and copy number changes in a target region / homologous region, aiming to achieve a one-time solution for different mutation types to save time and economic costs. The method uses a first primer pair and a second primer pair to perform PCR amplification on a template respectively, sequences the amplification products, and analyzes to obtain the sequence changes and copy number changes in the target region / homologous region of the template; the homologous region has a homologous region marker site; the homologous region marker site is a conserved site verified as a homologous region sequence, and the mutation ratio of the site can be used to judge the copy number of the homologous region.
[0005] The first primer pair contains semi-specific primers, such that when the target region exists in the template, it mainly amplifies the sequence in the target region, and when the target region does not exist in the template but the homologous region exists, it amplifies the sequence in the homologous region; the second primer pair amplifies both the sequence in the target region and the sequence in the homologous region when both the target region and the homologous region exist in the template; the semi-specific primers can specifically bind to the sequence in the target region, but there are 2 non-complementary nucleotides when binding to the sequence in the homologous region.
[0006] The analysis to obtain the sequence changes and copy number changes in the target region / homologous region includes the following steps:
[0007] S1. Determine the situation: According to the site ratio of the target region obtained by sequencing the PCR amplification product of the first primer pair falling greater than 0.95 or less than 0.05, it is judged that the target region is a situation without heterozygous sites, otherwise it is judged that the target region is a situation with heterozygous sites; the situation with heterozygous sites in the target region is further divided into a first sub-situation and a second sub-situation; the first sub-situation is the situation where the site ratio of the target region obtained by sequencing the PCR amplification product of the first primer pair is 0.5±0.05 and the site ratio of the PCR amplification product of the second primer pair is 0.5±0.05; the second sub-situation is the situation where the site ratio of the target region obtained by sequencing the PCR amplification product of the first primer pair is 0.33±0.05 or 0.67±0.05 and there is no site ratio of 0.5±0.05.
[0008] S2. Under the situation determined in step S1, further determine the copy number, including Method I and Method II.
[0009] Method I: According to the site ratio of the PCR amplification product of the second primer pair and the site ratio of the homologous region marker site, the corresponding mixed copy number and homologous region copy number are obtained, and the target region copy number is obtained by subtracting the homologous region copy number from the mixed copy number.
[0010] The second method: Based on the site ratio of the PCR amplification product of the second pair of primers and the site ratio of the PCR amplification product of the first pair of primers, the mixed copy number and the target region copy number corresponding to the site ratio are obtained, and the homologous region copy number is obtained by subtracting the target region copy number from the mixed copy number.
[0011] In some embodiments, in the case of no heterozygous sites, according to the first method, it is obtained that:
[0012] If the site ratio of the PCR amplification product of the second pair of primers falls within 0.2 - 0.3 or 0.7 - 0.8, the mixed copy number is 4. If the site ratio of the homologous region marker site falls within 0.45 - 0.55, the homologous region copy number is 2. Therefore, the target region copy number is 2.
[0013] If the site ratio of the PCR amplification product of the second pair of primers falls within 0.2 - 0.3 or 0.7 - 0.8, the mixed copy number is 4. If the site ratio of the homologous region marker site falls within 0.7 - 0.8, the homologous region copy number is 3. Therefore, the target region copy number is 1.
[0014] If the site ratio of the PCR amplification product of the second pair of primers falls within 0.3 - 0.4 or 0.6 - 0.7 and there is no polymorphic site falling within 0.45 - 0.55, and the site ratio of the homologous region marker site falls within 0.61 - 0.71, the mixed copy number is 3, the homologous region copy number is 2. Therefore, the target region copy number is 1.
[0015] If the site ratio of the PCR amplification product of the second pair of primers falls within 0.3 - 0.4 or 0.6 - 0.7 and there is no polymorphic site falling within 0.45 - 0.55, and the site ratio of the homologous region marker site falls within 0.3 - 0.38, the mixed copy number is 3, the homologous region copy number is 1. Therefore, the target region copy number is 2.
[0016] If the site ratio of more than 1 / 6 of the polymorphic sites of the PCR amplification product of the second pair of primers falls within 0.1 - 0.2 or 0.8 - 0.9 and there is a polymorphic site falling within 0.45 - 0.55, and the site ratio of the homologous region marker site falls within 0.62 - 0.72, the mixed copy number is 6, the homologous region copy number is 4. Therefore, the target region copy number is 2.
[0017] If the site ratio of the PCR amplification product of the second pair of primers falls within 0.35 - 0.45 or 0.55 - 0.65 and there is no polymorphic site falling within 0.45 - 0.55, and the site ratio of the homologous region marker site falls within 0.75 - 0.85, the mixed copy number is 5, the homologous region copy number is 4. Therefore, the target region copy number is 1.
[0018] The site ratio of the PCR amplification product of the second pair of primers falls within 0.35 - 0.45 or 0.55 - 0.65 and there is no polymorphic site falling within 0.45 - 0.55. If the site ratio of the homologous region marker site falls within 0.55 - 0.65, the mixed copy number is 5, the homologous region copy number is 3, and thus the target region copy number is 2;
[0019] In the case of no heterozygous site, according to Method II, it is obtained that:
[0020] If more than 80% of the site ratio in the PCR amplification product of the second pair of primers falls within 0.45 - 0.55, the mixed copy number is 2 and the target region copy number is 1.
[0021] In some embodiments, in the first sub - case, according to Method I, it is obtained that:
[0022] If the site ratio of the homologous region marker site ≥ 0.9, the mixed copy number is 2, the homologous region copy number is 2, and thus the target region copy number is 0;
[0023] If the site ratio of the homologous region marker site ≤ 0.1, the mixed copy number is 2, the homologous region copy number is 0, and thus the target region copy number is 2;
[0024] In the first sub - case, according to Method II, it is obtained that:
[0025] If the site ratio of the PCR amplification product of the second pair of primers falls within 0.2 - 0.3 or 0.7 - 0.8, the mixed copy number is 4 and the target region copy number is 2;
[0026] If the site ratio of the PCR amplification product of the second pair of primers falls within 0.28 - 0.38 or 0.61 - 0.71, the mixed copy number is 3 and the target region copy number is 2;
[0027] If there is no site ratio of the PCR amplification product of the second pair of primers falling within 0.2 - 0.3 or 0.7 - 0.8 and there is a site ratio of the PCR amplification product of the second pair of primers falling within ≤ 0.2 or ≥ 0.8, then the mixed copy number > 4 and the target region copy number is 2.
[0028] In some embodiments, in the second sub - case, according to Method I, it is obtained that: If the site ratio of the homologous region marker site falls within ≤ 0.1, the mixed copy number is 3, the homologous region copy number is 0, and thus the target region copy number is 3;
[0029] In the second sub - case, according to Method II, it is obtained that:
[0030] If the site ratio of the PCR amplification product of the second pair of primers falls within 0.2 - 0.3 or 0.7 - 0.8, the mixed copy number is 4 and the copy number in the target region is 3;
[0031] If the site ratio of the PCR amplification product of the second pair of primers falls within 0.55 - 0.65, the mixed copy number is 5 and the copy number in the target region is 3;
[0032] If the site ratio of the PCR amplification product of the second pair of primers is ≤ 0.2 or ≥ 0.8, the mixed copy number is 6 and the copy number in the target region is 3.
[0033] In some embodiments, the target region is the CYP21A2 + TNXB exon 32 - 44 interval (CYP21A2 - TNXB exon 32 - 44); the homologous region is CYP21A1P - TNXA; the semi - specific primer is SEQ ID NO:1.
[0034] In some embodiments, the CYP21A2 + TNXB exon 32 - 44 interval is the chr6:32006083 - 32013949 interval of the Hg19 reference sequence; CYP21A1P - TNXA is the chr6:31973359 - 31981050 interval of the Hg19 reference sequence.
[0035] In some embodiments, the CYP21A2 + TNXB exon 32 - 44 interval is the chr6:32006089 - 32013689 interval of the Hg19 reference sequence and / or the chr6:32006158 - 32013689 interval of the Hg19 reference sequence; CYP21A1P - TNXA is the chr6:31973424 - 31980835 interval of the Hg19 reference sequence and / or the chr6:31973355 - 31980835 of the Hg19 reference sequence.
[0036] In some embodiments, both the first primer pair and the second primer pair contain a common reverse primer SEQ ID NO:2; the second primer pair also contains a non - specific amplification forward primer SEQ ID NO:3. That is, in this embodiment, the first primer pair is SEQ ID NO:1 and SEQ ID NO:2; the second primer pair is SEQ ID NO:3 and SEQ ID NO:2.
[0037] In some embodiments, the target region sequence fragment amplified by SEQ ID NO:1 and SEQ ID NO:2 is the interval of chr6:32006089-32013689 of the Hg19 reference sequence; the target region sequence fragment amplified by SEQ ID NO:1 and SEQ ID NO:2 is the interval of chr6:31973355-31980835 of the Hg19 reference sequence; the homologous region sequence fragment amplified by SEQ ID NO:3 and SEQ ID NO:2 is the interval of chr6:31973424-31980835 of the Hg19 reference sequence; the target region sequence fragment amplified by SEQ ID NO:3 and SEQ ID NO:2 is the interval of chr6:32006158-32013689 of the Hg19 reference sequence.
[0038] In some embodiments, in the PCR reaction, the template amount used for the first primer pair is 1.5 - 3 times that of the second primer pair.
[0039] In some embodiments, the first primer pair and the second primer pair are each independently or non - independently subjected to PCR amplification under the following PCR reaction program:
[0040]
[0041] In some embodiments, two kinds of PCR amplification products obtained by the first primer pair and the second primer pair for each sample are sequenced. When the target average depth and data volume are reached, all sequencing results are forcibly matched with the target region reference sequence to calculate all sites different from the target region reference sequence, and then the site ratio is calculated; theoretically, when the sequenced segment is uniform, the value of the site ratio itself has a theoretical correlation with copy number variation, as shown in Table 7; the target region reference sequence is Hg19:chr6:32006083 - 32013949; the forced matching means that even when partial mutations occur in the target region product, it is still forcibly matched to the target region. If the forced matching with the target region of the reference sequence is not performed, in the general process, the homologous region or the target region will be matched according to the similarity of the measured sequence results. In this way, when partial mutations occur in the target region product, it may be matched to the homologous region by the standard process, so forced matching with the target region is required.
[0042] In some embodiments, the sequencing is second - generation sequencing. Further, the second - generation sequencing is second - generation sequencing using an Illumina sequencing platform.
[0043] On the other hand, the present invention also discloses a computer - readable storage medium, including a program, which can be executed by a processor to implement the method as described above.
[0044] The present invention can comprehensively cover various mutation situations [single nucleotide polymorphisms (SNPs), SNVs, small insertions, duplications, and deletions (Smallinsertion, duplication / deletion), copy number variants (Copy number varaint)] in the CYP21A2+TNXB exon 32-44 region and its homologous region (CYP21A1P-TNXA) through one set of experimental procedures plus an innovative bioinformatics analysis logic algorithm, and can effectively distinguish the distribution or presence of these different types of mutations or polymorphisms in the CYP21A2+TNXB exon 32-44 region and CYP21A1P-TNXA, without always having to adopt the time-consuming, laborious, and high-cost existing technologies (MLPA+LR-PCR / mSanger-seq) to conduct differential studies on the sequence changes of all samples. The present invention has innovations in the primer design concept for the target region, experimental methods, and the design of bioinformatics analysis methods, which are described as follows:
[0045] 1. In the primer design concept for the target region (CYP21A2+TNXB exon 32-44 region), the regions and methods used by the present invention to design primers are innovative segments and experimental methods not considered in previous literature. In the primer design concept of the present invention, in order to analyze various CNV situations together with SNPs, SNVs, and other small insertions / duplications / deletions, a design concept of a combination of semi-specific primers and non-specific primers is adopted, and a design concept of one "semi-specific" primer and primers that are completely homologous (CYP21A1P-TNXA) in two regions (referring to the target region and the homologous region) is designed ( Figure 5 ). The "semi-specific" designed primer refers to a primer designed at a position in the 5' end of CYP21A2 that has only two nucleotide differences from the 5' end of CYP21A1P, that is, Figure 5 CYP21A2_F (5'-G T ACCTGA A GGTGGGGTC-3', SEQ ID NO:1) in Figure 5The CYP21A1P / A2_UTR_F forward primer (5'-GAGCTATAAGTGGCACCTCAG-3', SEQ ID NO: 3) located at the 5' end of the two intervals and the TNXA / TNXB-31in_R reverse primer (5'-CTGCTCCTCATTACTCGCTG-3', SEQ ID NO: 2) at the 3' end in the middle. In theory, when both the target region sequence (the CYP21A2 + TNXB exon 32-44 interval) and the homologous region sequence (CYP21A1P-TNXA) are present in the sample, the PCR reaction of the combination of CYP21A2_F and TNXA / TNXB-31in_R will have a great bias towards amplifying the target region (the CYP21A2 + TNXB exon 32-44 interval). Therefore, in the case where the target region sequence exists in the genome, this primer pair (the first primer pair) will not amplify the homologous region sequence (CYP21A1P-TNXA); while the PCR reaction of the combination of CYP21A1P / A2_UTR_F and TNXA / TNXB-31in_R (the second primer pair) will simultaneously produce the products of the target region and the homologous region defined by the present invention. However, if the target region sequence is completely absent, the PCR products of the combination of CYP21A2_F and TNXA / TNXB-31in_R and the combination of CYP21A1P / A2_UTR_F and TNXA / TNXB-31in_R will be exactly the same, and the results of whether the target region or the homologous region is absent can be determined by five biomarker sites in the homologous region after performing experiments and bioinformatics data analysis (such as Figure 7 shown in Standard Sample 6 in the middle). Therefore, in the concept of this primer design, there are innovative points for detecting various genomic sequence situations and structures in this region.
[0046] 2. Traditional methodologies need to use the multi-probe kit of MLPA [P050-B2 CAH or P050-C1 CAH kit (MRC Holland, Amsterdam, The Netherlands)], and only amplify the partial target region of CYP21A2-TNXB exon 35-44 (since the 35th exon of TNXA is 120 bp less than that of TNXB, most traditional methodologies or current second-generation sequencing will design highly specific primers in this 120 bp sequence, but this will cause the sequence of the 32-34th exon region of TNXB to be ignored) to detect copy number variation and sequence difference type variation respectively by PCR multi-primers ( Figures 1-3)。However, if highly specific primers are used to amplify the long - fragment PCR products of this target region for performing second - generation sequencing, current literature also requires 4 long - fragment PCR reactions for sequencing
Gangodkar P, Khadilkar V, Raghupathy P, Kumar R, Dayal AA, Dayal D, Ayyavoo A, Godbole T, Jahagirdar R, Bhat K, Gupta N, KamalanathanS, Jagadeesh S, Ranade S, Lohiya N, Oke RL, Ganesan K, Khatod K, Agarwal M, Phadke N, Khadilkar A. Clinical application of anovel next generation sequencing assayfor CYP21A2 gene in 310 cases of 21 - hydroxylase congenital adrenalhyperplasia from India. Endocrine. 2021 Jan;71(1):189 - 198. doi:10.1007 / s12020 - 020 - 02494 - z. Epub 2020Sep 18. PMID:32948948.
[0047] 3. After the data is downloaded, the data generated by the target region product (primer pairs containing semi-specific primers, such as SEQ ID NO:1 and SEQ ID NO:2) and the target / homologous region mixed product are respectively compared with the target region chr6:32006083-32013949 of the Hg19 reference sequence as the only reference sequence for alignment. The Hg19 reference sequence of the homologous region (chr6:31973359-31981050) is excluded to generate VCF and Bam files. Through the sequencing results of each sample in the two PCR products, the copy number changes of the target region and the homologous region itself can be judged by the distribution of the site ratios (call ratio) generated by the two PCR products ( Figures 6-8 ), and the confidence interval of the result is judged by the formula data converted from the distribution of each site in the peak region; in addition to the copy number change, the matching difference between the detected sequence in the target region and the corresponding reference sequence chr6:32006083-32013949 on Hg19 can generate the changes of the actual sites in the sample target region [including SNPs, SNVs, small deletions / duplications / insertions (small deletion / duplication / insertion)].
[0048] Through the above innovation from primer design concept to experimental process and process optimization, plus a special bioinformatics analysis process designed for the data result distribution, the present invention designs an experimental scheme and a set of bioinformatics processes to analyze the sequence changes and copy number changes of samples at one time. Traditionally, to detect the actual polymorphic sites, point mutations, small deletions / duplications / insertions in the target region, long-chain polymerase reaction amplification is required, and then segmental Sanger sequencing is carried out to detect such variant sites, but this method cannot detect copy number variations ( Figure 3 ). To perform copy number change analysis, it is necessary to rely on multiplex ligation-dependent probe amplification (MLPA) to perform it alone, but it cannot detect various site changes in the target region. Therefore, traditionally, these two methodologies are required to detect the site changes and copy number changes in the entire target region ( Figure 1 and Figure 2 ), and for simple second-generation sequencing, since the genomic sequence is directly fragmented, the target region and the homologous region sequences cannot be distinguished, so the sequence source and copy number situation of the site changes themselves cannot be effectively distinguished ( Figure 4 ). Therefore, the method of the present invention uses a method of designing special primers to amplify the target region and the target / homologous region ( Figure 5) By combining with the methodology of next-generation sequencing, it can determine the region where the locus change occurs and the detection of copy number variation at one time, and allows multiple samples to be pooled in a single library construction. Through analysis, the data of each sample can be separated. Based on the ratio value of the variant sites of each sample, the copy number variation change is distinguished, and a method for detecting different types of variations (locus change and copy number) at one time is achieved through the designed bioinformatics logic. Figures 6-8 ) The verification results of the present invention illustrate actual cases. Figures 9A-9F ) Thus, it solves the problem of time-consuming and laborious of the traditional methodology and overcomes the problem of inaccurate measurement of simple next-generation sequencing.
[0049] As used herein, the terms "target region" and "target area" have the same meaning and can be used interchangeably.
[0050] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the drawings to fully understand the purpose, features and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It shows that CYP21A2-TNXB and CYP21A1P-TNXA are homologous and there is a possibility of recombination and substitution, as well as the traditional methodology.
[0052] Figure 2 It shows various copy number differences and the traditional multiplex ligation-dependent probe amplification technology.
[0053] Figure 3 It shows the types of small-scale sequence differences and the traditional method of using long-range PCR + segmental Sanger detection.
[0054] Figure 4 It shows the process differences and result comparisons between the present invention and traditional next-generation sequencing.
[0055] Figure 5 It shows the special primer design and PCR product classification.
[0056] Figure 6 It shows the copy number judgment in the case of no heterozygous sites in the target region.
[0057] Figure 7 It shows the copy number judgment in the case where the product site ratios in the target region all fall within 0.5 ± 0.05.
[0058] Figure 8 It shows the copy number judgment in the case where the product site ratios in the target region all fall within 0.33 ± 0.05 or 0.67 ± 0.05.
[0059] Figures 9A-9F It shows the graphical case data of the standard sample site ratio (call ratio) determination method. Detailed implementation manners
[0060] In order to make the technical means, creative features, achieved purposes and functions of the invention easy to understand, the present invention will be further described below in conjunction with specific drawings. However, the present invention is not limited to the following implemented cases.
[0061] It should be noted that the structures, proportions, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Therefore, they do not have technical substantive significance. Any modification of the structure, change of the proportional relationship or adjustment of the size, without affecting the functions that the present invention can produce and the purposes that can be achieved, should still fall within the scope that can be covered by the technical content disclosed by the present invention.
[0062] The specific implementation manner of the present invention is to verify by purchasing standard samples of known sequences, and it is confirmed that the innovative concept in the above theory has the possibility of being practiced. The specific implementation method is as follows, and the list of standard samples used and their genotypes are shown in Table 1. The standard samples are human leukocyte line samples of known sequences purchased from the Coriell Institute of Medical Research.
[0063] Table 1. Sequence differences and copy number differences of selected known standard samples
[0064]
[0065] Experimental design and step description:
[0066] First, use the following combination of 3 primers for the above standard samples and perform PCR with a high-fidelity rapid DNA polymerase at a high annealing temperature (65 - 75 °C) to respectively amplify Figure 5 the 2 kinds of products described in, use primers 1 and 2 to amplify specific products (when there is a target region in the genome, this primer combination will mainly amplify the target region products, but if there is no target region, it will amplify homologous region products), use primers 3 and 2 to amplify non-specific products (products that can amplify the target region and homologous segments in the genome in proportion), and for the use of templates, the template amount of primer 1 + 2 is 1.5 - 3 times that of the primer 2 + 3 reaction.
[0067] Primer 1 is the sequence shown in SEQ ID NO:1; Primer 2 is the sequence shown in SEQ ID NO:2; Primer 3 is the sequence shown in SEQ ID NO:3. The target region sequence fragment amplified by SEQ ID NO:1 and SEQ ID NO:2 is the interval of chr6:32006089-32013689 of the Hg19 reference sequence; the target region sequence fragment amplified by SEQ ID NO:1 and SEQ ID NO:2 is the interval of chr6:31973355-31980835 of the Hg19 reference sequence; the homologous region sequence fragment amplified by SEQ ID NO:3 and SEQ ID NO:2 is the interval of chr6:31973424-31980835 of the Hg19 reference sequence; the target region sequence fragment amplified by SEQ ID NO:3 and SEQ ID NO:2 is the interval of chr6:32006158-32013689 of the Hg19 reference sequence.
[0068] Primer:
[0069] 1. CYP21A2_F (SEQ ID NO:1): 5’-G T ACCTGA A GGTGGGGTC-3’
[0070] 2. TNXA / TNXB-31in_R (SEQ ID NO:2): 5’-CTGCTCCTCATTACTCGCTG-3’
[0071] 3. CYP21A1P / A2_UTR_F (SEQ ID NO:3): 5’-GAGCTATAAGTGGCACCTCAG-3’
[0072] Reagents and equipment (Tables 2 and 3):
[0073] Table 2. Reagents
[0074]
[0075] Table 3. Equipment
[0076]
[0077] Experimental procedures:
[0078] I. Long-fragment PCR:
[0079] 1. Prepare the reaction premix according to the following method:
[0080] The PCR reaction systems corresponding to 2 primer pairs are shown in Tables 4 and 5 respectively.
[0081] Table 4. Long fragment PCR reaction system for primer CYP21A1P / A2_UTR_F / TNXA / TNXB-31in_R
[0082] PCR Reactants Volume (μL) 2X PCR Buffer Matched with KOD FX Neo 12.5 2 mM dNTPs 5 CYP21A1P / A2_UTR_F (10 μM) 0.375 TNXA / TNXB-31in_R (10 μM) 0.375 KOD FX Neo (1 U / μL) 0.5 Nuclease-Free Water 5.25 Total 24
[0083] Table 5. Long fragment PCR reaction system for primer CYP21A2_F / TNXA / TNXB-31in_R
[0084] PCR Reactants Volume (μL) 2X PCR Buffer Matched with KOD FX Neo 12.5 2 mM dNTPs 5 CYP21A2_F (10 μM) 0.375 TNXA / TNXB-31in_R (10 μM) 0.375 KOD FX Neo (1 U / μL) 0.5 Nuclease-Free Water 3.75 Total 22.5
[0085] 2. Add 1 μL or 2.5 μL of sample DNA (10 - 50 ng) as a template into the PCR tube according to different primers, and the total reaction system is 25 μL.
[0086] 3. Place the PCR tube on the PCR instrument and run the PCR program (Table 6).
[0087] Table 6. PCR program
[0088]
[0089] 4. Detect the fragment size of the PCR product by agarose gel electrophoresis.
[0090] 5. Purify the PCR product using 1 volume of nucleic acid purification magnetic beads.
[0091] 6. Quantify the purified PCR product using the Qubit dsDNA kit.
[0092] II. Construction of sequencing library
[0093] 1. Use 10 ng of the purified PCR product and construct the sequencing library using the Hieff NGS OnePot Pro DNA Library Prep Kit for Illumina according to the official instructions.
[0094] 2. Measure the library concentration using the Qubit dsDNA kit and measure the fragment length of the library using the Agilent D1000 reagent.
[0095] III. Library sequencing
[0096] Sequence the constructed library on the Illumina NovaSeq 6000 sequencer in the PE150 mode according to the official operation instructions.
[0097] IV. Actual execution steps of the bioinformatics process:
[0098] Two amplification products of each sample are sequenced (simultaneously if possible). When the target average depth and data volume are reached, all sequencing results are forcibly matched with the reference sequence of the target region to calculate all sites that are different from the sequence of the target region (Hg19: chr6: 32006083-32013949), and then calculate their variant ratio (call ratio). Theoretically, when the sequenced segments are uniform, the value of the ratio of the measured sites itself has a theoretical correlation with copy number variation (see Table 1 in Nucleic Acids Research, 2015, Vol.43, No.14 e90 doi: 10.1093 / nar / gkv319 Figure 1 )(organized as shown in Table 7).
[0099] Table 7. Distribution of copy number changes and allele sites in Nucleic Acids Research, 2015, Vol.43, No.14 e90
[0100]
[0101] This table refers to the information in Table 1 published in Nucleic Acids Research, 2015, Vol, 43, No.14 e90 doi: 10.1093 / nar / gkv319
[0102] When the amplified products of the target region of each sample meet the range of ±0.05 in the values in the parentheses in the above table (1 copy number is no het), the copy numbers of the target region and the target + homologous region can be judged respectively through Figures 6-8 logic, and the theoretical copy number can be corresponded back based on the ±0.05 algorithm of the proportion of each site (Table 9). In the case of no heterozygosity in some target regions, the total copy number of the target homologous mixed product is deduced by using the proportion (call ratio) of the sites (which can be any SNP or SNV appearing in the product) of the target / homologous mixed sequence [for example, in Standard Sample 2, the proportion of sites in the target + homologous region appears at 0.33 ± 0.05 or 0.67 ± 0.05, so it can be speculated that the probability of no heterozygosity in the target region is that there is only 1 copy number in the target region. Because in most cases, for most people (more than 80%), the proportion of sites in the two allelic target regions will be 0.5 ± 0.05]. Even in the case of two target regions with exactly the same sequences, the present invention can also use the consistency of the proportion of 5 marker sites in the homologous region (Table 8) to judge the proportion of the homologous region product in the target / homologous mixed amplification product, and then reverse-deduce whether the actual copy number of the target region is 1 or 2 copies with exactly the same sequences ( Figure 6 ). Therefore, using the above Figures 6-8The actual copy number of the target region can be determined by the logic, and since the semi-specific primers can bias the amplification of the target region sequence when the target region exists, the point of the target region can be directly measured [for example, c.293-13C>G (or chr6:32006858C>G) of standard sample 4]. Other sites that can be distinguished and measured in the target region or homologous region are listed in Table 10. Based on the case of these 6 standard samples, it is proved that the present invention can detect the existence of the target region sites and the target / homologous mixed region sites of multiple samples from the second generation sequencing results of the two amplification products of one sample at a time, and through the proportion of these sites in the sequencing results, the copy number of the target region can be further analyzed through the designed bioinformatics analysis process, or in specific cases, the copy number of the target region can be reversely estimated by comparing the (specific) variation ratio of the target region and the target / homologous mixed product [for example, when the two allele sequences are completely consistent, or when there is no target region (standard sample 6)].
[0103] Table 8. Marker sites of 5 homologous regions
[0104]
[0105] Table 9. Data results can accurately judge the copy number difference of standard samples
[0106]
[0107] Table 10. Sites of target region and homologous region that can be distinguished (known sites located in the interval, but not limited to the sites in this table)
[0108]
[0109]
[0110]
[0111]
[0112] Bioinformatics analysis process Figures 6-8 ):
[0113] S1. Determine the situation: If the site ratio of the target region obtained by sequencing the PCR amplification product of the first primer pair falls greater than 0.95 or less than 0.05, it is determined that the target region is a situation without heterozygous sites; otherwise, it is determined that the target region is a situation with heterozygous sites. The situation with heterozygous sites in the target region is further divided into a first sub - situation and a second sub - situation. The first sub - situation is the situation where the site ratio of the target region obtained by sequencing the PCR amplification product of the first primer pair is 0.5 ± 0.05 and the site ratio of the PCR amplification product of the second primer pair is 0.5 ± 0.05. The second sub - situation is the situation where the site ratio of the target region obtained by sequencing the PCR amplification product of the first primer pair is 0.33 ± 0.05 or 0.67 ± 0.05 and there is no site ratio of 0.5 ± 0.05.
[0114] S2. Under the situation determined in step S1, further determine the copy number, including Method I and Method II;
[0115] Method I: Based on the site ratio of the PCR amplification product of the second primer pair and the site ratio of the homologous region marker site, obtain the corresponding mixed copy number and homologous region copy number, and subtract the homologous region copy number from the mixed copy number to obtain the target region copy number;
[0116] Method II: Based on the site ratio of the PCR amplification product of the second primer pair and the site ratio of the PCR amplification product of the first primer pair, obtain the corresponding mixed copy number and target region copy number, and subtract the target region copy number from the mixed copy number to obtain the homologous region copy number.
[0117] In the situation without heterozygous sites ( Figure 6 ), according to Method I, it is obtained that:
[0118] If the site ratio of the PCR amplification product of the second primer pair falls between 0.2 - 0.3 or 0.7 - 0.8, the mixed copy number is 4. If the site ratio of the homologous region marker site falls between 0.45 - 0.55, the homologous region copy number is 2, so the target region copy number is 2;
[0119] If the site ratio of the PCR amplification product of the second primer pair falls between 0.2 - 0.3 or 0.7 - 0.8, the mixed copy number is 4. If the site ratio of the homologous region marker site falls between 0.7 - 0.8, the homologous region copy number is 3, so the target region copy number is 1;
[0120] The site ratio of the PCR amplification product of the second pair of primers falls within 0.3 - 0.4 or 0.6 - 0.7 and there is no polymorphic site falling within 0.45 - 0.55. If the site ratio of the homologous region marker site falls within 0.61 - 0.71, the mixed copy number is 3, the homologous region copy number is 2, so the target region copy number is 1;
[0121] The site ratio of the PCR amplification product of the second pair of primers falls within 0.3 - 0.4 or 0.6 - 0.7 and there is no polymorphic site falling within 0.45 - 0.55. If the site ratio of the homologous region marker site falls within 0.3 - 0.38, the mixed copy number is 3, the homologous region copy number is 1, so the target region copy number is 2;
[0122] For the PCR amplification product of the second pair of primers, if the site ratio of more than 1 / 6 of the polymorphic sites falls within 0.1 - 0.2 or 0.8 - 0.9 and there is a polymorphic site falling within 0.45 - 0.55, and if the site ratio of the homologous region marker site falls within 0.62 - 0.72, the mixed copy number is 6, the homologous region copy number is 4, so the target region copy number is 2;
[0123] The site ratio of the PCR amplification product of the second pair of primers falls within 0.35 - 0.45 or 0.55 - 0.65 and there is no polymorphic site falling within 0.45 - 0.55. If the site ratio of the homologous region marker site falls within 0.75 - 0.85, the mixed copy number is 5, the homologous region copy number is 4, so the target region copy number is 1;
[0124] The site ratio of the PCR amplification product of the second pair of primers falls within 0.35 - 0.45 or 0.55 - 0.65 and there is no polymorphic site falling within 0.45 - 0.55. If the site ratio of the homologous region marker site falls within 0.55 - 0.65, the mixed copy number is 5, the homologous region copy number is 3, so the target region copy number is 2;
[0125] In the case of no heterozygous sites, according to Method II, it is obtained that:
[0126] If more than 80% of the site ratio of the PCR amplification product of the second pair of primers falls within 0.45 - 0.55, the mixed copy number is 2 and the target region copy number is 1.
[0127] In the first sub - case ( Figure 7 ), according to Method I, it is obtained that:
[0128] If the site ratio of the homologous region marker site ≥ 0.9, the mixed copy number is 2, the homologous region copy number is 2, so the target region copy number is 0;
[0129] If the site ratio of the homologous region marker site ≤ 0.1, the mixed copy number is 2, the copy number of the homologous region is 0, and thus the copy number of the target region is 2;
[0130] In the first sub-case, according to Method II, it is obtained that:
[0131] If the site ratio of the PCR amplification product of the second pair of primers falls within 0.2 - 0.3 or 0.7 - 0.8, the mixed copy number is 4 and the copy number of the target region is 2;
[0132] If the site ratio of the PCR amplification product of the second pair of primers falls within 0.28 - 0.38 or 0.61 - 0.71, the mixed copy number is 3 and the copy number of the target region is 2;
[0133] If there is no site ratio of the PCR amplification product of the second pair of primers falling within 0.2 - 0.3 or 0.7 - 0.8 and there is a site ratio of the PCR amplification product of the second pair of primers falling within ≤ 0.2 or ≥ 0.8, then the mixed copy number > 4 and the copy number of the target region is 2.
[0134] In the second sub-case ( Figure 8 ), according to Method I, it is obtained that: if the site ratio of the homologous region marker site falls within ≤ 0.1, the mixed copy number is 3, the copy number of the homologous region is 0, and thus the copy number of the target region is 3;
[0135] In the second sub-case, according to Method II, it is obtained that:
[0136] If the site ratio of the PCR amplification product of the second pair of primers falls within 0.2 - 0.3 or 0.7 - 0.8, the mixed copy number is 4 and the copy number of the target region is 3;
[0137] If the site ratio of the PCR amplification product of the second pair of primers falls within 0.55 - 0.65, the mixed copy number is 5 and the copy number of the target region is 3;
[0138] If the site ratio of the PCR amplification product of the second pair of primers falls within ≤ 0.2 or ≥ 0.8, the mixed copy number is 6 and the copy number of the target region is 3.
[0139] The above has described in detail the preferred specific embodiments of the present invention. It should be understood that those of ordinary skill in the art can make many modifications and variations according to the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the protection scope determined by the claims.
Claims
1. A method for detecting sequence variations and copy number variations in a target region / homologous region for non-diagnostic and non-therapeutic purposes, characterized in that, The template is subjected to PCR amplification using a first primer pair and a second primer pair respectively, and the amplified products are sequenced to analyze the sequence changes and copy number changes in the target region / homologous region of the template; the homologous region is provided with a homologous region marker site; the homologous region marker site is a conserved site verified as the homologous region sequence, and the variation ratio of the site can be used to judge the copy number of the homologous region; The first primer pair contains a semi-specific primer such that when the target region exists in the template, the sequence in the target region is mainly amplified, and when the target region does not exist in the template but the homologous region exists, the sequence in the homologous region is amplified; the second primer pair amplifies both the sequence in the target region and the sequence in the homologous region when both the target region and the homologous region exist in the template; the semi-specific primer can specifically bind to the sequence in the target region, but there are 2 nucleotides that cannot be complementary paired when binding to the sequence in the homologous region; In the PCR reaction, the amount of template used for the first primer pair is 1.5 - 3 times the amount of template used for the second primer pair; The analysis of obtaining the sequence changes and copy number changes in the target region / homologous region includes the following steps: S1. Determine the situation: According to the site ratio of the target region obtained by sequencing the PCR amplification product of the first primer pair falling greater than 0.95 or less than 0.05, it is judged that the target region is a situation without heterozygous sites, otherwise it is judged that the target region is a situation with heterozygous sites; the situation with heterozygous sites in the target region is further divided into a first sub-situation and a second sub-situation; the first sub-situation is the situation where the site ratio of the target region obtained by sequencing the PCR amplification product of the first primer pair is 0.5 ± 0.05 and the site ratio of the PCR amplification product of the second primer pair is 0.5 ± 0.05; the second sub-situation is the situation where the site ratio of the target region obtained by sequencing the PCR amplification product of the first primer pair is 0.33 ± 0.05 or 0.67 ± 0.05 and there is no site ratio of 0.5 ± 0.05; S2. Under the situation determined in step S1, determine the copy number again, including Method I and Method II; Method I: According to the site ratio of the PCR amplification product of the second primer pair and the site ratio of the homologous region marker site, the corresponding mixed copy number and homologous region copy number are obtained, and the target region copy number is obtained by subtracting the homologous region copy number from the mixed copy number; Method II: According to the site ratio of the PCR amplification product of the second primer pair and the site ratio of the PCR amplification product of the first primer pair, the corresponding mixed copy number and target region copy number are obtained, and the homologous region copy number is obtained by subtracting the target region copy number from the mixed copy number; Sequencing is performed on the two PCR amplification products obtained for each sample using the first primer pair and the second primer pair. When the target average depth and data volume are reached, all the sequencing results are forced to match the target region reference sequence to calculate all the sites that are different from the target region reference sequence, and the site ratio is calculated. Theoretically, when the sequenced segments are uniform, there is a theoretical correlation between the value of the site ratio itself and copy number variation, and the correlation is as follows: The forced matching means that even when partial mutations occur in the target region products, they are still forced to match the target region.
2. The method for detecting sequence variation and copy number variation of a target region / homologous region for non-diagnostic and non-therapeutic purposes as claimed in claim 1, wherein In the case of no heterozygous sites, according to Method I, it is obtained that: If the site ratio of the PCR amplification product of the second primer pair falls within 0.2 - 0.3 or 0.7 - 0.8, the mixed copy number is 4. If the site ratio of the homologous region marker sites falls within 0.45 - 0.55, the homologous region copy number is 2. Therefore, the target region copy number is 2. If the site ratio of the PCR amplification product of the second primer pair falls within 0.2 - 0.3 or 0.7 - 0.8, the mixed copy number is 4. If the site ratio of the homologous region marker sites falls within 0.7 - 0.8, the homologous region copy number is 3. Therefore, the target region copy number is 1. If the site ratio of the PCR amplification product of the second primer pair falls within 0.3 - 0.4 or 0.6 - 0.7 and there are no polymorphic sites falling within 0.45 - 0.55, and the site ratio of the homologous region marker sites falls within 0.61 - 0.71, the mixed copy number is 3, and the homologous region copy number is 2. Therefore, the target region copy number is 1. If the site ratio of the PCR amplification product of the second primer pair falls within 0.3 - 0.4 or 0.6 - 0.7 and there are no polymorphic sites falling within 0.45 - 0.55, and the site ratio of the homologous region marker sites falls within 0.3 - 0.38, the mixed copy number is 3, and the homologous region copy number is 1. Therefore, the target region copy number is 2. If more than 1 / 6 of the polymorphic site ratios of the PCR amplification product of the second primer pair fall within 0.1 - 0.2 or 0.8 - 0.9 and there are polymorphic sites falling within 0.45 - 0.55, and the site ratio of the homologous region marker sites falls within 0.62 - 0.72, the mixed copy number is 6, and the homologous region copy number is 4. Therefore, the target region copy number is 2. If the site ratio of the PCR amplification product of the second primer pair falls within 0.35 - 0.45 or 0.55 - 0.65 and there are no polymorphic sites falling within 0.45 - 0.55, and the site ratio of the homologous region marker sites falls within 0.75 - 0.85, the mixed copy number is 5, and the homologous region copy number is 4. Therefore, the target region copy number is 1. If the site ratio of the PCR amplification product of the second primer pair falls within 0.35 - 0.45 or 0.55 - 0.65 and there are no polymorphic sites falling within 0.45 - 0.55, and the site ratio of the homologous region marker sites falls within 0.55 - 0.65, the mixed copy number is 5, and the homologous region copy number is 3. Therefore, the target region copy number is 2. In the case of no heterozygous sites, according to Method II, it is obtained that: If the proportion of sites in the PCR amplification product of the second primer pair exceeding 80% falls within 0.45 - 0.55, then the mixed copy number is 2 and the copy number in the target region is 1.
3. The method for detecting sequence variation and copy number variation of a target region / homologous region for non-diagnostic and non-therapeutic purposes as claimed in claim 1, wherein In the first sub - case, according to Method I, it is obtained that: If the proportion of the homologous region marker sites ≥ 0.9, then the mixed copy number is 2, the copy number in the homologous region is 2, so the copy number in the target region is 0; If the proportion of the homologous region marker sites ≤ 0.1, then the mixed copy number is 2, the copy number in the homologous region is 0, so the copy number in the target region is 2; In the first sub - case, according to Method II, it is obtained that: If the proportion of sites in the PCR amplification product of the second primer pair falls within 0.2 - 0.3 or 0.7 - 0.8, then the mixed copy number is 4 and the copy number in the target region is 2; If there is no proportion of sites in the PCR amplification product of the second primer pair falling within 0.2 - 0.3 or 0.7 - 0.8 and there is a proportion of sites in the PCR amplification product of the second primer pair falling within ≤ 0.2 or ≥ 0.8, then the mixed copy number > 4 and the copy number in the target region is 2.
4. The method for detecting sequence variations and copy number variations in a target region / homologous region for non-diagnostic and non-therapeutic purposes as claimed in claim 1, wherein In the second sub - case, according to Method I, it is obtained that: if the proportion of the homologous region marker sites falls within ≤ 0.1, then the mixed copy number is 3, the copy number in the homologous region is 0, so the copy number in the target region is 3; In the second sub - case, according to Method II, it is obtained that: If the proportion of sites in the PCR amplification product of the second primer pair falls within 0.2 - 0.3 or 0.7 - 0.8, then the mixed copy number is 4 and the copy number in the target region is 3; If the proportion of sites in the PCR amplification product of the second primer pair falls within 0.55 - 0.65, then the mixed copy number is 5 and the copy number in the target region is 3; If the proportion of sites in the PCR amplification product of the second primer pair falls within ≤ 0.2 or ≥ 0.8, then the mixed copy number is 6 and the copy number in the target region is 3.
5. A computer-readable storage medium, characterized in that, It includes a program that can be executed by a processor to implement the method according to any one of claims 1 - 4.
6. The computer-readable storage medium according to claim 5, wherein The target region is the interval of exon 32 - 44 of CYP21A2+TNXB; the homologous region is CYP21A1P - TNXA; the semi - specific primer is SEQ ID NO:
1.
7. The computer-readable storage medium according to claim 6, wherein Both the first primer pair and the second primer pair contain a common reverse primer SEQ ID NO: 2; the second primer pair also contains a non - specific amplification forward primer SEQ ID NO:
3.
8. The computer-readable storage medium according to claim 7, wherein Both the first primer pair and the second primer pair are independently or non - independently subjected to PCR amplification under the following PCR reaction program: Pre - denaturation at 94°C for 2 min ; Incubation at 4°C.
9. The computer-readable storage medium according to claim 7, wherein The homologous region marker sites are shown as follows; ; The reference sequence of the target region is Hg19:chr6:32006083 - 32013949.
Citation Information
Patent Citations
Kit, reaction system and method for detecting gene copy number variation in sample
CN113736865A
Methods and systems for identifying recombinant variants
WO2022261010A1