Target region amplification method suitable for long-read-length three-generation sequencing
By combining long-read third-generation sequencing with multiple pairs of overlapping amplification primers, the complexity and high cost of single-gene genetic disease detection in existing technologies have been solved, enabling efficient and accurate detection in the absence of probands or incomplete family pedigrees.
Patent Information
- Application Number
- CN202511550129.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-01-13
AI Technical Summary
Existing methods for detecting single-gene genetic diseases, such as STR technology, SNP chips, and second-generation sequencing technology, are complex to operate, costly, and have limited accuracy. They are particularly difficult to complete when family samples are incomplete or there is no proband. Furthermore, haplotype linkage analysis must be performed separately, resulting in long testing times and high costs.
By employing long-read third-generation sequencing combined with multiple pairs of overlapping amplification primers, an amplification method covering the target region was designed. Through two-stage long-fragment PCR and third-generation sequencing, information on single-gene genetic diseases and haplotypes was directly obtained, reducing sequencing costs and analysis time.
It significantly reduces sequencing costs and analysis time, improves detection accuracy, and enables single-gene disease detection in the absence of probands or incomplete family histories, thus shortening the detection cycle.
Smart Images

Figure CN121320501A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biological detection technology, and in particular to a method for amplifying target regions suitable for long-read third-generation sequencing. Background Technology
[0002] Single-gene genetic diseases, also known as monogenic diseases, are hereditary diseases caused by mutations in one or both alleles of a single gene; they are also called Mendelian genetic diseases. Depending on the chromosome in which the disease-causing gene is located, they can be divided into autosomal and sex-linked genetic diseases. Currently, there are over 6,000 types of single-gene genetic diseases with clearly defined phenotypes and molecular pathogenic mechanisms. These diseases often have fatal, teratogenic, or disabling characteristics, and only about 5% are treatable, with expensive treatment costs. Therefore, prenatal diagnosis or preimplantation genetic testing (PGT) is crucial for the prevention of single-gene genetic diseases.
[0003] In the field of assisted reproduction, preimplantation genetic testing (PGT) is an extremely important diagnostic tool. It allows for genetic testing of embryos before implantation into the uterus, selecting normal or low-risk embryos for transfer and providing a strong guarantee for improving pregnancy success rates. Among these, preimplantation genetic testing for single-gene diseases (PGT-M) involves biopsiing and testing embryos from families with known pathogenic mutations in single-gene inherited diseases. Embryos without the pathogenic mutation are selected for transfer, thereby blocking the inheritance of single-gene diseases and resulting in healthy offspring.
[0004] PGT-M testing includes the detection of pathogenic mutation sites, linkage analysis of family sample genes based on SNP sites, and analysis of embryonic chromosome aneuploidy. Currently, there are three common linkage analysis methods: short tandem repeats (STRs), SNP microarrays, and next-generation sequencing technology. However, all three methods have limitations to varying degrees.
[0005] STR technology is relatively complex and costly, and STR and SNP chips can only analyze known sites, thus missing some unknown potential SNP sites that can be used for haplotype identification, limiting the accuracy of detection. While next-generation sequencing (NGS) technology has high throughput and resolution, it can only detect DNA fragments of about 300 bases, while the average spacing of SNP sites in the human genome is about 1000 bases. Therefore, finding continuously usable SNPs is difficult, limiting the widespread use of this method.
[0006] The three methods described above typically require complete family pedigrees to obtain haplotypes associated with the variant sites. This is not only complex and costly, but also necessitates obtaining samples from both parents and the proband. However, in clinical practice, incomplete family samples or the absence of a proband often make these methods difficult to complete. Secondly, current haplotype linkage analysis can only utilize known loci or select only a limited number of loci to construct haplotypes, resulting in the waste of a large number of SNP loci. Furthermore, in current clinical practice, single-gene genetic disease testing and haplotype analysis are generally performed separately, leading to lengthy and costly preimplantation genetic testing (PGT), placing a greater burden on patients and their families. Summary of the Invention
[0007] Based on this, the present invention overcomes the shortcomings and deficiencies of the prior art and provides a method for amplifying target regions suitable for long-read third-generation sequencing to obtain single-gene genetic disease information and haplotype information of the target region.
[0008] The above-mentioned objective of this invention is achieved through the following technical solution: A method for amplifying target regions suitable for long-read third-generation sequencing includes the following steps: S1. Design at least one pair of amplification primers so that the amplification primers can cover the target region; If the length of the target region is less than or equal to 20kb, then design a pair of amplification primers; If the length of the target region is greater than 20kb, multiple pairs of amplification primers are designed. Adjacent amplification primers have overlapping regions. The multiple pairs of amplification primers are divided into two groups, and the coverage regions of each amplification primer in each group do not overlap. S2. The gDNA of the test sample is amplified using the amplification primers from step S1 to obtain the test product. S3. Library construction and third-generation sequencing of the test product from step S2; S4. Perform single-gene genetic disease detection and / or haplotype analysis on the sequencing data to obtain single-gene genetic disease information and / or haplotype information of the target gene.
[0009] Compared with existing technologies, this invention significantly reduces sequencing costs (equivalent to a reduction of more than 45 times the original 70-100G sequencing cost) and analysis time without affecting the accuracy of the test results, and shortens the testing cycle. It can be widely used in multiple fields such as single-gene disease detection, preimplantation genetic diagnosis, and prenatal diagnosis.
[0010] Furthermore, when the length of the target region is greater than 20kb, step S2 includes: The two sets of amplification primers from step S1 were used to perform the first stage of long fragment amplification of the gDNA of the test sample to obtain the intermediate amplification product of the target gene. The intermediate amplification product was diluted and used as an amplification template. The amplification template was then subjected to the second stage of long fragment amplification using the two sets of amplification primers from step S1, resulting in two sets of amplification products. The two sets of amplification products were purified and recovered, and then mixed in equal amounts to obtain the test product.
[0011] Furthermore, the conditions for long fragment amplification in the first stage are: pre-denaturation at 94℃ for 5 min, denaturation at 94℃ for 30 s, extension at 68℃ for 12 min, 10-20 cycles, extension at 68℃ for 7 min, and then cooling to 4℃.
[0012] Furthermore, the conditions for long fragment amplification in the second stage are as follows: pre-denaturation at 94°C for 5 min, denaturation at 94°C for 30 s, extension at 68°C for 12 min, 20-30 cycles, extension at 68°C for 7 min, and then cooling to 4°C.
[0013] Furthermore, in step S2, the intermediate amplification product is diluted by a factor of 40-60.
[0014] Furthermore, when the length of the target region is less than or equal to 20kb, the conditions for long fragment amplification in step S2 are: pre-denaturation at 94℃ for 5 min, denaturation at 94℃ for 30 s, extension at 68℃ for 10 min or 12 min, 30-40 cycles, extension at 68℃ for 7 min, and then cooling to 4℃.
[0015] Furthermore, the duration of the extension is 30-60 s / kb.
[0016] Furthermore, when detecting information on single-gene genetic diseases, the target region is the full-length sequence of the target gene; When detecting only haplotypes or simultaneously detecting single-gene genetic diseases and haplotypes, the target region is a region at least 500kb upstream and downstream of the target gene locus.
[0017] Furthermore, each pair of amplification primers can cover a region of up to 20kb; adjacent amplification primers have an overlap region of 100bp-5kb.
[0018] Furthermore, the amplification primers meet the following requirements: GC content of 40%-60%, primer annealing temperature of 66-68℃, and primer length of 25-30bp. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the target region amplification method for long-read third-generation sequencing provided by the present invention.
[0020] Figure 2 As in Example 1 HBB Electrophoresis diagram of gene amplification products.
[0021] Figure 3 As in Example 1 KCNV2 Electrophoresis diagram of gene amplification products.
[0022] Figure 4 As in Example 1 PROM1 Electrophoresis diagram of gene amplification products.
[0023] Figure 5 This is a diagram of the third-generation sequencing results of sample M1 from Example 2.
[0024] Figure 6 This is a diagram of the third-generation sequencing results of sample M2 from Example 2.
[0025] Figure 7 This is a diagram of the third-generation sequencing results of sample M31 from Example 2.
[0026] Figure 8 This is a diagram of the third-generation sequencing results of sample M32 from Example 2.
[0027] Figure 9 This is a diagram of the third-generation sequencing results of sample M33 from Example 2.
[0028] Figure 10 This is a diagram of the third-generation sequencing results of sample M4 from Example 3.
[0029] Figure 11 This is a diagram of the third-generation sequencing results for Example 4. Detailed Implementation
[0030] Long-read sequencing, also known as third-generation sequencing, has the ability to read longer DNA fragments. The average length of the human genome is approximately 2500 bp, and long-read sequencing not only completes sequencing quickly but also detects more complete genomic information, effectively reducing subsequent genome assembly work. Based on these characteristics, this invention considers utilizing the long-read features of third-generation sequencing to directly perform third-generation sequencing on target genes / regions in both partners of the patient. By genotyping the SNP sites in the sequencing results, haplotype analysis can be completed without a proband or incomplete family pedigree, for the detection of preimplantation genetic disorders.
[0031] Currently, haplotype analysis based on third-generation sequencing mainly employs whole-genome sequencing protocols. This approach avoids PCR amplification, has higher reagent and sequencing costs, and involves more complex data analysis, all of which limit its widespread clinical application.
[0032] This invention attempts to perform third-generation sequencing on the products amplified by PCR to reduce the complexity of data analysis. However, conventional PCR has limited amplification capacity, amplifying fragments up to 2-3 kb at most. Long-fragment PCR can amplify fragments up to 20 kb in length. Combining long-fragment PCR with third-generation sequencing helps reduce the amount of data, but the obtained fragment length is still relatively short, making it difficult to obtain enough SNP sites for haplotype analysis.
[0033] To address the aforementioned issues, this invention designs multiple pairs of amplification primers using an overlapping method to cover the target gene locus and its upstream and downstream regions of at least 500 kb, thereby ensuring sufficient SNP loci are obtained for haplotype analysis. See also... Figure 1 The present invention provides a method for amplifying target regions suitable for long-read third-generation sequencing, comprising the following steps: S0, Extract gDNA from the test sample.
[0034] S1. Design primers for amplifying the target region.
[0035] If the sequence length of the target region is less than or equal to 20kb, design a pair of amplification primers and then proceed with steps S2 to S4 in sequence. If the sequence length of the target region is greater than 20kb, design multiple pairs of amplification primers so that the multiple pairs of amplification primers can cover the target region and that adjacent amplification primers have overlapping regions. Then, divide the multiple pairs of amplification primers into two groups, and ensure that the coverage regions of each amplification primer in each group do not overlap. Then, proceed with steps S2 to S4 in sequence.
[0036] When only single-gene genetic diseases are detected, the target region is the full-length sequence of the target gene.
[0037] When haplotype analysis is required (either haplotype analysis only or haplotype analysis and single-gene genetic disease detection are performed simultaneously), the target region is a region at least 500kb upstream and downstream of the target gene locus.
[0038] S2. Using the amplification primers from step S1, the gDNA from step S0 is amplified to obtain the amplification product of the target gene. The amplification product is purified and recovered to obtain the product to be tested.
[0039] S3. Library construction and third-generation sequencing of the test product from S2.
[0040] S31. The amplification product from step S2 is subjected to end repair and "A" is added to obtain the product with end repair and "A" added.
[0041] S32, end repair plus purification of product "A".
[0042] S33. The purified end-repaired product with "A" is ligated to a sequencing adapter to obtain the ligation product.
[0043] S34. Purification of the ligation product.
[0044] S35. Perform third-generation sequencing on the purified ligation product.
[0045] S4. Perform single-gene genetic disease detection and / or haplotype analysis on the sequencing data to obtain single-gene genetic disease information and / or haplotype information.
[0046] Because the target gene locus and its upstream and downstream sequences are too long, long-fragment PCR is prone to problems such as low yield and low mismatch rate, leading to reduced sequencing accuracy. To improve sequencing accuracy, this invention further improves step S2: When the sequence length of the target region is less than or equal to 20kb, the gDNA in step S0 is amplified once for a long fragment, with a cycle number of 30-40 cycles, i.e., steps S2, S3, and S4 are performed sequentially. When the sequence length of the target region is greater than 20kb, the gDNA in step S0 is amplified in two stages. The first stage has 10-20 cycles, and the second stage has 20-30 cycles, i.e., steps S21, S22, S23, S3, and S4 are performed sequentially.
[0047] S21. First stage of long fragment amplification: The gDNA from step S0 is amplified using the two sets of amplification primers from step S1 to obtain the intermediate amplification product of the target gene.
[0048] S22. Second stage of long fragment amplification: The intermediate amplification product of step S21 is diluted and used as an amplification template. The amplification template is then amplified with long fragments using the two sets of amplification primers from step S1 to obtain two sets of amplification products.
[0049] S23. The two groups of amplification products were purified and recovered separately to obtain two groups of purified products. After the purified products passed the quality inspection, they were mixed in equal amounts to obtain the product to be tested.
[0050] Compared to conventional long-fragment PCR, the two-stage + dilution method achieves higher yield, lower mismatch rate, and higher uniformity. In the first stage of long-fragment amplification, a low cycle number is used, allowing enzymes and primers to be preferentially consumed by the long fragment amplification, reducing over-amplification of non-specific fragments and primer dimers. Diluting the intermediate amplification products from the first stage reduces the concentration of non-specific amplification and primer dimers, minimizing their inhibitory effects in subsequent amplification rounds. The second stage of long-fragment amplification introduces new enzymes and primers, re-entering multi-cycle exponential amplification, resulting in more efficient amplification of the long fragment.
[0051] Based on the above method, the present invention adopts HBB , KCNV2 , PROM1 , PHEX Using four genes as examples, this invention introduces a method for amplifying target regions suitable for long-read third-generation sequencing. The method is then validated using positive samples verified by Sanger sequencing. The following detailed description is provided in conjunction with examples.
[0052] Unless otherwise specified, the reagents, methods, and equipment used in this invention are conventional reagents, methods, and equipment in this technical field. Test methods in the examples that do not specify specific experimental conditions are generally performed under conventional experimental conditions or according to the manufacturer's recommended experimental conditions. Unless otherwise specified, the reagents and raw materials used in this invention are commercially available.
[0053] Example 1 This embodiment is... HBB , KCNV2 , PROM1 The three genes were amplified to their full length using the following methods: S0. Extract gDNA from the test sample. Specifically, gDNA was extracted from the test sample using the MagPure Blood DNA LQ Kit from Meiji Biotechnology, and the concentration was detected using Qubit.
[0054] S1. Design amplification primers for the target region. The target region is the full-length sequence of the target gene.
[0055] The full-length sequences of the three target genes were downloaded from the NCBI database, using the human reference genome GRCh37 / hg19 as a benchmark. HBB The full-length sequence of the gene includes chr11:524694-5248301, and the gene length is 9019 bp. KCNV2 The full-length sequence of the gene includes chr9:2717510-2730037, and the gene length is 12.5 kb. PROM1 The full-length sequence of the gene includes chr4:15969851-16085646, and the gene length is 115.8 kb.
[0056] Primers are designed based on different gene lengths. For genes with a length of 20kb or less, a pair of amplification primers is designed to cover the target region.
[0057] For genes longer than 20kb, multiple pairs of amplification primers need to be designed until the entire target region is covered. Adjacent primers need to overlap to ensure that the entire gene is effectively amplified.
[0058] In this embodiment, each pair of amplification primers covers a region of 15kb-20kb, and two adjacent pairs of amplification primers have an overlap region of 100bp-3kb. The size of the overlap region is determined according to the gene sequence and primer design.
[0059] Specifically, under the premise of meeting the primer design requirements (i.e., primer GC content of 40%-60%, primer annealing temperature of 66-68℃, primer length of 25-30bp, and primer specificity), the size of the overlapping region should be minimized as much as possible. When the sequence of some regions cannot meet the primer design requirements, such as excessive GC content, insufficient specificity, or difficulty in finding suitable matching primers, the overlapping region should be expanded to meet the primer design requirements.
[0060] After the primers were designed, their specificity was evaluated using NCBI. Once the evaluation was successful, primer synthesis proceeded. Primer synthesis was performed by Sangon Biotech. Before use, the primers must be diluted to 10 μM.
[0061] because HBB Since the gene length is less than 20kb, a pair of amplification primers were designed using Primer Premier5 software. The primer sequences are shown in SEQ ID No.1 and SEQ ID No.2, covering a region of 9964bp.
[0062] Table 1 HBB Gene amplification primer sequences
[0063] because KCNV2 Since the gene length is less than 20kb, a pair of amplification primers were designed, the sequences of which are shown in SEQ ID No. 83 and SEQ ID No. 84, covering a region of 15.4kb.
[0064] Table 2 KCNV2 Gene amplification primer sequences
[0065] because PROM1 Since the gene length is greater than 20kb, multiple pairs of amplification primers were designed using Primer Premier5 software. Each pair of primers covers a region of 15kb-20kb, and adjacent primers have an overlap of 100bp-3kb. In this embodiment, PROM1 The gene has 7 pairs of amplification primers. PROM1-1F / 1R to PROM1-4F / 4R , PROM1-22F / 22R to PROM1-24F / 24RThe sequences are shown in SEQ ID No. 149-SEQ ID No. 156 and SEQ ID No. 191-SEQ ID No. 196, covering a total area of 119kb. The coverage areas of the above 7 pairs of amplification primers are 18.9kb, 16.5kb, 18.9kb, 18.3kb, 17.4kb, 16.2kb, and 17.6kb, respectively, with an overlap of 400bp-3kb between adjacent primers.
[0066] Table 3 PROM1 Gene amplification primer sequences
[0067] S2. The amplification primers from step S1 are used to amplify the gDNA from step S0 into a long fragment to obtain the amplification product.
[0068] The reaction system for long fragment amplification is shown in Table 4. The amount of gDNA used was 200 ng, and the long fragment amplification reagent used was Novizan 2×Vazyme LAmp Master Mix. The reaction program for long fragment amplification is shown in Table 5. The extension time in the reaction program met the requirement of 30-60 s / kb. HBB The gene extension time is 10 minutes. KCNV2 Genes and PROM1 The gene extension time is 12 minutes.
[0069] Table 4 Reaction system for long fragment amplification
[0070] Table 5 Reaction Procedures for Long Fragment Amplification
[0071] PCR products were purified and recovered using 1.0× purification magnetic beads (Kesheng). HBB The gene amplification product was named LS01-1. KCNV2 The amplification product of the gene was named LS02-1. PROM1 Seven gene amplification products were obtained, named LS03-1, LS03-2, LS03-3, LS03-4, LS03-5, LS03-6, and LS03-7. The recovered purified products were analyzed by agarose gel (0.6%) electrophoresis. The electrophoresis results are shown below. Figure 2-4 As shown, the length of the target band obtained is as expected, and the product band is uniform.
[0072] The concentration of the purified product was detected using Qubit. As shown in Table 6, the concentration of the amplified product is within the normal range and meets the sequencing requirements.
[0073] Table 6. Concentration of Amplified Products
[0074] Example 2 This embodiment includes HBB Positive samples M1 and M2 containing gene mutation sites, and samples containing KCNV2 Samples M31, M32, and M33, which showed positive results for gene mutation sites, were used as test samples to introduce the detection method for single-gene genetic diseases (gene length ≤20kb). Among them, the mutation site in positive sample M1 was... HBB: CD41-42:c.126_129delCTTT, chromosomal location chr11:5247993-5247996, the variant site of positive sample M2 is... HBB: IVS-II-654:c.316-197C>T, chromosomal location chr11:5247153, the variant sites in positive samples M31, M32, and M33 are all... KCNV2 :c.523G>A, both chromosomes are located at chr9:2718262.
[0075] S0, Extract gDNA from the test sample.
[0076] S1. Design amplification primers for the target region. The target region is the full-length sequence of the target gene. HBB The primer sequences for gene amplification are shown in SEQ ID No. 1 and SEQ ID No. 2. KCNV2 The primer sequences for gene amplification are shown in SEQ ID No. 83-SEQ ID No. 84.
[0077] S2. Using the amplification primers from step S1, perform a long-fragment amplification of the gDNA from step S0 to obtain the amplification product. The long-fragment amplification method is the same as step S2 in Example 1, and will not be described again.
[0078] The amplification products were purified and recovered using 1.0× ratio purification magnetic beads (Kesheng). The concentrations of the purified products were detected using a Qubit analyzer. The concentrations of M1, M2, M31, M32, and M33 were 86.8 ng / μL, 80.53 ng / μL, 90.27 ng / μL, 78.1 ng / μL, and 92.9 ng / μL, respectively, all within the normal range and meeting sequencing requirements.
[0079] S3. Library construction and third-generation sequencing of the purified products. The amplified products were prepared into a library using the CycloneSEQ universal library preparation kit, and then third-generation sequencing was performed on the BGI CycloneSEQ series sequencing platform using the CycloneSEQ adaptation sequencing kit and sequencing chip.
[0080] S31. End Repair with "A": Add 1 μg of purified product to a PCR tube, and bring the volume to 45 μL with nuclease-free water. Add the end repair with "A" reaction solution to the purified product at a ratio of 15 μL of end repair with "A" reaction solution to 1 μg of purified product. Mix well and react according to the conditions in Table 7 to obtain the end repair with "A" product. The end repair with "A" reaction solution includes end repair buffer and end repair enzyme at a volume ratio of 4:1.
[0081] Table 7. End-stage repair with "A" reaction conditions
[0082] S32. Purification of the end-repair plus "A" product: Pipette 60 μL of magnetic beads (DNA Clean Beads) into a new 1.5 mL low-adsorption tube, add the end-repair plus "A" product, and mix well. Fix the low-adsorption tube on a rotary mixer and rotate slowly, incubating at room temperature for 10 minutes. After brief centrifugation, place on a magnetic rack and let stand for 2-5 minutes until the liquid is clear. Aspirate and discard the supernatant.
[0083] Keep the low-adsorption tube on the magnetic rack, add 200 μL of freshly prepared 80% ethanol to rinse the magnetic beads and tube walls, carefully aspirate the supernatant and discard it, and repeat this step. Then, aspirate the remaining liquid at the bottom of the low-adsorption tube, open the tube cap, and dry at room temperature until the surface of the magnetic beads is no longer reflective and cracked.
[0084] Remove the low-adsorption tube from the magnetic rack, add 62 μL of nuclease-free water for DNA elution, mix thoroughly, centrifuge briefly for 1 second, and then incubate at room temperature for 10 min. After the reaction, place it on the magnetic rack and let it stand for 2-5 minutes until the liquid is clear. Transfer 61 μL of the supernatant to a new low-adsorption tube to obtain the purified end-repair plus "A" product. Take 1 μL of this product for concentration determination.
[0085] S33. Sequencing adapter ligation: Prepare the ligation reaction solution according to the ratio of 25 μL ligation buffer: 10 μL ligase: 2.5 μL nuclease-free water. Add 2.5 μL of sequencing adapter to the sample from step S32, mix well, add 37.5 μL of ligation reaction solution, mix thoroughly, and incubate the ligation reaction at 25°C for 30 min to obtain the ligation product.
[0086] S34. Purification of ligation product: Add 40 μL of magnetic beads (DNA Clean Beads) to the ligation product and mix well. Fix the mixture on a rotary mixer and incubate at room temperature for 10 min. Then place it on a magnetic rack and let it stand for 2-5 min until the liquid is clear. Then aspirate and discard the supernatant.
[0087] Add 150 μL of long fragment cleaning solution (WB for Long Fragment) and mix well. Let it stand on a magnetic rack for 2-5 minutes until the liquid is clear. Then aspirate and discard the supernatant. Repeat this step and aspirate any remaining liquid at the bottom of the tube.
[0088] Add 17 μL of elution buffer and mix well. Incubate at room temperature for 10 min, then place on a magnetic rack and let stand for 2-5 min until the liquid becomes clear. Transfer 16 μL of the supernatant to a new PCR tube to obtain the purified ligation product. Take 1 μL of this product for concentration determination.
[0089] S35. On the BGI CycloneSEQ series sequencing platform, the ligation products were sequenced using the CycloneSEQ compatible sequencing kit and sequencing chip for third-generation sequencing.
[0090] S4. Data analysis to identify single-gene genetic diseases of the target gene.
[0091] The data after sequencing were filtered for quality and compared. The results are shown in Table 8. Figure 5-9 As shown, positive sites in all 5 groups of samples can be accurately detected, requiring only 1-2G of data and a sequencing depth of over 6000×, far exceeding the sequencing depth requirement for target site detection (normal sequencing depth requirement is >200×).
[0092] Table 8. Third-generation sequencing data output
[0093] Whole-genome third-generation sequencing was performed directly on the gDNA of the test samples M1, M2, M31, M32, and M33 as comparative example 1. The sequencing data was then filtered for quality and compared with the data in example 2.
[0094] As shown in Table 8, the depth of the target site in Comparative Example 1 is less than 100×, which cannot meet the normal sequencing depth requirements. In contrast, the method in Example 2 significantly reduces the amount of sequencing data and significantly increases the average depth of the target site.
[0095] Example 3 This embodiment includes PROM1 The positive sample M4 at the gene mutation site is used as the test sample to introduce the detection method for single-gene genetic diseases (gene length > 20kb).
[0096] Among them, the mutation site of positive sample M4 is: PROM1 :c.1090C>T, chromosome location is chr4:16014922.
[0097] S0, Extract gDNA from the test sample.
[0098] S1. Design amplification primers for the target region. The target region is the full-length sequence of the target gene. PROM1 The primer sequences for gene amplification are shown in SEQ ID No. 149-SEQ ID No. 156 and SEQ ID No. 191-SEQ ID No. 196.
[0099] The seven pairs of amplification primers were divided into two groups, with the coverage areas of multiple pairs of amplification primers in each group not overlapping. In this embodiment, the non-overlapping amplification primers... PROM1-1F / 1R to PROM1-4F / 4R Mixed into a set (M4-1), non-overlapping amplification primers PROM1-22F / 22R to PROM1-24F / 24R Mix them into one group (M4-2).
[0100] S2. Using the amplification primers from step S1, perform two-stage long fragment amplification of the gDNA from step S0 to obtain... PROM1 The amplification product of a gene.
[0101] S21. The two sets of amplification primers from step S1 were used to perform the first stage of long fragment amplification of the gDNA from M4 in step S0. Amplification primers M4-1 and M4-2 were used to amplify the gDNA, and the reaction system is shown in Table 4, while the reaction program is shown in Table 9, yielding two sets of intermediate amplification products. The extension time was 12 min, and the number of cycles was 15.
[0102] Table 9. Reaction Procedure for Long Fragment Amplification in Stage 1
[0103] S22. After diluting the intermediate amplification product, use it as an amplification template. Then, use the two sets of amplification primers from step S1 to perform the second stage of long fragment amplification on the amplification template to obtain two sets of amplification products. The reaction system for the second stage of long fragment amplification is shown in Table 10, and the reaction procedure is shown in Table 11.
[0104] Table 10 Reaction system for the second stage of long fragment amplification
[0105] Table 11 Reaction Procedure for Long Fragment Amplification in Stage Two
[0106] S23. The two sets of amplification products were purified and recovered using 1.0× purification magnetic beads, resulting in two sets of purified products. The concentration of the purified products was detected using a Qubit analyzer, and the concentrations were 98.25 ng / μL and 90.6 ng / μL, respectively, which are within the normal range and meet the sequencing requirements. After the two sets of purified products passed quality control, equal volumes of the purified products from M4-1 and M4-2 were mixed to obtain the product to be tested.
[0107] S3. The test product from S23 is prepared into a library and subjected to third-generation sequencing. The library preparation and third-generation sequencing methods are the same as in Example 2, and will not be repeated in this example.
[0108] S4. Analysis of single-gene hereditary diseases.
[0109] The data after sequencing were filtered and compared for quality. Figure 10 Showing positive sites ( PROM1 The sequencing depth of the target site was 2795.5×, which is 1.35G and 1090C>T. It can be accurately detected. The average sequencing depth of the target site is 2795.5×, which far exceeds the sequencing depth requirement for target site detection (normal sequencing depth requirement is >200×).
[0110] Whole-genome third-generation sequencing of the gDNA in the test samples was used as Comparative Example 2. The sequenced data underwent quality filtering and was compared with the data from Example 3. The results showed that Comparative Example 2 had a data volume of 95.33 G, which was more than 70 times larger than that of Example 3. Furthermore, the average depth of the target loci in Comparative Example 2 was only 30.82 × 10⁻⁶, which does not meet the requirements for normal sequencing depth.
[0111] Example 4 This embodiment uses family samples that have been verified by microarray detection as test samples to introduce a method for simultaneously performing single-gene genetic disease detection and haplotype analysis.
[0112] The proband of the tested sample carried a heterozygous variant. PHEX c.2148-1G>A (chrX:22265967), inherited from the father, the mother does not carry this variant.
[0113] S0, Extract gDNA from the test sample.
[0114] S1. Design primers for amplifying the target region.
[0115] PHEXThe full-length sequence of the gene includes chrX:22050443-22269427, and the gene length is 219 kb. Forty-one pairs of amplification primers were designed using Primer Premier 5 software to cover the target region (i.e., the positive variant site). PHEX The sequences, as shown in SEQ ID No. 231-SEQ ID No. 312, cover a region of 691kb (chrX:21916266-22607263). Each pair of amplification primers covers a region of 15kb-20kb, and adjacent primers have an overlap of 100bp-3kb.
[0116] The 41 pairs of amplification primers were divided into two groups, with the coverage areas of multiple pairs of amplification primers in each group not overlapping. In this embodiment, the non-overlapping amplification primers... PHEX-1F / 1R to PHEX-21F / 21R (Sequences shown in SEQ ID No. 231-SEQ ID No. 272) are mixed into a group (Tube01), with non-overlapping amplification primers. PHEX-22F / 22R to PHEX-41F / 41R (The sequences shown in SEQ ID No. 273-SEQ ID No. 312) are mixed into one group (Tube02).
[0117] When only single-gene genetic diseases are being detected, only 14 pairs of amplification primers are needed. PHEX-5F / 5R to PHEX-11F / 11R , PHEX-25F / 25R to PHEX-31F / 31R This can cover the entire gene length, with primer sequences shown in SEQ ID No. 239-SEQ ID No. 252 and SEQ ID No. 279-SEQ ID No. 292. These are non-overlapping amplification primers. PHEX-5F / 5R to PHEX- 11R / 11R Mixed into a set (Tube01), non-overlapping amplification primers PHEX-25F / 25R to PHEX-31R / 31R Mixed into one group (Tube02).
[0118] S2, two-stage long fragment amplification.
[0119] S21. The gDNA from step S0 is amplified in the first stage using the two sets of amplification primers Tube01 and Tube02 from step S1, respectively.
[0120] S22. After diluting the intermediate amplification product, use it as an amplification template and use the two sets of amplification primers Tube01 and Tube02 from step S1 to perform the second stage of long fragment amplification on the amplification template.
[0121] S23. The two sets of amplification products were purified and recovered using 1.0× purification magnetic beads to obtain two sets of purified products. After passing quality inspection, the two sets of purified products were mixed in equal volumes to obtain the product to be tested.
[0122] The methods of steps S21-S23 are the same as those of steps S21-S23 in Embodiment 3, and will not be repeated in this embodiment.
[0123] Table 12 PHEX Gene amplification primer sequences
[0124] S3. The test product from S23 is prepared into a library and subjected to third-generation sequencing. The library preparation and third-generation sequencing methods are the same as in Example 1, and will not be repeated in this example.
[0125] S4. Single-gene genetic disease detection and haplotype analysis.
[0126] Single-gene disease detection: The data from steps S2-S3, which involved amplification and sequencing using 14 pairs of amplification primers (primer sequences shown in SEQ ID No. 239-SEQ ID No. 252 and SEQ ID No. 279-SEQ ID No. 292), underwent quality filtering. The results are as follows: Figure 11 As shown. Sequencing data volume 1.49G, positive sites PHEX :c.2148-1G>A can be accurately detected, and the average depth is 3570×.
[0127] Table 13 SNP typing results
[0128] Haplotype analysis: Haplotype analysis was performed on the data from steps S2-S3, which were amplified and sequenced using 41 pairs of amplification primers (primer sequences are shown in SEQ ID No. 231-SEQ ID No. 312). The SNP typing results are shown in Table 13. There are at least 30 SNP sites available upstream and downstream of the pathogenic site.
[0129] In the table, HAP1 and HAP2 represent two haplotypes, and the SNP information corresponds to the positive strand of the reference genome. The coordinate 22265967 in the table corresponds to... PHEX The c.2148-1G>A site is shown in the table. It can be seen from the table that HAP2 is the haplotype containing the pathogenic mutation. Underlined and bold text in the table indicates the pathogenic mutation.
[0130] Table 14 Results of Chip-Based Family Linkage Analysis
[0131] Refer to Table 14, which shows the results of the pedigree linkage analysis using the microarray method. F-HAP represents the male's haplotype (pathogenic chain), M-HAP1 and M-HAP2 represent the two haplotypes of the female, and P-HAP1 and P-HAP2 represent the two haplotypes of the proband, where P-HAP2 is the pathogenic chain. SNP information corresponds to the positive strand of the reference genome. Underlined bold text in the table indicates pathogenic mutations.
[0132] As can be seen from Tables 13 and 14, the pathogenic strand HAP2 of the proband directly assembled by the third-generation sequencing method is completely consistent with the pathogenic strand P-HAP2 of the proband in the pedigree linkage analysis by the microarray method.
[0133] Example 5 This embodiment uses family samples 01-03, whose pathogenic variants have been validated by Sanger sequencing, as the test samples to introduce a method for simultaneously detecting single-gene genetic diseases and haplotypes. The method in this embodiment is the same as in Embodiment 3, except that the target region is a region at least 500kb upstream and downstream of the target gene locus.
[0134] The tested samples were from individuals without a family history or proband. Sanger sequencing of family sample 01 confirmed that the male partner was a carrier. HBB :IVS-II-654:c.316-197C>T heterozygous variant, female carries it HBB The :CD41–42:c.126_129delCTTT heterozygous variant was observed, and both males and females underwent third-generation sequencing.
[0135] The Sanger sequencing results of family sample 02 showed that the woman was a carrier. KCNV2 The :c.523G>A heterozygous variant was present; the male did not carry it, there was no proband, and the female underwent next-generation sequencing.
[0136] The Sanger sequencing results of family sample 03 showed that the male was a carrier. PROM1 The c.1090C>T heterozygous variant (autosomal dominant / recessive inheritance) was not carried by the female partner, there was no proband, and the male partner underwent third-generation sequencing.
[0137] Refer to Table 1-3. HBB The 41 pairs of amplification primer sequences for the gene are shown in SEQ ID No. 1-SEQ ID No. 82, covering a region of 680 kb (chr11:4912800-5592907). The amplification primers are non-overlapping. HBB-1F / 1R to HBB-21F / 21R Mixed into a set (Tube01), non-overlapping amplification primers HBB-22F / 22R to HBB-41F / 41R They are mixed into one group (Tube02) to ensure that there is no interference from small fragments and to simplify the amplification process.
[0138] KCNV2 The 33 pairs of amplification primer sequences for the gene are shown in SEQ ID No. 83-SEQ ID No. 148, covering a region of 535.5 kb (chr9:2443372-2978898). The amplification primers are non-overlapping. KCNV2-1F / 1R to KCNV2-17F / 17R Mixed into a set (Tube01), non-overlapping amplification primers KCNV2-18F / 18R to KCNV2-32F / 32R Mixed into one group (Tube02).
[0139] PROM1 The 41 pairs of amplification primer sequences for the gene are shown in SEQ ID No. 149-SEQ ID No. 230, covering a region of 682 kb (chr4:15969851-16085646). These are non-overlapping amplification primers. PROM1-1F / 1R to PROM1-21F / 21R Mixed into a set (Tube01), non-overlapping amplification primers PROM1-22F / 22R to PROM1-41F / 41R Mixed into one group (Tube02).
[0140] Each pair of amplification primers covers a region of 14kb-20kb, and adjacent amplification primers have an overlap region of 100bp-4kb.
[0141] The sequencing data were subjected to haplotype analysis, and the SNP genotyping results are shown in Tables 15-17. There are at least 30 SNP sites available upstream and downstream of the pathogenic site.
[0142] Table 15 Haplotype Analysis Results of Family Sample 01
[0143] Refer to Table 15. Bold text with an underline indicates pathogenic mutations, and - indicates a base deletion. In the table, F-HAP1 and F-HAP2 represent the two haplotypes in the male, and M-HAP1 and M-HAP2 represent the two haplotypes in the female. SNP information corresponds to the negative strand of the reference genome. The coordinates 5247153 in the table correspond to... HBB :IVS-II-654:c.316-197C>T site, coordinates of 5247993 and 5247995 correspond HBBThe table shows the CD41–42:c.126_129delCTTT locus. It can be seen from the table that the male's HAP1 is the haplotype containing the pathogenic mutation, and the female's HAP1 is also the haplotype containing the pathogenic mutation. The results in Table 15 are consistent with the Sanger sequencing verification, indicating that the long fragment amplification + third-generation sequencing method in this embodiment can effectively and accurately obtain SNP locus information for the detection of preimplantation genetic disorders.
[0144] Table 16 Haplotype Analysis Results of Family Sample 02
[0145] Refer to Table 16. Underlined bold text indicates pathogenic mutations. HAP1 and HAP2 represent two haplotypes, and SNP information corresponds to the positive strand of the reference genome. The coordinates 2751726 in the table correspond to... KCNV2 The c.523G>A site is shown in the table. It can be seen from the table that HAP1 is the haplotype containing the pathogenic mutation. The results in Table 16 are consistent with the Sanger sequencing verification, indicating that the long fragment amplification + third-generation sequencing method in this embodiment can effectively and accurately obtain SNP site information for the detection of preimplantation genetic disorders.
[0146] Table 17 Haplotype Analysis Results of Family Sample 03
[0147] Refer to Table 17. Underlined bold text indicates pathogenic mutation sites. HAP1 and HAP2 represent two haplotypes, and SNP information corresponds to the negative strand of the reference genome. The coordinate 16014922 in the table corresponds to... PROM1 The c.1090C>T site is shown in the table. It can be seen from the table that HAP1 is the haplotype containing the pathogenic mutation. The results in Table 17 are consistent with the Sanger sequencing verification, indicating that the long fragment amplification + third-generation sequencing method in this embodiment can effectively and accurately obtain SNP site information for the detection of preimplantation genetic disorders.
[0148] In summary, the method provided by this invention has the following advantages: (1) By designing primers to amplify long fragments, the target gene with a length of 20kb can be amplified. The amplified product can be sequenced in the third generation and can be used for the detection of single-gene genetic diseases and the differentiation of compound heterozygous variants.
[0149] (2) For target genes with a full length > 20kb, primers were designed by overlapping, and amplification was performed by two-stage + dilution PCR. After amplification, the products were mixed in equal amounts and then sequenced for the detection of pathogenic mutations in single genes.
[0150] (3) Based on the amplification of the full length of the target gene, primers are designed upstream and downstream of the target gene locus. PCR amplification is performed by a two-stage + dilution method so that the amplification product covers a region of at least 500kb upstream and downstream of the target locus. SNP haplotype analysis is performed on the sequencing results to effectively and accurately obtain SNP locus information for the detection of preimplantation genetic diseases.
[0151] (4) The method of the present invention can also be used to obtain information on single-gene genetic diseases and / or haplotype information of other target genes.
[0152] (5) Without affecting the accuracy of the test results, the method of the present invention significantly reduces the sequencing cost (compared to the original sequencing cost of 70-100G, which is equivalent to a cost reduction of more than 45 times) and analysis time, and shortens the detection cycle. It can be widely used in multiple fields such as detection of single-gene diseases, preimplantation genetic diagnosis and prenatal diagnosis.
[0153] This invention is not limited to the above-described embodiments. If any modifications or variations to this invention do not depart from the spirit and scope of this invention, and if such modifications and variations fall within the scope of the claims and equivalent technologies of this invention, then this invention also intends to include such modifications and variations.
Claims
1. A method for amplifying target regions suitable for long-read third-generation sequencing, characterized in that, Includes the following steps: S1. Design at least one pair of amplification primers so that the amplification primers can cover the target region; If the length of the target region is less than or equal to 20kb, then design a pair of amplification primers; If the length of the target region is greater than 20kb, multiple pairs of amplification primers are designed. Adjacent amplification primers have overlapping regions. The multiple pairs of amplification primers are divided into two groups, and the coverage regions of each amplification primer in each group do not overlap. S2. The gDNA of the test sample is amplified using the amplification primers from step S1 to obtain the test product. S3. Library construction and third-generation sequencing of the test product from step S2; S4. Perform single-gene genetic disease detection and / or haplotype analysis on the sequencing data to obtain single-gene genetic disease information and / or haplotype information of the target gene.
2. The target region amplification method according to claim 1, characterized in that, When the length of the target region is greater than 20kb, step S2 includes: The two sets of amplification primers from step S1 were used to perform the first stage of long fragment amplification of the gDNA of the test sample to obtain the intermediate amplification product of the target gene. The intermediate amplification product was diluted and used as an amplification template. The amplification template was then subjected to the second stage of long fragment amplification using the two sets of amplification primers from step S1, resulting in two sets of amplification products. The two sets of amplification products were purified and recovered, and then mixed in equal amounts to obtain the test product.
3. The target region amplification method according to claim 2, characterized in that, The conditions for long fragment amplification in the first stage are: pre-denaturation at 94℃ for 5 min, denaturation at 94℃ for 30 s, extension at 68℃ for 12 min, 10-20 cycles, extension at 68℃ for 7 min, and then cooling to 4℃.
4. The target region amplification method according to claim 3, characterized in that, The conditions for long fragment amplification in the second stage are: pre-denaturation at 94℃ for 5 min, denaturation at 94℃ for 30 s, extension at 68℃ for 12 min, 20-30 cycles, extension at 68℃ for 7 min, and then cooling to 4℃.
5. The target region amplification method according to claim 4, characterized in that, In step S2, the intermediate amplification product is diluted by a factor of 40-60.
6. The target region amplification method according to claim 2, characterized in that, When the length of the target region is less than or equal to 20kb, the conditions for long fragment amplification in step S2 are: pre-denaturation at 94℃ for 5 min, denaturation at 94℃ for 30 s, extension at 68℃ for 10 min or 12 min, 30-40 cycles, extension at 68℃ for 7 min, and then cooling to 4℃.
7. The target region amplification method according to claim 6, characterized in that, The extension time is 30-60 s / kb.
8. The target region amplification method according to any one of claims 1-7, characterized in that, When detecting information on single-gene genetic diseases, the target region is the full-length sequence of the target gene; When detecting only haplotypes or simultaneously detecting single-gene genetic diseases and haplotypes, the target region is a region at least 500kb upstream and downstream of the target gene locus.
9. The target region amplification method according to claim 8, characterized in that, Each pair of amplification primers can cover a region of up to 20kb; adjacent amplification primers have an overlap region of 100bp-5kb.
10. The target region amplification method according to claim 8, characterized in that, The amplification primers meet the following requirements: GC content of 40%-60%, primer annealing temperature of 66-68℃, and primer length of 25-30bp.
Citation Information
Cited By
Kidd blood group system genotyping amplification primer group, amplification system, amplification method, library building method and sequencing method based on long-reading long nanopore sequencing
CN122279025A