Human whole exome sequencing probe group and use thereof
By providing a whole-exome sequencing probe set containing multiple probe sets, the problem of inability to detect multiple genomic variants at the same time in the prior art is solved, and higher detection accuracy and lower cost are achieved, and clinical detection needs are met.
Patent Information
- Application Number
- PCT/CN2024/138051
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-12-10
- Publication Date
- 2025-06-19
AI Technical Summary
Existing all-exome sequencing products cannot efficiently detect deep intron variants, mitochondrial variants, small fragment CNVs and large fragment CNVs at the same time, resulting in low detection accuracy and high cost, low efficiency and long time.
A human whole exome sequencing probe set is provided, including a basic whole exome capture probe set, mitochondrial full-length genome probe set, copy number variant (CNV) backbone probe set, deep intron pathogenic variant capture probe set and intragenic pathogenic CNV capture probe set. Through the combination of these probe sets, a wider genomic region is covered and detection accuracy is improved.
It realizes efficient detection of deep introns, mitochondria, genome-wide CNV and small fragments of CNV in genes, improves detection accuracy, reduces cost and detection time, and meets clinical needs.
Smart Images

Figure PCTCN2024138051-FTAPPB-I100001 
Figure PCTCN2024138051-FTAPPB-I100002 
Figure PCTCN2024138051-FTAPPB-I100003
Abstract
Description
A human whole exome sequencing probe set and its application Technical Field
[0001] The present invention belongs to the field of biotechnology. Specifically, the present invention relates to a human whole exome sequencing probe set and its application.
[0002] Background of the Invention
[0003] In recent years, with the rapid advancement of genomics and sequencing technologies, sequencing costs have decreased exponentially, and gene sequencing technology has been widely used in the exploration of genetic causes and prenatal diagnostic screening. The human genome consists of approximately 3 billion base pairs and 22,000 genes. Exons, which encode proteins, account for only 1% of the genome. Exons are only about 30M in size, but contain approximately 85% of pathogenic variants. Whole Exome Sequencing (WES) is a technical method that uses probe hybridization to enrich DNA sequences in exon regions, followed by high-throughput sequencing, primarily to identify and study genetic variants in coding regions associated with disease. WES achieves high detection efficiency at a low cost and is currently the most cost-effective solution for genetic disease diagnosis.
[0004] Currently, there are many whole-exome capture kits on the market, mainly including Agilent's SureSelect, IDT's xGen, and Twist's Human Exome. Agilent's SureSelect Human Whole Exome, represented by V6, captures approximately 60Mb of protein-coding regions from known genes from RefSeq, GENCODE, CCDS, HGMD, and OMIM. IDT's xGen Exome Research Panel v1 covers 34Mb of the human genome. This probe capture reagent reduces exon flanking regions, making capture more efficient and cost-effective. Twist's Human Comprehensive Exome offers excellent coverage, covering 36.8Mb of the human protein-coding region. Its probe design is based on the latest published databases (RefSeq, CCDS, GenCode, Clinvar, ACMG73, etc.). Twist utilizes unique oligonucleotide synthesis technology, achieving optimal sequencing coverage uniformity. Its design reduces invalid data coverage in non-target regions, improving data utilization, and allows for flexible probe addition, making it suitable for personalized customization. All three probes only capture the protein coding region and its flanking sequences, which has certain limitations in clinical use.
[0005] Twist WES (Human Comprehensive Exome, https: / / www.twistbioscience.com / products / ngs) covers 36.8M protein-coding regions and is already a mature commercial product, but it has certain limitations that can be improved through customized probes. The detection area of Twist WES is limited, and some known pathogenic variants are located in deep intronic regions. Twist WES can only detect coding regions and intronic regions adjacent to about 15bp. No probes are laid in deep intronic regions, resulting in a lower positive detection rate. Mitochondria are organelles that provide energy in the human body. They are only 16569bp in size, but contain 37 genes. Diseases caused by mitochondrial mutations are also common genetic metabolic diseases, but Twist WES does not lay probes to capture the mitochondrial genome. Copy number variation (CNV) is the deletion or duplication of genomic fragments caused by genomic rearrangements. Twist WES only lays probes in the exon region and lacks CNV probes covering the entire genome, resulting in insufficient CNV detection accuracy, easy missed detection, and failure to meet clinical needs; some diseases are caused by CNVs within genes, and Twist WES probes have difficulty detecting exon-sized CNVs, which are easily missed.
[0006] Existing whole-exome sequencing products cannot simultaneously solve the problems of detecting deep intronic variations, mitochondrial variations, small-fragment CNVs, and the accuracy of detecting larger-fragment CNVs. Their usage scenarios are greatly limited and must be supplemented by other technical means such as mitochondrial panels, chromosome chip analysis (CMA), Sanger, etc., which are costly, inefficient, and time-consuming.
[0007] Therefore, there is still a need in the art for improved whole exome sequencing probe combinations, corresponding kits, and methods of use.
[0008] Summary of the Invention
[0009] In order to solve the problems existing in the prior art, the present invention provides an improved whole exome sequencing probe combination, a corresponding kit and a method of use.
[0010] The present invention includes but is not limited to the following embodiments:
[0011] Embodiment 1. A human whole exome sequencing probe set comprising:
[0012] 1) Basic whole-exon capture probe set;
[0013] 2) Mitochondrial full-length genome probe set;
[0014] 3) copy number variation (CNV) backbone probe set;
[0015] 4) Deep intronic pathogenic variant capture probe sets; and
[0016] 5) Intragenic pathogenic CNV capture probe set.
[0017] Embodiment 2. The human whole exome sequencing probe set according to embodiment 1, wherein the basic whole exome capture probe set is the Human Comprehensive Exome probe set from Twist Bioscience.
[0018] Embodiment 3. The human whole exome sequencing probe set according to embodiment 1 or 2, wherein the mitochondrial full-length genome probe set is the Mitochondrial Panel from Twist Bioscience.
[0019] Embodiment 4. The human whole exome sequencing probe set according to any one of embodiments 1-3, wherein the CNV backbone probe set is a probe set obtained based on one probe per 100 kb of the human whole genome.
[0020] Embodiment 5. The human whole-exome sequencing probe set of embodiment 4, wherein the CNV backbone probe set is the RnD_backbone_100K_optPLnF probe set from Twist Bioscience (Catalog #106984; Twist Custom Panel, ID: TE-96905785, RnD_backbone_100K_optPLnF).
[0021] Embodiment 6. The human whole exome sequencing probe set of any one of embodiments 1-5, wherein the deep intronic pathogenic variant capture probe set is obtained by the following method:
[0022] a) Screening variants with a pathogenicity rating of two stars or higher and a P or LP in the Clinvar (v202106) database, removing variants located in the probes of the comprehensive whole-exon capture probe set, and obtaining pathogenic variants in deep intronic regions of Clinvar (v202106);
[0023] b) screening variants recorded as DM and DM? in the HGMD (v202106) database, removing variants located in probes of the comprehensive whole-exon capture probe set, and obtaining pathogenic variants in deep intronic regions in HGMD (v202106); and
[0024] c) Design probes based on the positional information of pathogenic variants in deep intronic regions obtained in a) and b)
[0025] Embodiment 7. A human whole exome sequencing probe set according to embodiment 6, wherein the deep intronic pathogenic variant capture probe set comprises the probes shown in Table 19.
[0026] Embodiment 8. A human whole exome sequencing probe set according to any one of embodiments 1-7, wherein the intragenic pathogenic CNV capture probe set comprises probes for the following 60 genes: ABCD1, ACVRL1, ENG, ANO5, ATP7A, ATP7B, BTK, CAPN3, COL4A5, CREBBP, DMD, F8, F9, FANCA, FBN1, FH, UBE3A, GBA, GLDC, KCNQ1, KCNH2, KCNQ2, LDLR, LMNA, MEN 1. MYBPC3, NF1, NPHP1, PAH, PCCA, PCCB, PKD1, PKD2, PLP1, PRRT2, SCN1A, SCN5A, SLC22A5, SOX9, TH, TSC1, TSC2, BRCA1, BRCA2, CHD7, EXT1, EXT2, MECP2, RB1, SHANK3, GCK, HNF1A, HNF1B, HNF4A, ARID1A, ARID1B, EHMT1, USH2A, ADGRV1 and PMP2.
[0027] Embodiment 9. The human whole exome sequencing probe set according to embodiment 8, wherein a probe is designed every 1 kb within the 60 genes.
[0028] Embodiment 10. The human whole exome sequencing probe set according to embodiment 8, wherein the intragenic pathogenic CNV capture probe set comprises the probes shown in Table 20.
[0029] Embodiment 11. A method for performing whole exome sequencing on a human sample, the method comprising:
[0030] i) extracting genomic DNA from the human sample;
[0031] ii) fragmenting the genomic DNA, performing end-repair on the obtained DNA fragments, and adding A nucleotides to the ends of the DNA fragments;
[0032] iii) adding a universal sequencing adapter to the DNA fragments obtained in step ii);
[0033] iv) amplifying the DNA fragments obtained in step iii) to obtain a genomic DNA sequencing prelibrary of the human sample;
[0034] v) mixing the genomic DNA sequencing prelibrary with the probes of the probe set according to any one of embodiments 1 to 10, so that the probes capture corresponding target DNA fragments by hybridization and remove non-target DNA fragments that are not captured;
[0035] vi) amplifying the captured target DNA fragments, thereby obtaining a captured DNA library; and
[0036] vii) performing high-throughput sequencing on the captured DNA library.
[0037] Embodiment 12. The method according to embodiment 11, wherein steps ii) to iv) are performed using the Library Preparation Enzymatic Fragmentation (EF) Kit 2.0 and the Universal Adapter System from Twist Bioscience.
[0038] Embodiment 13. The method according to embodiment 12, wherein the enzymatic fragmentation reaction of step ii) is carried out at about 37°C for about 10 minutes.
[0039] Embodiment 14. The method according to embodiment 12 or 13, wherein step ii) is performed in a PCR instrument with a heated lid temperature of about 70°C.
[0040] Embodiment 15. A method according to any one of embodiments 12-14, wherein the ratio of the volume of 1) basic whole exon capture probe group, 2) mitochondrial full-length genome probe group, 3) copy number variation (CNV) backbone probe group, 4) deep intron pathogenic variation capture probe group and 5) intragenic pathogenic CNV capture probe group in the human whole exome sequencing probe group is 2μl:8μl:4μl:4μl:4μl; and the concentration of each probe is as follows: 206.80nM, 0.07nM, 14.92nM, 0.93nM, 4.00nM, respectively.
[0041] Embodiment 16. The method according to any one of embodiments 11 to 15, wherein steps v) to vi) are performed using the Fast Hybridization and Wash Kit from Twist Bioscience.
[0042] Embodiment 17. A method according to any one of embodiments 11-16, wherein the high-throughput sequencing is performed using BGI's second-generation sequencing platform.
[0043] Embodiment 18. A method according to any one of embodiments 11-17, wherein the human sample is selected from amniotic fluid, tissue (such as abortion tissue, skin tissue), cell line, blood (such as peripheral blood, dried blood spots, cord blood).
[0044] Embodiment 19. A kit for performing whole-exome sequencing on a human sample, the kit comprising the human whole-exome sequencing probe set according to any one of embodiments 1-10.
[0045] Embodiment 20. A kit according to embodiment 19, for performing whole exome sequencing on a human sample by the method according to any one of embodiments 11-18.
[0046] BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 shows the sample library construction flow chart.
[0048] Figure 2 shows the bioinformatics analysis flow chart.
[0049] Figure 3 shows the detection of large CNV fragments in 15 samples.
[0050] Figure 4 shows the CNV detection results of 10 samples.
[0051] Figure 5 shows IGV maps of two reported non-coding region variants from parallel experiments with three reagents.
[0052] Figure 6 shows that Omniseek has good coverage of the mitochondrial genome and has good detection ability for mitochondrial variants.
[0053] Detailed Description of the Invention
[0054] In one aspect, the present invention provides a human whole exome sequencing probe set comprising:
[0055] 1) Basic whole-exon capture probe set;
[0056] 2) Mitochondrial full-length genome probe set;
[0057] 3) copy number variation (CNV) backbone probe set;
[0058] 4) Deep intronic pathogenic variant capture probe sets; and
[0059] 5) Intragenic pathogenic CNV capture probe set.
[0060] In some specific embodiments, the basic whole exon capture probe set is the Human Comprehensive Exome probe set from Twist Bioscience, with product number 102033, for example.
[0061] In some specific embodiments, the mitochondrial full-length genome probe set is the Mitochondrial Panel from Twist Bioscience, with a product number of, for example, 102040.
[0062] The "Copy Number Variation (CNV) Backbone Probe Set" described herein is a probe set for genome-wide CNV detection, designed to overcome the low CNV detection accuracy of existing techniques that only use probes in exon regions. The Copy Number Variation (CNV) Backbone Probe Set includes probes spaced at intervals across the genome.
[0063] In some specific embodiments, the CNV backbone probe set is a probe set obtained based on approximately one probe per 100 kb of the entire human genome.
[0064] In some embodiments, the CNV backbone probe set is the RnD_backbone_100K_optPLnF probe set from Twist Bioscience (Catalog #106984 Twist Custom Panel, ID: TE-96905785, RnD_backbone_100K_optPLnF).
[0065] Existing basic whole-exon capture probe sets can only detect coding regions and the adjacent intronic regions of approximately 15 bp, but cannot detect deep intronic regions. To overcome this limitation of the existing technology, the inventors screened for pathogenic variants located in deep intronic regions of the genome and set probes based on these variant locations, thus creating a deep intronic pathogenic variant capture probe set.
[0066] In some embodiments, the deep intronic pathogenic variant capture probe set is obtained by:
[0067] a) Screening variants with a pathogenicity rating of two stars or higher and a P or LP in the Clinvar (v202106) database, removing variants located in the probes of the comprehensive whole-exon capture probe set, and obtaining pathogenic variants in deep intronic regions of Clinvar (v202106);
[0068] b) screening variants recorded as DM and DM? in the HGMD (v202106) database, removing variants located in probes of the comprehensive whole-exon capture probe set, and obtaining pathogenic variants in deep intronic regions in HGMD (v202106); and
[0069] c) Design probes based on the positional information of pathogenic variants in deep intronic regions obtained in a) and b)
[0070] In some embodiments, the deep intronic pathogenic variant capture probe set comprises the probes listed in Table 19.
[0071] Existing basic whole-exome capture probe sets typically only detect large CNVs, potentially missing CNVs within genes or at the exon level. To address this, the inventors creatively screened 60 genes harboring intragenic pathogenic CNVs and developed encrypted probes targeting these genes to achieve higher detection accuracy. This results in the present intragenic pathogenic CNV capture probe set.
[0072] In some embodiments, the intragenic pathogenic CNV capture probe set comprises probes for the following 60 genes: ABCD1, ACVRL1, ENG, ANO5, ATP7A, ATP7B, BTK, CAPN3, COL4A5, CREBBP, DMD, F8, F9, FANCA, FBN1, FH, UBE3A, GBA, GLDC, KCNQ1, KCNH2, KCNQ2, LDLR, LMNA, MEN1, MYBPC3, NF1, N PHP1, PAH, PCCA, PCCB, PKD1, PKD2, PLP1, PRRT2, SCN1A, SCN5A, SLC22A5, SOX9, TH, TSC1, TSC2, BRCA1, BRCA2, C HD7, EXT1, EXT2, MECP2, RB1, SHANK3, GCK, HNF1A, HNF1B, HNF4A, ARID1A, ARID1B, EHMT1, USH2A, ADGRV1 and PMP2.
[0073] In some embodiments, wherein a probe is designed approximately every 1 kb within the 60 genes. In some specific embodiments, the intragenic pathogenic CNV capture probe set comprises the probes shown in Table 20.
[0074] In another aspect, the present invention provides a method for performing whole exome sequencing on a human sample, the method comprising:
[0075] i) extracting genomic DNA from the human sample;
[0076] ii) fragmenting the genomic DNA, performing end-repair on the obtained DNA fragments, and adding A nucleotides to the ends of the DNA fragments;
[0077] iii) adding a universal sequencing adapter to the DNA fragments obtained in step ii);
[0078] iv) amplifying the DNA fragments obtained in step iii) to obtain a genomic DNA sequencing prelibrary of the human sample;
[0079] v) mixing the genomic DNA sequencing prelibrary with the probes of the probe set according to any one of embodiments 1 to 10, so that the probes capture corresponding target DNA fragments by hybridization and remove non-target DNA fragments that are not captured;
[0080] vi) amplifying the captured target DNA fragments, thereby obtaining a captured DNA library; and
[0081] vii) performing high-throughput sequencing on the captured DNA library.
[0082] In some embodiments, the fragmentation of step ii) is enzymatic fragmentation.
[0083] In some embodiments, steps ii) to iv) are performed using Library Preparation Enzymatic Fragmentation (EF) Kit 2.0 (Product No. 104207) and Universal Adapter System (Product No. 101308) from Twist Bioscience.
[0084] When using commercial kits or systems, the steps of the method of the present invention can generally be performed according to the manufacturer's instructions. However, the present inventors surprisingly found that when a heated lid temperature of about 70°C is used in a thermal cycler and the enzyme fragmentation is performed at about 37°C for about 10 minutes, significantly better results can be obtained than the recommended heated lid temperature of 105°C.
[0085] Therefore, in some preferred embodiments, the enzymatic fragmentation reaction of step ii) is carried out at about 37°C for about 10 minutes. In some more preferred embodiments, step ii) is carried out in a thermal cycler (PCR instrument) with a heated lid temperature of about 70°C.
[0086] The ratio between different subgroups in the probe group will also significantly affect the final test results. In some preferred embodiments, the ratio of 1) basic whole exon capture probe group, 2) mitochondrial full-length genome probe group, 3) copy number variation (CNV) backbone probe group, 4) deep intron pathogenic variation capture probe group and 5) intragenic pathogenic CNV capture probe group in the human whole exome sequencing probe group is 2μl: 8μl: 4μl: 4μl: 4μl. In some embodiments, the concentrations of each probe are as follows: 206.80nM, 0.07nM, 14.92nM, 0.93nM, 4.00nM.
[0087] In some embodiments, steps v) to vi) are performed using the Fast Hybridization and Wash Kit (Product No. 104181) from Twist Bioscience.
[0088] In some specific embodiments, the method of the present invention is performed as described in Example 2 of the present application.
[0089] In some embodiments, the high-throughput sequencing is performed using BGI's next-generation sequencing platform.
[0090] In some embodiments, the method of the present invention further comprises analyzing the obtained sequencing data by bioinformatics. For example, the sequencing data can be analyzed by the Genetic Disease High-Throughput Sequencing Data SNV+Indel Analysis System (Software Copyright Registration Number: 2018SR845218) and the Genetic Disease High-Throughput Sequencing Data CNV Analysis System (Software Copyright Registration Number: 2018SR847212). In some embodiments, the sequencing results can be analyzed by the process shown in Figure 2.
[0091] In another aspect, the present invention provides use of the human whole exome sequencing probe set of the present invention in preparing a kit for performing whole exome sequencing on a human sample.
[0092] In another aspect, the present invention provides a kit for performing whole-exome sequencing on a human sample, wherein the kit comprises the human whole-exome sequencing probe set of the present invention.
[0093] In some embodiments, the kit is used to perform whole exome sequencing on a human sample using the method described herein.
[0094] In some embodiments of the various aspects of the invention, the human sample is selected from amniotic fluid, tissue (eg, abortion tissue, skin tissue), cell line, blood (eg, peripheral blood, dried blood spots, cord blood).
[0095] Advantages of the present invention include:
[0096] 1. This invention adds probes for deep intronic pathogenic regions, which can detect known pathogenic variants in deep introns;
[0097] 2. The present invention adds a probe for the mitochondrial genome, which can detect variations in the mitochondrial genome;
[0098] 3. This invention adds CNV probes covering the entire genome, which increases the accuracy of CNV detection and meets clinical needs;
[0099] 4. This invention encrypts probes for 60 genes and can detect small CNVs within genes;
[0100] 5. The present invention optimizes the wet experiment process and significantly improves the quality of library construction and sequencing by optimizing the hot cover temperature and probe ratio. Example
[0101] The following examples are only for better illustrating the present invention and are not intended to limit the scope of the present invention.
[0102] Example 1. Design of human whole exome sequencing probe set
[0103] To address the problem of missing deep intron pathogenic regions, the present invention adds probes that capture pathogenic variants in deep intron regions based on the position information of deep intron variants known in relevant databases.
[0104] To address the problem of missing mitochondrial genome variations, the present invention adds a full-length mitochondrial genome probe.
[0105] To address the problem of lack of CNV probes covering the entire genome, the present invention lays a probe at intervals of a certain length throughout the genome to improve CNV detection accuracy.
[0106] To address the problem of being unable to detect CNVs within genes, the present invention screens out the main genes where CNVs within genes cause diseases based on the database, and additionally lays probes on these genes to improve the CNV detection capability, enabling it to detect CNVs at the exon level.
[0107] Based on Twist WES, deep intron and mitochondrial genome probes are added, as well as whole-genome CNV backbone probes. Probes for some genes are encrypted, which can expand the full-range detection range and improve the CNV detection capability, making it more in line with actual clinical detection needs, reducing costs and saving detection time.
[0108] The specific probe design is as follows:
[0109] 1. Screen variants with a pathogenicity rating of two stars or above and a P or LP in the Clinvar (v202106) database, remove variants located in the Twist WES probe, and the remaining variants are pathogenic variants in the deep intronic regions of Clinvar (v202106).
[0110] 2. Variants recorded as DM and DM? in the HGMD (v202106) database were screened, and variants located in the Twist WES probe were removed. The remaining variants were pathogenic variants in the deep intronic regions of HGMD (v202106).
[0111] 3. Design probes based on the location information of pathogenic variants in deep intronic regions in Clinvar (v202106) and HGMD (v202106).
[0112] 4. Design probes based on the mitochondrial genome location information on NCBI (https: / / www.ncbi.nlm.nih.gov / nuccore / NC_012920.1).
[0113] 5. Lay one probe every 100 kb across the entire human genome to construct a CNV backbone probe.
[0114] 6. Based on experience, the inventors manually selected 60 key genes with intragenic pathogenic CNVs from the database, namely ABCD1, ACVRL1, ENG, ANO5, ATP7A, ATP7B, BTK, CAPN3, COL4A5, CREBBP, DMD, F8, F9, FANCA, FBN1, FH, UBE3A, GBA, GLDC, KCNQ1, KCNH2, KCNQ2, LDLR, LMNA, MEN1, MYBPC3, NF1 , NPHP1, PAH, PCCA, PCCB, PKD1, PKD2, PLP1, PRRT2, SCN1A, SCN5A, SLC22A5, SOX9, TH, TSC1, TSC2, BRCA1, BRCA2, CHD7, EXT1, EXT2, MECP2, RB1, SHANK3, GCK, HNF1A, HNF1B, HNF4A, ARID1A, ARID1B, EHMT1, USH2A, ADGRV1, PMP2. Within the 60 genes, the present invention encrypted the 60 genes by laying one probe per 1 kb.
[0115] Specifically, the composition of the probe set (OmniSeek) of the present invention is shown in Table 1 below:
[0116] Table 1. OmniSeek probe sets
[0117] Note:
[0118] 1) TE-98723561 (AddSNV), this part of the probe targets currently known pathogenic intronic variants, enabling the detection of such pathogenic intronic variants that may occur in diseases, and is a customized probe for the present invention;
[0119] 2) TE-93715706 (AddCNV), a probe designed for microdeletion / microduplication variants in 60 genes of key concern identified by the inventors, to detect microdeletion / microduplication variants in these 60 genes that may be associated with disease. This is a custom probe for the present invention.
[0120] 3) TE-96905785 (RnD_backbone_100K_optPLnF), this part is a probe already available from Twist, designed for copy number variations (CNVs) that may occur within the genome, namely the CNV backbone probe in this article.
[0121] Example 2, wet test technical process
[0122] The specific experimental process is shown in Figure 1 and is described in detail as follows:
[0123] 1. Concentration determination:
[0124] use dsDNA was quantitatively detected and analyzed using a fluorescence instrument.
[0125] 2. DNA fragmentation, end repair and A addition
[0126] 2.1 Thaw the Frag / AT Enzymes, Frag / AT Buffer, and gDNA sample on ice in advance, then flick the tube to mix.
[0127] 2.2 Take 100 ng DNA and dilute to 40 μl with nuclease-free water;
[0128] 2.3 Start the PCR instrument and set the PCR program according to the table below (Table 2) (heated cover 70℃), and pre-cool in advance;
[0129] Table 2. DNA fragmentation end filling plus A-PCR conditions
[0130] 2.4 Prepare the reaction solution according to the table below (Table 3) (operate on ice), gently pipette and mix the reagent system, place it in the pre-cooled PCR instrument, and start the program;
[0131] Table 3. Volume of DNA fragment end filling plus A-reagent
[0132] 3. Connector ligation and purification
[0133] After the 3.1 procedure is complete, add 5 μl of Universal Adapters to each well, gently pipette until completely mixed, and place on ice;
[0134] 3.2 Add 20 μl of Ligation Master Mix by inversion, gently pipette to mix, and incubate in a PCR machine at 20°C for 15 min. Do not use the heated cover mode.
[0135] 3.3 Magnetic bead method to purify samples
[0136] 3.3.1 Remove DNA Purification Beads from the 4°C refrigerator and place at room temperature for at least 30 minutes.
[0137] 3.3.2 Add 60 μl (0.8x) of mixed magnetic beads to the sample tube, vortex, and incubate at room temperature for 5 minutes.
[0138] 3.3.3 Place the reaction tube on the magnetic rack and let it stand for 3-5 minutes. When the liquid is clear, discard the supernatant, taking care not to touch the magnetic beads.
[0139] 3.3.4 Add 200 μl of freshly prepared 80% ethanol to each tube without breaking up the magnetic beads. Discard the supernatant after 1 minute.
[0140] Repeat step 3.3.5 once, washing the beads twice in total;
[0141] 3.3.6 After a brief centrifugation, remove any remaining liquid with a 10 μl pipette tip. Dry at room temperature for 3–5 minutes, taking care not to overdry the beads.
[0142] 3.3.7 Add 17 μl of nuclease-free water to each sample, vortex, and incubate at room temperature for 2 min.
[0143] 3.3.8 After a brief centrifugation, place the sample tube back on the magnetic rack and let it sit for 2-3 minutes until the liquid becomes clear.
[0144] 3.3.9 Pipette 15 μl of liquid into a new PCR tube and discard the magnetic beads.
[0145] 4. PCR amplification, purification and QC
[0146] 4.1 Set up the PCR reaction program according to the table below (Table 4):
[0147] Table 4. Amplification of "Adapter-Ligated" DNA Fragments - PCR Conditions
[0148] 4.2 Add 10 μl of UDIPrimers to each gDNA library and mix gently by pipetting;
[0149] 4.3 Add 25 μl of Equinox Library Amp Mix (2x) to the library and mix gently by pipetting;
[0150] 4.4 PCR instrument performs amplification process;
[0151] 4.5 Use 50μl (1x) magnetic beads to purify the amplified sample. For specific steps, refer to "3.3 Purification of samples by magnetic beads", add 22μl
[0152] Elute DNA fragments with nuclease-free water.
[0153] 5. Quality identification and quantification of pre-library after PCR amplification
[0154] 5.1 Use an Agilent Bioanalyzer 2200 (TapeStation) instrument to test the quality and concentration of the pre-library. The peak shape must be normal. The average fragment length should be 300-400 bp.
[0155] 5.2 Use The pre-library was detected by fluorescence instrument (HS kit), and the concentration should be ≥50ng / μl.
[0156] 6. Library Hybridization
[0157] 6.1 For each group of 8 pre-libraries, take 200 ng of each and place them into a new 1.5 ml centrifuge tube. Mix gently by pipetting. Add the following reagents to the mixed library according to the table below (Table 5, Table 6);
[0158] Table 5. Probe Mixture-Reagent Volumes
[0159] Table 6. Probe mixing ratio
[0160] 6.2 After gently mixing the premixed library, seal the hole at the tube with sealing film and evaporate the sample to dryness using a vacuum concentrator (<30°C).
[0161] 6.3 Preheat the Hybridization Mix at 65°C for 10 minutes, resuspend the dried premixed library with 20 μl of Hybridization Mix, and equilibrate at room temperature for 5 minutes.
[0162] 6.4 Pre-set the PCR sequence according to the table below (Table 7) and start preheating (heat cover 85°C);
[0163] Table 7. Hybridization capture reaction-PCR conditions
[0164] 6.5 Centrifuge the hybridization mixture to ensure that no bubbles remain and transfer it to a 0.2 ml tube;
[0165] 6.6 Add 30 μl of Hybridization Enhancer on top of the hybridization solution;
[0166] 6.7 Transfer the cross-mix solution to the PCR instrument and start the program.
[0167] 7. Target Sequence Enrichment
[0168] 7.1 Place Fast Wash Buffer 1 in a 68°C metal bath (approximately 450 μl per sample) and Wash Buffer 2 in a 48°C metal bath (approximately 700 μl per sample) and set aside.
[0169] 7.2 Remove the Binding Beads in advance and equilibrate them at room temperature for 30 minutes, then vortex mix.
[0170] 7.3 For each reaction, pipette 100 μl of Binding Beads into a 1.5 ml LoBind tube, add 200 μl of Binding buffer, vortex to mix, place on a magnetic rack until clear, and discard the supernatant.
[0171] Repeat 7.4 twice, washing the Binding Beads three times in total, and then resuspend the magnetic beads in 200 μl of Binding Buffer;
[0172] 7.5 After incubating the hybridization mixture at 60°C for approximately 4-17 hours, aspirate the hybridization solution directly from the PCR instrument and add it to the magnetic bead solution, gently pipetting to mix thoroughly.
[0173] 7.6 Place the tube on a rotator and incubate at room temperature for 30 minutes to ensure that the magnetic beads bind to the target sequence;
[0174] 7.7 After removing the magnetic bead mixing tube from the rotator, centrifuge briefly and place it on a magnetic rack for 3-5 minutes. After the liquid is clear, remove the supernatant.
[0175] 7.8 Wash the magnetic beads twice with Fast Wash buffer 1:
[0176] 7.8.1 Add 200 μl of Fast Wash buffer 1 preheated at 68°C, mix thoroughly by gentle pipetting, and centrifuge briefly.
[0177] 7.8.2 Incubate in a 68°C metal bath for 5 minutes, centrifuge briefly, place the tube on a magnetic rack, and discard the supernatant, avoiding contact with the magnetic beads.
[0178] 7.9 Transfer the second incubation with Wash buffer 1 to a new 1.5ml centrifuge tube, place it on a magnetic rack for 1 minute, and remove the supernatant.
[0179] 7.10 Wash the magnetic beads three times with Wash buffer 2:
[0180] 7.10.1 Add 200 μl of 48°C preheated Wash buffer 2, mix thoroughly by gentle pipetting, and centrifuge briefly.
[0181] 7.10.2 Incubate in a 48°C metal bath for 5 minutes, centrifuge briefly, place the tube on a magnetic rack, and discard the supernatant, avoiding contact with the magnetic beads.
[0182] 7.11 Add 45 μl of nuclease-free water to resuspend the magnetic beads. Store at -20°C.
[0183] 8. Target Sequence Amplification
[0184] 8.1 Prepare the reaction solution according to the following table (Table 8) (operate on ice):
[0185] Table 8. Capture Sequence PCR Amplification - Reagent Volumes
[0186] 8.2 Perform amplification (7 cycles) according to the PCR conditions in the table below (Table 9):
[0187] Table 9. Capture sequence PCR amplification - PCR conditions
[0188] 8.3 Purify the amplified sample with 90 μl (1.8x) magnetic beads. For specific steps, refer to "3.3 Purification of Samples by Magnetic Beads" and add 32 μl nuclease-free water to elute the library.
[0189] 9. Quality Assay and Quantification of Hybrid Capture Libraries
[0190] 9.1 Use an Agilent Bioanalyzer 2200 (Tape Station) to test sample quality and concentration. Peaks must be normal. The average fragment length should be 310–425 bp.
[0191] 9.2 Use The library was detected by fluorescence instrument (HS kit). The detection concentration should be ≥10 ng / μl and the total library volume should be ≥300 ng.
[0192] The library was sequenced by the second-generation sequencing platform of BGI.
[0193] Example 3, bioinformatics analysis process
[0194] The software copyrights for the Genetic Disease High-Throughput Sequencing Data SNV+Indel Analysis System (Software Copyright Registration Number: 2018SR845218) and the Genetic Disease High-Throughput Sequencing Data CNV Analysis System (Software Copyright Registration Number: 2018SR847212) used in this invention have been obtained. The specific bioinformatics analysis process is as follows (Figure 2):
[0195] 1. The downloaded fastq data were processed by fastp (v0.23.1), mainly to remove the adapter sequence and filter the low-quality sequence;
[0196] 2. The processed fastq data were aligned to the reference genome using the alignment software bwa (v0.7.17). The aligned sam files were then converted to bam files and sorted using GATK4 (v4.1.9.0). The bam files were then marked for duplication and base quality correction using GATK4 (v4.1.9.0) to obtain the processed bam files.
[0197] 3. Use GATK4 (v4.1.9.0) to detect single nucleotide variants (SNVs) and small insertion-deletions (InDels) in the processed bam files to obtain vcf files, and then perform variant filtering and annotation;
[0198] 4. Use CNVKIT (v0.9.7), XHMM (v1.0) and DECoN to detect CNVs on the processed bam files, and then perform variation filtering and annotation.
[0199] Example 4: Performance Verification
[0200] 1. Verification of standard products
[0201] NA12878 is a sequencing standard commonly used to verify sequencing performance. After sequencing and analyzing NA12878 using the probe set and method described above (also referred to herein as OmniSeek), the variant calls were statistically analyzed. The true variant set for NA12878 is available from NCBI (https: / / www.ncbi.nlm.nih.gov / ).
[0202] TP: variants included in the true set that were also detected in the test sample NA12878; FN: variants included in the true set but not detected in the test sample NA12878; FP: variants not included in the true set but detected in the test sample NA12878.
[0203] According to the formula Recall = TP / (TP+FN), the calculated value is 98.91±0.09%; according to the formula Precision = TP / (TP+FP), the calculated value is 98.17±0.09%; according to the formula F1-Score = 2*(Recall*Precision) / (Recall+Precision), the calculated value is 98.54±0.07%.
[0204] The values of Recall, Precision and F1-Score are all greater than 98%, indicating that the present invention is reliable in detecting variations.
[0205] 2. Clinical performance verification
[0206] 2.1 SNV / Indel Detection Capabilities
[0207] A total of 73 Sanger-verified SNV / Indel samples were selected, and the comparison results showed a positive coincidence rate of 98.63%. One Indel was not detected, and it was located in a low-complexity region dominated by CT bases, suggesting that OmniSeek has a good ability to detect SNV / Indels in non-special regions.
[0208] 2.2 CNV detection capability
[0209] 2.2.1 Large CNV Detection
[0210] Fifteen samples (including amniotic fluid, abortion tissue, and peripheral blood) that had been previously analyzed using microarray (affy 750K), CNVseq, or WGS were selected. A total of 17 CNVs (13 duplications, 4 deletions; CNV coverage ≥ 2 genes, with one chromosomal aneuploidy and the remaining CNVs ranging in size from 160 kb to 1 Mb) were selected from this batch of samples. CNV analysis was performed in parallel using the three reagents, the same algorithm, and the same parameters. The results are shown in Figure 3 and the table below.
[0211] Table 10
[0212] The results of this batch suggest that compared with the other two reagents of the same type on the market (reagent 1 (IDT V1) and reagent 2 (Agilent V6)), OmniSeek's detection rate for large-segment CNV (the number of genes covered by CNV ≥ 2) can be as high as 100%, which can significantly reduce the missed detection rate of large-segment CNV (the number of genes covered by CNV ≥ 2) and has better detection effect.
[0213] 2.2.2 Exon-level CNV verification
[0214] Ten samples (amniotic fluid, abortion tissue, and peripheral blood) that had been previously analyzed using microarray (affy 750K), CNVseq, or WGS were selected. Ten CNVs (six duplications and four deletions; CNV sizes ranged from 4kb to 500kb, all involving only partial exons of a gene) that had been validated by other methods were selected from this batch of samples. CNV analysis was performed in parallel using the three reagents, the same algorithm, and the same parameters. The results are shown in Figure 4 and the table below:
[0215] Table 11
[0216] The results of this batch suggest that compared with the other two reagents of the same type on the market (reagent one (IDT V1) and reagent two (Agilent V6)), OmniSeek's detection rates for CNVs with less than 7 exons and less than 5 exons can be as high as 80.00% and 71.43% respectively (both uncalled samples were single exon heterozygous deletions), which can significantly reduce the exon-level CNV missed detection rate and has better detection effect.
[0217] 2.3 Detection capability of deep intronic variants
[0218] The detection capabilities of intronic regions were tested in parallel using this capture probe combination (OmniSeek) and two major exon capture kits on the market: reagent 1 (IDT V1) and reagent 2 (Agilent V6). The comparative results showed that OmniSeek significantly outperformed the other two reagents in terms of coverage and sequencing depth in deep intronic regions, indicating that OmniSeek can supplement the capture blind spots of conventional WES reagents.
[0219] Table 12
[0220] The results of the parallel experiments with three reagents on the IGV maps of two reported non-coding region variants are shown in Figure 5.
[0221] 2.4 Detection capability of small deletions and duplications within key genes
[0222] Small deletions and duplications within genes are common genetic causes of gene-related diseases, but conventional exon capture kits have extremely limited ability to detect CNVs at the exon level. Therefore, this capture probe combination (OmniSeek) has encrypted probes for 60 "high-risk" genes and tested them in parallel with two major exon capture kits on the market. Statistics show that the average sequencing depth and 20X coverage of the 60 genes in OmniSeek are significantly higher than those of reagents 1 (IDT V1) and 2 (Agilent V6), effectively increasing the ability to detect CNVs in these genes.
[0223] Table 13. Average sequencing depth and 20X coverage of 60 genes using different sequencing reagents
[0224] Table 14. Specific sequencing results of 60 genes in OmniSeek:
[0225] 2.5 Detection capability of mitochondrial variants
[0226] The capture probe combination (OmniSeek) was used to test peripheral blood DNA, abortion product DNA, and amniotic fluid DNA, respectively. The data volume was normalized to 10G. The analysis results (Figure 6) showed that Omniseek had good coverage of the mitochondrial genome and had good detection capabilities for mitochondrial mutations.
[0227] Example 5: Experimental process optimization
[0228] 3.1 Enzyme digestion hot cover temperature and enzyme digestion time
[0229] Different enzyme digestion times and different hot cover temperatures were tested. The results shown in the table below show that when the enzyme digestion time was 10 minutes and the hot cover temperature was 70°C, the pre-library fragment size met the expected requirements. The experiment was repeated under these conditions, and the pre-library fragment size was within the expected range. Therefore, the experiment was carried out under these conditions.
[0230] Table 15. Test of different enzyme digestion times and different hot cover temperatures
[0231] 3.2 Probe mixing ratio
[0232] The probes were tested in different mixing ratios, and the specific test ratios and performances are as follows: the concentrations of the comprehensive full-exo, mitochondrial, backbone, AddSNV, and AddCNV probes are as follows: 206.80 nM, 0.07 nM, 14.92 nM, 0.93 nM, and 4.00 nM.
[0233] Table 16. Probe ratio test
[0234] Table 17. Average sequencing depth of each capture probe at different mixing ratios
[0235] Table 18. 20X coverage of each capture probe at different mixing ratios:
[0236] Taking the comprehensive full-exoprobe input as the benchmark, the first and second batches of data showed that when the comprehensive full-exoprobe input was 4 μl, the input ratio of mitochondrial and customized capture probes (AddSNV, AddCNV, backbone) was adjusted, and there was no significant difference in the average sequencing depth and 20X coverage of each probe component, suggesting that the probe input may be excessive.
[0237] The third batch of tests used a reduced probe dosage, but maintained the same probe ratio as the second batch. Test data showed no significant difference in 20X coverage across all probe components compared to the second batch, although the average sequencing depth for each probe component decreased. With the exception of mitochondria, the average sequencing depth for all components remained above the minimum quality control threshold, meeting the requirements for subsequent clinical data analysis.
[0238] The fourth to sixth batches of tests optimized the mitochondrial probe input based on the third batch of tests, kept the input of other components unchanged, and increased the mitochondrial probe input. The test data showed that, except for mitochondria, the 20X coverage and average sequencing depth of each probe component did not show significant differences compared with the third batch of tests, but the average sequencing depth of mitochondria was significantly improved. In addition, the three batches of data used the same probe input, and the sequencing quality between batches was relatively stable. Therefore, subsequent experiments were carried out with this probe input ratio, and the customized capture probes (AddSNV, AddCNV, backbone) were combined into one tube.
[0239] Part of the probes / probe sets involved in the present invention
[0240] Table 19. Deep Intronic Pathogenic Variant Probe Sets
[0241] Table 20. Intragenic Pathogenic CNV Capture Probe Sets Note: In Tables 19 and 20, probes are displayed in a similar format as chr15:48786674-48786794: chr15 represents chromosome 15, 48786674 is the start nucleotide, and 48786794 is the end nucleotide. chr15:48786674-48786794 indicates that the probe sequence covers the 120 nucleotides from nucleotide 48786674 to nucleotide 48786794 on chromosome 15. Reference is made to the human genome sequence GRCh37 / hg19, published on February 27, 2009.
Claims
1. A human whole exome sequencing probe set, comprising: 1) Basic whole exon capture probe set; 2) Mitochondrial full-length genome probe set; 3) Copy number variation (CNV) backbone probe set; 4) Deep intronic pathogenic variant capture probe sets; and 5) Intragenic pathogenic CNV capture probe set.
2. The human whole exome sequencing probe set according to claim 1, wherein the basic whole exome capture probe set is the Human Comprehensive Exome probe set from Twist Bioscience. 3 . The human whole exome sequencing probe set according to claim 1 , wherein the mitochondrial full-length genome probe set is the Mitochondrial Panel from Twist Bioscience.
4. The human whole exome sequencing probe set according to any one of claims 1 to 3, wherein the CNV backbone probe set is a probe set obtained by designing one probe per 100 kb of the human whole genome. 5 . The human whole exome sequencing probe set according to claim 4 , wherein the CNV backbone probe set is the RnD_backbone_100K_optPLnF probe set from Twist Bioscience.
6. The human whole exome sequencing probe set according to any one of claims 1 to 5, wherein the deep intronic pathogenic variant capture probe set is obtained by the following method: a) screening variants with a pathogenicity rating of two stars or above and P or LP in the Clinvar database, removing variants located in the probes of the comprehensive whole exon capture probe set, and obtaining pathogenic variants in deep intronic regions in Clinvar; b) screening variants recorded as DM and DM? in the HGMD database, removing variants located in the probes of the comprehensive whole exon capture probe set, and obtaining pathogenic variants in deep intronic regions in HGMD; and c) Design probes based on the positional information of pathogenic variants in deep intronic regions obtained in a) and b).
7. The human whole exome sequencing probe set according to claim 6, wherein the deep intronic pathogenic variant capture probe set comprises the probes shown in Table 19.
8. The human whole exome sequencing probe set according to any one of claims 1-7, wherein the intragenic pathogenic CNV capture probe set comprises probes for the following 60 genes: ABCD1, ACVRL1, ENG, ANO5, ATP7A, ATP7B, BTK, CAPN3, COL4A5, CREBBP, DMD, F8, F9, FANCA, FBN1, FH, UBE3A, GBA, GLDC, KCNQ1, KCNH2, KCNQ2, LDLR, LMNA, MEN1, MYBPC3, NF1, NPHP1, PAH, PCCA, PCCB, PKD1, PKD2, PLP1, PRRT2, SCN1A, SCN5A, SLC22A5, SOX9, TH, TSC1, TSC2, BRCA1, B RCA2, CHD7, EXT1, EXT2, MECP2, RB1, SHANK3, GCK, HNF1A, HNF1B, HNF4A, ARID1A, ARID1B, EHMT1, USH2A, ADGRV1 and PMP2. 9 . The human whole exome sequencing probe set according to claim 8 , wherein a probe is designed every 1 kb within the 60 genes.
10. The human whole exome sequencing probe set according to claim 8, wherein the intragenic pathogenic CNV capture probe set comprises the probes shown in Table 20.
11. A method for whole exome sequencing of a human sample, the method comprising: i) extracting genomic DNA from the human sample; ii) fragmenting the genomic DNA, performing end-repair on the obtained DNA fragments, and adding A nucleotides to the ends of the DNA fragments; iii) adding a universal sequencing adapter to the DNA fragment obtained in step ii); iv) amplifying the DNA fragments obtained in step iii) to thereby obtain a genomic DNA sequencing prelibrary of the human sample; v) mixing the genomic DNA sequencing prelibrary with the probes of the probe set according to any one of claims 1 to 10, so that the probes capture corresponding target DNA fragments by hybridization and remove non-target DNA fragments that are not captured; vi) amplifying the captured target DNA fragments, thereby obtaining a captured DNA library; and vii) performing high-throughput sequencing on the captured DNA library.
12. The method according to claim 11, wherein steps ii) to iv) are performed using Library Preparation Enzymatic Fragmentation (EF) Kit 2.0 and Universal Adapter System from Twist Bioscience.
13. The method according to claim 12, wherein the enzyme fragmentation reaction of step ii) is carried out at about 37°C for about 10 minutes.
14. The method according to claim 12 or 13, wherein step ii) is performed in a PCR instrument with a heated lid temperature of about 70°C.
15. The method according to any one of claims 12-14, wherein the ratio of the volume of 1) basic whole exon capture probe group, 2) mitochondrial full-length genome probe group, 3) copy number variation (CNV) backbone probe group, 4) deep intron pathogenic variation capture probe group and 5) intragenic pathogenic CNV capture probe group in the human whole exome sequencing probe group is 2μl:8μl:4μl:4μl:4μl, and the concentration of each probe is as follows: 206.80nM, 0.07nM, 14.92nM, 0.93nM, 4.00nM, respectively.
16. The method according to any one of claims 11 to 15, wherein steps v) to vi) are performed using the Fast Hybridization and Wash Kit from Twist Bioscience.
17. The method according to any one of claims 11-16, wherein the high-throughput sequencing is performed using the second-generation sequencing platform of BGI.
18. The method according to any one of claims 11-17, wherein the human sample is selected from amniotic fluid, tissue (such as abortion tissue, skin tissue), cell line, blood (such as peripheral blood, dried blood spots, cord blood).
19. A kit for performing whole exome sequencing on a human sample, the kit comprising the human whole exome sequencing probe set according to any one of claims 1 to 10.
20. The kit according to claim 19, which is used for performing whole exome sequencing on a human sample by the method according to any one of claims 11-18.
Citation Information
Patent Citations
Capture of breast cancer related genes, and preparation method and application of probes
CN103757709A
Whole-exome sequencing data analysis method
CN105930690A
Capture probe and gene chip for screening recessive genetic diseases as well as method and application of capture probe and gene chip
CN115807078A
Human whole exome sequencing probe set and application thereof
CN118222694A
Methods for Diagnosis of Familial Hypercholesterolemia
KR1020170055306A
Cited By
Design method and application of probe for enhancing molecular residual focus detection
CN121075423A