Novel smMIP probe, targeted detection kit and application of novel smMIP probe
By optimizing the structure of the smMIP probe and the primer sequence, the compatibility issues of smMIP technology on the MGI sequencing platform were resolved, the utilization rate and sensitivity of the probe were improved, and efficient targeted capture and detection on the MGI platform were achieved.
Patent Information
- Application Number
- CN202610209311.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-13
- Publication Date
- 2026-04-10
AI Technical Summary
The smMIP technology suffers from compatibility issues on the MGI sequencing platform, with low probe utilization and sensitivity, hindering its application and development among MGI sequencing users.
An smMIP probe was designed, comprising a 5' linker primer, UMI molecular tag 1, a sequencing library universal backbone, UMI molecular tag 2, and 3' extension primers. The probe structure and primer sequences were optimized to adapt to the MGI sequencing platform. A kit for targeted detection and a method for constructing targeted capture libraries were also provided.
This solves the incompatibility issue between smMIP and MGI library structures, improves probe utilization and sensitivity, reduces dependence on the Illumina platform, and enables efficient targeted capture and detection on the MGI platform.
Smart Images

Figure CN121826154A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of molecular detection technology, specifically relating to a novel smMIP probe, a targeted detection kit, and their applications. Background Technology
[0002] With the development of DNA sequencing technology, high-throughput sequencing has been widely applied in various fields of life science research. Although the cost of sequencing technology is decreasing, whole-genome sequencing itself remains expensive. Targeted capture technology enriches the target region of interest before high-throughput sequencing. Commonly used targeted sequencing methods mainly fall into three categories: hybridization capture, multiplex PCR enrichment (amplicon sequencing), and molecular inversion probes (MIP), which is a hybrid of the two technologies.
[0003] Probe hybridization capture is based on the complementary base pairing principle of nucleic acid molecules, and probes (DNA or RNA) are designed. These probes can be partially or completely complementary to the target region, and the hydrogen bonding force of complementary base pairing is used to enrich the target region DNA molecules, thereby achieving targeted sequencing detection.
[0004] Multiplex PCR is a technique developed to address the need for multi-target detection. It involves adding multiple primer pairs to a single reaction system to amplify multiple target sequences. In the field of leukemia diagnosis, multiplex PCR has been widely used for gene mutation screening, fusion gene detection, and minimal residual disease (MRD) monitoring.
[0005] Molecular inversion probes (MIPs) are technology that uses specific probes designed and synthesized to capture, amplify, and detect target genes through hybridization. The amplified products can be directly used as next-generation sequencing libraries.
[0006] Single-molecule molecular inversion probes (smMIPs) are a high-performance targeted sequencing technology developed in recent years. Combining the advantages of molecular inversion probe (MIP) technology and single-molecule PCR, smMIPs enable efficient and accurate targeted enrichment and sequencing at the single-molecule level. By hybridizing specific probes with the target DNA region, followed by enzymatic reactions and PCR amplification, smMIPs achieve highly efficient enrichment of specific genomic regions.
[0007] Single-molecule inverted probe technology utilizes the inverted sequence at the probe's end to capture, amplify, and detect target genes through hybridization. Compared to hybridization capture, it eliminates the need for constructing complex DNA libraries, making it more efficient and economical. Compared to multiplex PCR, the capture sequences at both ends of the probe are physically linked, resulting in higher capture and amplification specificity. Furthermore, introducing a single-molecule tag (UMI) into the probe structure can eliminate the impact of PCR and sequencing error rates, significantly improving the sensitivity of variant detection. In specific targeted capture applications, such as leukemia (AML, CML) and pan-cancer low-frequency variant detection, single-molecule inverted probes offer significant technical advantages.
[0008] However, existing smMIP technologies are primarily designed based on sequencing platforms such as Illumina, exhibiting significant incompatibility with MGI library structures (such as adapter design, primer binding sites, and amplification conditions). This hinders the development and application of smMIP technology by MGI sequencing users and delays its development and widespread adoption. Furthermore, while the single UMI in traditional smMIP probes can eliminate PCR duplication and correct random sequencing errors, its low utilization rate leads to low sensitivity. Summary of the Invention
[0009] To address the incompatibility issue between smMIP and MGI library structures, improve the utilization and sensitivity of smMIP probes, and meet the needs of MGI sequencing users for smMIP technology, this invention provides the following technical solutions.
[0010] In a first aspect, the present invention provides an smMIP probe, wherein the smMIP probe comprises, from 5' to 3', a 5' linker primer, a UMI molecular tag 1, a sequencing library universal backbone, a UMI molecular tag 2, and a 3' extension primer, wherein the nucleotide sequence of the sequencing library universal backbone is shown in SEQ ID NO.1 or SEQ ID NO.2.
[0011] Specifically, the nucleotide sequence of the general backbone of the sequencing library is as follows: AAGTCGGATCGTAGCCATGTCGTTCTGTGAGCCAAGGAGTTGTTGTCTTCCTAAGACCGCTTGGCCTCCGACTT (SEQ ID NO. 1).
[0012] AAGTCGGATCGTAGCCATGTCGTTCTGACCGCTTGGCCTCCGACTT (SEQ ID NO. 2).
[0013] Preferably, the UMI molecular tag 1 and the UMI molecular tag 2 are N bases with a length of 6-12 bp.
[0014] Furthermore, N can be any one of A, T, G, and C.
[0015] Furthermore, the lengths of the UMI molecular tag 1 and the UMI molecular tag 2 are selected from any one of 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt or 12 nt, more preferably 6 nt.
[0016] Furthermore, the sequences of UMI molecular tag 1 and UMI molecular tag 2 are different.
[0017] Preferably, the 5' connector primer is 14-30 nt in length, for example: 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, 30 nt.
[0018] Furthermore, the 5' junction of the smMIP probe is phosphorylated.
[0019] Preferably, the length of the 3' extension primer is 14-30 nt, for example: 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, 30 nt.
[0020] Preferably, the 5' linker and the 3' extension are matched with the target DNA sequence.
[0021] Furthermore, the targeted DNA is a gene for disease diagnosis and typing.
[0022] Furthermore, the diseases mentioned include, but are not limited to, leukemia, lymphoma, liver cancer, lung cancer, stomach cancer, esophageal cancer, breast cancer, cervical cancer, nasopharyngeal carcinoma, colorectal cancer, prostate cancer, thyroid cancer, myeloma, hemangioma, etc.
[0023] In a second aspect, the present invention provides an smMIP probe set, which includes smMIP probes with sequences as shown in SEQ ID NO. 7~284 or smMIP probes with sequences as shown in SEQ ID NO. 285~1439.
[0024] Thirdly, the present invention provides a DNA sequencing composition comprising the smMIP probe described in the first aspect, the smMIP probe set described in the second aspect, and primers.
[0025] Preferably, the primers include UDB-primer1 and UDB-primer2 used in conjunction with the smMIP probe containing SEQ ID NO.1 or Amp-Primer1 and Amp-Primer2 used in conjunction with the smMIP probe containing SEQ ID NO.2; The sequence of UDB-primer1 is as follows: CTCTCAGTACGTCAGCAGTTNNNNNNNNNNCAACTCCTTGGCTCACAGAAC (SEQ ID NO.3); The sequence of UDB-primer2 is as follows: GCATGGCGACCTTATCAGNNNNNNNNNNTTGTCTTCCTAAGACCGCTTG (SEQ ID NO.4); The sequence of Amp-Primer1 is as follows: CTCTCAGTACGTCAGCAGTTNNNNNNNNNNCAACTCCTTGGCTCACAGAACGACATGGCTACGATCCGACTT (SEQ ID NO.5); The sequence of Amp-Primer2 is as follows: GCATGGCGACCTTATCAGNNNNNNNNNNTTGTCTTCCTAAGACCGCTTGGCCTCCGACTT (SEQ IDNO.6); The N mentioned above can be any one of A, T, G, and C.
[0026] Fourthly, the present invention provides a targeted detection kit, characterized in that the kit comprises the smMIP probe described in the first aspect, the smMIP probe set described in the second aspect, or the composition described in the third aspect.
[0027] Preferably, the kit further includes hybridization buffer, nick-filling reaction buffer, nick-filling reaction enzyme mixture, digestive enzyme reaction mixture, and PCR mixture.
[0028] Furthermore, the hybridization buffer contains a DNA ligase buffer.
[0029] Furthermore, the gap-filling reaction buffer contains dNTPs, betaine, nicotinamide adenine dinucleotide, and DNA ligase buffer.
[0030] Furthermore, the gap-filling reaction enzyme mixture contains DNA ligase and hot-start DNA high-fidelity polymerase.
[0031] Furthermore, the digestive enzyme reaction mixture contains exonuclease I and exonuclease III.
[0032] Fifthly, the present invention provides a method for constructing a targeted capture library, comprising the steps of probe hybridization, nick filling, exonuclease digestion, PCR amplification, and purification.
[0033] Preferably, the reaction time for probe hybridization is 2.0-2.2 h.
[0034] Preferably, the probe hybridization reaction system consists of smMIP probes, target genomic DNA, hybridization buffer, and deionized water free of DNase / RNase.
[0035] Furthermore, the reaction conditions for probe hybridization are: 98℃ for 3-5 min; gradient cooling from 98℃ to 56℃ (0.1℃ / sec); 56℃ for 100-130 min; and holding at 56℃.
[0036] Preferably, the gap-filling reaction system consists of a gap-filling reaction buffer, a gap-filling reaction enzyme mixture, and deionized water free of DNA / RNA enzymes.
[0037] Preferably, the reaction conditions for filling the gap are: 56℃ for 3-7 min; 72℃ for 3-5 min; and 4℃ for heat preservation.
[0038] Preferably, the digestive enzyme reaction mixture of the exonuclease digestion contains exonuclease I and exonuclease III.
[0039] Preferably, the reaction time for the exonuclease digestion is 15-20 min, for example: 15 min, 16 min, 17 min, 18 min, 19 min, 20 min.
[0040] Preferably, the reaction conditions for the exonuclease digestion are: 37°C for 10-12 min; 95°C for 5-8 min; and 4°C for incubation.
[0041] Preferably, the PCR amplification reaction system consists of PCR mixture, amplification primers, exonuclease digestion products, and deionized water free of DNase / RNase.
[0042] The amplification primers are either UDB-amplification primers or Amp-amplification primers.
[0043] Furthermore, the UDB amplification primers include UDB-primer1 and UDB-primer2; the Amp amplification primers include Amp-Primer1 and Amp-Primer2.
[0044] Preferably, the PCR amplification reaction conditions are: 98℃, 30 s; 98℃, 5 s; 60℃, 10 s; 72℃, 15 s; 72℃, 2 min; and incubation at 4℃.
[0045] Preferably, in the purification step, magnetic beads are used to concentrate the PCR amplification product in two steps.
[0046] Furthermore, the volume ratio of the magnetic beads concentrated in the first step to the amplification product is 0.6-0.8, for example: 0.6, 0.65, 0.7, 0.75, 0.8.
[0047] Furthermore, the volume ratio of the magnetic beads concentrated in the second step to the amplification product is 0.1-0.3, for example: 0.1, 0.15, 0.2, 0.25, 0.3.
[0048] In a sixth aspect, the present invention provides a genome sequencing method based on the MGI sequencing platform, which includes the step of constructing a targeted capture library according to the method described in the fifth aspect.
[0049] In a seventh aspect, the present invention provides the use of the smMIP probe of the first aspect, the smMIP probe set of the second aspect, the composition of the third aspect, or the kit of the fourth aspect in constructing a targeted capture library.
[0050] Preferably, the targeted capture library is based on the MGI sequencing platform.
[0051] Eighthly, the present invention provides the application of the smMIP probe of the first aspect, the smMIP probe set of the second aspect, the composition of the third aspect, or the kit of the fourth aspect in disease diagnosis.
[0052] Preferably, the diseases include, but are not limited to, leukemia, lymphoma, liver cancer, lung cancer, stomach cancer, esophageal cancer, breast cancer, cervical cancer, nasopharyngeal carcinoma, colorectal cancer, prostate cancer, thyroid cancer, myeloma, or hemangioma.
[0053] The beneficial effects of this invention are: This invention constructs an smMIP probe for the MGI sequencing platform, which is used for target DNA genome sequencing. It solves the problem of incompatibility between smMIP and MGI library structure, improves the utilization and sensitivity of smMIP probe, and reduces dependence on specific platforms (such as Illumina). Attached Figure Description
[0054] Figure 1 The diagram shown is a schematic of the smMIP probe structure of the present invention. Figure 2 The diagram shown is a schematic of smMIP probe hybridization capture. Figure 3 The diagram shows the primer capture at the end of the smMIP probe. Figure 4 The capture performance of smMIP probes with different skeletons is shown in Figure A. A represents the long / short skeleton smMIP probe structure; B represents the capture performance of the long / short skeleton smMIP probe. Data is extracted from 3M PEread / 0.9 Gb data. Figure 5 The image shows the capture performance of the smMIP probe and the MIPgen probe; A represents the annealing temperature and amplicon length distribution of the smMIP probe set's linker and extension sequences; B represents the annealing temperature and amplicon length distribution of the MIPgen probe set's linker and extension sequences. p The value represents the significance level of the difference in the T-test. Figure 6 The figure shows the performance difference between the two probe sets at a probe concentration of 40 pM, with data extracted from 2M PEread / 0.6 Gb data. Figure 7 The figure shows the effect of two probe sets on capture performance under different probe deployment amounts; the data is extracted from 2 MPEreads / 0.6 Gb data. Figure 8 The results show the optimization of the capture process and the comparison of capture performance; A is the process flow; B is the target hit rate and coverage of different capture processes. p The value represents the significance level of the T-test; C represents the effect of different gDNA input amounts on the capture performance of the smMIP probe; data were extracted from 1.3M PEread / 0.4 Gb data. Figure 9 The image shows the detection results of the uniformity of the target area captured by the smMIP probe, with data extracted from 5M PEread / 1.5 Gb data. Figure 10 The figure shows the percentage of target regions captured at sequencing depths greater than 100X; data is extracted from 5M PEreads / 1.5 Gb. Detailed Implementation
[0055] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments and accompanying drawings. The described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the present invention... The embodiments described herein, and other embodiments obtained by those skilled in the art without inventive effort, are all within the scope of protection of this invention.
[0056] It should be noted that, unless otherwise specified, the experimental methods and reagents used in the embodiments of the present invention are all conventional experimental methods and reagents in the art.
[0057] Some of the reagents used in the examples are as follows: Table 1. Reagents used in the examples Example 1: Detection of the capture performance of single-molecule inverted probes with different backbone lengths 1.1 Design and preparation of smMIP probes and primers 1.1.1 According to Figure 1 The smMIP probe structure shown includes, from 5' to 3', a 5' linker, UMI molecular tag 1, a universal sequencing library backbone, UMI molecular tag 2, and a 3' extension. Two universal backbone sequences are designed: a 74 nt long backbone and a 46 nt segment backbone. Specific smMIP probes are shown below. Figure 4 As shown in Figure A.
[0058] The long skeleton sequence (74 nt) is as follows: AAGTCGGATCGTAGCCATGTCGTTCTGTGAGCCAAGGAGTTGTTGTCTTCCTAAGACCGCTTGGCCTCCGACTT (SEQ ID NO. 1).
[0059] The short skeleton sequence (46 nt) is as follows: AAGTCGGATCGTAGCCATGTCGTTCTGACCGCTTGGCCTCCGACTT (SEQ ID NO. 2).
[0060] The primers for PCR amplification after capture are as follows: The primer sequences used for the long backbone probe are: UDB-primer1: CTCTCAGTACGTCAGCAGTTNNNNNNNNNNCAACTCCTTGGCTCACAGAAC (SEQ ID NO.3); UDB-primer2: GCATGGCGACCTTATCAGNNNNNNNNNNTTGTCTTCCTAAGACCGCTTG (SEQ ID NO.4); The short backbone primer sequence is Amp-Primer1: CTCTCAGTACGTCAGCAGTTNNNNNNNNNNCAACTCCTTGGCTCACAGAACGACATGGCTACGATCCGACTT (SEQ ID NO.5); Amp-Primer2: GCATGGCGACCTTATCAGNNNNNNNNNNTTGTCTTCCTAAGACCGCTTGGCCTCCGACTT (SEQ ID NO. 6).
[0061] All of the above primers are 5' end phosphorylation modified primers, where NNNNNNNNNN is the sample barcode label.
[0062] 1.1.2 Based on the extended and connecting end sequences of 278 molecular inversion probes disclosed in the literature (BIEZUNER T, BRILON Y, ARYE AB, et al. An improved molecular inversion probe based targeted sequencing approach for low variantallele frequency [J]. NAR Genom Bioinform, 2022, 4(1): lqab125.), long backbone probes or short backbone sequences were connected to paired 6nt single-molecule tags (dUMIs) to form the long backbone (74 nt) probe set and the short backbone (46 nt) probe set tested in this embodiment (see Table 2). The only difference between the two probe sets is the backbone.
[0063] Table 2 Short-skeleton smMIP probe set 1.1.3 Preparation of probe working solution The two probe sets designed in 1.1.2 were synthesized into 5' phosphorylated ssDNA probes by Changzhou Xinyisheng Life Technology Co., Ltd., and then prepared into working solutions with final concentrations of 0.04 pM and 0.4 pM, respectively. The sample was NA12878 genomic DNA, and the initial input was 200 ng.
[0064] 2. Probe annealing 2.1 Remove the hybridization buffer and probe working solution and place them on an ice box to thaw. Gently vortex to mix, briefly centrifuge, and then place on an ice box for later use. Then prepare the PCR reaction solution on an ice box according to the table below, vortex to mix, and briefly centrifuge.
[0065] Table 3. Formulation of PCR reaction solution 2.2 Place the PCR tube containing the PCR reaction solution in the PCR instrument and start the following program. The hot cap temperature is 105℃ and the volume is 9.5 μL. Perform the PCR amplification reaction according to the reaction program in the table below.
[0066] Table 4 PCR reaction procedure 3. Filling the gap 3.1 Remove the gap-filling reaction buffer and place it on an ice box to thaw. Gently vortex to mix, briefly centrifuge, and then place it on an ice box for later use. Prepare the reaction solution on an ice box according to the table below, vortex to mix, and briefly centrifuge.
[0067] Table 5 Formulation of Gap Filling Reaction Solution 3.2 Remove the PCR tube from the PCR instrument, briefly centrifuge for 3 seconds, and place at room temperature. Then, take 6 μL of the reaction solution from 3.1 and add it to the room-temperature PCR tube. Gently vortex to mix, centrifuge for 5 seconds, tap to remove air bubbles, and centrifuge for another 5 seconds. Place the PCR tube back in the PCR instrument and start the following program: hot cap temperature 105℃, volume 15.5 μL: Table 6 Gap-Filling PCR Reaction Procedure 4. Exonuclease digestion Remove the PCR tube from the PCR instrument, centrifuge briefly for 3 seconds, and place at room temperature. Then, transfer 3 μL of the digestive enzyme reaction mixture to the room-temperature PCR tube, gently vortex to mix, centrifuge for 5 seconds, tap to remove air bubbles, and centrifuge for another 5 seconds. Place the PCR tube back into the PCR instrument and start the following program: hot cap temperature 105 °C, volume 18.5 μL: Table 7 Exonuclease Digestion PCR Reaction Procedure 5. PCR amplification 5.1 Thaw the amplification primers at room temperature, gently vortex to mix, briefly centrifuge, and then store at room temperature. Remove the PCR mixture, gently vortex to mix, briefly centrifuge, and then place on an ice pack. Then remove the PCR tube from the PCR instrument, briefly centrifuge for 3 seconds, and place on an ice pack.
[0068] 5.2 Prepare the reaction solution on an ice box according to the table below, vortex to mix, and centrifuge briefly.
[0069] Table 8 Formulation of Amplification PCR Reaction Solution Note: The primers for amplification of long backbone smMIP probes are UDB-primer1 and UDB-primer2; the primers for amplification of short backbone smMIP probes are Amp-Primer1 and Amp-Primer2.
[0070] 5.3 Place the PCR tube containing the amplification reaction solution in the PCR instrument and start the following program, with the hot cap temperature set to 105℃ and the volume to be 42.5 μL.
[0071] Table 9 Amplification PCR Reaction Procedure 6. Purification 6.1 Remove the purified magnetic beads (brand: Novizan, catalog number: N411-03) and equilibrate at room temperature for 30 min. Then mix 42.5 µL of amplification product with 34 µL of purified magnetic beads (0.8x) in a PCR tube, vortex for 10 s, centrifuge briefly for 3 s, and incubate at room temperature for 5 min. Transfer the supernatant to a new PCR tube.
[0072] 6.2 Add 8.5 µL of purified magnetic beads (0.2x) to a PCR tube containing supernatant, vortex to mix, and incubate at room temperature for 5 min. Briefly centrifuge the PCR tube, place it on a magnetic rack, and let it stand for 2–5 min until the liquid becomes clear. Carefully aspirate and discard the supernatant, being careful not to aspirate the magnetic beads. Then, slowly add 160 µL of 80% ethanol along the side away from the magnetic beads, and aspirate the supernatant with a pipette, being careful not to aspirate the magnetic beads. Add ethanol again and aspirate the supernatant.
[0073] 6.3 Remove the PCR tube from the magnetic rack, centrifuge briefly for 3 seconds, place it back on the magnetic rack, aspirate all the supernatant with a 10 µL pipette, open the tube cap, and let it sit at room temperature for about 2 minutes until the magnetic beads are dry (the surface of the magnetic beads should be rough and dull).
[0074] 6.4 Remove the PCR tube from the magnetic rack, add 17 µL of DNase / RNase-free deionized water, vortex thoroughly for 1 min, centrifuge gently, and incubate at room temperature for 5 min. Then place the PCR tube on the magnetic rack, allow it to absorb for 1 min, aspirate 15 µL of the supernatant into a new centrifuge tube, and label it with a library tag. This is the final library.
[0075] 6.5 Use the Qubit dsDNA HS Assay Kit (brand: NeoGene, catalog number: LS-DY-R-00004) for concentration quality control. A library is considered acceptable if the detected concentration is ≥1 ng / µL, and can then proceed to the next step of sequencing.
[0076] 7. High-throughput sequencing and data analysis Sequencing in PE150 mode was performed using an MGI G99 sequencer. PE reads were filtered using SOAPnuke (version 1.5.6) to remove low-quality data, allowing for the retention of duplicates and data containing library adapters. The filtered data was then assembled into a single sequence using PEAR (version 0.9.11). Adapter sequences were removed during assembly, resulting in a clean dataset. The obtained sequences contained single molecular tags (UMIs) at both ends, which were used to remove PCR duplicates. The assembled sequences were aligned to a reference sequence using Bowtie2 (version 2.4.1). The alignment results were used to calculate the target hit rate and UMI utilization (i.e., the percentage of non-duplicated data). After UMI deduplication, coverage, uniformity (coverage at 0.2X average sequencing depth), >100X coverage, and >1000X coverage were further evaluated to assess the smMIP probe capture performance.
[0077] like Figure 4 As shown in Figure B, there was no significant difference in the molecular inversion probe capture performance between long and short backbone probes. The sequencing performance of both backbones was comparable, with a 0.4 pM probe concentration showing better performance than a 0.04 pM backbone. The long backbone probe exhibited a higher target rate, while at a 0.4 pM concentration, the short backbone probe showed better uniformity and >100x coverage than the long backbone probe. Considering probe compatibility, the long backbone probe was chosen for subsequent experiments.
[0078] Example 2: Comparison of capture performance between smMIP probe and MIPgen probe This example demonstrates the difference in capture performance between the long skeleton smMIP probe and the MIPgen probe prepared in Example 1.
[0079] like Figure 3As shown, the smMIP probe captures double-stranded DNA from the target region in the sample, and its terminal primer binding region employs an alternating positive and negative strand design. The annealing temperature of the linker primer is 3-8°C higher than that of the extension primer, and the extension primer has no non-specific binding sites throughout the genome. The annealing temperature of the probe terminal primer is controlled above 50°C, and the hairpin folding free energy is higher than -5 kcal / mole to prevent the formation of a stable hairpin structure with the probe and enhance the binding efficiency to the target site. The design prioritizes ensuring the characteristics of the paired-end primers, and the amplicon length is strictly lower than the cumulative read length of sequencing (73-200 bp) to achieve complete sequencing throughput. Annealing temperature and folding free energy were calculated using the NN model (see BORER PN, DENGLER B, TINOCO I, et al. Stability of ribonu gap-filling eic acid double-stranded helices [J]. Journal of Molecular Biology, 1974, 86(4): 843-53) and the ViennaRNA tool (see LORENZ R, BERNHARTS H, HöNER ZU SIEDERDISSEN C, et al. ViennaRNA Package 2.0 [J]. Algorithms for Molecular Biology, 2011, 6(1): 26). Based on the above strategies, a set of 1155 long backbone smMIP probes was finally designed (see Table 10). The preparation of the probe working solution was the same as in Example 1.
[0080] MIPgen probes are designed using the MIPgen design tool (see BOYLE EA, O'ROAK BJ, MARTIN BK, et al. MIPgen: optimized modeling and design of molecular inversion probes for targeted resequencing [J]. Bioinformatics, 2014, 30(18): 2670-2), with default parameters.
[0081] The specific steps are the same as in Example 1: probe annealing, gap filling, exonuclease digestion, PCR amplification, and purification. The concentrations of the probe working solutions were 0.4 pM, 4 pM, and 40 pM, respectively.
[0082] Table 10 Long-skeleton smMIP probe set The smMIP probe design strategy of this invention prioritizes annealing temperature (Tm) to meet the molecular dynamics requirements of the probe's paired primer sequences, while the MIPgen probe prioritizes amplicon length.
[0083] like Figure 5 As shown in Figures AB, the annealing temperature Tm of the smMIP probe's connecting end is 62℃, and the annealing temperature Tm of the extension end is 56℃. In contrast, the annealing temperature Tm of the MIPgen probe's connecting end is 65℃, and the annealing temperature Tm of the extension end is 55℃. This indicates that the temperature fluctuation of the smMIP probe primers is lower than that of the MIPgen probe. The peak length distribution of the target region captured by the smMIP probe is 120 nt, while the target region length captured by the MIPgen probe is more concentrated around 120 nt.
[0084] like Figure 6As shown, the hit rate of the MIPgen probe was 74.44%, while that of the smMIP probe was 89.09%. Compared to the MIPgen probe, the smMIP probe showed a significantly improved hit rate at 40 pM and a higher 1000X coverage, indicating that the smMIP probe has superior capture performance.
[0085] like Figure 7 As shown, the smMIP probe outperforms the MIPgen probe in terms of target hit rate and 1000X coverage under different probe concentration conditions.
[0086] Example 3: Effects of reagent composition and process optimization on smMIP probe capture performance This embodiment demonstrates the optimization of the molecular inverted probe process and tests the effective range of genome input under the new process conditions.
[0087] In this embodiment, the original process referenced for optimization is ( Figure 8 A) refers to the molecular inversion probe capture (iMIP) process disclosed by Biezuner et al. in 2022 (see BIEZUNER T, BRILON Y, ARYE AB, et al. An improved molecular inversion probe based targeted sequencing approach for low variantallele frequency [J]. NAR Genom Bioinform, 2022, 4(1): lqab125.). The optimized process mainly includes: probe annealing, nick filling, exonuclease digestion, PCR amplification, and purification. The specific operations are the same as in Example 1, and the long backbone smMIP probe is the same as in Example 2.
[0088] The final concentration of the smMIP probe working solution in the hybridization system was 40 pM. The library used was NA12878 genomic DNA, and the amounts of genomic DNA used for genomic testing were 10 ng, 20 ng, 50 ng, 100 ng, 200 ng, 300 ng, 500 ng, 750 ng, 1000 ng, and 2000 ng, respectively.
[0089] like Figure 8 As shown in Figure A, compared to the iMIP process, the optimized process reduces the targeted capture sample processing time to approximately 3 hours, a 25% saving compared to the iMIP process's 4 hours. The required number of components in the gap-filling step is reduced from 6 to 2, decreasing the number of reagent components and manual operation time, thus reducing the risk of human error and better meeting the needs of automated workstations.
[0090] like Figure 8As shown in B, to reduce probe contamination of the data (target fragment is small, <200nt), the optimized process of this invention employs a two-step purification method (0.8x / 0.2x magnetic beads). The target hit rate of the optimized process of this invention is 93.67%, significantly higher than the target hit rate of 89.11% of the one-step purification process. Compared with the one-step purification (0.75x magnetic beads), the coverage of >1000X of the process of this invention is significantly improved.
[0091] like Figure 8 As shown in Figure C, when the gDNA input is greater than or equal to 100 ng, the optimized process exhibits good performance, namely, a target hit rate greater than 50%, uniformity (0.2x coverage) greater than 85%, and a >100x coverage reaching 97.97%, indicating that the process of this invention can effectively detect 100 ng genomic samples. Furthermore, the overall capture performance of the smMIP probe improves with increasing input volume.
[0092] Example 4: Determination of sensitivity and accuracy of variation detection in standard DNA This embodiment is used to verify the detection performance of the smMIP probe and optimized process for target mutations under low starting amount conditions, thereby evaluating the detection sensitivity, repeatability and quantitative accuracy of the sequencing method of the present invention in standards.
[0093] The pan-tumor DNA standard (GW-OGTM800) produced by Jingliang Gene Technology Co., Ltd. was used as the test sample. This standard contains the NRAS gene Q61K mutation (variant allele frequency of 1.04%), the FLT3 gene ΔI836 mutation (variant allele frequency of 2.14%), and the KIT gene D816V mutation (variant allele frequency of 2%), and all of these mutation sites are located within the target region covered by the method of this invention. The mutations represent different genes, different mutation types, and different levels of variant allele frequency. Using the long backbone smMIP probe from Example 2, the standard was tested with DNA input amounts of 30 ng, 50 ng, 100 ng, and 140 ng, respectively, with two independent replicate tests performed under each DNA input condition. The specific operating steps were the same as in Example 1, except that the purification step after library construction used a two-step magnetic bead purification method (0.8x / 0.2x).
[0094] To ensure the comparability of test results under different DNA input conditions, a uniform 5 MPE read / 1.5 GB was extracted for analysis during downstream analysis. The results are shown below. Figure 7 .
[0095] like Figure 9-10As shown, under the above sequencing data volume conditions, all samples obtained good sequencing quality, with the target region coverage exceeding 99% (>100×) and coverage uniformity (0.2x coverage) greater than 95%. This indicates that even under low sequencing data volume conditions, the sequencing method of this invention can still obtain stable and uniform target region coverage, providing a foundation for reliable detection of low-frequency mutations.
[0096] Under the above sequencing quality and coverage conditions, the detection results of the target mutation in the standard were analyzed, and the results are shown in the table below.
[0097] Table 11. Allele frequencies of NRAS gene Q61K, FLT3 gene ΔI836, and KIT gene D816V variants in the standards. As shown in Table 11, under the conditions of DNA input of 30 ng, 50 ng, 100 ng, and 140 ng, the sequencing method of the present invention can simultaneously detect three mutations: NRAS Q61K, FLT3ΔI836, and KIT D816V. Furthermore, consistent detection results were obtained in two repeated tests, with a detection accuracy of 100% and a coefficient of variation (CV) of less than 20% for the variant allele frequency (VAF). Notably, the NRAS gene Q61K mutation can still be accurately quantified even with a variant allele frequency of approximately 1% (mean VAF measurement 1.1%, CV 15%).
[0098] The above results demonstrate that the method of the present invention can achieve stable detection of multiple genes, different mutation types, and low mutation frequency mutations under conditions of low DNA input (30 ng) and limited sequencing data (5 MPEreads / 1.5 Gb), and has high detection sensitivity, good repeatability, and quantitative accuracy.
[0099] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural transformations made based on the content of the present invention specification, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. An smMIP probe, characterized in that, The smMIP probe includes, from 5' to 3', a 5' linker primer, UMI molecular tag 1, a sequencing library universal backbone, UMI molecular tag 2, and a 3' extension primer. The nucleotide sequence of the sequencing library universal backbone is shown in SEQ ID NO.1 or SEQ ID NO.
2.
2. The smMIP probe according to claim 1, characterized in that, The UMI molecular tag 1 and the UMI molecular tag 2 are each 6-12 bp in length, with N being any one of A, T, G, and C; the sequences of the UMI molecular tag 1 and the UMI molecular tag 2 are different.
3. The smMIP probe according to claim 1, characterized in that, The 5' linker primer and the 3' extension primer are 14-30 nt in length, and the 5' linker is phosphorylated.
4. An smMIP probe set, characterized in that, The smMIP probe set includes smMIP probes with sequences as shown in SEQ ID NO.7~284 or smMIP probes with sequences as shown in SEQ ID NO.285~1439.
5. A DNA sequencing composition, characterized in that, The composition comprises any of the smMIP probes of claims 1-3 or the smMIP probe set and primers of claim 4.
6. The composition according to claim 5, characterized in that, The primers include UDB-primer1 and UDB-primer2 used in conjunction with the smMIP probe containing SEQ ID NO.1 or Amp-Primer1 and Amp-Primer2 used in conjunction with the smMIP probe containing SEQ ID NO.2; The sequence of UDB-primer1 is as follows: CTCTCAGTACGTCAGCAGTTNNNNNNNNNNCAACTCCTTGGCTCACAGAAC; The sequence of UDB-primer2 is as follows: GCATGGCGACCTTATCAGNNNNNNNNTTGTCTTCCTAAGACCGCTTG; The sequence of Amp-Primer1 is as follows: CTCTCAGTACGTCAGCAGTTNNNNNNNNNNCAACTCCTTGGCTCACAGAACGACATGGCTACGATCCGACTT; The sequence of Amp-Primer2 is as follows: GCATGGCGACCTTATCAGNNNNNNNNNNTTGTCTTCCTAAGACCGCTTGGCCTCCGACTT; N can be any one of A, T, G, and C.
7. A kit for targeted detection, characterized in that, The kit comprises any of the smMIP probes of claims 1-3, the smMIP probe set of claim 4, or any of the compositions of claims 5-6.
8. A method for constructing a targeted capture library, comprising the steps of probe hybridization, nick-filling, exonuclease digestion, PCR amplification, and purification, characterized in that, The reaction system for the gap-filling step consists of a gap-filling reaction buffer, a gap-filling reaction enzyme mixture, and deionized water free of DNA / RNA enzymes.
9. The construction method according to claim 8, characterized in that, The reaction time for probe hybridization is 2.0-2.2 h; and / or The reaction time for the exonuclease digestion is 15-20 min; and / or The purification process employs a two-step magnetic bead concentration method. In the first step, the volume ratio of magnetic beads to amplification products is 0.6-0.8, and in the second step, the volume ratio is 0.1-0.
3.
10. The use of any of the smMIP probes of claims 1-3, the smMIP probe set of claim 4, any of the compositions of claims 5-6, or the kit of claim 7 in constructing a targeted capture library.