A double-stranded consensus sequencing method based on hierarchical targeted enrichment and a special reagent combination
Patent Information
- Application Number
- CN202611098642.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-08-21
AI Technical Summary
(1)接头连接效率受限:DS技术采用随机UMI接头,易发生自连和接头二聚体形成,导致大量原始DNA分子在连接阶段即告丢失,传统DS技术通常需要约1μg的起始DNA输入量
(1)双链X-96 UMI接头并非仅依据计算所得的编辑距离选择UMI或仅凭样本索引读数分布进行经验性筛选,而是在随机接头池实际连接表现的基础上,按连接效率筛选,并结合汉明距离≥3、无连续≥3个相同碱基、GC平衡等序列设计原则共同确定。由此兼顾了双链UMI接头的实际连接效率、降低UMI碰撞率,从而提高链分组及单链/双链共有序列(SSCS/DCS)构建的准确性以及提高分子回收率。与现有商品化 Duplex 测序接头(如TS-DUMI)相比,在相同实验条件下,采用本发明X-96 UMI集合所构建的文库,其接头连接效率(CV35.83% vs 72.64%),及分子回收率显著提高。与采用较短 UMI 的方案(如 CN113913495B的 5~6nt UMI)相比,本发明 UMI 为8nt UMI,候选空间为48=65,536种,远大于现有5~6nt UMI的4,096种(CN113913495B);且96种UMI的两端组合多样性达96×96=9,216种,约为现有64种UMI(64×64=4,096种)的2.25倍,可有效降低高深度测序场景下的UMI碰撞率;进一步采用复合分子标识符,以UMI序列与插入片段端点处8bp基因组序列共同构成标识符,总有效标识符空间扩展为9,216×B(B为断点多样性),相比仅用UMI使碰撞概率再降低;且该2.25倍的相对优势不随断点多样性或输入量变化。
Smart Images

Figure CN122609690A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular detection technology, specifically to a double-stranded consensus sequencing method based on hierarchical targeted enrichment and a combination of dedicated reagents. Background Technology
[0002] Ultra-low frequency DNA mutation detection has significant clinical value in areas such as early tumor diagnosis, liquid biopsy, and non-invasive prenatal testing. However, existing detection technologies still have significant limitations in terms of sensitivity, sample requirements, and detection efficiency.
[0003] The baseline error rate of conventional next-generation sequencing (NGS) is approximately 10. - ² to 10 - ³, which is much higher than the actual abundance of ultra-low frequency mutations in cfDNA (cell-free DNA), and the real mutation signal is easily masked by background noise.
[0004] Duplex Sequencing (DS) technology reduces the background error rate to <5×10⁻⁶ by independently labeling and sequencing complementary Watson and Crick strands, utilizing the two-strand consensus principle. -7 However, DS technology still faces the following key bottlenecks: (1) Limited adapter ligation efficiency: DS technology uses random UMI adapters, which are prone to self-ligation and adapter dimer formation, resulting in the loss of a large number of original DNA molecules during the ligation stage. Traditional DS technology usually requires about 1 μg of initial DNA input.
[0005] (2) The DS technology uses hybridization capture to enrich the target region, which has the following inherent defects: First, the hybridization capture system is usually optimized for large panels of hundreds of kb to several Mb. When the target region is reduced to the exon region of a single or a few genes (such as only covering TP53 or EGFR hotspot exons, the target region accounts for 10% of the genome). -6 ~10 -7 First, the probe-target hybridization efficiency is significantly reduced, with a large number of target molecules being irreversibly lost during hybridization and elution, resulting in low overall molecule utilization. Second, it usually requires two rounds of overnight hybridization, with the entire process taking 2 to 3 days. Third, the cost of probes is based on the entire panel, and even if only a few sites are detected, the cost is close to that of a complete panel, making it difficult to flexibly customize for individualized MRD tracking sites.
[0006] SaferSeqS technology uses strand-specific nested PCR instead of hybridization capture to achieve targeted enrichment, but its core enrichment strategy relies on multiple rounds of exponential PCR amplification, which brings two inherent limitations: (1) Amplification bias leads to a competitive disadvantage for low-abundance mutant templates. Low-abundance mutant templates are at a natural disadvantage in competition with high-copy wild-type templates and are easily suppressed or even lost by dominant templates; (2) SaferSeqS uses strand-specific nested PCR to directly target and enrich the whole genome library. The specificity and amplification efficiency of the primers are the only factors that determine the success of the enrichment. This type of pure nested PCR enrichment method has the following inherent limitations: its enrichment depends entirely on the exponential amplification of the target region by the inner and outer nested primers, so the specificity of the primers is extremely high. In the case of insufficient primer specificity or complex adjacent sequences of the target region, non-specific amplification products are easily generated, leading to enrichment failure or increased background. Since not all genomic regions can be designed with nested primer pairs that meet such high specificity requirements, the applicable genomic regions of this method are significantly limited. It is difficult to cover target regions located near repetitive sequences, with extreme GC content, or lacking ideal primer design windows, thus limiting its versatility in detecting a wide range of gene loci. In addition, when expanding to large detection panels for multiple genes and multiple loci, mutual interference and non-specific amplification can easily occur between a large number of inner and outer primers, further restricting its multiplexing capability.
[0007] In addition, another similar technique to SaferSeqS, a strand-separated double-strand sequencing method (see publication number CN110520542B, hereinafter referred to as the prior art) enriches the two strands separately by dividing the library into tubes and using primers specific to different adapter sequences in each tube. The prior art also describes several alternative schemes to improve enrichment efficiency, including: (1) a primer site disruption scheme, which involves using methylation primers in conjunction with methylation-dependent restriction enzymes (such as MspJI, FspEI) to disrupt one side of the adapter primer site; and (2) a biotin affinity enrichment scheme, which involves introducing biotin onto the target primer and then enriching the target product with streptavidin.
[0008] While the comparative document mentions adding a linear amplification step to each tube after segmentation to aid enrichment, these linear amplification steps are all placed after segmentation and primarily serve to facilitate the installation of artificial tags that disrupt primer sites or to reduce off-target effects associated with biotin enrichment. Furthermore, the comparative document does not provide any experimental data demonstrating the extent to which the described linear amplification step improves transformation efficiency. The comparative document reports a free DNA transformation efficiency of approximately 1% at baseline.
[0009] Therefore, existing technologies address the loss of target molecules caused by splitting by "rescuing after splitting," which involves complex workflows, relies on additional chemical steps such as methylation modification or biotin enrichment, and still has significant room for improvement in conversion efficiency.
[0010] In summary, while existing hybridization capture enrichment methods have a wide applicability, they suffer from low molecule recovery rates and a sharp drop in efficiency in small target regions. These methods also rely on customized biotin-labeled capture probes, require lengthy hybridization processes, involve numerous steps, have high reagent costs, and have detection cycles lasting several days, resulting in high economic and time costs. Existing pure nested PCR enrichment methods (such as SaferSeqS) improve molecule recovery rates, but their complete reliance on nested primers for enrichment specificity leads to extremely high primer specificity requirements, limited applicability to different genomic regions, and difficulties in multiplexing. Currently, there is a lack of a targeted enrichment method that offers high molecule recovery rates, broad applicability to different genomic regions, flexible expansion capabilities, and effectively reduces economic and time costs.
[0011] To address the issues of unverifiable and inefficient connection of random UMIs, two fixed UMI connector solutions have emerged in the existing technology, but both have shortcomings: The authorized patent CN113913495B provides 64 types of fixed Duplex UMI connectors with a UMI length of 5~6nt and an edit distance of ≥2 between any two sequences. This design solves the unverifiable problem of random UMIs to some extent, but still has the following limitations: (1) The UMI length is relatively short (5~6 nt), and the sequence candidate space is limited (4 nt). 6 =4,096), the combinatorial diversity of UMI is only 64×64=4,096, and the probability of UMI collision is relatively high in high-depth sequencing scenarios with high DNA input.
[0012] (2) The edit distance is only ≥2, and a single base sequencing error can shorten the distance to 1, so the error correction tolerance space is limited; (3) The frame sequence of this adapter is designed only for the MGI sequencing platform and is not applicable to the Illumina sequencing platform; (4) The hybridization capture method is inefficient in enriching the target gene.
[0013] TS-DUMI adapters use predefined fixed UMI sequences instead of random UMIs, reducing the required DNA input from ~1 μg in traditional DS to ~100 ng. However, TS-DUMI exhibits poor ligation uniformity; the Bioscience coefficient of variation (CV) for the ligation efficiency of 96 adapters is as high as 72.64% (95% CI: 61.04%–83.59%), indicating significant differences in ligation efficiency among different UMI sequences, with some sequences exhibiting preferential ligation or ligation inhibition. This high dispersion in ligation efficiency leads to a reduction in the number of usable effective UMIs, a decrease in equivalent molecular resolution, and a still limited molecular recovery rate (the proportion of double-stranded consistency reads to the theoretical number of input molecules, a comprehensive indicator reflecting the retention efficiency of raw DNA molecules throughout the entire process from adapter ligation, enrichment to sequencing).
[0014] Neither the TS-DUMI nor the CN113913495B adapters were screened for UMI sequences based on a systematic evaluation of ligation efficiency in real sequencing data. This lack of synergistic optimization of ligation efficiency, base balance, and sequence discriminability results in significant room for improvement in ligation uniformity and molecule recovery efficiency. Furthermore, both employed relatively inefficient hybridization capture methods for enrichment.
[0015] For ultra-low frequency mutation detection technology to achieve clinical translation in the field of liquid biopsy, it must simultaneously possess an ultra-low error rate (<10). -7 High molecular weight recovery and compatibility with large-scale personalized mutation tracking panels are also crucial. There is an urgent need to develop a novel detection method that can overcome these bottlenecks while maintaining the ultra-high specificity of double-stranded sequencing. Summary of the Invention
[0016] To address the problems existing in the prior art, one of the objectives of this invention is to provide a hierarchical targeted enrichment-based double-stranded consensus sequencing method, the specific steps of which are as follows: (1) DNA fragmentation: Extract DNA from the sample to be tested and break the DNA into small fragments of 100-300bp.
[0017] (2) End repair: The small fragments obtained in step (1) are dephosphorylated and flattened.
[0018] (3) Adding adapters: Add double-stranded X-96 UMI adapters to both ends of the DNA fragments that have undergone end repair to obtain DNA fragments with adapters, and then purify them with magnetic beads to obtain magnetic bead purified products.
[0019] (4) The purified magnetic bead product was subjected to PCR amplification to obtain the initial library construction product.
[0020] (5) Using the initial library construction product obtained in step (4) as a template, linear pre-amplification is performed using LP primers.
[0021] (6) Divide the library product obtained in step (5) into two parts and perform the first round of amplification of Watson and Crick chains respectively. Watson chain amplification is performed using GSP1 primer and universal amplification primer P7; Crick chain amplification is performed using GSP1 primer and universal amplification primer P5, to obtain Watson and Crick chains amplified in the first round respectively.
[0022] (7) The Watson and Crick chains obtained in the first round of amplification obtained in step (6) are subjected to a second round of Watson and Crick chain amplification, respectively. The second round of Watson chain amplification is performed using GSP2 Watson primers and universal primer Index-P7; the second round of Crick chain amplification is performed using GSP2 Crick primers and universal primer Index-P5, respectively, to obtain the second round of amplified Watson and Crick chains.
[0023] Preferably, the sequence of primer P7 is GTGACTGGAGTTCAGACGTGTGCTCTTCCGATC; and the sequence of primer P5 is ACACTCTTTCCCTACACGACGCTCTTCCGATCT.
[0024] Preferably, the sequence of the universal primer Index-P7 is CAAGCAGAAGACGGCATACGAGAT-Index sequence-GTGACTGGAGTTCAGACGTGT; the sequence of the universal primer Index-P5 is AATGATACGGCGACCACCGAGATCTACAC-Index sequence-ACACTCTTTCCCTACACGAC.
[0025] Preferably, LP primers, GSP1 primers, GSP2 Watson primers, and GSP2 Crick primers are designed and screened according to the target region to be detected, in order to improve the enrichment efficiency and amplification specificity of the target fragment, and are designed according to the following principles: A. Design candidate primers within a range of 50–200 bp upstream of the target detection region, with a candidate primer length of 16–25 bp; B. Avoid the formation of obvious primer dimers and hairpin structures in the primer sequence; C. The annealing temperatures of the candidate primers are kept similar; D. The GC content in the primer sequence is controlled between 40% and 60%; E. Perform whole-genome alignment of candidate primer sequences and select primers that specifically match the target region and have no significant homology with other non-target regions.
[0026] The second objective of this invention is to provide an X96 UMI set, which is derived from a predefined set of specific sequences, each 8 bases in length. This set contains 96 sequences (referred to as the "X-96 UMI set"). Unlike existing non-random UMIs (such as 6-7 base vNRUMIs) constructed algorithmically based solely on sequence edit distance and of variable length, and Duplex adapters using degenerate / semi-degenerate base random tags, the X-96 UMI set of this invention has a uniform length of 8 bases and is systematically obtained through a combination of experimental data-driven active screening and multiple sequence constraints. The nucleic acid sequences of the X-96 UMI set are as follows: TTCGTCCA, TTCCGAGT, TGCTTAAC, TTGTCATG, TGTGATAA, TGCTATGT, TCTTAGAC, TTAGATCG, TGACATCA, TGCAGGCT, TCTCGATC, TGGCTAAG, TCTAGCCA, TGAATTAT, GTACATCC, TGAACGAG, TAGTGCCA, TATGCGGT, GGTGCAAC, TCGTAAGG, GTCGAGAA, TATAGATT, GGTCCTCC, TCGATGAG, GGATACTA, TAGCTGTT, GACTAGGC, TCCGGATG, GGAACATA, GTCTGGAT, GACACTAC, TATGTCCG, GCGTATCA, GTCCACAT, CTTAGGCC, GGCATCCG, GCCAATGA, GCAGCATT, CTGTAGCC, GGCAATTG, GATTCGGA, GACGAGTT, CGATAATC, GCTGGCTG, GATCTAGA, GAAGCGAT, CCTGCCTC, GCGCGGAG, GAGACGTA, CGGCCATT, CCGCTTCC, GCACCTTG, GAAGGCGA, CGAGTCCT, CCAAGAGC, CTTCATCG, CTTAGCGA, CACTGCTT, CAGACATC, CTCCTAAG, CGACTTGA,ATGACCAT,CACCGTTC,CTACCTAG,CGACGCCA,AGATTGAT,ATTGTCTC,CCGAGACG,CCTGTGCA,ACTTGGAT,ATGGACAC,CCATTCCG, CAGCCAGA, ACTATCGT, ATGCGCTC, CAGCAATG, ATTGCAGA, ACGGATGT, ATATGACC, CAATTGTG, ATGGCTTA, ACTCGGT, ATAAGGAC, CAACTTCG, ATCTCTGA, AAGGCCTT, AGTAGTGC, ATTATCCG, AGAGACGA, AACTTACT, AGGACAGC, AATTATAG, AAGAGCGA, AACGTGGT, ACCTCTAC, AACATGCG.
[0027] A third objective of this invention is to provide a dual-chain X-96 UMI connector, the connector comprising a frame sequence and one sequence from the X-96UMI set.
[0028] Preferably, the dual-chain X-96 UMI connector is a Y-type dual-link connector. One chain in the dual-link connector has a structure of frame 1 sequence - a sequence of the X-96 UMI set, and the other chain has a structure of X-96 UMI set sequence - T - frame 2 sequence, wherein the sequence of frame 1 is ACACTCTTTCCCTACACGACGCTCTTCCGATC; and the sequence of frame 2 is GATCGGAAGAGCACACGTCTGAACTCCAGTCAC.
[0029] The fourth objective of this invention is to provide a dedicated reagent set for a hierarchical targeted enrichment-based double-stranded consensus sequencing method. The dedicated reagent set includes a double-stranded X-96 UMI adapter and a set of nested primers targeting the target region to be detected, including LP primers, GSP1 primers, GSP2 Watson primers, and GSP2 Crick primers. The binding sites of the LP primers, GSP1 primers, and GSP2 primers are nested inward in sequence, forming three targeting recognition levels with different sequences and binding sites.
[0030] Preferably, the special reagent combination further includes universal amplification primers that specifically bind to the adapter sequence, including universal amplification primer P7, universal amplification primer P5, universal primer Index-P7, and universal primer Index-P5.
[0031] Preferably, LP primers, GSP1 primers, GSP2 Watson primers, and GSP2 Crick primers are designed and screened according to the target region to be detected, in order to improve the enrichment efficiency and amplification specificity of the target fragment, and are designed according to the following principles: A. Design candidate primers within a range of approximately 50–200 bp upstream of the target detection region, with the preferred primer length being 16–25 bp; B. Avoid the formation of obvious primer dimers and hairpin structures in the primer sequence; C. The annealing temperatures of the candidate primers are kept similar; D. The GC content in the primer sequence is controlled between 40% and 60%; E. Perform whole-genome alignment of candidate primer sequences and select primers that specifically match the target region and have no significant homology with other non-target regions.
[0032] The fifth objective of this invention is to provide a specific reagent combination for detecting EGFR gene mutations in concrete applications. When detecting EGFR gene Exon19 deletion mutations, the LP primer sequence in the dedicated reagent kit is GTGTCCCTCACCTTCGG; the GSP1 primer sequence is GTGTCCCTCACCTTCGG; the GSP2 Watson primer sequence is AATGATACGGCGACCACCGAGATCTACACAGACTCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTGTGCATCGCTGGTAAC; and the GSP2 Crick primer sequence is CAAGCAGAAGACGGCATACGAGATTGGTGGAAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGTGCATCGCTGGTAAC. When detecting the Exon21 L858R point mutation in the EGFR gene, the sequence of the LP primer in the special reagent kit is GGATCAGTAGTCACTAAC; the sequence of the GSP1 primer is CTAACGTTCGCCAGCCAT; the sequence of the GSP2 Watson primer is AATGATACGGCGACCACCGAGATCTACACAGACTCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTCAGCCATAAGTCCTCGACGT; and the sequence of the GSP2 Crick primer is CAAGCAGAAGACGGCATACGAGATTGGTGGAAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTCAGCCATAAGTCCTCGACGT. When detecting the Exon20 T790M point mutation in the EGFR gene, the LP primer sequence in the dedicated reagent kit is TCCAGGAAGCCTACGTGATG; the GSP1 primer sequence is TGATGGCCAGCGTGGAC; the GSP2 Watson primer sequence is AATGATACGGCGACCACCGAGATCTACACAGACTCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTGCGTGGACAACCCCCAC; and the GSP2 Crick primer sequence is CAAGCAGAAGACGGCATACGAGATTGGTGGAAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGCGTGGACAACCCCCAC. When detecting the Exon20 C797S point mutation in the EGFR gene, the LP primer sequence in the dedicated reagent kit is TCCAGGAAGCCTACGTGATG; the GSP1 primer sequence is TGATGGCCAGCGTGGAC; the GSP2 Watson primer sequence is AATGATACGGCGACCACCGAGATCTACACAGACTCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTGCGTGGACAACCCCCAC; and the GSP2 Crick primer sequence is CAAGCAGAAGACGGCATACGAGATTGGTGGAAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGCGTGGACAACCCCCAC.
[0033] The fifth objective of this invention is to provide a dedicated reagent combination for detecting TP53 gene mutations. When detecting Exon4 deletions or mutations in the TP53 gene, the LP primer sequence in the dedicated reagent combination is GCAGCCTCTGGCATTCT; the GSP1 primer sequence is GCAGCCTCTGGCATTCT; the GSP2 Watson primer sequence is AATGATACGGCGACCACCGAGATCTACACAGACTCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTGGCATTCTGGGAGCTTCAT; and the GSP2 Crick primer sequence is CAAGCAGAAGACGGCATACGAGATTGGTGGAAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGGCATTCTGGGAGCTTCAT. When detecting Exon8 deletion or mutation in the TP53 gene, the LP primer sequence in the dedicated reagent kit is AACTGCACCCTTGGTCTC; the GSP1 primer sequence is CTTGGTCTCCTCCACC; the GSP2 Watson primer sequence is AATGATACGGCGACCACCGAGATCTACACAGACTCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTCCTCCACCGCTTCTTGTC; and the GSP2 Crick primer sequence is CAAGCAGAAGACGGCATACGAGATTGGTGGAAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTCCTCCACCGCTTCTTGTC.
[0034] Compared with the prior art, the present invention provides the following beneficial effects: (1) The double-stranded X-96 UMI adapters are not selected solely based on the calculated edit distance or empirically based solely on the sample index reading distribution. Instead, they are selected based on the actual ligation performance of the random adapter pool, according to ligation efficiency, and combined with sequence design principles such as Hamming distance ≥3, no consecutive ≥3 identical bases, and GC balance. This balances the actual ligation efficiency of double-stranded UMI adapters and reduces the UMI collision rate, thereby improving the accuracy of strand grouping and single-stranded / double-stranded common sequence (SSCS / DCS) construction and improving molecule recovery. Compared with existing commercially available Duplex sequencing adapters (such as TS-DUMI), under the same experimental conditions, the library constructed using the X-96 UMI set of this invention shows significantly improved adapter ligation efficiency (CV 35.83% vs 72.64%) and molecule recovery. Compared with schemes using shorter UMIs (such as the 5-6 nt UMI of CN113913495B), the UMI of this invention is 8 nt UMI, with a candidate space of 4 8 =65,536 species, far exceeding the 4,096 species of existing 5-6 nt UMIs (CN113913495B); and the end-coupled diversity of the 96 UMIs reaches 96×96=9,216 species, which is about 2.25 times that of the existing 64 UMIs (64×64=4,096 species), which can effectively reduce the collision rate of UMIs in high-depth sequencing scenarios; furthermore, a composite molecular identifier is adopted, which is composed of the UMI sequence and the 8 bp genome sequence at the end of the inserted fragment to form the identifier, and the total effective identifier space is expanded to 9,216×B (B is the breakpoint diversity), which further reduces the collision probability compared to using only UMIs; and this 2.25-fold relative advantage does not change with breakpoint diversity or input volume.
[0035] (2) This invention proposes a hierarchical targeted enrichment-based dual-strand consensus sequencing method, namely, a progressive linear pre-amplification enrichment strategy (Linear-enhanced Duplex Sequencing, LinDS). This invention introduces a linear pre-amplification step of the target region before strand-specific nested PCR enrichment, forming a hierarchical enrichment system of "linear pre-amplification + strand-specific nested exponential PCR". The linear pre-amplification uses primers specific to different adapter sequences, paired with gene-specific primers (GSP1, GSP2) nested inside the linear pre-amplification primers for amplification, replacing traditional hybridization capture. This hierarchical system avoids the sharp drop in efficiency caused by an excessively small target region in hybridization capture. This characteristic makes this invention naturally suitable for clinical scenarios such as minimal residual disease (MRD) monitoring and companion diagnostics, where only a few known mutation sites need to be detected. Specifically, the target-specific linear pre-amplification before separation can significantly improve… The conversion efficiency of target molecules to double-stranded common sequences is improved, and by increasing the copy number of the two strands of under-amplified rare molecules and reducing the probability of strand loss during separation, it is particularly beneficial for the recovery of low-frequency mutations in circulating tumor DNA. However, linear amplification placed after separation cannot recover lost strands. At the same time, the linear pre-amplification primer, GSP1, and GSP2 form a three-level nested hierarchy, requiring the amplification product to pass through three independent target recognitions consecutively, which can more effectively suppress off-target products and improve the on-target rate. The above characteristics enable the present invention to significantly improve the molecule recovery rate and reduce time and economic costs. Specifically, the present invention adopts a three-primer progressive architecture, with LP (Linear Primer), GSP1 (Gene-Specific Primer 1), and GSP2 (Gene-Specific Primer 2) working together: In the LP linear pre-amplification stage, the LP primer targets and binds to the outer position of the target region, performing multiple rounds of linear amplification on the library template after adapter ligation in a single primer manner. This stage accumulates the target strand copy number linearly. Since only one-sided primers participate in amplification, non-target molecules cannot be exponentially amplified, thus achieving initial targeted enrichment. The GSP1 progressive linear amplification stage: GSP1 primers bind to the outer side of the target region and are nested inside the LP primers, amplifying in conjunction with the adapter primers. Molecules containing the target region can be simultaneously bound by GSP1 and the adapter primers, resulting in exponential amplification. Non-target molecules are only linearly amplified by the adapter primers. Because the GSP1 binding site is located inside the LP amplicon, only amplified products truly containing the target region can be effectively extended by GSP1, further improving enrichment specificity and excluding non-specific background amplification.Experiments confirmed that nested amplification of LP and GSP1 improves enrichment efficiency. In the GSP2 exponential amplification stage, the GSP2 primer is further nested inside the GSP1 primer, working in conjunction with the adapter primer to perform final exponential PCR amplification of the products pre-enriched by two rounds of linear enrichment via LP and GSP1. The GSP2 primer carries the adapter sequence required for sequencing; only target fragments containing its binding site can obtain the adapter at the gene-specific end, thus forming a complete library with both ends ready for sequencing together with the adapter primer. Off-target fragments remaining from the LP and GSP1 stages lack GSP2 binding sites and cannot obtain the adapter at that end, resulting in an incomplete library at both ends, and therefore not being sequenced. Thus, the final amplification is completed against a highly specific background where the target molecule is already highly enriched, effectively avoiding non-specific products occupying sequencing throughput. Since the target molecules are highly enriched after the first two rounds of amplification, the exponential amplification in this stage is performed under a highly specific background, effectively avoiding the waste of sequencing throughput by non-specific products. The key to the progressive three-primer architecture is that enrichment is achieved entirely through the ladder design of primer positions and linear amplification itself, rather than relying on physical capture methods. This design allows the amplification products from each round to proceed to the next reaction without additional purification or selection steps, thereby maximizing the preservation of molecular diversity in the library. Experimental verification shows that the molecular recovery rate achieved by the LinDS method of this invention can reach up to 81.19%, significantly higher than that of existing double-stranded sequencing methods using physical capture enrichment (such as CRISPR-DS, which achieve 6%–12%). The improvement in molecular recovery rate (approximately 6–10 times) is a technical effect that cannot be reasonably expected by those skilled in the art based on existing technologies.
[0036] (3) This invention introduces a linear pre-amplification step, decomposing the enrichment process into a hierarchical system of "linear pre-amplification for initial enrichment + two-phase chain-specific nested PCR for fine enrichment". In the linear pre-amplification stage, a single LP primer is used for mild initial enrichment of the target region. Based on this, subsequent nested PCR is only performed against the background of the pre-enriched product, thus significantly reducing the dependence on the specificity of a single nested primer. Since the enrichment specificity is handled by the linear pre-amplification and nested PCR hierarchically, rather than entirely by the nested primer, the method of this invention can be applied to a wider range of genomic regions. Furthermore, when expanding to large detection panels with multiple genes and multiple loci, the method of this invention exhibits higher tolerance to interference between primers and stronger multiplexing capabilities. This technical effect is unpredictable by those skilled in the art based on existing pure nested PCR enrichment techniques.
[0037] (4) The present invention further provides a combination of targeted primers for the EGFR gene and the TP53 gene. The primer combination is obtained through large-scale experimental screening and is specifically adapted to the progressive linear pre-amplification enrichment strategy. In the prior art, primer design for the above target genes is usually based on sequence parameters predicted by bioinformatics tools (such as Primer3, NCBI Primer-BLAST, etc.). There are few systematic experimental evaluations combined with the special requirements of Duplex sequencing library construction (including SMI / UMI integrity preservation, amplification specificity of each strand, progressive amplification compatibility, etc.).
[0038] (5) This invention designs and synthesizes a large number of candidate primers for the targeted mutation regions of the EGFR gene (including key exons / mutation sites related to the sensitivity and resistance of targeted drugs) and the functional regions of the TP53 gene. Attached Figure Description
[0039] Figure 1 It is a Y-type double-chain X-96 UMI connector structure.
[0040] Figure 2 The product prepared by the Y-type double-chain X-96 UMI linker was identified by Qseq bio-fragment analyzer.
[0041] Figure 3 Figure A compares the connection uniformity of the Y-type double-chain X-96 UMI connector and the commercial TS-DUMI connector; Figure B is a bar chart comparing the connection uniformity of the Y-type double-chain X-96 UMI connector (CV=35.83%) and the commercial TS-DUMI connector (CV=72.64%).
[0042] Figure 4 A comparison of molecular recovery rates between the Y-type double-chain X-96 UMI connector and the commercially available TS-DUMI connector.
[0043] Figure 5 To verify the quantitative accuracy of the Y-type double-stranded X-96 UMI linker in the Spike-in experiment; Figure A shows the consistency analysis of two technical repeatability tests; Figure B shows the spike-in experimental design; Figure C shows the consistency between observed and theoretical values confirmed by the spike-in experiment (R2=0.9949, P<0.0001); Figure D shows an example at the TP53:7674828 site, where the theoretical value (E, EXPECTED) and the observed value (O, OBSERVED) are consistent.
[0044] Figure 6To improve the diversity of molecular species and the on-target ratio in pre-linear amplification; Figure A shows the number of SSCS reads in the hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) and SaferSeqS method described in this invention; Figure B shows the proportion of reads in the target region in the hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) and SaferSeqS method described in this invention.
[0045] Figure 7 The effect of a three-level progressive nested amplification system on the enrichment efficiency of the target region and the inhibition effect of non-specific amplification.
[0046] Figure 8 This figure compares the performance of the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS), the hybridization capture method based on Y-type double-stranded X-96 UMI adapter, and the SaferSeqS method described in this invention at an input volume of 10 ng. Figure A shows the molecular recovery efficiency of the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS), the hybridization capture method based on Y-type double-stranded X-96 UMI adapter, and the SaferSeqS method in fresh tumor tissue. Figure B shows the minimum detection frequency of the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS), the hybridization capture method based on Y-type double-stranded X-96 UMI adapter, and the SaferSeqS method in fresh tumor tissue. Figure C shows the molecular recovery efficiency of the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS), the hybridization capture method based on Y-type double-stranded X-96 UMI adapter, and the SaferSeqS method in urinary ctDNA. Figure D shows the molecular recovery efficiency of the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS), the hybridization capture method based on Y-type double-stranded X-96 UMI adapter, and the SaferSeqS method in urinary ctDNA. Minimum detection frequency of UMI adapter hybridization capture method and SaferSeqS method in urinary ctDNA.
[0047] Figure 9 The recovery rate of the EGFR probe molecule is given.
[0048] Figure 10 This invention validates the quantitative accuracy of spike-in sequencing using the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) at an input volume of 10 ng. Figure A shows the spike-in design, where sample A is gradient-mixed into sample B at proportions of 0.5%, 0.1%, 0.01%, and 0.001% to construct the spike-in model. Figure B shows a high linear correlation between the theoretical expected mutation frequency and the actual observed value for all mutation sites (R² = 0.96). Figure C shows an example of the consistency between the theoretical and observed values for a single mutation site in the EGFR gene. Figure D shows a high degree of consistency between two technical replicates of the tumor sample.
[0049] Figure 11 The figures show the error correction performance of SSCS vs. DCS; Figure A represents the error correction performance of SSCS after single-chain consensus correction; Figure B represents the error correction performance of DCS after dual-chain consensus correction.
[0050] Figure 12 The invention demonstrates that the experimental operation time of the hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in the invention has been reduced from 3 days to 1 day.
[0051] Figure 13 This invention validates the clinical application of the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) in bladder cancer ctDNA. Figure A shows the frequency distribution of all mutations in tumor tissue, with the vertical axis representing mutation frequency (logarithmic scale). Each point represents one mutation, and the mutation frequency in the tumor spans four orders of magnitude (0.0001-1), including high-frequency driver mutations and low-frequency subclonal mutations. Figure B compares the overall detection rate of the three detection methods. Figure C shows the correlation analysis between urinary ctDNA and tumor tissue mutation frequency.
[0052] Figure 14 Figure 1 shows the dynamic monitoring results of ctDNA in neoadjuvant therapy patients; Figure A represents the sampling time point for neoadjuvant therapy patients; Figure B represents the dynamic monitoring results of tumor major efficiency mutations in ctDNA in two NAC1 MIBC patients before and after three treatment cycles; Figure C represents the dynamic monitoring results of tumor major efficiency mutations in ctDNA in two NAC2 MIBC patients before and after three treatment cycles.
[0053] Figure 15 This diagram illustrates the dynamic tracking of TP53 mutations in urinary ctDNA using the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention. Figure A shows the sampling time; Figure B shows the gradual decrease in TP53 mutation frequency as treatment progresses. Detailed Implementation
[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0055] The reagent kits, manufacturers, and uses used in the examples are shown in Table 1.
[0056] Table 1 Overview of the reagent kits used in the examples Example 1 Primer design In this embodiment, a large number of candidate primers were designed and synthesized targeting mutation regions of the EGFR gene (including key exons / mutation sites related to targeted drug sensitivity and resistance) and functional regions of the TP53 gene, and systematically screened using the following criteria: (1) Progressive amplification compatibility: The candidate primers must be compatible with the LP / GSP1 / GSP2 three-primer progressive architecture of the present invention, that is, the binding sites of LP, GSP1 and GSP2 must be arranged in a ladder, and the extension products of each primer can effectively enter the next step of amplification. (2) Target enrichment efficiency: The enrichment efficiency of each candidate primer combination on the target region was evaluated by the measured on-target rate, and primer combinations with significantly better enrichment efficiency than other candidates were screened. (3) UMI integrity retention: After amplification with the primers, the X-96 double-stranded UMI must be completely retained in the amplification product to ensure the accuracy of subsequent double-stranded identity sequence (DCS) analysis; (4) Non-specific amplification inhibition: The selected primers must maintain high specificity in the complex background of the Duplex sequencing library (containing a large number of non-target molecules and adapter dimers), and the proportion of non-specific amplification products must be lower than the preset threshold. (5) Double-strand amplification uniformity: To meet the requirements of accurate reconstruction of double-strand information in duplex sequencing, the double-strand separation and amplification system established in this invention should ensure that the target DNA double strands maintain uniform amplification efficiency during subsequent amplification processes. This avoids introducing additional strand bias due to differences in tube separation operations, adapter structures, or universal primers, which could lead to an imbalance in double-strand coverage depth and affect the construction efficiency of the double-strand consensus sequence (DCS) and the accuracy of low-frequency mutation detection. By evaluating the family number, sequencing depth, and pairing efficiency of the positive and reverse strands in the double-strand amplification products, the established system was verified to have good double-strand amplification uniformity. Through the above large-scale experimental screening, this invention finally determined the appropriate primer combinations for EGFR and TP53 gene targeting, as shown in Table 2.
[0057] Table 2 Overview of primers used in the examples Note: Index sequences are a standard design in this field. The core of selecting index sequences is to ensure that the indexes of different samples can be accurately distinguished without affecting sequencing quality. Commercial index sequences can be used directly.
[0058] Example 2 Design of X-96 UMI assembly and fabrication of Y-type double-chain X-96 UMI connector The following principles were followed for the design and screening of 8bp UMIs: (a) 8bp fixed nucleotide sequences, 96 in each group, with base balance in each adapter to ensure that the proportion of A, T, C, and G bases at each position (positions 1–8) of each UMI is close to 25%, effectively reducing PCR amplification bias, as shown in Table 3; (b) the minimum edit distance between any two UMIs in the set is not less than 3; (c) the number of consecutive identical bases in each UMI does not exceed 2, and the design results are shown in Table 4.
[0059] Table 3 Table 4X-96 UMI ensemble design results Based on the sequence characteristics of the Illumina sequencing platform, the universal framework sequence 1 was designed as follows: ACACTCTTTCCCTACACGACGCTCTTCCGATC-X96UMI-T; The general framework sequence 2 is designed as follows: X96UMI-T-GATCGGAAGAGCACACGTCTGAACTCCAGTCAC The UMI sequence and the framework sequence were assembled to obtain each single strand of the synthesized double-stranded X-96 UMI connector, as shown in Table 5.
[0060] Table 5. Sequence of each single chain of the double-chain X-96UMI connector. Connection of 96 pairs of Y-type double-chain X-96 UMI connectors (1) Adapter dissolution and mixing: Each single-strand sequence of the synthesized double-stranded X-96 UMI adapter (as shown in Table 5) needs to be dissolved using Duplex-buffer (IDT, cat 11-05-01-03) (100 mM potassium acetate; 30 mM HEPES, pH 7.5). The concentration of the adapter after dissolution should be adjusted to 100 µM to ensure sufficient reaction volume for subsequent annealing. The dissolved single strands are then mixed in equal proportions.
[0061] (2) Annealing process: The dissolved single strand was placed in the PCR instrument, the temperature was set to 95°C, and annealing was performed for 5 minutes. After that, the instrument was turned off and left to stand for 1 hour. The Y-type double-stranded X-96 UMI adapter was then obtained.
[0062] (3) Disassembly and storage of connectors: The Y-type double-chain X-96 UMI connectors obtained after annealing need to be disassembled into multiple small containers. The disassembled connectors should be stored immediately in a -80°C freezing environment to ensure the stability and long-term storage of the connectors.
[0063] Effect Analysis The structure of the 96 pairs of Y-type double-chain X-96 UMI connectors prepared in this embodiment is as follows: Figure 1 As shown, the Y-type double-chain X-96 UMI connector was identified using the Qseq bio-fragment analyzer, and the results are as follows. Figure 2 As shown, from Figure 2 It can be seen that the Y-type double-chain X-96 UMI linker has only a single main peak, indicating that its preparation efficiency is very high.
[0064] Example 3 Comparison of connection uniformity between Y-type double-chain X-96 UMI connector and commercial TS-DUMI fixed connector 200 ng of human genomic DNA was taken and used with the Y-type double-stranded X-96 UMI adapter described in this invention and the commercial TS-DUMI fixed adapter (TwinStrand Biosciences) under the same conditions for end repair, adapter ligation and library amplification. Sequencing was performed on the Illumina platform, and the proportion of molecules ligated by each UMI to the total number of molecules was counted. The coefficient of variation (CV) of the ligation ratio among the 96 UMIs was calculated. The lower the CV, the more uniform the ligation efficiency between different UMI sequences. The higher the ligation uniformity, the closer the actual usage frequency of each UMI is to the theoretical equal probability (1 / 96), and the higher the effective utilization of the UMI combinatorial space, thereby reducing the probability of molecular collisions.
[0065] The samples used in Example 3 included 293T cell line and SW480 cell line, and the library was constructed according to the following steps: (1) DNA was extracted from fresh 293T cells using the Cell Genomic DNA Extraction Kit (DC102) from Novizan. The specific operation steps were performed according to the kit instructions.
[0066] (2) The DNA obtained in step (1) was fragmented by ultrasonic fragmentation (Covaris M220) to obtain DNA fragments of 100-300 bp.
[0067] (3) The DNA fragments obtained in step (2) were subjected to end repair treatment using the Novizan VAHTS Universal DNA Library Prep Kit for Illumina V4 kit to obtain end repair products. The end repair system is shown in Table 6 (X in Table 6 needs to be calculated according to the actual situation. For example, if 33ng DNA needs to be input, if the DNA of a certain sample is 11ng / ul, then add 3ul). The reaction procedure is shown in Table 7.
[0068] Table 6 End-of-life repair system Table 7 Reaction Procedure (4) Connect the 96 pairs of Y-type double-chain X-96UMI connectors prepared in Example 2 to the end repair product obtained in step (3) to obtain the Adapter Ligation. The connector connection system is shown in Table 8 and the connection procedure is shown in Table 9.
[0069] Table 8 Connector Connection System Table 9 Connection Program (5) After the reaction was completed, 88 μL of VAHTS DNA Clean Beads was added to 110 μL of AdapterLigation product to adsorb the target fragment. The magnetic beads were washed twice with freshly prepared 80% ethanol to remove excess adapters and reaction solution. The DNA was dissolved in 25 μL of ddH2O solution. The purified magnetic bead product was subjected to initial library finite-cycle amplification using the VAHTS HiFi Amplification Mix kit. The library amplification system is shown in Table 10, and the library amplification program is shown in Table 11. Table 10 Library amplification system Table 11 Library Expansion Procedures Results Analysis: The proportion of 96 Y-type double-stranded X-96 UMI adapters in the sequencing data was statistically analyzed. Ideally, each type should account for approximately 1 / 96 ≈ 1.04%. Figure 3 (A). The Y-type double-stranded X-96 UMI linker exhibits better base composition balance and sequence differentiation, effectively reducing cross-reactions and secondary structure formation tendencies between linkers, resulting in a more uniform distribution of the connection efficiency of the 96 Y-type double-stranded X-96 UMI linkers. Figure 3 The A-type double-stranded X-96 UMI connector avoids preferential ligation or suppression of specific sequences; the coefficient of variation (CV) of the Y-type double-stranded X-96 UMI connector is 35.83% (95% CI: 31.58%–39.94%), while the CV of the TS-DUMI connector is 72.64% (95% CI: 61.04%–83.59%), indicating that the latter has a significantly higher degree of dispersion in ligation efficiency (p<0.0001). Figure 3 In the middle B. X96 UMI, 99.07% of the total data volume was accounted for. Figure 3 (B)
[0070] Note: This indicator is crucial. Although the amount of connector added is excessive, it is only possible if all 96 UMIs are present in equal amounts in the mixing pool. If one UMI is too much or too little, it will lead to uneven distribution, which is equivalent to reducing the number of usable UMIs and increasing the disadvantage of molecular collision rate.
[0071] Example 4 Comparison of molecular recovery rates between Y-type double-chain X-96 UMI connectors and commercially available TS-DUMI connectors Genomic DNA from human lung cancer tissue, adjacent normal tissue, and 293T cell line was collected (200 ng each). Y-type double-stranded X-96 UMI adapters and TS-DUMI fixed adapters (TwinStrand Biosciences) were used. Under the same DNA input volume (200 ng), the same library construction method (end repair → adapter ligation → library amplification), the same hybridization capture probe and capture procedure, and the same sequencing data volume, the molecular recovery rate of the two adapters was calculated. The molecular recovery rate refers to the proportion of independent molecules that finally obtain DCS support to the total number of theoretical original molecules of input DNA, which comprehensively reflects the molecular retention efficiency of the entire process from adapter ligation, enrichment to sequencing.
[0072] 200 ng of human genomic DNA corresponds to approximately 60,600 haploid genome equivalents (200 ng ÷ 3.3 pg / genome), meaning that theoretically, there are approximately 60,600 independent original DNA molecules at any given genomic locus. After library construction (using the same method as in Example 2), hybridization capture, and sequencing, the number of independent molecule families supported by the final double-stranded consensus sequence (DCS) at that locus is counted. The molecule recovery rate = number of DCS families ÷ theoretical input molecule number × 100%. First, the library is constructed according to the method described in Example 3 to obtain the library construction product. Then, the library construction product is subjected to two rounds of hybridization capture using the Novozymes hybridization capture kit (NC103). The library after the first round of hybridization capture is amplified for 19 cycles, and the library after the second round of hybridization capture is amplified for 10 cycles. The capture probe covers the exon region of the TP53 gene. The specific steps of the two-round hybridization capture process are as follows: (1) The library construction product was mixed with Cot-1 DNA. The library mixing system is shown in Table 12. The reaction was allowed to stand at room temperature for 10 min. After the reaction was completed, the magnetic beads were washed twice with freshly prepared 80% ethanol to remove excess Cot-1 DNA.
[0073] Table 12 (2) Prepare the hybridization capture solution as shown in Table 13.
[0074] Table 13 The sequence of the TP53 gene probe is shown below (biotin is attached to the 5' end of the sequence): AAAAAAAAAAAAAAAAGAAAAGCTCCTGAGGTGTAGACGCCAACTCTCTCTAGCTCGCTAGTGGGTTGCAGGAGGTGCTTACGCATGTTTGTTTCTTTGCTGCCGTCTTCCAGTTGCTTT、ATCTGTTCACTTGTGCCCTGACTTTCAACTCTGTCTCCTTCCTCTTCCTACAGTACTCCCCTGCCCTCAACAAGATGTTTTGCCAACTGGCCAAGACCTGCCCTGT GCAGCTGTGGGTTG,ATTCCACACCCCCGCCCGGCACCCGCGTCCGCGCCATGGCCATCTACAAGCAGTCACAGCACATGACGGAGGTTGAGGCGCTGCCCCCACCATGAGCGCTGCTCAGATAGC GATGGTG and CCTAGGTTGCCTCTGACTGTACCACCATCCACTACAACTACATGTGTAACAGTTCCTGCATGGGCGGCATGAACCGGAGGCCCATCCTCACCATCATCACACTGGAAGACTCCAGGTCAG.
[0075] (3) Add 19µL of the prepared hybridization capture solution to the magnetic beads washed with 80% ethanol to dissolve the DNA. Then take 17µL for capture. The capture reaction procedure is shown in Table 14. After the hybridization capture is completed, use CA-28 Streptavidin Beads of the hybridization capture kit (NC103) of Novizan to elute the captured target fragment.
[0076] Table 14 (4) The captured target fragment was amplified by PCR. The reaction system is shown in Table 15 and the reaction procedure is shown in Table 16. After the reaction was completed, PCR products were obtained. 40 µL of VAHTS DNA Clean Beads magnetic beads were added to 50 µL of library amplification products to adsorb the target fragment. The magnetic beads were washed twice with freshly prepared 80% ethanol to remove excess adapters and reaction solution. The DNA on the magnetic beads was dissolved with 35 µL of ddH2O.
[0077] Table 15 Table 16 (5) Take 1µL of the DNA obtained in step (4) and use a Qubit4 instrument to detect the concentration.
[0078] (6) Take 1µL of the DNA obtained in step (4) and use the Qsep1 biological fragment analyzer to detect the fragment size.
[0079] (7) Repeat the hybridization capture once more (i.e., the second hybridization capture). In the second hybridization capture, the target fragment after hybridization capture is amplified for 10 cycles (i.e., when performing PCR amplification in the second hybridization capture, the 19 cycles in Table 16 are set to 10 cycles), and the concentration and fragment size are detected.
[0080] Each sample underwent two rounds of hybridization capture. From library construction to assay, a total of three PCR amplifications were performed. The PCR products obtained from the above steps were subjected to paired-end sequencing on the Illumina platform. The sequencing data were quality controlled and filtered based on the characteristics of the Y-type double-stranded X-96 UMI adapter fixed sequence set: reads whose first 8 bases did not belong to the Y-type double-stranded X-96 UMI adapter fixed set and reads whose 9th base was not T were removed. Subsequently, the reads were grouped into molecular families based on the composite molecular identifier composed of UMI and 8bp endpoint sequence, and the number of double-stranded consistent sequences obtained by different methods was counted.
[0081] Effect Analysis: The Y-type double-stranded X-96 UMI adapter, through synergistic optimization of the adapter's terminal structure and ligase affinity, significantly improves the overall ligation efficiency and substantially increases the molecule recovery rate compared to TS-DUMI. Figure 4 Molecular recovery rate refers to the proportion of independent molecules ultimately supported by a double-stranded consensus sequence (DCS) out of the total theoretical number of original molecules in the input DNA. This metric comprehensively reflects the retention efficiency of original DNA molecules throughout the entire process from adapter ligation and enrichment to sequencing.
[0082] Example 5 Y-type double-chain X-96 UMI connector Spike-in quantitative accuracy verification Genomic DNA was extracted from SW480 and 293T cell lines. Using SW480 cell line genomic DNA as a background, 293T cell line genomic DNA was mixed in at mass ratios of 50%, 10%, 1%, and 0.1%. Figure 5 The blue text represents SW480 cell line genomic DNA, and the red text represents 293T cell line genomic DNA. The total input amount was 200 ng. The 293T cell line genomic DNA was used to construct a library according to the method described in Example 2, and hybridization and capture were performed according to the method described in Example 3. The data were obtained after sequencing on the Illumina platform and PIPELINE analysis with double-strand error correction.
[0083] The 293 and SW480 cell lines each carry a known set of characteristic mutations, and their mutation profiles do not overlap. By performing deep sequencing on both cell lines separately beforehand, the actual allele frequencies of their respective mutation sites were determined. When the two cell lines are mixed at a known mass ratio, the theoretical expected frequency of each mutation site can be precisely calculated based on the actual frequency of the mutation in the source cell line and the mixing ratio. For example, if the actual frequency of a mutation in SW480 is f, when SW480 and 293 are mixed at a ratio of 1:99, the theoretical mutation frequency of this site in the mixed sample is f ÷ (1 + 99) = f / 100. By constructing gradient mixed samples of 1:1, 1:10, 1:100, and 1:1000, a series of mutation standards covering high to very low frequencies can be generated in the same experiment. Comparing the observed mutation frequencies with the theoretical expected frequencies, the linear correlation (R²) reflects the quantitative accuracy, and the lowest theoretical frequency that can be reliably detected is the lower limit of detection sensitivity (LOD).
[0084] like Figure 5 As shown in Figure A, the detection results from the two technical replications were highly consistent (R² = 0.9938, P < 0.0001), and the detected mutation sites were completely identical, indicating that the detection system has good repeatability. Comparing the actual observed mutation frequencies with the theoretical expected frequencies for all data points, the two were highly consistent (R² = 0.9949, P < 0.0001). Figure 5 (C and D) indicate that the detection system based on the Y-type double-stranded X-96 UMI linker maintains excellent quantitative accuracy across a mutation frequency range spanning three orders of magnitude. When the desired number of mutant molecules is ≥1 (desired mutation frequency ≥5×10⁻⁶), the accuracy is significantly improved. -4 At this point, the Obs / Exp ratio remained stable between 0.5 and 2.0, indicating good quantitative accuracy. Even when the expected mutation frequency decreased further and the number of corresponding expected mutant molecules was less than one, the system could still detect rare mutations, for example, when the theoretical expected mutation frequency was as low as 6.6 × 10⁻⁶. -5 The mutation at the specified site was still successfully detected, corresponding to a single-molecule detection event, indicating that this detection system has the ability to detect rare mutations down to the single-molecule level.
[0085] Example 6 To verify the effectiveness of the target enrichment strategy (LP linear pre-amplification) described in this invention, linear pre-amplification enrichment systems and nested GSP enrichment systems using the same DNA samples were constructed respectively. The experimental group employed the gene-specific single-primer linear amplification strategy before tube separation as described in this invention; the control group employed the nested gene-specific exponential amplification strategy of SaferSeqS. Except for the target enrichment method, all other experimental conditions, including DNA input volume, adapter ligation, double-strand separation, sequencing, and data analysis procedures, remained consistent. To verify the effectiveness of the gene enrichment method (LP linear pre-amplification) described in this invention, 33 ng of fresh lung cancer tissue DNA and adjacent normal morphological tissue DNA were used as samples, and comparisons were made by detecting the TP53 gene and exon4 gene in the samples.
[0086] First, the library was constructed according to Example 3. Then, the library was pre-linearly amplified and the product was tailed to meet the requirements for sequencing. The enrichment effect was compared with that of the library without linear amplification (i.e., the SaferSeqS method).
[0087] The enrichment of the target gene using the hierarchical targeted enrichment double-stranded consensus sequencing method (LinDS) described in this invention is as follows: (1) Using a DNA library with Y-type double-stranded X-96 UMI adapter as a template, linear pre-amplification of the target region was performed using the primers shown in Table 2 (SEQ ID NO:05). The amplification system is shown in Table 17, and the amplification program is shown in Table 18.
[0088] Table 17 Amplification System Table 18 Amplification Procedure (2) After linear pre-amplification, the product was used directly as a template and primers (LP binding adapter tail sequence AATGATACGGCGACCACCGAGATCTACACAGACTCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTGCAGCCTCTGGCATTCT) were used to amplify the product with SEQ ID NO:28 in Table 2. The amplification system is shown in Table 19 and the amplification program is shown in Table 20 to obtain a library product that meets the requirements of Illumina sequencing.
[0089] Table 19 Amplification System Table 20 Amplification Procedure After PCR, the product was purified using 0.8×AMP magnetic beads to obtain the final sequencing library.
[0090] The SaferSeqS method flow is as follows: Using a DNA library ligated with a Y-type double-stranded X-96 UMI adapter as a template, two rounds of exponential amplification were performed directly without linear pre-amplification: First round of PCR: The DNA library ligated with the Y-type double-stranded X-96 UMI adapter was divided into two equal parts, and the Watson and Crick strands were amplified separately. Watson strand reaction: Amplification was performed using primers shown in SEQ ID NO:05 and SEQ ID NO:14 in Table 2, respectively. Crick strand reaction: Amplification was performed using primers shown in SEQ ID NO:05 and SEQ ID NO:13 in Table 2, respectively. The amplification system is shown in Table 21, and the amplification program is shown in Table 22 (the number of cycles is kept consistent with SaferSeqS: 19X). This step achieves the physical separation of complementary double strands, laying the foundation for subsequent double-stranded consensus analysis.
[0091] Table 21 Amplification System Table 22 Amplification Procedure Second-round PCR: The amplification products from the previous round were subjected to further nested PCR, amplifying the Watson and Crick chains separately. The Watson chain reaction was performed using primers shown in Table 2 (SEQ ID NO: SEQ ID NO: 23 and SEQ ID NO: 28). The Crick chain reaction was performed using primers shown in Table 2 (SEQ ID NO: 24 and SEQ ID NO: 27). The amplification system is shown in Table 23. The amplification program is shown in Table 24 (the cycle number remains consistent with SaferSeqS: 17X). This step further enriches the target region while adding complete adapters to the molecule to meet the Illumina sequencing requirements.
[0092] Table 23 Amplification System Table 24 Amplification Procedure After PCR, the product was purified using 0.8×AMP magnetic beads to obtain the final sequencing library.
[0093] PCR products obtained by both methods were subjected to paired-end sequencing on the Illumina platform. Sequencing data were quality-controlled based on the characteristics of the X96 fixed sequence set: reads whose first 8 bases did not belong to the X96 fixed set and reads whose 9th base was not T were removed. Subsequently, reads were grouped into molecular families based on a composite molecular identifier composed of the UMI and the 8bp endpoint sequence, and the number of single-stranded consistent sequences obtained by different methods (SSCS reads) and the proportion of reads in the target region (on-target ratio) were calculated.
[0094] Technical effects: Molecular diversity increased by 4-5 times: The SSCS reads of the hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention are 19,989-29,287, while SaferSeqS only has 3,892-8,382 reads. Figure 6 (A) The on-target ratio is improved by 6-7 times: LinDS is 25.01%~42.23%, and SaferSeqS is 3.55%~6.09% ( Figure 6 (B). The above results indicate that the hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention effectively improves the molecular diversity and enrichment efficiency of the target fragment through pre-linear amplification, which is significantly better than the SaferSeqS method.
[0095] Example 7 Construction and Validation of a Three-Level Progressive Nested Amplification System The role of linear preamplification primers as an independent nested level was verified by detecting point mutations in TP53 gene Exon8, EGFR gene Exon21 L858R, and EGFR gene Exon20 T790M. Two experimental groups were set up: the control group used the same primer for both linear preamplification and the first-round chain-specific nested PCR (i.e., the linear preamplification primer was GSP1 primer, second-level nesting); the experimental group used three primers with different sequences and binding sites for linear preamplification, GSP1, and GSP2 (i.e., the linear preamplification primer was not equal to GSP1 primer, third-level nesting, with the linear preamplification primer located on the outermost side), and all other conditions were kept the same. The targeting enrichment efficiency of the two groups was compared.
[0096] The results showed that the target enrichment efficiency of the experimental group (linear pre-amplification primers are not equal to GSP1 primers, three-level nesting) was significantly higher than that of the control group (linear pre-amplification primers are not equal to GSP1 primers, two-level nesting), confirming that setting the linear pre-amplification primers as the outermost nesting level independent of GSP1 can improve the target enrichment efficiency.
[0097] like Figure 7As shown, compared with the two-primer amplification system, the three-level progressive amplification system constructed in this invention significantly improves the enrichment efficiency of the target region, reduces the proportion of non-specific amplification products, and increases the proportion of effective sequencing data and the effective molecule recovery rate.
[0098] Example 8 Performance validation of the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention at an extremely low input level of 10 ng, and comparison of molecule recovery rates between the hybridization capture method based on Y-type double-stranded X-96 UMI adapter and the SaferSeqS method. The amount of extractable DNA from clinical liquid biopsy samples (such as plasma, urine, cerebrospinal fluid, etc.) is typically only a few nanograms to tens of ng, far lower than the 100 ng to 1 μg required for traditional double-stranded sequencing. Therefore, efficient recovery of original molecules under extremely low input volumes is a key prerequisite for the clinical translation of the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention. This embodiment uses 10 ng as a uniform input volume to evaluate the molecular recovery rate and detection sensitivity of the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) under this condition, and compares it with the hybridization capture method based on the Y-type double-stranded X-96 UMI adapter and the SaferSeqS method based on nested PCR.
[0099] Experimental samples included DNA from fresh lung cancer tissue and ctDNA from urine of bladder cancer patients; The target genes detected in this embodiment are: TP53 gene Exon4, TP53 gene Exon8, EGFR gene Exon19 deletion mutation, EGFR gene Exon20 T790M point mutation, and EGFR gene Exon20 C797S point mutation. The specific experimental procedure is as follows: (1) First, construct the library according to the method in Example 3, and divide the constructed library into 3 parts; One of the constructed libraries was hybridized and captured according to the method described in Example 4 to obtain hybridization capture sequencing samples (hybridization capture method based on Y-type double-stranded X-96 UMI adapter); one of the constructed libraries was divided into two equal parts and amplicon sequencing samples were obtained according to the method described in the second step (2) of Example 6, SaferSeqS method; the last constructed library was ligated according to the hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) of this invention, i.e., Y-type double-stranded X-96 UMI adapter, and then enriched by linear pre-amplification + strand-specific nested PCR. The specific method is as follows: (1) Using a DNA library with Y-type double-stranded X-96 UMI adapter as a template (prepared according to the method in Example 3), the target region was linearly pre-amplified using one of the primers in Table 2, SEQ ID NO:01~SEQ ID NO:06. The amplification system is shown in Table 25, and the amplification program is shown in Table 26.
[0100] Table 25 Amplification System Table 26 Amplification Procedure (2) The obtained linear pre-amplified product was divided into two parts and amplified by Watson and Crick chains respectively. Watson chain reaction: amplification was performed using primers shown in Table 2 (SEQ ID NO:07~SEQ ID NO:12) and SEQ ID NO:14 respectively. Crick chain reaction: amplification was performed using primers shown in Table 2 (SEQ ID NO:07~SEQ ID NO:12) and SEQ ID NO:13 respectively. The amplification program is shown in Table 27 and Table 28. This step realizes the physical separation of complementary double strands, laying the foundation for subsequent double-strand consensus analysis.
[0101] Table 27 Amplification System Table 28 Amplification Procedure (3) The amplification products from the previous round were subjected to further nested PCR, with amplification of the Watson and Crick chains respectively. Watson chain reaction: primers shown in Table 2 (SEQ ID NO:15, SEQ ID NO:17, SEQ ID NO:19, SEQ ID NO:21, SEQ ID NO:23, SEQ ID NO:25) were reacted with primers shown in SEQ ID NO:28 respectively. Crick chain reaction: primers shown in Table 2 (SEQ ID NO:16, SEQ ID NO:18, SEQ ID NO:20, SEQ ID NO:22, SEQ ID NO:24, SEQ ID NO:26) were reacted with primers shown in SEQ ID NO:27 respectively. The amplification system is shown in Table 29. The amplification program is shown in Table 30. This step further enriches the target region and adds complete adapters to the molecule to meet the Illumina sequencing conditions.
[0102] Table 29 Amplification System Table 30 Amplification Procedure After PCR, the product was purified using 0.8×AMP magnetic beads to obtain the final sequencing library.
[0103] The final PCR products were subjected to paired-end sequencing on the Illumina platform. Quality control filtering was performed based on the X96 fixed set characteristics: reads whose first 8 bases did not belong to the X96 set, and reads whose 9th base was not T, were removed. Molecular families were grouped based on the composite molecular identifier formed by the UMI + 8bp endpoint sequence, and a double-stranded consensus sequence was constructed, retaining only variations that were consistent in both the Watson and Crick strand families.
[0104] Technical effects: Regarding molecular recovery rate, in urine ctDNA ( Figure 8 The average molecule recovery rate of the hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention is 43.18% (n=49, range 23.47%~75.70%), compared to only 1.232% (n=35, range 0.01%~2.68%) for the hybridization-capture-based Y-type double-stranded X-96 UMI adapter method and 11.59% (n=8, range 2.27%~21.05%) for the SaferSeqS method. The hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) achieves approximately 35-fold improvement over the hybridization-capture-based Y-type double-stranded X-96 UMI adapter method and approximately 3.7-fold improvement over the SaferSeqS method. In fresh tumor tissue ( Figure 8 The average recovery rate of the hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention is 58.46% (n=27, range 27.0%~81.19%), compared to 6.06% (n=24) for the hybridization-capture-based Y-type double-stranded X-96 UMI adapter method and 18.85% (n=3) for the SaferSeqS method. The hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention represents an improvement of approximately 9.6 times over the hybridization-capture-based Y-type double-stranded X-96 UMI adapter method.
[0105] Regarding detection sensitivity ( Figure 8 (B and D) 10 ng of human genomic DNA corresponds to approximately 3,030 genomic equivalents, meaning that theoretically there are approximately 3,030 raw DNA molecules available for detection at each target site. Under ideal conditions (100% molecule recovery), the highest sensitivity is 1 / 3,030 ≈ 3.3 × 10⁻⁶. -4 In actual testing, the number of molecules that can be recovered directly determines the level of sensitivity that can be achieved. In urine ctDNA ( Figure 8The hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention recovered an average of approximately 1,308 DCS molecules from 3,030 input molecules (recovery rate 43.18%), with a maximum recovery of approximately 2,293 molecules, close to the theoretical input amount. This means that the hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention can detect an average of approximately 1 / 1,308 ≈ 7.6 × 10⁻⁶ molecules. -4 The low-frequency mutations have a maximum sensitivity of 1 / 2, 293≈4.36×10 -4 This is close to the theoretical upper limit of 10 ng input. In contrast, the SaferSeqS method recovered approximately 351 DCS molecules from the same 3,030 input molecules (recovery rate 11.59%), with sensitivity limited to approximately 1 / 351 ≈ 2.85 × 10⁻⁶. - ³; X96-DS recovered only about 37 DCS molecules (recovery rate 1.232%), and its sensitivity was limited to approximately 1 / 37 ≈ 2.7 × 10³. -2 In fresh tumor tissue ( Figure 8 According to the present invention, the hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) recovers an average of approximately 1,771 DCS molecules (average recovery rate of 58.46% (27%-80.11%), with a maximum recovery of approximately 2,427 molecules, and an average sensitivity of 1 / 1,771 ≈ 5.6 × 10⁻⁶). -4 The highest sensitivity reaches 1 / 2427 ≈ 4.12 × 10⁻⁶. -4 The SaferSeqS method recovered approximately 571 molecules (recovery rate 18.85%), with sensitivity limited to approximately 1 / 571 ≈ 1.75 × 10⁻⁶. - ³; The hybridization-capture-based Y-type double-stranded X-96 UMI linker method recovered only about 183 molecules (recovery rate 6.06%), with sensitivity limited to approximately 1 / 183 ≈ 5.46 × 10³. - ³.
[0106] The above results demonstrate that the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention can successfully convert approximately 43% (urine ctDNA) to 58% (fresh tissue) of the original input molecules into double-stranded consensus sequences with an extremely low input volume of only 10 ng. The maximum DCS depth can reach 2,293 × (urine ctDNA) and 2,427 × (fresh tissue), and the detection sensitivity is close to the theoretical upper limit of 10 ng input volume (1 / 3,030 ≈ 3.3 × 10⁻⁶). -4 The invention describes a molecular recovery rate based on hierarchical targeted enrichment-based double-stranded consensus sequencing (LinDS), regardless of whether the tissue is fresh (…). Figure 8 In (A and B) or in ctDNA ( Figure 8 (C and D) and the lowest detectable frequency ( Figure 8 Both B and D yielded significantly higher results than the hybridization-capture-based Y-type double-stranded X-96UMI adapter method and the SaferSeqS method. This high molecular recovery rate enables the invention's graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) to reliably detect low-frequency mutations and dynamically monitor MRDs in minute clinical samples such as urine ctDNA, overcoming the dependence of traditional double-stranded sequencing on high DNA input volumes.
[0107] In addition, the invention's graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) reduces economic and time costs. Compared to hybridization capture, the enrichment method for the target gene changes from hybridization capture enrichment to PCR, shortening the operation time from 3 days to 1 day. Figure 12 Compared to patent CN201880020286 (Method for Targeted Nucleic Acid Sequence Enrichment and its Application in Error-Correcting Nucleic Acid Sequencing), this method does not rely on primer site disruption, biotin affinity enrichment, or adapter methylation modification. It simplifies the workflow while significantly improving the conversion efficiency of target molecules to double-stranded consensus sequences. The double-stranded consensus sequencing method (LinDS) based on hierarchical targeted enrichment described in this invention reduces economic and time costs and is easier to implement.
[0108] Example 9 Dedicated probe assemblies for dynamic monitoring of EGFR-targeted drug resistance mutations and their applications A set of gene-specific primers was designed specifically for monitoring EGFR-targeted drug resistance, covering the following sites: Exon19 deletion mutation (19del), Exon21 L858R point mutation, Exon20 T790M point mutation, and Exon20 C797S. For each site, a linear preamplification primer (LP), a first gene-specific primer (GSP1), and a second gene-specific primer (GSP2) were designed, with their sequences and binding sites nested sequentially inwards (the linear preamplification primer being the outermost). The sequences of each primer are shown in Sequence Listing 2.
[0109] It is worth noting that in the EGFR targeting primer set used in this embodiment, (LP primer SEQ ID NO: 01 + GSP1 primer SEQ ID NO: 07 + GSP2 Watson chain primer SEQ ID NO: 15, Crick chain primer SEQ ID NO: 16) covers EGFR Exon 19; (LP primer SEQ ID NO: 03 + GSP1 primer SEQ ID NO: 09 + GSP2 Watson chain primer SEQ ID NO: 19, Crick chain primer SEQ ID NO: 20) covers Exon 20; and (LP primer SEQ ID NO: 02 + GSP1 primer SEQ ID NO: 08 + GSP2 Watson chain primer SEQ ID NO: 17, Crick chain primer SEQ ID NO: 18) covers Exon 21 and other core decision regions for NSCLC clinical targeted therapy. Exon 19: Covers the Exon 19 deletion hotspot region (represented by p.E746_A750del). The in-frame deletion mutation represented by p.E746_A750del (ΔELREA) is a core sensitive mutation of first-generation EGFR-TKIs such as gefitinib, erlotinib, and icotinib; second-generation EGFR-TKIs such as afatinib and dacomitinib; and third-generation EGFR-TKIs such as osimertinib, ametinib, and vormetinib. Exon 20: T790M mutation is the most common acquired resistance mutation after first- and second-generation EGFR-TKI treatment, and it is also a predictive / indicative site for the efficacy of third-generation EGFR-TKIs (osimertinib, ametinib, vormetinib); C797S mutation is a key acquired resistance site after third-generation EGFR-TKI treatment.
[0110] The Exon 21: L858R point mutation is one of the two classic EGFR activating mutations, along with the Exon 19 deletion, and is also a core sensitive mutation of the first to third generation EGFR-TKIs mentioned above.
[0111] The above-mentioned sites have all been listed as essential sites for NSCLC targeted therapy decision-making (including first-line drug selection, drug resistance mechanism identification and subsequent regimen adjustment) by authoritative domestic and international guidelines such as NCCN and CSCO.
[0112] Following the procedure in Example 8, urinary ctDNA was detected, with an average molecular recovery rate of approximately 55.47% (n=18, range 33.46%-75.7%). See [link to Example 8]. Figure 9 .
[0113] To evaluate the detection capability of the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) for low-frequency mutations at clinically targeted EGFR sites in a liquid biopsy scenario, a clinical scenario of detecting tumor mutations from blood was simulated. Fresh tumor tissue DNA and paired normal blood DNA were taken from the same lung cancer patient. The tumor DNA was mixed into the normal blood DNA in a gradient of 50% (AB1), 10% (AB2), 1% (AB3), and 0.1% (AB4), with a total input of 10 ng. The library was constructed using steps (2) to (5) of Example 3 and the method described in Example 8. Since the target gene detected in this example is EGFRExon19, the primers used in the library construction process are the corresponding primers shown in Table 2.
[0114] Double-stranded consensus sequences were constructed after sequencing using the Illumina platform. Two independent technical replicates were performed on tumor tissue samples. The mutation frequencies of EGFR Exon 19 E746-A750del (chr7:55174771, c.2235_2249del) were 0.0845 and 0.0685, respectively (mean 0.0765, not detectable in normal blood). The correlation R² between the two replicates for the mutation frequencies of all detected sites was 0.99. Figure 10 (D), demonstrating the reproducibility of Lin-DS detection. The tumor carries a 15bp in-frame deletion (c.2235_2249del, p.Glu746_Ala750del, i.e., the classic E746-A750del / ΔELREA) in EGFR Exon 19 (chr7:55174771), with a mutation frequency of 0.0765 in tumor tissue (0.0845 and 0.0685 in two technical replicates, respectively), and this mutation is not present in normal blood.
[0115] E746-A750del is the most common EGFR Exon 19 activating mutation in non-small cell lung cancer (NSCLC), accounting for approximately 60–70% of all Exon 19 deletions. It is a core predictive biomarker for the efficacy of first- to third-generation EGFR-TKIs such as gefitinib, erlotinib, afatinib, dacomitinib, and osimertinib, and has been listed as a mandatory site for first-line targeted therapy of NSCLC by authoritative domestic and international guidelines such as NCCN and CSCO.
[0116] Using this mutation as a tracking target, the invention's graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) was evaluated to assess whether it could reliably detect the mutation from 10 ng of mixed DNA under different tumor DNA incorporation ratios.
[0117] Technical effects: A 10 ng DNA input corresponds to approximately 3,030 genomic equivalents. At various gradients, the DCS depth of this target site ranged from 1,640 to 1,791×. The detection results are as follows: AB1 (50% tumor DNA incorporation): Expected mutation frequency 0.03825, expected number of mutant molecules approximately 63. Observed frequency 0.0488, Obs / Exp=1.28, accurate quantification, successfully detected. Figure 10 ).
[0118] AB2 (10% tumor DNA incorporation): Expected mutation frequency 0.00765, expected number of mutant molecules approximately 13. Observed frequency 0.00796, Obs / Exp=1.04, highly accurate quantification, successfully detected.
[0119] AB3 (1% tumor DNA incorporation): Expected mutation frequency 0.000765, expected number of mutant molecules approximately 1.4. Observed frequency 0.00223, Obs / Exp = 2.92. Successful detection even with an expected number of only about one mutant molecule demonstrates near-single-molecule detection capability.
[0120] AB4 (0.1% tumor DNA incorporation): Expected mutation frequency 0.0000765, expected number of mutated molecules approximately 0.13 (less than 1 molecule). Not detected, in line with expectations.
[0121] The above results demonstrate that the hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention, with an input of only 10 ng DNA, can achieve sequencing speeds down to 7.65 × 10⁻⁶ DNA molecules. -4 Tumor-specific mutations with a mutation frequency (1% tumor DNA incorporation) can still be successfully detected and accurately quantified at the 10% incorporation level, demonstrating near-single-molecule-level detection capabilities. Given that the E746-A750del used in this validation is a core decision site for EGFR-TKI targeted therapy in NSCLC, the above detection performance can directly support key clinical scenarios such as postoperative MRD monitoring in early-stage lung cancer, early warning of TKI resistance in advanced patients, and plasma genotyping when tissue samples are unavailable.
[0122] In addition, the denoising and low-frequency mutation fidelity of SSCS (single-stranded shared sequence) and DCS (double-stranded shared sequence) were compared. Figure 11 ); Figure 11 In the case of SSCS single-chain consensus correction, a large amount of background noise signals are widely distributed, which masks the ultra-low frequency real mutation signals in 0.1% of the samples. Figure 11 After DCS dual-chain consensus correction, background noise is effectively removed, and the real mutation signals in 50% and 1% of the samples are completely preserved, achieving the dual goals of "denoising" and "fidelity preservation".
[0123] Example 10 Application verification of the hybridization capture method based on Y-type double-stranded X-96 UMI linker described in this invention in clinical sample detection and neoadjuvant therapy efficacy monitoring in bladder cancer. To verify the effectiveness of the hybridization capture method based on the Y-type double-stranded X-96 UMI linker in the clinical detection of bladder cancer, this embodiment tested urine, blood, and tumor tissue samples from bladder cancer patients. Furthermore, dynamic monitoring of ctDNA was performed on bladder cancer patients receiving neoadjuvant therapy before, during, and after treatment. The test results were compared with imaging examinations and postoperative pathological results.
[0124] 1. Detection of tumor-specific mutations in clinical bladder cancer samples using the Y-type double-stranded X-96 UMI linker. The samples used in this embodiment include urine samples, blood samples, and tumor tissue DNA samples from bladder cancer patients. Adapters were prepared according to the method described in Example 2, libraries were constructed according to the method described in Example 3, and hybridization capture and sequencing analysis were performed according to the method described in Example 4 to obtain double-stranded identical sequences, i.e., DCS.
[0125] The test results showed that the Y-type double-stranded X-96 UMI adapter can effectively capture tumor-specific mutations in urinary ctDNA. Figure 13 Further comparison of the mutation frequencies of corresponding mutations in urinary ctDNA and tumor tissue DNA showed a correlation, suggesting that urinary ctDNA can reflect the mutation characteristics in tumor tissue to a certain extent.
[0126] Meanwhile, the mutation frequency distribution in tumor tissue is broad, covering both high-frequency driver mutations and low-frequency subclonal mutations. The Y-type double-stranded X-96 UMI linker has a high detection capability for low-frequency mutations, which can improve the detection sensitivity of low-abundance tumor-specific mutations in bladder cancer urine and blood ctDNA.
[0127] Figure 13 In the middle A, the frequency distribution of all mutations in the tumor tissue is shown. The vertical axis represents the mutation frequency (logarithmic scale). Each point represents a mutation. The mutation frequency in the tumor spans four orders of magnitude (0.0001-1), including high-frequency driver mutations and low-frequency subclonal mutations. Figure 13 B represents the overall detection rate of the three detection methods; Figure 13 C represents the correlation analysis between urinary ctDNA and the mutation frequency in tumor tissue.
[0128] 2. Y-type double-chain X-96 UMI connector for dynamic monitoring of the efficacy of neoadjuvant therapy for bladder cancer. Further, Y-type double-stranded X-96 UMI adapters were used to dynamically monitor ctDNA in bladder cancer patients receiving neoadjuvant therapy. Peripheral blood samples were collected from patients before treatment, during each treatment cycle, and after treatment to detect the mutational allele frequency (AF) of tumor-specific mutations, and the sum of AFs was calculated as a molecular tumor burden indicator.
[0129] The results showed that in NAC1 patients with a good treatment response, the pre-treatment absorptivity (AF) was 0.562, which decreased to 0.00047 after treatment, a reduction of approximately 99.9%; after the first cycle of treatment, AF decreased to 0.134, a reduction of approximately 76.2%. Imaging examinations showed that the patient's tumor shrank significantly from 7.2×3.7×4.1cm, and the pathological stage decreased from T4aN0M0 to ypT1, indicating a good response to neoadjuvant therapy.
[0130] In contrast, NAC2 patients with poor treatment response had an absorptivity (AF) level of 0.68 before treatment, which decreased to 0.43 after treatment, a reduction of approximately 36.8%; after the first cycle of treatment, AF only decreased to 0.64, a reduction of approximately 5.9%. Dynamic monitoring showed that while AF decreased during treatment, it rebounded during treatment intervals, suggesting that the tumor clone was not continuously and effectively eliminated. Imaging examinations showed that the tumor only shrank from 2.7 × 2.3 cm to 2.7 × 1.6 cm, and the pathological stage decreased from T3bN1M0 to pT3bN0M0, indicating limited downstaging.
[0131] Furthermore, the Y-type double-stranded X-96 UMI adapter can detect low-frequency residual mutations with an AF < 0.01. For example, residual mutations in the major clone can still be detected in NAC1 patients after treatment, with an AF of 0.00047, which is lower than the common detection threshold of conventional next-generation sequencing technology. This indicates that the Y-type double-stranded X-96 UMI adapter can identify tiny residual lesions that may be missed by conventional detection methods.
[0132] The above results indicate that the decrease in total ctDNA AF, changes in early treatment, and residual levels after treatment are consistent with the degree of tumor shrinkage on imaging and the results of postoperative pathological downstaging. These results demonstrate that the Y-type double-stranded X-96 UMI linker can be used for dynamic assessment, early prediction, and identification of minimal residual disease in neoadjuvant therapy for bladder cancer.
[0133] Figure 14The figures show the dynamic monitoring results of ctDNA in neoadjuvant therapy patients. Figure A represents the sampling time point for neoadjuvant therapy patients. Figure B shows the dynamic monitoring results of tumor major efficiency mutations in ctDNA in two NAC1 MIBC patients before and after three treatment cycles. Figure C shows the dynamic monitoring results of tumor major efficiency mutations in ctDNA in two NAC2 MIBC patients before and after three treatment cycles. The blue background indicates the treatment period. The red arrows indicate the clonal proliferation rebound after the emergence of drug resistance. The purple dashed line shows the minimum detection threshold of conventional next-generation sequencing.
[0134] Example 11 Application verification of the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in the invention for dynamic tracking of TP53 mutations in urinary ctDNA. To verify the effectiveness of the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention in the detection of ctDNA in low-input urine, this embodiment dynamically monitors TP53 mutations in ctDNA in urine during the treatment of bladder cancer patients, and compares the results with the hybridization capture detection results based on the Y-type double-stranded X-96 UMI adapter and the results of clinical imaging assessment.
[0135] The samples used in this embodiment included urine, blood, and tumor tissue DNA from bladder cancer patients. The LinDS (Limited Injection Sequencing) method based on graded targeted enrichment was used for adapter and probe preparation according to Examples 1-2, library construction according to Example 3, target gene enrichment according to Example 8, and analysis according to Example 8 to obtain DCS (Digital Candidate Sequencing Scheme). The hybridization capture method based on Y-type double-stranded X-96 UMI adapters was used for adapter preparation according to Example 2, library construction according to Example 3, target gene enrichment according to Example 4, and analysis according to Example 4.
[0136] The results show that ( Figure 15 The TP53 mutation frequency detected by the hierarchical targeted enrichment-based double-stranded consensus sequencing method (LinDS) of this invention at a DNA input of 10 ng was highly consistent with the detection results of the hybridization capture method based on Y-type double-stranded X-96 UMI adapter at a DNA input of 200 ng. Figure 15 This indicates that the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in the invention has a high molecular recovery rate and can achieve stable detection under low sample input conditions. Further dynamic monitoring showed that as treatment progressed, the frequency of TP53 mutations in urinary ctDNA gradually decreased, consistent with the tumor shrinkage trend shown in imaging.
[0137] The above results indicate that the graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) described in this invention can achieve real-time dynamic evaluation of bladder cancer treatment efficacy through non-invasive urine ctDNA detection under extremely low DNA input conditions, and is especially suitable for clinical samples with limited ctDNA content.
[0138] Figure 15 LinDS is used for dynamic tracking of TP53 mutations in urinary ctDNA. Figure A shows the sampling time; Figure B shows the gradual decrease in TP53 mutation frequency as treatment progresses, and the detection results of the invention's graded targeted enrichment-based double-stranded consensus sequencing method (LinDS) are consistent with those of the hybridization capture method based on the Y-type double-stranded X-96 UMI adapter.
[0139] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A double-stranded consensus sequencing method based on hierarchical targeted enrichment, characterized in that: The specific steps are as follows: (1) DNA fragmentation: Extract DNA from the sample to be tested and break the DNA into small fragments of 100-300bp; (2) End repair: The small fragments obtained in step (1) are dephosphorylated and flattened. (3) Adding adapters: Add double-stranded X-96 UMI adapters to both ends of the DNA fragments that have undergone end repair to obtain DNA fragments with adapters, and then purify them with magnetic beads to obtain magnetic bead purified products; (4) The purified magnetic bead product was subjected to PCR amplification to obtain the initial library construction product; (5) Using the initial library construction product obtained in step (4) as a template, linear pre-amplification is performed using LP primers, which are specific to a single target gene, before the tubes are separated. (6) Divide the library product obtained in step (5) into two parts and perform the first round of amplification of the Watson chain and Crick chain respectively. The Watson chain is amplified using gene-specific GSP1 primers and adapter sequence-specific universal amplification primers P7; the Crick chain is amplified using gene-specific GSP1 primers and adapter sequence-specific universal amplification primers P5, to obtain the Watson chain and Crick chain amplified in the first round respectively; the GSP1 is nested inside the LP binding site; (7) The Watson and Crick chains obtained in the first round of amplification obtained in step (6) are subjected to a second round of Watson and Crick chain amplification, respectively. The second round of Watson chain amplification is performed using GSP2 Watson primers and universal primer Index-P7; the second round of Crick chain amplification is performed using GSP2 Crick primers and universal primer Index-P5, respectively, to obtain the second round of amplified Watson and Crick chains; the GSP2 is nested inside the GSP1 binding site.
2. The double-stranded consensus sequencing method based on hierarchical targeted enrichment according to claim 1, characterized in that: The double-stranded X-96 UMI adapter mentioned in step (3) is a Y-type double-stranded adapter. One strand of the double-stranded X-96 UMI adapter has the structure of: frame 1 sequence - a sequence of the X-96 UMI set, and the other strand has the structure of: a sequence of the X-96 UMI set - T - frame 2 sequence, where the sequence of frame 1 is: ACACTCTTTCCCTACACGACGCTCTTCCGATC; the sequence of frame 2 is: GATCGGAAGAGCACACGTCTGAACTCCAGTCAC; the X-96 UMI set consists of 96 nucleic acid sequences, each with a length of 8 bases, and the nucleic acid sequences are: TTCGTCCA, TTCCGAGT, TGCTTAAC, TTGTCATG, TGTGATAA, TGCTATGT, TCTTAGAC, TTAGATCG, TGACATCA, TGCAGGCT, TCTCGATC, TGGCTAAG, TCTAGCCA, TGAATTAT, GTACATCC, TGAACGAG, TAGTGCCA, TATGCGGT, GGTGCAAC, TCGTAAGG, GTCGAGAA, TATAGATT, GGTCCTCC, TCGATGAG, GGATACTA, TAGCTGTT, GACTAGGC, TCCGGATG, GGAACATA, GTCTGGAT, GACACTAC, TATGTCCG, GCGTATCA, GTCCACAT, CTTAGGCC, GGCATCCG, GCCAATGA, GCAGCATT, CTGTAGCC, GGCAATTG, GATTCGGA, GACGAGTT, CGATAATC, GCTGGCTG, GATCTAGA, GAAGCGAT, CCTGCCTC, GCGCGGAG, GAGACGTA, CGGCCATT, CCGCTTCC, GCACCTTG, GAAGGCGA, CGAGTCCT, CCAAGAGC, CTTCATCG, CTTAGCGA, CACTGCTT, CAGACATC, CTCCTAAG, CGACTTGA, ATGACCAT, CACCGTTC, CTACCTAG, CGACGCCA, AGATTGAT, ATTGTCTC, CCGAGACG, CCTGTGCA, ACTTGGAT, ATGGACAC, CCATTCCG, CAGCCAGA, ACTATCGT, ATGCGCTC, CAGCAATG, ATTGCAGA, ACGGATGT, ATATGACC, CAATTGTG, ATGGCTTA, ACCTCGGT, ATAAGGAC, CAACTTCG, ATCTCTGA, AAGGCCTT, AGTAGTGC, ATTATCCG, AGAGACGA, AACTTACT, AGGACAGC, AATTATAG, AAGAGCGA, AACGTGGT, ACCTCTAC, AACATGCG。 3. The double-stranded consensus sequencing method based on hierarchical targeted enrichment according to claim 1, characterized in that: The sequence of the primer P7 is GTGACTGGAGTTCAGACGTGTGCTCTTCCGATC; the sequence of the primer P5 is ACACTCTTTCCCTACACGACGCTCTTCCGATCT.
4. The hierarchical targeted enrichment-based double-stranded consensus sequencing method according to claim 1, characterized in that: The sequence of the universal primer Index-P7 is CAAGCAGAAGACGGCATACGAGAT-Index sequence-GTGACTGGAGTTCAGACGTGT; the sequence of the universal primer Index-P5 is AATGATACGGCGACCACCGAGATCTACAC-Index sequence-ACACTCTTTCCCTACACGAC.
5. The double-stranded consensus sequencing method based on hierarchical targeted enrichment according to claim 1, characterized in that: LP primers, GSP1 primers, GSP2 Watson primers, and GSP2 Crick primers were designed and screened based on the target region to be detected, and were designed according to the following principles: A. Design candidate primers within a range of 50–200 bp upstream of the target detection region, with a candidate primer length of 16–25 bp; B. Avoid the formation of obvious primer dimers and hairpin structures in the primer sequence; C. The annealing temperatures of the candidate primers are kept similar; D. The GC content in the primer sequence is controlled between 40% and 60%; E. Perform whole-genome alignment of candidate primer sequences and select primers that specifically match the target region and have no significant homology with other non-target regions.
6. A specific reagent combination for use in the method of claim 1, characterized in that: The specialized reagent kit includes a double-stranded X-96 UMI adapter, LP primers, GSP1 primers, GSP2 Watson primers, and GSP2 Crick primers.
7. The special reagent combination according to claim 6, characterized in that: The specialized reagent set also includes universal amplification primers P7, P5, Index-P7, and Index-P5.
8. The application of the special reagent combination according to claim 6 in the detection of EGFR gene mutations, characterized in that: When detecting EGFR gene Exon19 deletion mutations, the LP primer sequence in the dedicated reagent kit is GTGTCCCTCACCTTCGG; the GSP1 primer sequence is GTGTCCCTCACCTTCGG; the GSP2 Watson primer sequence is AATGATACGGCGACCACCGAGATCTACACAGACTCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTGTGCATCGCTGGTAAC; and the GSP2 Crick primer sequence is CAAGCAGAAGACGGCATACGAGATTGGTGGAAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGTGCATCGCTGGTAAC. When detecting the Exon21 L858R point mutation in the EGFR gene, the sequence of the LP primer in the special reagent kit is GGATCAGTAGTCACTAAC; the sequence of the GSP1 primer is CTAACGTTCGCCAGCCAT; the sequence of the GSP2 Watson primer is AATGATACGGCGACCACCGAGATCTACACAGACTCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTCAGCCATAAGTCCTCGACGT; and the sequence of the GSP2 Crick primer is CAAGCAGAAGACGGCATACGAGATTGGTGGAAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTCAGCCATAAGTCCTCGACGT. When detecting the Exon20 T790M point mutation in the EGFR gene, the LP primer sequence in the dedicated reagent kit is TCCAGGAAGCCTACGTGATG; the GSP1 primer sequence is TGATGGCCAGCGTGGAC; the GSP2 Watson primer sequence is AATGATACGGCGACCACCGAGATCTACACAGACTCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTGCGTGGACAACCCCCAC; and the GSP2 Crick primer sequence is CAAGCAGAAGACGGCATACGAGATTGGTGGAAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGCGTGGACAACCCCCAC. When detecting the Exon20 C797S point mutation in the EGFR gene, the LP primer sequence in the dedicated reagent kit is TCCAGGAAGCCTACGTGATG; the GSP1 primer sequence is TGATGGCCAGCGTGGAC; the GSP2 Watson primer sequence is AATGATACGGCGACCACCGAGATCTACACAGACTCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTGCGTGGACAACCCCCAC; and the GSP2 Crick primer sequence is CAAGCAGAAGACGGCATACGAGATTGGTGGAAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGCGTGGACAACCCCCAC.
9. The application of the special reagent combination according to claim 6 in the detection of TP53 gene mutations, characterized in that: When detecting Exon4 deletion or mutation in the TP53 gene, the LP primer sequence in the special reagent kit is GCAGCCTCTGGCATTCT; the GSP1 primer sequence is GCAGCCTCTGGCATTCT. The sequences of the GSP2 Watson primers are: AATGATACGGCGACCACCGAGATCTACACAGACTCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTGGCATTCTGGGAGCTTCAT; the sequences of the GSP2 Crick primers are: CAAGCAGAAGACGGCATACGAGATTGGTGGAAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGGCATTCTGGGAGCTTCAT. When detecting Exon8 deletion or mutation in the TP53 gene, the LP primer sequence in the dedicated reagent kit is AACTGCACCCTTGGTCTC; the GSP1 primer sequence is CTTGGTCTCCTCCACC; the GSP2 Watson primer sequence is AATGATACGGCGACCACCGAGATCTACACAGACTCCTACACTCTTTCCCTACACGACGCTCTTCCGATCTCCTCCACCGCTTCTTGTC; and the GSP2 Crick primer sequence is CAAGCAGAAGACGGCATACGAGATTGGTGGAAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTCCTCCACCGCTTCTTGTC.
Citation Information
Patent Citations
Methods for targeted nucleic acid sequence enrichment with applications to error corrected nucleic acid sequencing
CN110520542A
Method for enrichment of targeted nucleic acid sequences and application in error-corrected nucleic acid sequencing
CN110520542B
Duplex UMI adapter and sequencing method
CN113913495B