A pancreatic cancer early screening kit based on structural variation chimera in peripheral blood free DNA fragments
Patent Information
- Application Number
- CN202611255790.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-25
AI Technical Summary
[0018]与现有技术相比,本发明具有如下有益效果:本发明通过将“双末端锚定探针捕获”与“选择性环化-滚环扩增”相结合,首次实现了在cfDNA中无需预先知道断点序列即可高灵敏度、高特异性地检测结构变异嵌合体。
Smart Images

Figure CN122811369A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of genetic engineering technology, and in particular relates to a pancreatic cancer early screening kit based on chimeras with structural variations in cell-free DNA fragments in peripheral blood. Background Technology
[0002] Cell-free DNA (cfDNA) is cell-free DNA fragments found in human peripheral blood plasma, primarily ranging in length from 150 to 200 bp, originating from apoptotic or necrotic cells. In cancer patients, approximately 0.1%–10% of cfDNA originates from tumor cells and is termed circulating tumor DNA (ctDNA). Liquid biopsy based on cfDNA has become an important tool for early cancer screening, molecular subtyping, treatment monitoring, and prognostic assessment due to its advantages such as being non-invasive, repeatable, and capable of reflecting tumor heterogeneity.
[0003] Currently, cfDNA detection technology mainly focuses on two types of molecular markers: (1) point mutations (such as driver mutations in genes like KRAS, EGFR, and TP53), detected by digital PCR (ddPCR) or targeted sequencing panels; and (2) changes in DNA methylation patterns, detected by sequencing after bisulfite conversion. However, neither of these two types of markers can directly reflect information on genomic structural variations (SVs).
[0004] Structural variations, including translocations, inversions, deletions, duplications, chromosome fragmentation, and breakage-fusion-bridge cycles, are important driving events in tumorigenesis and development. Chimeric DNA fragments generated by structural variations carry junction sequences of two or more different genomic regions, exhibiting tumor specificity and making them ideal biopsy markers. However, detecting chimeric structural variations in cfDNA faces three major technical bottlenecks: (a) cfDNA fragments are extremely short (150-200 bp), with chimeric junctions spanning only a few to tens of bp, making it difficult for conventional whole-genome sequencing (WGS) reads to cover them; (b) the abundance of chimeras in cfDNA is extremely low (VAF often <1%), with signals drowned out by background noise at conventional sequencing depths; and (c) the breakpoint locations of structural variations are highly heterogeneous, with significant differences in breakpoint locations between different patients and even different subclones of the same tumor, making it impossible to cover unknown structural variations using fixed panels or PCR primers designed based on known breakpoints.
[0005] Existing structural variation detection technologies mainly include fluorescence in situ hybridization (FISH) and immunohistochemistry (IHC), PCR / qPCR / ddPCR detection based on known fusion designs, targeted capture sequencing, whole-genome sequencing (WGS) and whole-exome sequencing (WES), and circularization-mediated amplification techniques. However, all of these methods have significant limitations. For example, FISH and immunohistochemistry are only applicable to tissue samples and cannot be used for liquid biopsies; their resolution is also limited, making it impossible to accurately determine breakpoint sequences. PCR / qPCR / ddPCR detection based on known fusion designs includes methods such as RT-PCR detection of EML4-ALKV1 / V2 / V3 and qPCR detection of BCR-ABL1. These methods have high sensitivity (ddPCR detection limit can reach 0.01% VAF), but require prior knowledge of the fusion partner combination and the precise breakpoint location, and are ineffective for novel or rare fusions. Targeted capture sequencing includes methods such as RNA-based fusion gene detection panels (e.g., Archer FusionPlex) or DNA breakpoint capture panels. RNA-based methods are highly susceptible to sample quality and RNA degradation, making them unsuitable for plasma cfDNA. DNA-based methods typically employ a "whole exon + intron" coverage strategy, requiring a massive number of probes (>100,000), resulting in high costs, and still primarily cover known fusion regions. Whole genome sequencing and whole exon sequencing: WGS can detect unknown structural variations, but cfDNA samples require extremely high sequencing depths (>100×, even >1000×) to detect low-frequency chimeras, leading to extremely high costs. WES does not cover intron regions, and since most fusion breakpoints are located within introns, it is unsuitable. Circulation-mediated amplification techniques: such as CIRCLE-seq and Padlock probe circularization. These methods amplify signals by circularizing linear DNA followed by rolling circle amplification (RCA), but current techniques usually require prior knowledge of the breakpoint sequence to design Padlock probes, and lack selectivity for circularization of linear cfDNA, resulting in high background noise and insufficient specificity.
[0006] Therefore, how to detect chimeras with genomic structural variations in cfDNA samples with extremely low abundance and high background noise has become a fundamental problem that urgently needs to be solved in this field. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to provide a pancreatic cancer early screening kit based on structural variation chimeras in cell-free DNA fragments in peripheral blood.
[0008] To achieve the above-mentioned objectives, the present invention provides the following technical solution: This invention provides a method for detecting structurally variant chimeric fragments in cell-free DNA from peripheral blood, comprising the following steps: (1) Extract cfDNA from the sample; (2) Perform probe hybridization on cfDNA to obtain hybridization products; (3) Add magnetic beads to the hybridization product, incubate to allow the hybridization product to bind with the magnetic beads, wash, elute and purify to obtain enriched DNA; (4) Add end-repair and A-tailing premixed enzyme to the enriched DNA, incubate, purify, and obtain end-repaired DNA; (5) Dilute the end-repaired DNA, perform circularization, and purify it to obtain purified circularized DNA; (6) Add end repair and A-tailing premixed enzyme to the circularized DNA, incubate, add adaptor mixture, incubate, and obtain the circularized product; (7) Perform rolling circle amplification on the cyclized product to obtain the PCA product; (8) Perform PCR amplification on the PCA product to obtain the PCR product; (9) Purify the PCR product to obtain the target library, quantify the target library and analyze its fragment distribution to obtain libraries with different indices; (10) According to the requirements of the sequencing platform, the libraries of different indices are mixed in equal molar amounts and paired-end 150bp sequencing is performed. The analysis is carried out through bioinformatics process to identify the chimeric fragments.
[0009] Preferably, the probe is designed based on knowledge of human genomics and predefined tens of thousands of "potential rearrangement hotspots", wherein the 5' end of each probe is labeled with biotin and the probe length is 80-120 nt.
[0010] Preferably, the incubation in step (4) is performed at 20°C for 30 minutes, followed by incubation at 65°C for 30 minutes.
[0011] This invention provides a kit for detecting chimeric fragments with structural variations in cell-free DNA from peripheral blood, comprising a cfDNA extraction and purification module, a chimeric targeted enrichment module, a circularization transformation and amplification module, a sequencing library preparation module, and a quality control module.
[0012] Preferably, the cfDNA extraction and purification module includes plasma lysis buffer A, binding buffer B, silanol magnetic bead suspension, washing buffer C, elution buffer D, and proteinase K.
[0013] Preferably, the chimeric targeted enrichment module includes a smart anchoring probe panel, a probe reconstitution solution, a 2X hybridization buffer, streptavidin magnetic beads, strict washing solution I, strict washing solution II, and elution buffer E.
[0014] Preferably, the circularization transformation and amplification module includes an end-repair and A-tailing premixed enzyme, a specific adaptor mixture with index, a T4 DNA ligase, a 5X circularization reaction buffer, a Phi29 DNA polymerase and a 10X reaction buffer, a dNTP mixture, and rolling circle primers. The sequence of the rolling circle primer is shown in SEQ ID No. 1.
[0015] Preferably, the sequencing library preparation module includes 2X high-fidelity PCR premix, universal PCR primer mixture, and PCR purification magnetic beads; The universal PCR primer mixture contains a P5 forward primer and a P7 reverse primer. The P5 primer sequence is shown in SEQ ID No. 2, and the P7 primer sequence is shown in SEQ ID No. 3.
[0016] Preferably, the quality control module includes positive control DNA, negative control DNA, DNA quantitative standards, library quantitative qPCR premix, and primers and probes.
[0017] The present invention also provides the application of the kit in the preparation of in vitro diagnostic products for tumor liquid biopsy, early screening, molecular typing, dynamic monitoring or prognostic assessment.
[0018] Compared with the prior art, the present invention has the following beneficial effects: By combining "double-end anchoring probe capture" with "selective circularization-rolling circle amplification", the present invention achieves for the first time the detection of structural variant chimeras in cfDNA with high sensitivity and high specificity without prior knowledge of the breakpoint sequence.
[0019] (1) The dual-end anchoring probe system of the present invention: Instead of targeting known breakpoint sequences, it is based on the distribution patterns of tumor structural variations in databases such as TCGA and ICGC, and designs biotinylated LNA modified probe sets at both ends of potential rearrangement hotspot regions (oncogene introns, tumor suppressor gene break clusters, genomic fragile sites, etc.). When the cfDNA molecule carries both A and B region end sequences simultaneously (i.e., structural variation chimera), both ends are captured simultaneously and enriched by streptavidin magnetic beads; while normal linear cfDNA can only bind to a single probe set and is eluted under strict washing conditions. This strategy breaks through the dependence of fixed panels on known breakpoints and realizes the simultaneous detection of "known + unknown" structural variations.
[0020] (2) Selective circularization and rolling circle amplification of the present invention: After end repair, the enriched chimeric DNA undergoes intramolecular circularization in the presence of extremely low concentrations (<1 nM) and a molecular crowding reagent (PEG 8000) to form covalently closed circular DNA. This circularization step is selective—only blunt-ended DNA after end repair can be circularized efficiently, while normal linear DNA has extremely low circularization efficiency due to the lack of complementary sequences at both ends. The circularized DNA is then subjected to rolling circle amplification (RCA) with Phi29 DNA polymerase, and a single target molecule can produce tandem repeat products of >70 kb, with a signal amplification of approximately 700 times, completely overcoming the limitation of insufficient cfDNA starting amount.
[0021] (3) Dual detection pathway: The kit of the present invention provides a rapid detection pathway based on PCR (qPCR / ddPCR) (no sequencer required, suitable for routine clinical screening) and a precise verification pathway based on high-throughput sequencing (used for the first discovery of unknown linkage points or confirmation of PCR positive results), taking into account both detection efficiency and discovery capability.
[0022] (4) The kit of this invention is a complete, end-to-end biochemical solution designed to specifically capture, enrich, transform, and prepare tumor-derived cfDNA chimeric fragment libraries for high-throughput sequencing from complex human peripheral blood plasma samples. Its core value lies in combining highly innovative probe design with an ingenious molecular biology transformation process, solving the fundamental problem of detecting rare and sequence-unknown genomic structural variant chimeras in cfDNA samples with extremely low abundance and high background noise. As a standalone "wet assay" product, this kit is compatible with commonly used sequencing platforms and bioinformatics analysis services on the market, possessing strong scalability and practicality. Attached Figure Description
[0023] Figure 1 The following are one-dimensional scatter plots of ddPCR for representative VAF gradient samples in Example 4, where A is a one-dimensional scatter plot of ddPCR with 0.1% VAF, B is a one-dimensional scatter plot of ddPCR with 0.05% VAF, C is a one-dimensional scatter plot of ddPCR with 0.02% VAF, and D is a one-dimensional scatter plot of ddPCR for the negative control. Figure 2 This is a bar chart comparing the chimeric recovery rates of the dual-end capture group and the single-end capture group in Example 5. P < 0.0001 (one-way ANOVA, Tukey multiple comparison test; comparison of the two-end capture group with the one-end capture group A and the one-end capture group B respectively), ns indicates that there is no statistically significant difference between the one-end capture group A and the one-end capture group B (P > 0.05). Detailed Implementation
[0024] This invention provides a method for detecting structurally variant chimeric fragments in cell-free DNA from peripheral blood, comprising the following steps: (1) Extract cfDNA from the sample; (2) Perform probe hybridization on cfDNA to obtain hybridization products; (3) Add magnetic beads to the hybridization product, incubate to allow the hybridization product to bind with the magnetic beads, wash, elute and purify to obtain enriched DNA; (4) Add end-repair and A-tailing premixed enzyme to the enriched DNA, incubate, purify, and obtain end-repaired DNA; (5) Dilute the end-repaired DNA, perform circularization, and purify it to obtain purified circularized DNA; (6) Add end repair and A-tailing premixed enzyme to the circularized DNA, incubate, add adaptor mixture, incubate, and obtain the circularized product; (7) Perform rolling circle amplification on the cyclized product to obtain the PCA product; (8) Perform PCR amplification on the PCA product to obtain the PCR product; (9) Purify the PCR product to obtain the target library, quantify the target library and analyze its fragment distribution to obtain libraries with different indices; (10) According to the requirements of the sequencing platform, the libraries of different indices are mixed in equal molar amounts and paired-end 150bp sequencing is performed. The analysis is carried out through bioinformatics process to identify the chimeric fragments.
[0025] In this invention, the probe is designed based on tens of thousands of predefined "potential rearrangement hotspot regions" that cover human genomics knowledge. These hotspot regions are selected based on the distribution patterns of tumor structural variations in the TCGA and ICGC databases, including (a) introns and flanking regions (containing repetitive sequences and fragile sites, easily broken) of known oncogenes (such as ALK, ROS1, KRAS, MYC, EGFR, MET), such as ALK (intron 19-20), ROS1 (intron 31-32), RET (intron 11-12), NTRK1 (intron 8-9), NTRK2 (intron 15-16), NTRK3 (intron 13-14), MET (intron 13-14), NRG1 (intron 1-2), EGFR (intron 24-25), KRAS (intron 2-3), BRAF (intron 7-8), PIK3CA ... (a) Known fusion partner genes: EML4 (intron 13), KIF5B (intron 15), CCDC6 (intron 1), TFG (intron 5), SLC34A2 (intron 4), CD74 (intron 6), TRIM33 (intron 3), KCNQ5 (intron 1); (b) Known break cluster regions of tumor suppressor genes (such as TP53, SMAD4, CDKN2A) are chromosome fragmentation hotspots, such as TP53 (exon 4-8 break cluster), SMAD4 (intron 8-9), CDKN2A (intron 1-2), APC (intron 14-15), RB1 (intron 1-2), etc. 17-18); (c) Fragile sites in the genome (FRA3B corresponds to the FHIT gene, FRA16D corresponds to the WWOX gene) preferentially break under replication pressure, such as FRA3B (chr3: 60.2-60.4 Mb, FHIT gene region) and FRA16D (chr16: 79.6-79.8 Mb, WWOX gene region); (d) Boundaries of highly homologous repeat sequence regions (such as Alu / LINE boundaries, which are prone to MMEJ due to microhomologous sequences), such as Alu-rich regions. (chr2: 29.4-29.6 Mb, dense region of Alu elements upstream of ALK locus), LINE-rich region (chr6: 117.2-117.4 Mb); (e) Boundaries of genomic segments commonly amplified in tumors (e.g., 8q24 containing MYC and 11q13 containing CCND1 are high-incidence regions of ecDNA breakpoints), such as 8q24 (MYC locus, chr8: 127.2-127.4 Mb) and 11q13 (CCND1 locus, chr11: 69).4-69.6Mb)。.
[0026] In this invention, for each selected hotspot region "A", a series (typically 50-100 probes per end) of overlapping oligonucleotide probes, each 80-120 nt in length, are designed at its ends (e.g., within a 1kb upstream and 1kb downstream range). All probes are labeled with biotin at their 5' ends. Similarly, B-left-end and B-right-end probe groups are designed for hotspot region "B", where hotspot region "A" and hotspot region "B" are symmetrically located on opposite sides of the same hotspot region. Each probe in this invention is labeled with biotin at its 5' end, and the probe length is 80-120 nt. Probes that are too short (<60 nt) have insufficient hybridization stability, while probes that are too long (>150 nt) have high synthesis costs and are prone to secondary structure formation. The 5' biotin label is spaced 10-15 nt from the hybridization region by a spacer sequence (such as a C6 spacer or TEG spacer) to avoid steric hindrance. Each end has 50-100 overlapping probes, with adjacent probes overlapping by 20-30 nt to ensure gapless coverage. The probe Tm value is 65-75℃, and the GC content is 40-60%. Preferably, 2-3 locked nucleic acid (LNA) or peptide nucleic acid (PNA) monomers are embedded at key positions in each probe sequence (such as the central region or a location expected to mismatch with the target sequence), preferably 3 locked nucleic acid (LNA) monomers. These nucleic acid analogs have extremely high hybridization affinity with DNA / RNA (Tm value can be increased by 2-8℃ / monomer) and significantly improve the ability to distinguish single-base mismatches. This allows hybridization to be performed at higher temperatures (e.g., 72℃) and more stringent salt concentrations, greatly reducing non-specific binding background caused by partial sequence homology. The locked-loop structure of LNA restricts the conformational freedom of the sugar ring, allowing the LNA-DNA double strand to adopt a helical structure closer to the A-form, increasing base stacking forces and hydrogen bond stability. When a single base mismatch occurs, the conformational distortion of LNA-DNA double strands is much greater than that of conventional DNA-DNA double strands, resulting in a 5-15°C decrease in the Tm value of the mismatched double strands (compared to only a 2-5°C decrease for conventional DNA). This "mismatch sensitivity amplification effect" means that under hybridization conditions at 72°C, only one base mismatched non-specific binding double strand completely dissociates, while the perfectly complementary target double strand remains stable. Universal blocking sequences: A large amount of excess "blocking oligonucleotides" complementary to highly repetitive sequences in the human genome (such as Alu and LINE) are added to the probe pool. These blockers preferentially bind to repetitive sequences in the sample, preventing non-specific consumption of the probe, thereby improving the probe's effective capture efficiency for target low-copy-number chimeric sequences. Approximately 50% of the human genome consists of repetitive sequences (Alu accounts for ~10%, LINE for ~17%), and these repetitive sequences are extremely abundant in cfDNA. The blocking oligonucleotides are designed as full-length complementary sequences to repetitive sequences such as Alu and LINE, with the 3' end modified with a C3 spacer or dideoxy terminator to prevent DNA polymerase elongation.The concentration of the blocking oligonucleotide is 10-50 times the total concentration of the probe pool, preferably 20-30 times, and even more preferably 25 times; through competitive hybridization, the probe preferentially occupies the repetitive sequence, so that the probe can only bind to the specific target.
[0027] In this invention, the LNA modification site is one or a combination of the following two categories: (i) the last three bases at the 3' end of the probe; (ii) the central region of the probe, i.e., bases 45-52 from the 5' end of the probe; the interval between two adjacent LNA monomers is 1-2 nt; no LNA modification is introduced within the first 10 nt after biotin labeling at the 5' end of the probe. The LNA modification increases the probe Tm value by 6-8 °C, thereby supporting rigorous hybridization at 65-72 °C.
[0028] In this invention, when a cfDNA molecule is a chimera resulting from the rearrangement of genomic regions "A" and "B", its junction carries the terminal sequence of region A on one side and the terminal sequence of region B on the other. During hybridization, the molecule will be stably hybridized and bound simultaneously to multiple probes from both the A-terminal and B-terminal probe sets. Since all probes are biotinylated, the chimeric DNA molecule will be efficiently "pulled down" by streptavidin magnetic beads. Normal, unrearranged linear cfDNA fragments, on the other hand, can only be bound by one end of a single hotspot region's probe set at most, exhibiting weak binding and lacking paired-end specificity, and will be effectively removed in subsequent rigorous washing steps. The specificity of paired-end capture stems from "geometric constraint" and "thermodynamic synergy": the chimeric DNA, carrying both A-terminal and B-terminal sequences, simultaneously forms hybrid double strands with both the A-terminal and B-terminal probe sets in three-dimensional space, producing a "molecular bridging" effect that significantly enhances the binding stability with magnetic beads. Normal linear DNA fragments can only bind to a single probe set. Under strict washing conditions (low salt, high temperature), the dissociation rate constant (koff) of a single probe-target double strand is significantly higher than that of the synergistic binding of two probes, and therefore it is easily eluted.
[0029] In this invention, different premixed probe panels can be designed for different cancer types (such as "Lung Cancer Panoramic Chimera Detection Panel", "Sarcoma-Related Rearrangement Panel", and "Hematologic Malignancy Fusion Gene Expansion Panel"). Each panel contains tens of thousands to hundreds of thousands of pairs of "paired-end anchored" probes, comprehensively covering known and potential rearrangement-related genomic regions for that cancer type. The panel design incorporates redundancy to ensure that even if the breakpoint occurs within a hotspot region rather than at the precise end, there are still enough probes to cross-cover, guaranteeing a high capture success rate. The modular design of the probe panels is based on the structural variation spectrum characteristics of different cancer types: the lung cancer panel covers introns of genes such as ALK, ROS1, RET, NTRK1 / 2 / 3, MET, and NRG1, as well as common fusion partner genes (EML4, KIF5B, CCDC6); the sarcoma panel covers EWSR1, FLI1, SS18, and FUS; and the hematologic malignancy panel covers BCR-ABL, PML-RARA, and MLL rearrangements. Each panel follows a dual strategy of "known + unknown": it covers both known fusion partner combinations and genomic regions that may be involved in new fusions, enabling simultaneous detection of known and unknown structural variations.
[0030] In this invention, a high-fidelity end-repair enzyme mixture (containing T4 DNA polymerase, T4 polynucleotide kinase, and Klenow fragment) is used to repair the sticky ends of enriched DNA into blunt ends, ensuring 5' phosphorylation. T4 DNA polymerase possesses both 5'→3' polymerase and 3'→5' exonuclease activities, simultaneously repairing both 5' protruding ends (removal) and 3' protruding ends (filling), producing blunt ends. T4 polynucleotide kinase (PNK) catalyzes the transfer of the γ-phosphate group of ATP to the 5' hydroxyl group of DNA, causing 5' phosphorylation. This is a necessary prerequisite for the ligation reaction catalyzed by T4 DNA ligase (T4 DNA ligase only catalyzes the formation of the phosphodiester bond between the 5'-phosphate and 3'-hydroxyl groups). The end-repaired DNA is diluted with nuclease-free water to a concentration <1 nM (approximately 20-50 pg / μL), 25 μL of 5X cyclization reaction buffer (containing PEG 8000) and 2.5 μL of T4 DNA ligase (high concentration) are added, and water is added to a final volume of 125 μL. The circularization program was run at 22°C for 30 minutes (initial ligation), followed by a programmed cooling to 16°C (1°C decrease per hour) for a total of 6 hours. At low DNA concentrations (<1 nM) and in the presence of PEG 8000 molecular crowding reagent, the ends of DNA fragments were catalyzed by T4 DNA ligase to form covalently closed circular DNA molecules due to spatial proximity and random collisions. Under normal high-concentration conditions, intermolecular linkages mainly occurred, forming polymers. Circulation is based on the Jacobson-Stockmayer theory: the probability of intramolecular circularization is proportional to the 3 / 2 power of the fragment length, and the probability of intermolecular linkage is proportional to the square of the concentration. Intramolecular circularization is absolutely dominant under <1 nM conditions. PEG 8000 increases the effective viscosity of the solution through size exclusion, compressing the DNA gyration radius and increasing the frequency of encounters between the two ends.
[0031] In this invention, 10 μL of end-repair and A-tailing premixed enzyme (containing Klenowexo- fragment and dATP) is added to 20 μL of circularized DNA. The mixture is incubated at 20°C for 30 minutes, then at 65°C for 30 minutes, to add a poly(dA) tail to the gap in the circular DNA. 2.5 μL of adaptor mixture (with index, containing a poly(dT) sequence complementary to the poly(dA) tail), 25 μL of 5X ligation buffer, and 2.5 μL of T4 DNA ligase are added directly, and the mixture is incubated at 22°C for 30 minutes. The poly(dT) sequence in the adaptor anneals to the poly(dA) tail of the circular DNA, forming a phosphodiester bond catalyzed by T4 DNA ligase, thus ligating the sequencing adapter (P5 / P7 / Index) to the circular DNA. The Klenow fragment (exo-) lacks 3'→5' exonuclease activity, allowing only unidirectional dNTP addition without excising the synthesized poly(dA) tail, ensuring uniform tail length. The poly(dA) tail length is 20-50 nt: too short (<10 nt) leads to insufficient hybridization stability with the linker poly(dT) and reduced cyclization efficiency; too long (>100 nt) increases molecular flexibility and promotes intermolecular linkages rather than intramolecular cyclization.
[0032] In this invention, Phi29 DNA polymerase is used. This enzyme has extremely strong strand displacement activity and continuous synthesis capacity (>70kb) with high fidelity. Using circularized DNA as a template, a rolling circle primer is designed using the universal sequence on the adaptor. Phi29 polymerase is initiated with this primer and performs isothermal (30°C) rolling circle replication along the circular template for 12-16 hours, preferably 13-15 hours, and more preferably 14 hours, producing an ultra-long single-stranded DNA product composed of hundreds of template inverted repeat sequences tandemly. Technical advantages: (a) Exponential signal amplification: A single target molecule is transformed into a huge DNA cluster, completely overcoming the limitation of insufficient starting material. (b) Error dilution: Isothermal amplification avoids the problem of error accumulation during PCR cycles. (c) Solidification of linkage site information: The linkage site sequence of the chimera is "locked" in each repeat unit, ensuring the reliability of subsequent sequencing detection. Phi29 DNA polymerase, derived from Bacillus subtilis phage Phi29, possesses sustained synthesis capacity (>70kb) and strong strand substitution activity. Its 3'→5' exonuclease proofreading activity imparts extremely high fidelity (error rate approximately 10%). -6 / base). RCA primers are designed in the universal P5 / P7 sequence region of the adaptor. Phi29 polymerase synthesizes unidirectionally and continuously along the circular template using dNTPs as raw materials. Because the template is circular, the polymerase returns to the starting point after each cycle, pushing the previously synthesized strand apart through strand displacement activity to form a tandem repeat sequence (concatenator). After 12-16 hours of RCA, a 100bp circular template produces a >70kb tandem product containing approximately 700 repeat units, each retaining the original chimeric junction sequence. The RCA product is used directly as a PCR template: specific primer pairs that cross the junction point are designed (F in region A, R in region B), amplifying only when the junction point is present; or semi-universal primer pairs (anchor region primer + adaptor P7 universal primer) are designed for initial screening.
[0033] This invention provides a kit for detecting chimeric fragments with structural variations in cell-free DNA from peripheral blood, comprising a cfDNA extraction and purification module, a chimeric targeted enrichment module, a circularization transformation and amplification module, a sequencing library preparation module, and a quality control module.
[0034] In this invention, the cfDNA extraction and purification module includes plasma lysis buffer A, binding buffer B, silanol magnetic bead suspension, washing buffer C, elution buffer D, and proteinase K. Plasma lysis buffer A, using 10 mM Tris-HCl (pH 8.0) and 1 mM EDTA as solvents, is composed of guanidine hydrochloride at a final concentration of 4-6 M and proteinase K at a final concentration of 0.5-2 mg / mL. Preferably, the final concentration of guanidine hydrochloride is 5 M, and the final concentration of proteinase K is 1 mg / mL, meaning each 1 mL of lysis buffer contains approximately 33 μL of 30 mg / mL proteinase K stock solution. In this invention, plasma lysis buffer A lyses proteins in plasma and releases cfDNA through strong denaturation, while proteinase K digests proteins to eliminate DNA-protein cross-links. Binding buffer B, using 10 mM Tris-HCl (pH 8.0) as the solvent, consists of 40%-60% isopropanol and a specific salt (sodium acetate or sodium chloride). When the specific salt is sodium acetate, the concentration is 2.5-4 M, preferably 3 M; when the specific salt is sodium chloride, the concentration is 1-2 M, preferably 1.5 M. The volume molar ratio of isopropanol to the specific salt (calculated as sodium acetate) is 1:0.05-1:0.08, preferably 1:0.06, meaning that each 1 mL of binding buffer contains approximately 50 μL of 3 M sodium acetate. Binding buffer B in this invention promotes the binding of DNA to silanol magnetic beads by adjusting the ionic strength and dielectric constant of the solution. A silanol magnetic bead suspension, with a concentration of 10 mg / mL, was suspended in a preservation solution containing 20% ethanol, with a total volume of 1 mL. The magnetic beads had a particle size of approximately 1-2 μm and were surface-modified with silanol groups (Si-OH). Under high-salt, low-pH conditions, they bound to DNA through hydrogen bonds and electrostatic interactions. Washing buffer C is an 80% aqueous ethanol solution (v / v) containing 10 mM Tris-HCl (pH 8.0), used to remove residual binding buffer and protein impurities from the surface of magnetic beads; Elution buffer D is 10 mM Tris-HCl, pH 8.5, used to disrupt the binding of DNA to magnetic beads and release the purified cfDNA; Proteinase K (lyophilized powder) activity ≥30U / mg. Reconstitute with nuclease-free water to a stock solution of 30mg / mL before use.
[0035] In this invention, the chimeric targeted enrichment module includes a smart anchoring probe panel, a probe reconstitution solution, a 2X hybridization buffer, streptavidin magnetic beads, strict washing solution I, strict washing solution II, and elution buffer E. The smart anchoring probe panel includes tens of thousands of premixed biotinylated LNA-modified probes targeting specific cancer types. The probes are custom-synthesized by commercial synthesis services (such as IDT, Twist Bioscience, and Agilent), purified by HPLC or PAGE, lyophilized and stored at -20°C, and reconstituted with the probe reconstitution solution to a working concentration of 0.1-1 μM / probe, preferably 0.5 μM, before use. The probe reconstitution solution is DNase / RNase-free water used to reconstitute the lyophilized probe panels and prepare the probe mixture. The concentration of the reconstituted probe mixture is 0.1-1 μM / probe, preferably 0.5 μM, with a total volume of approximately 100-200 μL (depending on the number of probes and the amount synthesized). 5 μL of the reconstituted probe mixture contains approximately 0.5-5 pmol of each probe, which is sufficient to cover one hybridization reaction. 2X hybridization buffer, using nuclease-free water as solvent, consists of high-concentration SSC, Denhardt's solution, EDTA, and a surfactant. The high-concentration SSC is 3×-6× SSC, preferably 4× SSC; the Denhardt's solution is 1×-2×, preferably 1×; the EDTA concentration is 1-5 mM, preferably 2 mM; and the surfactant is sodium dodecyl sulfate (SDS) or polyethylene glycol octylphenyl ether (Triton X-100). When the surface activity is equivalent to the concentration of SDS, the concentration is 0.1%-0.5%, preferably 0.2%; when the surface activity is equivalent to the concentration of Triton X-100, the concentration is 0.1%-0.5%, preferably 0.2%. Example volume ratio of components: 4× SSC: 1× Denhardt's: 2 mM EDTA: 0.2% SDS = 100:10:1:1 (volume ratio). Streptavidin magnetic beads: binding capacity >500 pmol biotin / mg, concentration 10 mg / mL, suspended in PBS containing 0.05% Tween-20 and 0.1% BSA, total volume 500 μL. The magnetic beads have a particle size of approximately 1-2.8 μm and are covalently coupled with streptavidin on the surface for specific capture of biotinylated probe-target DNA complexes. Rigorous Wash Buffer I: Low-salt SSC buffer containing low-concentration SDS, preheated to 65°C before use. The concentration of the low-salt SSC buffer is 0.1×-0.5×SSC, preferably 0.2×SSC, and the concentration of low-concentration SDS is 0.1%-0.5%, preferably 0.2%. The solvent is nuclease-free water, and the total volume is 10 mL. Preheat to 65°C before use. Used to remove non-specifically bound DNA and probes. Rigorous Wash Buffer II: A lower-salt SSC buffer, used at room temperature. The concentration of the lower-salt SSC buffer is 0.05×-0.1×SSC, preferably 0.1×SSC. It is SDS-free, and the solvent is nuclease-free water. The total volume is 10 mL. Used at room temperature. For further removal of residual salt ions and weakly bound impurities. Elution buffer E (low salt TE): 10 mM Tris-HCl, 0.1 mM EDTA, pH 8.0, used for high-temperature denaturation elution of target DNA while maintaining DNA stability.
[0036] In this invention, the circularization transformation and amplification module includes an end-repair and A-tailing premixed enzyme, an index-specific adaptor mixture, T4 DNA ligase, 5X circularization reaction buffer, Phi29 DNA polymerase and 10X reaction buffer, a dNTP mixture, and rolling circle primers; wherein the end-repair and A-tailing premixed enzyme contains 0.5-2 U / μL of T4 DNA polymerase, preferably 1 U / μL; 5-20 U / μL of T4 PNK, preferably 10 U / μL; 2-10 U / μL of Klenow fragment (exo-), preferably 5 U / μL; and corresponding buffers (10× end-repair buffer: 500 mM Tris-HCl pH 8.0, 100 mM MgCl2, 100 mM DTT, 1 mg / mL BSA, 1 mM dNTPs), dNTPs at a concentration of 2.5 mM, and dATP at a concentration of 10 mM, with a total volume of 200 μL. Example of component volume ratio: T4 DNA polymerase: T4 PNK: Klenow exo-: 10× buffer: dNTPs: dATP = 1:1:1:2:1:1; A mixture of index-specific adaptors: adaptors containing different index sequences, at a concentration of 15 μM, with a total volume of 20 μL. The adaptors are double-stranded DNA, formed by annealing two complementary oligonucleotides. Example adaptor sequence (using the Illumina platform as an example): Long adaptor chain (5'→3'): 5'-AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTCCGATCT-[Index]-TTTTTTTTTT-3' (SEQ ID No. 4); The adaptor short strand (5'→3'): 5'-AAAAAAAAAA-[Index complementary]-AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGTAGATCTCGGTGGTCGCCGTATCATT-3' (SEQ ID No. 5); where [Index] is a 6-8 nt sample tag sequence (e.g., Index 1: ATCACG; Index 2: CGATGT, etc.), and the 3' end poly(dT) sequence (10 Ts) is complementary to the poly(dA) tail of the circular DNA. The 5' end of the short strand is phosphorylated to facilitate ligation catalyzed by T4 DNA ligase. The actual adaptor sequence is designed based on the official Illumina standard adapter sequence, and the Index sequence can be selected from the Illumina Index sequence library; T4 DNA ligase, at a concentration of 400 U / μL, in a total volume of 50 μL. Derived from *E. coli* infected with T4 bacteriophage, it catalyzes the formation of a phosphodiester bond between the 5'-phosphate and 3'-hydroxyl groups of DNA. 5X cyclization reaction buffer: Contains specific ionic components (Tris-HCl 250mM pH 7.5, MgCl2 50mM, DTT 25mM, ATP 5mM) and PEG 8000 (concentration 15%-25%, preferably 20%), in nuclease-free water, with a total volume of 1 mL. Example volume ratio of ionic components to PEG 8000: Tris-HCl:MgCl2:DTT:ATP:PEG 8000 = 50:10:5:1:200 (volume ratio, based on stock solution concentration). PEG 8000 acts as a molecular crowding reagent, promoting intramolecular cyclization of DNA molecules through size exclusion effect. Phi29 DNA polymerase and 10X reaction buffer: Phi29 DNA polymerase concentration 10 U / μL, total volume 100 μL; 10X reaction buffer: 330 mM Tris-HAc pH 7.5, 100 mM MgCl2, 660 mM KAc, 1% Tween-20, 10 mM DTT, total volume 100 μL; dNTP mixture (10mM each): containing 10mM each of dATP, dCTP, dGTP, and dTTP, with a total volume of 100μL; Rolling circle primers (100 μM): sequence 5'-AGATCGGAAGAGCGTCGTGTAGGGAAAGAG-3' (complementary to the P7 region of the adapter, SEQ ID No. 1), total volume 50 μL. These primers were annealed to the universal P7 sequence on the adapter to initiate Phi29 rolling circle amplification.
[0037] In this invention, the sequencing library preparation module includes 2X high-fidelity PCR premix, universal PCR primer mixture, and PCR purification magnetic beads; wherein the 2X high-fidelity PCR premix contains hot-start DNA polymerase (such as NEB Q5 or KAPA HiFi, concentration 1-2 U / μL), dNTPs (0.4 mM each), MgCl2 (2-4 mM), and optimized buffer, with a total volume of 1 mL; The universal PCR primer mixture contains a P5 forward primer and a P7 reverse primer. The P5 primer sequence is shown in SEQ ID No. 2, and the P7 primer sequence is shown in SEQ ID No. 3. Specifically: 5'-AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCCTTCCGATCT-3' (SEQID No. 2); 5'-CAAGCAGAAGACGGCATACGAGATCGGTCTCGGCATTCCTGCTGAACCGCTCTTCCGATCT-3' (SEQ ID No. 3); Both primers P5 and P7 are at a concentration of 100 μM, and the P5:P7 ratio in the mixture is 1:1 (volume ratio). Therefore, the mixture contains 50 μM of P5 and 50 μM of P7, with a total volume of 100 μL. For the PCR reaction, 2.5 μL of this mixture (containing approximately 125 pmol P5 and 125 pmol P7) is used for co-amplification with the RCA product. PCR purification magnetic beads: such as AMPure XP magnetic beads, concentration 10-20 mg / mL, total volume 5 mL. Used to purify PCR products and remove primer dimers and small fragments.
[0038] In this invention, the quality control module includes positive control DNA, negative control DNA, DNA quantitative standards, library-based quantitative qPCR premix, and primers and probes. The positive control (PC) DNA is a synthetically produced DNA mixture containing a known chimeric sequence (such as EML4-ALK V1) that simulates the fragmented state of cfDNA. The chimeric allele frequencies are 1%, 0.1%, and 0.01%, respectively, at 20 ng / μL. The synthetically produced EML4-ALK V1 chimeric DNA sequence, characterized by cfDNA fragmentation (150-200 bp), contains an EML4 intron 13 terminal sequence at its 5' end (designed based on the real GRCh38 sequence) and an ALK intron 19 terminal sequence at its 3' end (designed based on the real GRCh38 sequence), with a fusion junction in the middle. The positive control consisted of three concentration gradients: 1% VAF (high positive), 0.1% VAF (medium positive), and 0.01% VAF (low positive, close to the limit of detection), each at 20 ng / μL, used to verify the detection sensitivity and repeatability of the kit. Each concentration was aliquoted into three 10 μL tubes. During actual synthesis, the actual genomic sequences of EML4 and ALK were downloaded from the UCSC Genome Browser, and the corresponding regions were manually assembled. The negative control (NC) DNA was cfDNA extracted from mixed plasma of healthy donors, 20 ng / μL, 10 μL. It was used to assess the background noise and nonspecific amplification levels of the kit. DNA quantitative standards: used for calibration of fluorescence quantitative instruments; Library quantification qPCR premix and primers / probes include 2X qPCR premix (such as KAPA Library QuantKit), P5 / P7 primer pairs (10 μM each), and SYBR Green I fluorescent dye. qPCR primer sequences: P5 forward primer 5'-AATGATACGGCGACCACCGAGATCTACAC-3' (29 nt, SEQ ID No. 14), P7 reverse primer 5'-CAAGCAGAAGACGGCATACGAGAT-3' (24 nt, SEQ ID No. 15). Ensure equimolar amounts of the library are mixed for accurate quantification before sequencing.
[0039] The present invention also provides the application of the kit in the preparation of in vitro diagnostic products for tumor liquid biopsy, early screening, molecular typing, dynamic monitoring or prognostic assessment.
[0040] The technical solutions provided by the present invention will be described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.
[0041] Example 1: Probe Sequence Design
[0042] The probe and primer sequences of this invention are generated based on the design rules of this invention (length 80-12 nt, GC content 40-60%, Tm 65-75℃, 5' biotin labeling, LNA modification at key positions, etc.). The probes in the actual product are designed based on the real sequences of the GRCh38 reference genome, and their specificity is verified using bioinformatics tools (such as NCBI Primer-BLAST and UCSGCenome Browser). The design uses the EML4-ALK V1 fusion, the most common fusion in non-small cell lung cancer, as a model.
[0043] 1. Dual-end anchoring probe sequence
[0044] EML4 - Left Probe - 1 (covering the upstream end of EML4 intron 13): 5'-Biotin-AACTTTGAAGAAATTATGTGCATGCCTTCAAGACCCAGAGACCTAATCATAGCGCTCCTCATTTCGCTCATACGCATCTGGGACTTCGGCTTGAAATT+G+A+-3' (Length: 100 nt; GC content: 44.0%; + indicates LNA modification site, located in the last 3 bases of the 3' end, SEQ ID No. 6); EML4-Right-end Probe-1 (covering the downstream end of EML4 intron 13): 5'-Biotin-GGGCCAACCACGTGACTACTTCTACGAACCTATAAGATTGTCGTTC+G+C+GGATTACATTAAATAACATCGTTGTGGTAAGCGGGCAAACCATTTGGTGTCG-3' (Length: 100nt; GC content: 46.0%; + indicates LNA modification site, located at positions 45-47 in the central region, SEQ ID No. 7); ALK-Left-End Probe-1 (Covering the upstream end of ALK intron 19): 5'-Biotin-TAGAAATTTCGGTGATGAGCGCGGTTCTAACAAGTAATAATGATAAGCCTCTCGTCGCAAGAATCTCATCCTGCACATCAATCCTCTCGCAAGCAACT+C+T+-3' (Length: 100nt; GC content: 44.0%; + indicates LNA modification site, located in the last 3 bases of the 3' end, SEQ ID No. 8); ALK-Right-End Probe-1 (Covering the downstream end of ALK intron 19): 5'-Biotin-GGAAAGTACTGTACCACTTACGTTTAGATCGTCTAGAGTTGCCTTATGCCA+C+C+GCAACTCAAGCCGAGTCAGATCGACCACCGCGCTTGGTCGACCTGCG-3' (Length: 100nt; GC content: 54.0%; + indicates LNA modification site, located at positions 50-52 in the central region, SEQ ID No. 9) The above four probes are representative examples of probe sets in the EML4-ALK V1 detection panel. In actual panels, each end is designed with 50-100 overlapping probes (adjacent probes overlap by 20-30 nt), with a total of approximately 200-400 probes per fusion partner pair. All probes are labeled with biotin at their 5' ends, with 2-3 LNA monomers embedded at key locations, increasing the probe Tm value by 6-8 °C, allowing for rigorous hybridization at high temperatures of 65-72 °C. Probe length is controlled at 80-120 nt, and GC content is controlled at 40-60%.
[0045] The LNA modification of the probes in this invention follows a uniform rule: each probe contains 3-5 (preferably 4) LNA monomers; the LNA monomers are distributed within the anchoring binding region of the probe, with adjacent LNAs spaced 2-3 DNA nucleotides apart; no LNA modification is applied at least 2 nucleotides from the 5' and 3' ends of the probe to avoid affecting biotin labeling efficiency and hybridization kinetics. Modification according to this rule can increase the probe's Tm value by 6-8°C, allowing for rigorous hybridization at high temperatures of 65-72°C, and improving the capture specificity for chimeric fragments.
[0046] 2. qPCR / ddPCR primer pairs crossing the EML4-ALK V1 junction
[0047] Forward Primer (located at the end of EML4 intron 13): 5'-ACAGCAAGCCATTCCTTAGAGA-3' (Length: 22nt; GC content: 45.5%; Tm: 64℃, SEQ ID No. 10); Reverse Primer (located at the end of ALK intron 19): 5'-TGACGGTCGCCGTAACTTTCGT-3' (Length: 22nt; GC content: 54.5%; Tm: 68℃, SEQ ID No. 11); The target amplicon length is 100-120 bp. Primer pairs span the fusion junction and amplification is only possible when the EML4-ALK fusion chimera is present in the sample. Wild-type EML4 or ALK sequences alone cannot be amplified because the primer pairs are located in two different genomic regions and cannot coexist on the same DNA molecule. In practical applications, primer design is based on the actual GRCh38 sequence, and specificity is verified using NCBI Primer-BLAST.
[0048] 3. Semi-universal PCR primer pairs (for initial screening of unknown junction sites)
[0049] Semi-universal Forward Primer (located in the EML4 anchoring region): 5'-TACACTGACGTGTTGGGAAT-3' (Length: 20nt; GC content: 45.0%; Tm: 58℃, SEQ ID No. 12); Semi-universal Reverse Primer (universal P7 connector sequence): 5'-CAAGCAGAAGACGGCATACG-3' (Length: 20nt; GC content: 55.0%; Tm: 62℃; This sequence is the standard Illumina P7 connector sequence and is fixed, SEQ ID No. 13). The semi-universal forward primer is located within the anchor probe coverage of the potential hotspot region of EML4, while the semi-universal reverse primer is located within the universal sequence of the adaptor P7. This primer pair can only amplify when the EML4 region and the adaptor P7 are linked via a chimera. Positive amplification products are verified by HRM or Sanger sequencing, indicating the presence of a potential chimera, and the linkage site sequence can then be precisely identified through sequencing.
[0050] 4. Positive control DNA (PC) sequence description
[0051] The positive control DNA is a synthetically produced EML4-ALK V1 chimeric DNA sequence with cfDNA fragmentation characteristics (150-200 bp). Its 5' end contains the EML4 intron 13-terminal sequence (designed based on the real GRCh38 sequence), and its 3' end contains the ALK intron 19-terminal sequence (designed based on the real GRCh38 sequence), with a fusion junction in the middle. The positive control includes three concentration gradients: 1% VAF (high positivity), 0.1% VAF (medium positivity), and 0.01% VAF (low positivity, close to the detection limit), used to verify the kit's detection sensitivity and reproducibility. In actual synthesis, the real genomic sequences of EML4 and ALK were downloaded from the UCSC Genome Browser, and the corresponding regions were manually assembled. Specifically: EML4 (Ensembl accession number ENSG00000143924, RefSeq accession number NM_019063; GRCh38 chr2:42,169,312-42,332,548, positive strand) takes the intron sequence chr2:42,295,517-42,295,616 (100bp, positive strand) from sequence 13; ALK (Ensembl accession number ENSG00000171094, RefSeq accession number NM_004304; GRCh38 chr2:29,190,995-29,921,586, negative strand) takes the reverse complement (100bp) of the intron sequence chr2:29,223,529-29,223,628 from sequence 19. The two segments were spliced together to form a 200bp chimeric DNA sequence, with the fusion junction between nucleotides 100 and 101. The sequence (SEQ ID No. 16) is: 5'-GTACAGTATTCTTATATTAAACTCATTTCTGGTAATTCTCACATAGTACTCTTTCAGTCCCATCTCTTAGACCAGGAGAGAAAGAGCTGCAGTGTAACAAGGATATGGAGATCCAGGGAGGCTTCCTGTAGGAAGTGGCCTGTGTAGTGCTTCAAGGGCCAGGCTGCCAGGCCATGTTGCAGCTGACCACCCACCTGCAG-3'.
[0052] Example 2: A method for detecting structurally variant chimeric fragments in cell-free DNA from peripheral blood.
[0053] 1. Sample preparation and cfDNA extraction
[0054] (1) Take 2 mL of patient plasma and add an equal volume of plasma lysis buffer A (containing proteinase K), and incubate at 56 °C with shaking for 1 h. Guanidine hydrochloride in lysis buffer A destroys the plasma protein structure through strong denaturation, releasing cfDNA; proteinase K digests proteins and eliminates DNA-protein cross-links.
[0055] (2) Add 2 times the volume of binding buffer B and 20 μL of magnetic bead suspension, and bind for 15 min at room temperature. Add 2 times the volume (i.e., twice the plasma volume) of binding buffer B (containing 50% isopropanol and 3M sodium acetate) and 20 μL of silanol magnetic bead suspension (10 mg / mL), and bind for 15 min at room temperature (22℃). The isopropanol in binding buffer B lowers the dielectric constant of the solution, promoting hydrogen bonding and electrostatic interactions between DNA and the silanol (Si-OH) groups on the surface of the silanol magnetic beads; the high concentration of sodium acetate provides the necessary ionic strength to neutralize the negative charge of the DNA phosphate backbone, enhancing binding efficiency.
[0056] (3) Place on a magnetic rack and discard the supernatant. Wash twice with washing buffer C. Each time, use 200 μL and let stand at room temperature for 1 min before discarding the supernatant. Wash with ethanol to remove residual salt ions, protein debris and organic solvents.
[0057] (4) Open the lid and let it air dry for 5 minutes to remove residual ethanol. Avoid excessive drying, which may cause the magnetic beads to crack and make it difficult to wash out DNA.
[0058] (5) Add 50 μL of elution buffer D, incubate at 55 °C for 5 min to elute cfDNA, and transfer to a new tube. The low-salt buffer disrupts the binding of DNA to the magnetic beads, and 55 °C promotes the dissociation of DNA from the surface of the magnetic beads.
[0059] (6) Quantification should be performed using a fluorometer, with a total volume >10ng and an A260 / A280 ratio between 1.8 and 2.0. Quantification should be performed using a fluorometer (such as Qubit 4.0 with dsDNA HS Assay Kit). An A260 / A280 ratio <1.8 indicates protein contamination, and >2.0 indicates RNA contamination, both of which require repurification.
[0060] 2. Probe hybridization
[0061] (1) Take 20 ng cfDNA (sample, PC, NC) and add it to 50 μL with nuclease-free water.
[0062] (2) Add 50 μL of 2X hybridization buffer and 5 μL of reconstituted probe mixture. 2X hybridization buffer (containing 4×SSC, 1×Denhardt's solution, 2 mM EDTA and 0.2% SDS) and 5 μL of reconstituted probe mixture (0.5 μM / probe, containing 2.5 pmol of each probe). The probe mixture is obtained as follows: reconstitute the lyophilized probe panel with probe reconstitution solution (nuclease-free water) to 0.5 μM / probe, vortex to mix, briefly centrifuge, aliquot and store at -20℃, avoiding repeated freeze-thaw cycles.
[0063] (3) Run the PCR tester: 95℃ for 5 min (denaturation, causing the cfDNA double strand to dissociate into single strands), 65℃ for 16 h (hybridization, allowing the probe to anneal and bind to the target sequence). 65℃ is a strict hybridization temperature to ensure that only completely complementary or highly homologous sequences can bind stably.
[0064] 3. Magnetic bead capture and washing
[0065] (1) Transfer the hybridization reaction system to a new tube and add 50 μL of streptavidin magnetic beads that have been pre-equilibrated with hybridization buffer. The streptavidin magnetic beads (10 mg / mL) have been pre-equilibrated with 2X hybridization buffer. Equilibration procedure: Take 50 μL of streptavidin magnetic beads, wash twice with 200 μL of 2X hybridization buffer, and resuspend in 50 μL of 2X hybridization buffer. Equilibration removes BSA and Tween-20 from the magnetic bead preservation solution, preventing interference with subsequent hybridization reactions.
[0066] (2) Rotate and incubate at room temperature for 45 min to allow the biotinylated probe-target DNA complex to bind to the magnetic beads. Rotate and incubate at room temperature (22℃) for 45 min to allow the biotinylated probe-target DNA complex to bind to streptavidin on the surface of the magnetic beads. Streptavidin has a very high affinity for biotin (Kd≈10). -14 The binding is rapid and irreversible (under normal washing conditions). The biotinylated probe-target DNA complex forms naturally during hybridization (65°C for 16 h) and then binds to streptavidin magnetic beads.
[0067] (3) Place it on a magnetic rack and carefully remove the supernatant.
[0068] (4) Rigorous washing: Add 200 μL of preheated Rigorous Wash Buffer I, gently tumble to mix, and transfer to a 65°C metal bath to stand for 5 min. Return to the magnetic rack and discard the supernatant. Repeat this step once. Rigorous Wash Buffer I (0.2×SSC, 0.2% SDS). The high temperature and low salt conditions at 65°C cause the non-specifically bound DNA-probe double strands to dissociate, while the chimeric DNA anchored at both ends remains stably bound due to the "molecular bridging" effect.
[0069] (5) Second wash: Add 200 μL of strict wash buffer II, let stand at room temperature for 2 min, and discard the supernatant. Repeat once. Add strict wash buffer II (0.1×SSC), let stand at room temperature (22℃) for 2 min, and discard the supernatant. Repeat once. This further removes residual SDS and salt ions, preparing for subsequent elution.
[0070] 4. Target elution and purification
[0071] (1) Add 50 μL of elution buffer E and incubate at 95 °C for 5 min. The high temperature denatures and dissociates the DNA-probe hybrid. The LNA-modified probe is retained on the magnetic beads due to its binding with streptavidin, and the target DNA is released into the solution.
[0072] (2) Quickly transfer to a magnetic rack and carefully transfer the supernatant (50 μL) containing the enriched DNA to a new PCR tube. Avoid aspirating magnetic beads to prevent probe contamination.
[0073] (3) DNA purification beads (such as AMPure XP, at a volume ratio of 1.2) were used for concentration and buffer replacement, and finally eluted in 20 μL of EB buffer (10 mM Tris-HCl, pH 8.5). The purification beads removed residual salt ions, EDTA and probe degradation products to ensure the efficiency of subsequent enzyme reactions.
[0074] 5. End repair and blunt end preparation
[0075] (1) Add 10 μL of end-repair and A-tailing premixed enzyme to 20 μL of enriched DNA, incubate at 20℃ for 30 min, and then incubate at 65℃ for 30 min. The enzyme contains 1 U / μL T4 DNA polymerase, 10 U / μL T4 PNK, 5 U / μL Klenow exo-, 10× end-repair buffer, 2.5 mM each of dNTPs and 10 mM dATP. Incubate at 20℃ for 30 min (end repair, producing blunt ends), and at 65℃ for 30 min (enzyme heat inactivation to prevent interference with subsequent reactions).
[0076] (2) Purify the end-repair product using purifying magnetic beads (1.2 × volume ratio) and elute in 20 μL of nuclease-free water. Remove residual enzymes, dNTPs, and buffer components.
[0077] 6. Intramolecular cyclization under dilution conditions
[0078] (1) Dilute the end-repaired DNA with nuclease-free water to a concentration of <1 nM (20 pg / μL), add 25 μL of 5X cyclization reaction buffer (containing Tris-HCl 250 mM pH 7.5, MgCl 250 mM, DTT 25 mM, ATP 5 mM, PEG 8000 20%) and 2.5 μL of T4 DNA ligase (high concentration, 400 U / μL), and add water to a final volume of 125 μL.
[0079] (2) Run the circularization program: 22℃ for 30 min (initial ligation), then gradually cool down to 16℃ (1℃ per hour) for a total of 6 hours. Under low DNA concentration (<1 nM) and in the presence of PEG 8000 molecular crowding reagent, the ends of the DNA fragments are catalyzed by T4 DNA ligase to form covalently closed circular DNA molecules due to spatial proximity effect and random collisions; under normal high concentration conditions, intermolecular linkages mainly occur to form polymers.
[0080] (3) After cyclization, the cyclization product was purified using purification magnetic beads (1.2 × volume ratio) and eluted in 20 μL of nuclease-free water to remove uncyclized linear DNA, polymers, and reaction byproducts.
[0081] (4) Gap introduction: Add DNase I to the purified cyclized product to a final concentration of 0.005 U / μL, incubate at 37°C for 10 min, and then add EDTA to a final concentration of 20 mM to terminate the reaction. Limited DNase I digestion randomly introduces an average of 1-2 single-stranded gaps into the double-stranded circular DNA. The gaps provide free 3'-OH ends, serving as substrates for subsequent A-tailing. Furthermore, the ligation efficiency of the two strands differs during intramolecular ligation catalyzed by T4 DNA ligase, and the cyclized product also naturally contains circular DNA with single-stranded gaps, which, together with the gaps introduced in this step, serve as substrates for tailing.
[0082] The selectivity of the cyclization step in this invention is achieved through a triple mechanism: ① Concentration selection: When the DNA concentration is below 1 nM, the probability of spatial collisions between the two ends of the same molecule is much higher than that between different molecules, and the intramolecular linkage (cyclization) rate is much higher than that of intermolecular linkage. Normal linear DNA mainly forms ineffective multimers rather than effective loops; ② Molecular crowding promotion: PEG 8000 increases the local effective concentration at the DNA ends through the size exclusion effect, promoting intramolecular cyclization; ③ Secondary selection at the amplification level: Downstream Phi29 DNA polymerase rolling circle amplification only produces exponential tandem amplification for circular templates, and non-circularized linear DNA does not produce effective amplification signals. Preferably, the cyclization product can be digested with exonuclease I and exonuclease III at 37°C for 30 min to degrade the non-circularized linear DNA and multimers, further improving the purity and selectivity of the cyclization product.
[0083] 7. Add an A tail and connect it to the connector.
[0084] (1) Add 10 μL of end repair and A-tailing premixed enzyme (containing 5 U / μL of Klenow exo-fragment and 10 mM of dATP) to 20 μL of circularized DNA, incubate at 20℃ for 30 min, and incubate at 65℃ for 30 min. The Klenow exo-fragment uses the free 3'-OH end at the nick as a substrate to add a poly(dA) tail to the nick of the circular DNA. The success of tailing and adaptor ligation is verified as follows: Take a small amount of ligation product and perform qPCR using the universal P5 / P7 sequence of the adaptor as a primer. Use a known amount of circular positive template as a control. The difference in Ct value reflects the tailing-ligation efficiency. The adaptor contains a poly(dT) sequence. Only molecules that have successfully tailed can be annealed and then amplified by rolling circle amplification and library PCR. The amount of RCA product (>70 kb tandem) and the sequencing library yield (confirmed by capillary electrophoresis) can also be used as direct verification of successful tailing.
[0085] (2) Add 2.5 μL of adaptor mixture (with index, 15 μM, containing a poly(dT) sequence complementary to the poly(dA) tail), 25 μL of 5X ligation buffer (containing Tris-HCl 250 mM pH 7.5, MgCl2 50 mM, DTT 25 mM, ATP 5 mM), and 2.5 μL of T4 DNA ligase (high concentration, 400 U / μL), and incubate at 22 °C for 30 min. The poly(dT) sequence in the adaptor is annealed to be complementary to the poly(dA) tail of the circular DNA, and a phosphodiester bond is formed by the catalysis of T4 DNA ligase, which ligates the sequencing adapter (P5 / P7 / Index) to the circular DNA.
[0086] (3) Purify the ligation product using purifying magnetic beads (1.0 × volume ratio) and elute in 20 μL of nuclease-free water. Remove unligated adapters and reaction byproducts.
[0087] 8. Rolling circle amplification
[0088] (1) Take 50 μL of the cyclized product, add 10 μL of 10X Phi29 buffer, 4 μL of dNTPs, 2 μL of rolling circle primers, 2 μL of Phi29 polymerase, and add water to 100 μL. 10X Phi29 buffer (330 mM Tris-HAc pH 7.5, 100 mM MgCl2, 660 mM KAc, 1% Tween-20, 10 mM DTT), 4 μL of dNTPs (10 mM each), 2 μL of rolling circle primers (100 μM, sequence 5'-AGATCGGAAGAGCGTCGTGTAGGGAAAGAG-3'), 2 μL of Phi29 polymerase (10 U / μL), and add water to 100 μL.
[0089] (2) Incubate at 30℃ for 14 h, then heat at 85℃ for 10 min to inactivate the enzyme. The RCA product is an ultra-long single-stranded DNA tandem repeat sequence (concatenator), which can be directly used for subsequent PCR detection or sequencing library preparation.
[0090] 9. Sequencing Library PCR Amplification and Purification
[0091] (1) Take 5 μL of RCA product, add 25 μL of 2X high-fidelity PCR premix, 2.5 μL of universal PCR primers, and add water to 50 μL. RCA product (i.e., the product after enzyme inactivation at 85℃ in step 8), add 25 μL of 2X high-fidelity PCR premix (containing 1 U / μL of hot-start DNA polymerase, 0.4 mM of each dNTP, 4 mM of MgCl2 (i.e., the final concentration of 1X reaction is 2 mM)), 2.5 μL of universal PCR primer mixture (50 μM each of P5 and P7), and add water to 50 μL.
[0092] (2) PCR program: 98℃ for 30s; (98℃ for 10s, 60℃ for 30s, 72℃ for 30s) for 12 cycles; 72℃ for 5min. 98℃ for 30s (initial denaturation); (98℃ for 10s, 60℃ for 30s, 72℃ for 30s) for 12 cycles; 72℃ for 5min (final extension).
[0093] (3) Purify the PCR product using PCR purification magnetic beads at a ratio of 1.0X and elute in 20 μL of EB buffer. For example, use AMPure XP magnetic beads. Remove primer dimers (<100bp) and small fragments at a ratio of 1.0×, retaining the target library of 150-500bp.
[0094] (4) The library was accurately quantified and fragment distribution analyzed using qPCR and a high-sensitivity microarray electrophoresis system. The KAPA Library Quant Kit and Agilent Bioanalyzer 2100 were used. The target library with a main peak at 200-400 bp and a concentration ≥2 nM was required for sequencing.
[0095] 10. Sequencing and Data Analysis
[0096] (1) According to the requirements of the sequencing platform (such as Illumina NovaSeq), mix libraries of different indices in equimolar amounts. Illumina NovaSeq 6000. Use qPCR to accurately quantify the concentration of each library and calculate the mixing ratio according to the target data volume.
[0097] (2) Perform paired-end 150bp sequencing, with a recommended data volume of ≥50M reads per sample. Paired-end 150bp sequencing (PE150). The PE150 read length can cover approximately 75bp of sequence on both sides of the chimeric junction, which is sufficient to accurately locate the breakpoint.
[0098] (3) The raw sequencing data (FASTQ format files) are analyzed using a matching or general bioinformatics workflow to identify chimeric fragments. The analysis workflow includes: adapter sequence removal (cutadapt), low-quality read filtering (fastp), alignment with the human reference genome (GRCh38) (BWA-MEM), chimeric read identification (split-read and discordant-pair analysis), precise location annotation of junctions (ANNOVAR), and visualization (IGV). Chimeric read identification uses a junction detection algorithm based on split-read and discordant read pairs, which can be implemented using public tools such as circle_finder or a self-built workflow with the same judgment logic.
[0099] The rules for determining and filtering chimeric segments are as follows: (1) Number of supported reads: The same join point receives no less than 5 independent reads, and the reads are split-reads across join points or discordant reads from abnormal alignments. (1) Pair; where split-reads must be continuously aligned for at least 20 bp on each side of the linker point, and the alignment quality value MAPQ must be at least 30; (2) Deduplication: Remove PCR duplicate reads according to the read alignment start site or UMI tag; (3) False positive filtering: (i) Remove linker points where the sequence at either end of the linker point cannot be uniquely aligned to the reference genome (multi-mapping); (ii) Remove linker points that are also detected in the negative control sample; (iii) Remove micro-homology-mediated artifacts where the aligned sequences on both sides of the linker point have greater than 90% homology with the reference genome and the homology region length is not less than 30 bp; (iv) Remove linker points with fewer than 5 supporting reads or an estimated allele frequency of less than 0.02%; (4) Positive determination: Linker points that meet the above conditions and are reproduced in at least 2 independent replicate experiments are determined to be true chimeric fragments; among them, linker points with at least 20 supporting reads are determined to be high-confidence chimeras.
[0100] The determination of chimeric fragments requires the following conditions to be met simultaneously: ① Split-read support number ≥ 2, and the alignment length on each side of the breakpoint ≥ 30 bp; ② Discordant read pair support number ≥ 3 pairs; ③ Alignment quality MAPQ ≥ 30; ④ Anchor probe region overlap length ≥ 20 bp; ⑤ Negative control blacklist filtering: Candidate sites with a background frequency > 5 times and support reads ≥ 5 in the negative control (NC) sample are removed; ⑥ Repeatability determination: Candidate chimeras detected in ≥ 2 out of 3 independent replicate experiments are considered positive. After filtering according to the above rules, library preparation artifacts, alignment errors, and sample background noise can be effectively eliminated, achieving high specificity identification of chimeric fragments.
[0101] Example 3: PCR Detection Pathway (Rapid Screening)
[0102] Steps 9 and 10 in Example 2 describe a sequencing-based detection pathway. The core detection method of this invention also includes a low-cost, rapid PCR-based detection pathway: after the RCA step in step 8, the RCA product is directly used as a template for qPCR or ddPCR detection. For chimeras with known junction sites (such as EML4-ALK V1), pre-installed junction-specific primer pairs are used for direct detection; for unknown junction sites, semi-universal primer pairs (anchor region primer + adaptor P7 universal primer) are used for initial HRM screening. A positive result indicates the presence of a potential chimera, and then the sequencing verification process begins.
[0103] Example 4: Setting up gradient VAF simulated samples and validating the lowest detection limit (LoD)
[0104] (1) Preparation of gradient VAF simulated samples: The artificially synthesized EML4-ALK V1 chimeric DNA fragment (SEQ ID No. 16, 200bp) was incorporated into the cfDNA background of the mixed plasma of healthy donors at a mass ratio. The total amount of cfDNA in each sample was 50ng. The VAF gradient was set to 1%, 0.5%, 0.1%, 0.05%, 0.02%, 0.01% and 0% (negative control). The three gradients of 1%, 0.5% and 0.1% were each independently repeated 5 times, the three gradients of 0.05%, 0.02% and 0.01% which are close to the detection limit were each independently repeated 10 times, and the 0% gradient was repeated 5 times.
[0105] (2) Detection process: After each sample completes probe capture, circularization and rolling circle amplification according to steps 1-8 of Example 2, the ddPCR path described in Example 3 is used to perform qualitative detection with cross-connection point primer pairs (SEQ ID No. 10 / 11); each detection is simultaneously quality controlled with negative control (NC) and template-free control (NTC).
[0106] (3) Determination of the lowest detection limit: The lowest VAF with a qualitative detection rate of not less than 95% in independent repeated experiments is defined as the lowest detection limit.
[0107] (4) Results are shown Figure 1 .
[0108] Depend on Figure 1 It can be seen that the detection rate of 0.05% VAF is 10 / 10, the detection rate of 0.02% VAF is 7 / 10, and the detection rate of 0.01% VAF is 3 / 10. Based on this, the lowest detection limit of the method of the present invention is determined to be 0.05% (based on actual measurement) VAF (50ng cfDNA input).
[0109] Example 5: Individual Verification of Dual-End Anchoring Capture Effect
[0110] (1) Experimental grouping: Take 0.1% VAF EML4-ALK V1 simulated sample (50ng healthy donor cfDNA background) and divide it into four groups: double-end capture group (add complete EML4 left end, EML4 right end, ALK left end, ALK right end probes), single-end capture group A (add only EML4 end probe group), single-end capture group B (add only ALK end probe group), and blank group (no probe added). Each group is repeated 5 times independently.
[0111] (2) Experimental procedure: Each group performed probe hybridization, magnetic bead capture and strict washing according to steps 2-4 of Example 2. The elution product was used to quantify the chimeric copy number using cross-linking point qPCR primer pairs (SEQ ID No. 10 / 11), and the recovery rate was calculated (recovery rate = copy number in elution buffer / input copy number × 100%).
[0112] (3) Results are shown Figure 2 .
[0113] Depend on Figure 2 The results showed that the recovery rate of chimeras in the two-end capture group was 65.9% ± 6.6%; the recoveries in the single-end capture groups A and B were 3.5% ± 1.7% and 2.9% ± 1.2%, respectively; and no chimeras were detected in the blank group. These results indicate that chimera molecules can only withstand stringent washing conditions and be stably enriched when both ends are simultaneously anchored and captured, demonstrating the selectivity of the two-end anchoring capture strategy.
[0114] Example 6: Quantitative evaluation of reagent kit performance
[0115] (1) Capture efficiency: Chimeric simulated samples with known copy numbers (0.1% VAF, 50 ng background cfDNA) were captured according to steps 2-4 of Example 2. The target copy numbers added and eluted were quantified by qPCR. The capture efficiency = eluted copy number / added copy number × 100%, and the results were independently repeated 5 times. The capture efficiency was 64.7% ± 5.9%.
[0116] (2) Chimera detection sensitivity: determined according to Example 4, the lowest detection limit is 0.05% (subject to actual measurement) VAF.
[0117] (3) False positive rate: Plasma cfDNA samples from 10 healthy donors were processed according to the entire process in Example 2, and detected by ddPCR. The number of detected linkage points was counted according to the judgment and filtering rules described in step (10) of Example 2. Result: 0 false positive linkage points were detected in 10 samples, and the false positive rate was 0%.
[0118] (4) Repeatability: 0.1% VAF simulated samples were taken and tested in 3 batches, with 3 replicates per batch (9 times in total). The chimerism detection concordance rate and the coefficient of variation (CV) of the VAF measurement value were calculated. Results: The detection concordance rate was 9 / 9, and the CV of the VAF measurement value was 7.9%.
[0119] Application Example 1: Early Screening Detection of High-Risk Groups for Pancreatic Cancer
[0120] Sample: Plasma sample from a pancreatic cancer patient (Tianjin Cancer Hospital).
[0121] The kit used is: "Tumor Peripheral Blood cfDNA Panoramic Chimera Detection Kit (consisting of cfDNA extraction and purification module, chimera targeted enrichment module, circularization transformation and amplification module, sequencing library preparation module, and quality control module. The cfDNA extraction and purification module consists of 20 mL plasma lysis buffer A, 15 mL binding buffer B, 1 mL silanol magnetic bead suspension, 30 mL wash buffer C, 2 mL elution buffer D, and 2 mg proteinase K; the chimera targeted enrichment module consists of 1 tube of smart anchoring probe panel (lyophilized powder, 0.5~2 mg, probe panel V1.0), 1.5 mL probe reconstitution solution, 5 mL 2X hybridization buffer, 500 μL streptavidin magnetic beads, 10 mL strict wash buffer I, 10 mL strict wash buffer II, and 1 mL elution buffer E; the circularization transformation and amplification module consists of 200 μL end-repair and A-tailing premixed enzyme, 20 μL index-specific adaptor mixture, 50 μL T4 DNA ligase, and 1 mL..." The sequencing library preparation module consists of 1 mL of 2X high-fidelity PCR premix, 100 μL of universal PCR primer mixture, and 5 mL of PCR purification magnetic beads. The quality control module consists of 10 μL of positive control DNA, 10 μL of negative control DNA, DNA quantification standards, library quantification qPCR premix, and primers and probes. Its probe system specifically covers genomically unstable regions and driver genes associated with pancreatic cancer (such as flanking regions of genes like KRAS, TP53, SMAD4, and CDKN2A, as well as potential fusion partner gene regions like NRG1 and ALK).
[0122] Specifically, the probe panel for pancreatic cancer patients in this application example is designed as follows: Based on the pancreatic cancer genome structural variation spectrum (integrating rearrangement, deletion, and fusion events reported in the TCGA-PAAD and ICGC pancreatic cancer cohorts), potential rearrangement hotspot regions related to pancreatic cancer are predefined, covering the flanking regions of driver genes such as KRAS, TP53, SMAD4, and CDKN2A, as well as the intron regions of druggable fusion genes such as NRG1, NTRK1 / 2 / 3, ALK, ROS1, RET, BRAF, and FGFR2, and their common fusion partner gene regions; for the symmetrical ends on both sides of each hotspot region (within 1kb upstream and 1kb downstream), 50-100 probes with a length of 80-120nt and an overlap of 2 ms are designed. Overlapping oligonucleotide probes of 0-30 nt constitute the left and right probe sets of this hotspot region; all probes are biotin-labeled at the 5' end, with a 10-15 nt spacer sequence between the biotin label and the hybridization region; each probe contains 2-3 locked nucleic acid (LNA) monomers, modified at the last 3 bases of the probe's 3' end and / or the central region at positions 45-52 from the probe's 5' end, with a 1-2 nt interval between adjacent LNA monomers, and no LNA modification is introduced within the first 10 nt after biotin labeling at the 5' end; the probe Tm value is 65-75℃, and the GC content is 40-60%; blocking oligonucleotides complementary to highly repetitive sequences (Alu, LINE) in the human genome are added to the probe pool at 25 times the total molar concentration of the probe pool.
[0123] This application example targets the KRAS driver gene hotspot (exon 2, containing codons 12 and 13) and two constituent regions of the SND1-BRAF chimera in pancreatic cancer patient samples that need to be covered. Specifically, the following 6 target regions were selected and probe sets were constructed for each region, as shown in Table 1.
[0124] Table 1 shows the constructed probe sets.
[0125] The representative probe sequences (two for each probe group) and their parameters are shown in Table 2 below.
[0126] Table 2 Representative probe sequences for each probe group
[0127]
[0128] In each of the above probe groups, the backbone probes are arranged bidirectionally along the sense and antisense strands within the corresponding 1kb target region, with adjacent probes overlapping by 20-30 nt (stepping approximately 65 nt). Two LNA modification site variants are designed for each backbone sequence (variant 1: positions 86, 88, and 90, i.e., the last 3 bases at the 3' end; variant 2: positions 45, 47, and 49, i.e., the central region), resulting in 56 probes per probe group. The 5' biotinylate marker of each probe is separated from the hybridization region by a 12 nt poly-dT sequence (5'-TTTTTTTTTTTT-3'). The probe Tm value is 73.8-74.7℃, and the GC content is 40.0-42.2%. The table lists two representative probe sequences (SEQ ID No. 17-28) for each probe group, and the remaining probe sequences are determined one by one within the corresponding target region according to the above overlapping arrangement rules.
[0129] Operation: Strictly follow the 10-step process in Example 2.
[0130] The specific operation steps are as follows: (1) Take 2 mL of the patient's plasma sample, extract cfDNA and quantify it according to step (1) of Example 2, and take 50 ng (2) Following step (2) of Example 2, add the probe panel designed for pancreatic cancer patients, hybridize at 65°C for 16 hours to obtain hybridization products; (3) Following step (3) of Example 2, add streptavidin magnetic beads to the hybridization products, incubate, so that the hybridization products bind to the magnetic beads, wash sequentially with strict washing buffer I (65°C) and strict washing buffer II, elute, purify, and obtain enriched DNA; (4) Following step (4) of Example 2, add end repair and A-tail premixed enzyme to the enriched DNA, incubate, purify, and obtain end-repaired DNA; (5) Following step (5) of Example 2, dilute the end-repaired DNA to <1nM and perform circularization treatment, purify, and obtain purified circularized DNA; (6) Following step (6) of Example 2, add end repair and A-tail premixed enzyme to the circularized DNA, incubate, add an index-containing adapter mixture, incubate, and obtain the circularized product; (7) Following step (7) of Example 2, use Phi29 DNA polymerase was used to amplify the RCA product by rolling circle at 30°C for 14 hours; (8) The RCA product was used to qualitatively detect known targets by following the ddPCR path described in Example 3. On the other hand, PCR amplification, target library preparation and paired-end 150bp sequencing were performed according to steps (8)-(10) of Example 2; (9) The sequencing data were analyzed according to the judgment and filtering rules described in step (10) of Example 2: after removing duplicate reads, independent reads that cross the junction point were counted. The independent reads were required to be ≥5, the alignment length of the split-read on both sides of the junction point was ≥20bp and the MAPQ was ≥30. After filtering from four types of false positive sources, ≥2 independent experimental reproducibility, and ≥20 high-confidence junction point reads, chimeric fragments were detected.
[0131] Results: Bioinformatics analysis simultaneously identified two clinically significant molecular events in the individual's cfDNA: a low-frequency point mutation at codon 12 G12D of the KRAS gene, with an allele frequency (AF) of 0.5%; and a novel chimeric link supporting >20 reads, connecting intron 1 of the SND1 gene to exon 8 of the BRAF gene, forming a rare intergene rearrangement with potential driver function.
[0132] The KRAS G12D point mutation was detected by ddPCR (the mutant probe was labeled with FAM at the 5' end, and the wild-type probe was labeled with VIC at the 5' end), with a VAF of 0.5%, far exceeding the ddPCR detection limit (0.01%), confirming its detection. SND1-BRAF is a novel chimeric linker with an unknown linker sequence. It was precisely identified through sequencing library preparation in steps 9 and 10 and NovaSeq sequencing. Sequencing analysis confirmed that this linker connects intron 1 of the SND1 gene to exon 8 of the BRAF gene, supporting >20 reads. After confirming the linker sequence, specific qPCR / ddPCR primer pairs across the linker were designed to establish a subsequent monitoring protocol.
[0133] Conclusion: Using the method and kit of this invention, extremely low abundances of classical driver mutations and rare, structurally complex chimeric fragments were successfully and non-invasively detected simultaneously in the peripheral blood of high-risk individuals. These results suggest that the individual possesses clear clonal hematopoietic abnormalities or potential molecular characteristics of early pancreatic lesions, achieving very early risk warning and molecular stratification for pancreatic cancer. Based on these results, it is recommended that this individual undergo closer imaging (e.g., MRI / EUS) follow-up monitoring, potentially leading to early detection and intervention for pancreatic cancer.
[0134] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for detecting structurally variant chimeric fragments in cell-free DNA from peripheral blood, characterized in that, Includes the following steps: (1) Extract cfDNA from the sample; (2) Perform probe hybridization on cfDNA to obtain hybridization products; (3) Add magnetic beads to the hybridization product, incubate to allow the hybridization product to bind with the magnetic beads, wash, elute and purify to obtain enriched DNA; (4) Add end-repair and A-tailing premixed enzyme to the enriched DNA, incubate, purify, and obtain end-repaired DNA; (5) Dilute the end-repaired DNA, perform circularization, and purify it to obtain purified circularized DNA; (6) Add end repair and A-tailing premixed enzyme to the circularized DNA, incubate, add adaptor mixture, incubate, and obtain the circularized product; (7) Perform rolling circle amplification on the cyclized product to obtain the PCA product; (8) Perform PCR amplification on the PCA product to obtain the PCR product; (9) Purify the PCR product to obtain the target library, quantify the target library and analyze its fragment distribution to obtain libraries with different indices; (10) According to the requirements of the sequencing platform, the libraries of different indices are mixed in equal molar amounts and paired-end 150bp sequencing is performed. The analysis is carried out through bioinformatics process to identify the chimeric fragments.
2. The method according to claim 1, characterized in that, The probes are designed based on knowledge of human genomics and predefined tens of thousands of "potential rearrangement hotspots". Each probe has a biotinylated 5' end label and a probe length of 80-120 nt.
3. The method according to claim 1, characterized in that, The incubation in step (4) is to incubate at 20°C for 30 minutes, followed by incubation at 65°C for 30 minutes.
4. A kit for detecting structurally variant chimeric fragments in cell-free DNA from peripheral blood, characterized in that, It includes modules for cfDNA extraction and purification, chimera targeted enrichment, circularization transformation and amplification, sequencing library preparation, and quality control.
5. The reagent kit according to claim 4, characterized in that, The cfDNA extraction and purification module includes plasma lysis buffer A, binding buffer B, silanol magnetic bead suspension, washing buffer C, elution buffer D, and proteinase K.
6. The reagent kit according to claim 4, characterized in that, The chimeric targeted enrichment module includes an intelligent anchoring probe panel, probe reconstitution solution, 2X hybridization buffer, streptavidin magnetic beads, strict wash solution I, strict wash solution II, and elution buffer E.
7. The reagent kit according to claim 4, characterized in that, The circularization transformation and amplification module includes an end-repair and A-tail premixed enzyme, an index-specific adaptor mixture, a T4 DNA ligase, a 5X circularization reaction buffer, a Phi29 DNA polymerase and a 10X reaction buffer, a dNTP mixture, and rolling circle primers. The sequence of the rolling circle primer is shown in SEQ ID No.
1.
8. The reagent kit according to claim 4, characterized in that, The sequencing library preparation module includes 2X high-fidelity PCR premix, universal PCR primer mixture, and PCR purification magnetic beads; The universal PCR primer mixture contains a P5 forward primer and a P7 reverse primer. The P5 primer sequence is shown in SEQ ID No. 2, and the P7 primer sequence is shown in SEQ ID No.
3.
9. The reagent kit according to claim 4, characterized in that, The quality control module includes positive control DNA, negative control DNA, DNA quantitative standards, library quantitative qPCR premix, and primers and probes.
10. The use of the kit according to any one of claims 4 to 9 in the preparation of in vitro diagnostic products for tumor liquid biopsy, early screening, molecular typing, dynamic monitoring or prognostic assessment.