Method, system and kit for detecting genetic variation and chimera
By combining targeted enrichment and long-read sequencing with UMI technology, the problem of existing technologies being unable to effectively detect complex structural variations and low-frequency chimeric variations has been solved, achieving highly sensitive and accurate diagnosis of genetic diseases, which is suitable for routine clinical testing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ANZHEN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV
- Filing Date
- 2026-02-12
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies face challenges in effectively detecting complex structural variations and low-frequency mosaic variations when detecting complex genetic diseases. In particular, short-read sequencing cannot cross large fragment inversions, translocations, and rearrangements, MLPA cannot detect copy number neutral structural variations, and long-read sequencing is expensive and has a high error rate, resulting in a low diagnostic rate for genetic diseases.
An integrated approach combining targeted enrichment, long-read sequencing, and unique molecular identifier (UMI) technologies was employed to target and enrich high-molecular-weight DNA, and then combine high-fidelity long-read sequencing and UMI sequence markers for bioinformatics analysis to detect structural variations and low-frequency chimeric variations.
It achieves high sensitivity and high accuracy in detecting complex structural variations and low-frequency mosaic variations, significantly improving the diagnostic rate of difficult genetic diseases, reducing testing costs, and making it suitable for routine clinical testing.
Smart Images

Figure SMS_1
Abstract
Description
Technical Field
[0001] This invention relates to the fields of gene detection and molecular diagnostics, and in particular to a highly sensitive and accurate method for detecting complex structural variations, repetitive sequence abnormalities, and low-frequency mosaic variations in the human genome, which is especially suitable for the diagnosis of difficult cases of genetic diseases such as tuberous sclerosis (TSC). Background Technology
[0002] Currently, the genetic diagnosis of genetic diseases mainly relies on next-generation sequencing (NGS) technology, which is based on short-read sequencing, and multiplex linkage-dependent probe amplification (MLPA) technology, which is used to detect large copy number variations (CNVs).
[0003] Short-read NGS panel sequencing is the most widely used technique in clinical practice. This technique involves designing probes to capture or amplify the exons and adjacent splice regions of a target gene using multiplex PCR, followed by high-throughput sequencing (typically with read lengths of 150-300 bp). It excels in detecting single nucleotide variants (SNVs) and small insertions / deletions (Indels). MLPA technology, as a complement to NGS, effectively detects large exon-level deletions or duplications (CNVs) through the ligation and amplification of specific probes.
[0004] To detect low proportions of chimeric variants, existing techniques employ high-depth short-read sequencing (e.g., depth > 1000x) combined with unique molecular identifier (UMI) technology for error correction. This approach improves the detection capability of low-frequency SNVs and indels to some extent.
[0005] Long-read sequencing (LRS) technology, represented by PacBio and Oxford Nanopore, can generate reads ranging from thousands to tens of thousands of base pairs. Theoretically, it can overcome many limitations of short-read sequencing, especially in the detection of structural variations. There are already reports of using LRS for whole-genome sequencing to diagnose rare and difficult-to-diagnose diseases.
[0006] Despite the great success of the above technologies in clinical practice, there are still significant technical bottlenecks when dealing with complex genetic causes, resulting in approximately 10-15% of clinically diagnosed patients not being able to obtain a clear genetic diagnosis (known as "no mutation found", NMI).
[0007] The limitations of existing short-read NGS and MLPA technologies are mainly manifested in the following ways: they cannot effectively detect complex structural variants (SVs). Short reads cannot cross the breakpoints of SVs such as inversions, translocations, and complex rearrangements in large segments, and can only rely on indirect signals (such as read depth changes and segmentation alignment) for inference, resulting in high false positive and false negative rates, and they cannot resolve the precise structure of complex SVs. MLPA cannot detect copy-neutral SVs, such as inversions and balanced translocations. In addition, for highly repetitive sequence regions in the genome (such as tandem repeat amplification and pseudogene regions), short reads cannot be uniquely aligned, forming sequencing "black holes" and causing variants in these regions to be missed.
[0008] Although high-depth short-read sequencing combined with UMI can improve the detection of low-frequency SNVs / Indels, its short-read nature remains unchanged, and therefore it still cannot solve the diagnostic challenges caused by complex SVs and repetitive sequences.
[0009] Existing applications of long-read sequencing also have shortcomings: whole-genome long-read sequencing (WGS-LRS) is expensive and not suitable for large-scale cohort screening or routine clinical testing; while applying LRS to targeted sequencing faces the technical challenge of how to efficiently enrich high-quality long DNA fragments, and traditional hybridization capture methods are mainly designed for short DNA fragments, which is not very efficient; in addition, LRS itself has a certain error rate, and without sufficient error correction, the sensitivity and specificity of directly using it to detect low-frequency chimeric variants are limited.
[0010] Therefore, existing technologies still have significant shortcomings, and there is an urgent need to develop a new gene detection technology or method that can overcome the above-mentioned blind spots and is applicable to routine clinical diagnosis. Summary of the Invention
[0011] To address the shortcomings and deficiencies of existing technologies, this solution aims to provide a novel gene detection method that is highly sensitive, accurate, efficient, and cost-effective. This method innovatively integrates three technologies: targeted enrichment, long-read sequencing, and unique molecular identifiers (UMI). In a single test, it achieves both precise analysis of complex genetic variations and ultra-high sensitivity in identifying chimeric variants, ultimately significantly improving the diagnostic rate of intractable genetic diseases (such as NMI-TSC) and providing patients with clear etiological explanations and precise genetic counseling.
[0012] In one aspect, this invention provides a targeted long-read sequencing method integrating unique molecular identifiers (UMIs), the method comprising the following steps: (a) Extracting high molecular weight DNA from biological samples; (b) The high molecular weight DNA is labeled using a adapter containing a UMI sequence to form a UMI-labeled DNA library; (c) The UMI-labeled DNA library is enriched by using a probe targeting the target locus to obtain an enriched long-fragment DNA library. (d) Perform long-read sequencing on the enriched long-fragment DNA library to obtain sequencing data; (e) Perform bioinformatics analysis on the sequencing data, including error correction and variant detection based on UMI sequences; The method is used to simultaneously detect structural variations and low-frequency chimeric variations within a target genomic region.
[0013] In one embodiment of the present invention, the biological sample is selected from peripheral blood, saliva, amniotic fluid, chorionic villi, tumor tissue, skin tissue, semen, or a combination thereof.
[0014] In one embodiment of the present invention, the fragment length of the high molecular weight DNA is mainly distributed above 20kb, more preferably above 40kb.
[0015] In one embodiment of the present invention, in step (a), the extraction is performed using a magnetic bead method or a special kit to avoid mechanical shearing; and includes a quality control step to ensure that the DNA purity meets the requirements of A260 / 280>1.5 and A260 / 230>1.8.
[0016] In one embodiment of the present invention, in step (b), the UMI sequence is a random nucleotide sequence with a length of 10-30 bp, preferably 18 bp; the markers include end repair, dA tail addition, and adapter ligation.
[0017] In one embodiment of the present invention, in step (c), the probe is a biotinylated capture probe that covers the full-length region, introns, exons, and flanking regulatory regions of the target gene.
[0018] In one embodiment of the present invention, in step (c), the enrichment is carried out using liquid-phase hybridization capture or a targeted cleavage method based on the CRISPR-Cas system; preferably, the hybridization time is 2-24 hours, more preferably 4-16 hours.
[0019] In one embodiment of the present invention, in step (d), the long read sequencing uses a high-fidelity sequencing platform, including the PacBio or Oxford Nanopore system; the sequencing depth is >300x, preferably >500x; the sequencing produces sequences with read lengths of several thousand to tens of thousands of bases.
[0020] In one embodiment of the present invention, step (e) includes: grouping reads based on UMI sequences and performing multiple sequence alignment to generate molecularly consistent sequences; then detecting structural variations, SNVs, indels, and chimeric variations; preferably, using a statistical model to calculate the variation confidence level to support the detection of chimeric variations with allele frequencies as low as 0.5%.
[0021] In one embodiment of the present invention, the method further includes a false positive filtering step, which uses machine learning tools to score and verify candidate variants.
[0022] In one embodiment of the present invention, the target locus includes genes associated with genetic diseases, such as TSC1 / TSC2 or similar genes; the method is used to diagnose genetic diseases, including but not limited to rare diseases, tumor-related genetic variations, or mosaic diseases.
[0023] In a second aspect, the present invention provides a kit for using the sequencing method described above, the kit comprising: a high molecular weight DNA extraction reagent, an adapter containing a UMI sequence, a capture probe targeting a target locus, streptavidin magnetic beads, and sequencing reagents.
[0024] In one embodiment of the present invention, the kit further includes spectrophotometer reagents or electrophoresis reagents for quality control, as well as bioinformatics analysis software or scripts.
[0025] A third aspect of the present invention provides a system for targeted long-read sequencing that integrates unique molecular identifiers, the system comprising: DNA extraction module, used to extract high molecular weight DNA from biological samples; The library building module is used to label the DNA using UMI adapters; Enrichment module for targeted enrichment of DNA libraries with UMI markers; Sequencing module, used for long-read sequencing; Analysis module for UMI-based bioinformatics processing and variant detection.
[0026] In one embodiment of the present invention, the analysis module includes a processor and a storage medium, the storage medium storing instructions for performing UMI-based error correction, mutation detection, and statistical model calculation.
[0027] In one embodiment of the present invention, the system further includes a verification module for verifying candidate variants by ddPCR or targeted ONT sequencing.
[0028] In a fourth aspect, the present invention provides a computer-implemented bioinformatics analysis method for processing targeted long-read sequencing data, the method comprising: (a) Grouping sequencing reads based on UMI sequences; (b) Perform multiple sequence alignment on each UMI group to generate molecular consistency sequences; (c) Align the identical sequence to a reference genome; (d) Detecting structural variations, SNVs, indels, and chimeric variations; (e) Use statistical models to assess the confidence level of variants to support the detection of low-frequency chimeric variants.
[0029] In one embodiment of the present invention, the structural variant detection uses a combination of multiple algorithms, such as Sniffles, pbsv, or cuteSV; the chimeric variant detection is based on read counting supported by UMI, and a background error rate model is constructed using a binomial distribution or a Beta-binomial distribution to calculate the confidence level of low-frequency variants. In one embodiment of the present invention, the method further includes a machine learning filtering step to score candidate variants.
[0030] In a fifth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the above-described bioinformatics analysis method.
[0031] In a sixth aspect, the present invention provides the application of the above-described sequencing method in the preparation of reagents or devices for the diagnosis of genetic diseases.
[0032] In one embodiment of the present invention, the genetic disease includes, but is not limited to, tuberous sclerosis, cancer-related genetic variants, or mosaic-related diseases; the application enables the detection of complex structural variants and low-frequency mosaic variants.
[0033] Compared with the prior art, the present invention has the following beneficial effects: 1. Comprehensive variant detection capability: By adopting long-read sequencing, this approach can directly traverse and resolve complex structural variants, including large-fragment deletions / duplications, inversions, and translocations, and can directly detect high-GC-content and highly repetitive sequence regions, overcoming the technical "blind spots" of short-read sequencing.
[0034] 2. Extremely high detection sensitivity and accuracy: By combining UMI technology with high-fidelity long-read sequencing, this approach can efficiently correct errors in raw sequencing data, thereby detecting chimeric variants with allele frequencies as low as 1% or even lower with extremely high confidence, which is significantly better than conventional NGS methods.
[0035] 3. Multi-dimensional information integration: This approach can simultaneously obtain multi-dimensional genetic information such as SNV, Indel, SV, haplotype phase (due to the long read length which can cover multiple mutation sites) and DNA methylation (natively supported by LRS technology) in a single experiment, providing an integrated solution for elucidating complex pathogenic mechanisms.
[0036] 4. High efficiency and cost-effectiveness: By combining targeted enrichment technology, this approach focuses the advantages of long-read sequencing on specific target gene regions, which greatly reduces the detection cost compared to whole-genome long-read sequencing, making it more suitable for large-scale clinical cohort studies and applications. Detailed Implementation
[0037] The technical solution of the present invention will be further described in detail below with reference to specific embodiments. It should be understood that the following embodiments are merely illustrative and explanatory of the present invention, and should not be construed as limiting the scope of protection of the present invention. All technologies implemented based on the above content of the present invention are covered within the scope of protection intended by the present invention.
[0038] Unless otherwise stated, the raw materials and reagents used in the following examples are commercially available products or can be prepared by known methods.
[0039] Example 1: Library construction and sequencing workflow based on UMI-LRS-Panel The core of this approach is a targeted long-read sequencing (UMI-LRS-Panel) workflow that integrates UMI technology. This workflow comprises five key parts: sample processing, library construction, targeted enrichment, high-fidelity long-read sequencing, and customized bioinformatics analysis.
[0040] 1. Sample processing and high molecular weight (HMW) DNA extraction: Sample types: It can process a variety of biological samples, including but not limited to peripheral blood, saliva, amniotic fluid, chorionic villi, tumor tissue, skin tissue and semen, to support multi-tissue chimera analysis strategies.
[0041] Optimized magnetic bead methods or kits (such as the MagMAX™ HMW DNA Kit) are employed to avoid vigorous vortexing and centrifugation, minimizing mechanical shearing of DNA and ensuring that the extracted HMW DNA fragments are predominantly longer than 40 kb. Quality control of the DNA is performed using nanoliter spectrophotometers and pulsed-field gel electrophoresis to ensure purity meets sequencing requirements (A260 / 280 > 1.8, A260 / 230 > 2.0).
[0042] 2. UMI tagging and library construction: Design and synthesize double-ended DNA adapters containing a random nucleotide sequence (UMI, e.g., 18 bp, 5'-[UMI]-T-3'). The UMI sequence may contain specific patterns to avoid homopolymerization and aid in recognition. After end repair and dA tail addition to the HMW DNA, the UMI adapters are ligated to both ends of the DNA fragment using a ligase. This step ensures that each original DNA molecule is tagged with a unique UMI "barcode." In this example, NEBNext® Ultra was used. TM II. Use the DNA LibraryPrep Kit for end repair and ligation. Enzyme reaction conditions: 25℃ for 30 min (end repair), 16℃ for 1 h (ligation).
[0043] 3. Targeted enrichment optimized for long read lengths: Based on the principle of the patented technology (ZL 202111416309.5), a set of biotinylated capture probes targeting specific gene loci (such as TSC1 / TSC2) was designed and synthesized. This probe set covers the full-length region of the gene (including all exons and introns) and upstream and downstream flanking regulatory regions (e.g., 10kb each), and is optimized for high GC content and repetitive sequence regions to ensure capture uniformity. An optimized liquid-phase hybridization capture protocol was employed. A UMI-labeled DNA library was hybridized with an excess of capture probes for an extended period (e.g., 4-16 hours). Subsequently, streptavidin magnetic beads were used to capture the probe-bound DNA fragments, and after elution, an enriched long-fragment DNA library was obtained. In this embodiment, the hybridization buffer was 5×SSPE, 5× Denhardt's, and 0.1% SDS; the hybridization temperature was 65°C, and the time was 12 hours. Magnetic bead elution was performed using 0.1×TE buffer.
[0044] Alternative approach: CRISPR-Cas9-based targeted cleavage enrichment methods (such as PacBioPureTarget technology) can also be used. The Cas9 enzyme is guided to both sides of the target region by guide RNA for cleavage, and then a sequencing adapter is connected to achieve enrichment of the target region.
[0045] 4. High-fidelity long-read sequencing: Use a PacBio HiFi sequencing platform (such as the Revio or Sequel IIe system). HiFi sequencing generates high-precision congruent reads (>99.9% accuracy) through multiple circular sequencing of single molecules, which is crucial for accurate identification of SNVs and indels. Set a high target effective sequencing depth (e.g., >500x) to ensure sufficient raw reads to construct high-confidence UMI congruent sequences, thereby supporting the detection of low-frequency chimeric variants.
[0046] 5. Customized bioinformatics analysis workflow: Data preprocessing (UMI error correction): A dedicated script is developed to first group (cluster) the reads based on the UMI sequences at both ends. Reads with the same UMI sequence are considered to originate from the same original DNA molecule. Multiple sequence alignment is performed on all reads within each UMI group to generate a highly accurate molecular consistency read. This step reduces the original sequencing error and PCR amplification error rate to extremely low levels (e.g., one in a million), laying the foundation for subsequent low-frequency variant detection.
[0047] Variation detection: All generated molecularly consistent sequences are aligned to the reference genome.
[0048] Structural Variation (SV) Detection: Multiple SV detection algorithms optimized for long read data, such as Sniffles2, pbsv, and cuteSV, are used in parallel, and the results are intersected or integrated to improve the accuracy of SV detection.
[0049] SNV / Indel detection: High-precision algorithms such as DeepVariant are used for detection.
[0050] Chimeric variant detection: A statistical model is developed based on read counting supported by UMIs. For each candidate variant site, the number of independent UMIs supporting the reference allele and the variant allele is statistically calculated. By comparing with background error rate models (such as binomial distribution or beta-binomial distribution), the confidence level of the variant's presence is calculated, enabling highly sensitive detection of chimeric variants with allele frequencies (AF) as low as 1%.
[0051] False positive filtering and verification: Introducing machine learning-based SV filtering tools, such as Samplot-ML or SVDF, allows these tools to learn the characteristics of real SVs and false positives on alignment maps, scoring and automatically filtering candidate SVs, greatly improving analysis efficiency and accuracy.
[0052] All selected candidate pathogenic variants were ultimately validated using independent experimental techniques such as ddPCR or targeted ONT sequencing.
[0053] Example 2: Sensitivity verification of low-frequency chimerism mutation detection To verify the detection capability of this method for low-frequency variants, simulated chimeric samples were constructed. Positive DNA containing a known specific mutation (c.5021delC) in the TSC2 gene was mixed with wild-type DNA in different proportions to prepare test samples with theoretical mutation frequencies of 5%, 1%, 0.5%, and 0.1%, respectively.
[0054] The method described in this invention was used for detection, and the consistency between the detected mutation frequency and the expected frequency was statistically analyzed. The results are shown in Table 1 below: Table 1: Detection results of simulated chimeric samples at different frequencies
[0055] Conclusion: Experimental data show that, through UMI error correction technology, this method can effectively eliminate background noise and stably detect chimeric variants with a frequency as low as 0.5% at sequencing depths of approximately 800-1000×. Moreover, the measured frequency is highly consistent with the theoretical value, verifying the high sensitivity of this method in detecting low-frequency chimeras.
[0056] Example 3: Diagnostic Application of Clinically Difficult Cases (NMI) One TSC patient with typical clinical presentations (such as facial angiofibroma or subependymal giant cell astrocytoma) but negative in two short-read NGS Panel tests was selected.
[0057] Using the method of this invention for detection, data analysis revealed: 1. Detection of structural variations (SVs): A 2.5kb inversion variant was found in the intron region of the TSC2 gene. Because the breakpoint of this variant is located in a high-GC repetitive sequence region and its span exceeds the NGS read length, it was previously missed by routine testing.
[0058] 2. Detection of low-frequency chimeric mutations: A point mutation c.1177A>G was found in exon 12 of the TSC1 gene in another sample. The supporting reads after deduplication by UMI showed that the variant frequency (VAF) of this mutation was only 3.2%, which is lower than the confidence threshold of conventional NGS.
[0059] Validation: The above findings were validated using ddPCR. The validation results confirmed the existence of the variant. This indicates that the method of the present invention can effectively solve difficult cases that cannot be diagnosed by conventional techniques, significantly improving the diagnostic rate.
[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A targeted long-read sequencing method integrating unique molecular identifiers (UMIs), characterized in that: The method includes the following steps: (a) Extracting high molecular weight DNA from biological samples; (b) The high molecular weight DNA is labeled using a adapter containing a UMI sequence to form a UMI-labeled DNA library; (c) The UMI-labeled DNA library is enriched by using a probe targeting the target locus to obtain an enriched long-fragment DNA library. (d) Perform long-read sequencing on the enriched long-fragment DNA library to obtain sequencing data; (e) Perform bioinformatics analysis on the sequencing data, including error correction and variant detection based on UMI sequences; The method is used to simultaneously detect structural variations and low-frequency chimeric variations within a target genomic region.
2. The sequencing method according to claim 1, characterized in that: In step (a), the biological sample is selected from peripheral blood, saliva, amniotic fluid, chorionic villi, tumor tissue, skin tissue, semen, or a combination thereof; Preferably, the length of the high molecular weight DNA fragments is mainly distributed above 20kb, more preferably above 40kb; Preferably, the extraction is performed using the magnetic bead method or a dedicated kit; Preferably, a quality control step is also included to ensure that the DNA purity meets the requirements of A260 / 280>1.5 and A260 / 230>1.
8.
3. The sequencing method according to claim 1, characterized in that: In step (b), the UMI sequence is a random nucleotide sequence with a length of 10-30 bp, preferably 18 bp; the markers include end repair, dA tail addition, and adapter ligation.
4. The sequencing method according to claim 1, characterized in that: In step (c), the probe is a biotinylated capture probe that covers the full-length region, introns, exons, and flanking regulatory regions of the target gene. Preferably, the enrichment is performed using liquid-phase hybridization capture or a targeted cleavage method based on the CRISPR-Cas system; preferably, the hybridization time is 2-24 hours, more preferably 4-16 hours. In preferred step (d), the long read sequencing uses a high-fidelity sequencing platform, including the PacBio or Oxford Nanopore system; the sequencing depth is >300x, preferably >500x; the sequencing produces sequences with read lengths of several thousand to tens of thousands of bases.
5. The sequencing method according to claim 1, characterized in that: In step (e), the bioinformatics analysis includes: grouping reads based on UMI sequences and performing multiple sequence alignment to generate molecularly consistent sequences; then detecting structural variations, SNVs, indels, and chimeric variations; preferably, using a statistical model to calculate the variation confidence level to support the detection of chimeric variations with allele frequencies as low as 0.5%; Preferably, the method further includes a false positive filtering step, using machine learning tools to score and verify candidate variants; Preferably, the target locus includes genes associated with genetic diseases, such as TSC1 / TSC2 or similar genes; the method is used to diagnose genetic diseases, including but not limited to rare diseases, tumor-related genetic variations, or mosaic diseases.
6. A kit comprising: High molecular weight DNA extraction reagents, adapters containing UMI sequences, capture probes targeting specific loci, streptavidin magnetic beads, and sequencing reagents; Preferably, the kit also includes spectrophotometer reagents or electrophoresis reagents for quality control, as well as bioinformatics analysis software or scripts.
7. A system for targeted long-read sequencing that integrates unique molecular identifiers, said system comprising: DNA extraction module, used to extract high molecular weight DNA from biological samples; The library building module is used to label the DNA using UMI adapters; Enrichment module for targeted enrichment of DNA libraries with UMI markers; Sequencing module, used for long read sequencing; Analysis module for UMI-based bioinformatics processing and variant detection; Preferably, the analysis module includes a processor and a storage medium, the storage medium storing instructions for performing UMI-based error correction, mutation detection, and statistical model calculations; The system further includes a validation module for validating candidate variants via ddPCR or targeted ONT sequencing.
8. A computer-implemented bioinformatics analysis method for processing targeted long-read sequencing data, the method comprising: (a) Grouping sequencing reads based on UMI sequences; (b) Perform multiple sequence alignment on each UMI group to generate molecular consistency sequences; (c) Align the identical sequence to a reference genome; (d) Detecting structural variations, SNVs, indels, and chimeric variations; (e) Use statistical models to assess the confidence level of variants to support the detection of low-frequency chimeric variants; Preferably, the structural variation detection uses a combination of multiple algorithms, such as Sniffles, pbsv, or cuteSV; the chimeric variation detection is based on read counting supported by UMI, and uses a binomial distribution or Beta-binomial distribution to construct a background error rate model to calculate the confidence of low-frequency variations; Preferably, the method further includes a machine learning filtering step to score candidate variants.
9. A computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the bioinformatics analysis method of claim 8.
10. The use of the sequencing method according to any one of claims 1-5 and the kit according to claim 6 in the preparation of reagents or devices for the diagnosis of genetic diseases; Preferably, the genetic diseases include, but are not limited to, tuberous sclerosis, cancer-related genetic variations, or mosaic-related diseases; the application enables the detection of complex structural variations and low-frequency mosaic variations.