Single-ended targeting single-ended random molecular tagging method for quantitative analysis of polynucleic acid variation and application of single-ended targeting single-ended random molecular tagging method
By introducing random primers with adapters and UMIs during RNA reverse transcription and combining them with nested PCR amplification, the problems of insufficient sensitivity and large quantitative errors in existing fusion gene detection are solved, realizing the detection of unknown fusion events with high specificity and high sensitivity, which is suitable for high-throughput screening and dynamic monitoring of diseases such as tumors.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI JIAOTONG UNIV
- Filing Date
- 2025-12-04
- Publication Date
- 2026-04-14
AI Technical Summary
Existing fusion gene detection technologies suffer from problems such as insufficient sensitivity, large quantitative errors, and difficulty in identifying unknown fusion events, especially in low-abundance RNA samples or complex backgrounds.
The single-end targeted single-end random molecular tagging method is used to generate single-stranded or double-stranded cDNA with adapters and adapters by introducing random primers with adapters and UMIs during RNA reverse transcription. Combined with nested PCR amplification, this method enables the detection of fusion genes with high specificity and high sensitivity.
It improves the sensitivity and accuracy of fusion gene detection, can identify unknown fusion events, reduces costs and experimental errors, and is suitable for high-throughput screening and dynamic monitoring of diseases such as tumors.
Smart Images

Figure CN121852537A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fusion gene detection technology, and more specifically to a single-end targeted single-end random molecular tagging method for quantitative analysis of multiple nucleic acid variations and its application. Background Technology
[0002] Fusion genes are important drivers and specific molecular markers for various tumors, and are widely used in early tumor screening, subtyping diagnosis, therapeutic target identification, and efficacy monitoring. Currently, commonly used clinical methods for detecting fusion genes include FISH, RT-PCR, and next-generation sequencing. However, these methods generally suffer from insufficient sensitivity, large quantitative errors, and difficulty in identifying unknown fusion events, especially in low-abundance RNA samples or complex backgrounds (such as tumor heterogeneity or liquid biopsy), where they exhibit significant limitations.
[0003] On the other hand, existing RNA sequencing technologies generally lack a mechanism for introducing a unique molecular identifier (UMI) during library construction, making it difficult to accurately identify the original molecule. They are also susceptible to interference from PCR duplication and amplification bias, affecting the judgment of the true expression level and limiting their application in high-precision fusion gene detection.
[0004] Fusion genes, whether of known or unknown sequences, can drive various cancers. Their effects may include forming oncogenic fusion proteins and inactivating tumor suppressor factors. Many fusion genes include a known target gene, characterized by one or more unknown genes fusing to a target gene with a known function. Several methods are currently available for detecting such fusion events of known driver genes.
[0005] IDT's Archer FUSIONPlex product line uses RNA-targeting technology (such as...) Figure 1As shown in the figure, this method is used to identify tumor-related fusion genes, splicing variants, single nucleotide variants, insertions, deletions, and relative expression. Based on anchored multiplex PCR (AMP), this method combines gene-specific primers (GSP) to amplify and enrich target regions, and incorporates molecular barcodes (MBC) technology for error correction. The specific procedure is as follows: total RNA molecules are reverse transcribed to generate double-stranded cDNA; nested PCR using GSP generates an enriched library; each molecule is uniquely labeled with MBC; after multiple rounds of PCR amplification, sequencing is performed using a next-generation sequencing platform (NGS); the resulting FASTQ file is processed using Archer analysis software; after sequence deduplication, error correction, and data analysis, an analysis report including fusion genes is finally generated.
[0006] QIAGEN's QIAseq RNA Fusion XP series of products can also detect RNA fusions in a single sample (e.g., Figure 2 (As shown). This method combines single primer extension (SPE) chemistry to amplify the target region using a single gene-specific primer, and uses unique dual index (UDI) technology to correct sequencing errors. The specific procedure is as follows: total RNA molecules are reverse transcribed to generate double-stranded cDNA. After end repair and A-tailing of the double-stranded cDNA, it is ligated with an adapter containing a UMI. SPE is used to amplify the target region. After universal PCR amplification, QIAseq analysis software is used to analyze the raw sequencing results and generate analysis reports such as fusion genes and single-base mutations.
[0007] NuProbe USA's Random Reverse Primers-based fusion gene detection method (patent number: US20250011833A1) uses random primers to detect RNA fusion, single-base mutations, alternative splicing, etc. The specific process is as follows: total RNA molecules are reverse transcribed to generate double-stranded cDNA. Random primers are then combined with single-stranded or double-stranded cDNA through PCR denaturation and annealing processes. Targeted primers are then used to capture and amplify the random primers.
[0008] However, in general, the currently commonly used clinical methods for detecting fusion genes have the following shortcomings:
[0009] WGS sequences the entire genome in a sample at the same time. It has low sensitivity in detecting fusion genes, which limits its application in low-frequency detection. If the sensitivity is improved by increasing the sequencing depth, the cost will increase significantly, and the cycle will be long and require complex bioinformatics analysis.
[0010] RNA-seq performs sequencing analysis on the entire transcriptome in a sample, which has a smaller coverage than WGS. It only detects fusion genes at the transcriptional level, but because RNA samples are easily degraded, the preparation requirements are high, and it is difficult to accurately detect long fusion genes.
[0011] PCR detection of fusion genes uses paired-end primers to amplify specific regions of the fusion gene, which is relatively simple and fast. However, it can only detect fusion genes with known gene sequences and cannot detect unknown fusion genes.
[0012] Targeted detection methods, such as the Archer FUSIONPlex series and the QIAGEN QIAseq RNA Fusion XP series, can detect unknown fusions of specific target genes. Compared with other methods, they are less expensive. However, their disadvantage is that the experimental design includes an adapter ligation step, which has low ligation efficiency and cannot ensure that all sequences are successfully ligated, resulting in low yield and the possibility of missing fusion events.
[0013] NuProbe USA's detection method is based on random primers and adapters. While this avoids the inefficiency of the ligation step, its drawbacks include the addition of adapters after the synthesis of single-stranded or double-stranded cDNA. This results in biases introduced by the reverse transcription process, loss of quantitative information of the original RNA molecule, and the addition of an extra denaturation-annealing process, leading to more experimental steps and theoretically introducing more experimental biases.
[0014] Therefore, there is an urgent need to develop a fusion gene detection method that can solve the problems of insufficient sensitivity and large quantitative error that are common in existing fusion gene detection technologies. Summary of the Invention
[0015] The purpose of this invention is to provide a single-end targeted single-end random molecular tagging method for quantitative analysis of multiple nucleic acid variations and its application, thereby solving the problems of insufficient sensitivity, large quantitative error, and difficulty in identifying unknown fusion events that are common in existing fusion gene detection technologies.
[0016] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0017] According to a first aspect of the present invention, a single-end targeted single-end random molecular tagging method for quantitative analysis of multiple nucleic acid variations is provided, comprising the following steps:
[0018] 1) Sample RNA extraction: Extract total RNA from the sample to be tested;
[0019] 2) One-stranded cDNA synthesis: The total RNA was reverse transcribed using a one-stranded reverse transcriptase and a random primer RRP-UMI containing a linker and a UMI to generate single-stranded cDNA. The structure of the random primer RRP-UMI containing the linker and UMI from the 5' end to the 3' end is as follows: linker, 15-base UMI, 7-base specific sequence, and 5-40-base random sequence; the 7-base specific sequence is GTGCAAG.
[0020] 3) Two-stranded cDNA synthesis: The single-stranded cDNA obtained in step 2) is reverse transcribed using a two-stranded reverse transcriptase to generate double-stranded cDNA;
[0021] 4) First round of nested PCR amplification: Prepare the PCR system, add the outer PCR forward primer designed for the target gene and the universal outer PCR reverse primer, and perform PCR amplification using the double-stranded cDNA obtained in step 3) as a template;
[0022] 5) Second round of nested PCR amplification: Prepare the PCR system, add inner PCR forward primers designed for the target gene and universal inner PCR reverse primers, and perform highly specific PCR amplification using the PCR amplification product of step 4) as a template;
[0023] 6) Sequencing adapter ligation: Prepare the PCR system, add sequencing adapter primers compatible with the NGS sequencer, amplify the PCR amplification product from step 5), ligate the sequencing adapters, and obtain the library preparation product.
[0024] 7) NGS Sequencing and Analysis: The product was sequenced using an NGS sequencer to obtain FASTQ data, which was then subjected to bioinformatics analysis to screen for fusion gene detection results.
[0025] Preferably, in step 2), the length of the 5-40 base random sequence is 6-20 bases.
[0026] Preferably, in step 2), the modification of the linker portion 3' of RRP-UMI includes at least one of unmodified, complementary nucleotide, blocking probe with secondary structure, phosphorylation modification, and locked nucleic acid modification.
[0027] Preferably, the design of the outer PCR forward primer in step 4) and the inner PCR forward primer in step 5) are both based on one or more of the following parameters: primer binding free energy ΔG, primer self-complementarity, primer dimer formation probability, GC content, homopolymer length, and secondary structure stability.
[0028] Preferably, the design of the outer PCR forward primer in step 4) and the inner PCR forward primer in step 5) are both based on one or more of the following parameters: at 60 ℃, 0.18M Na + Under the given conditions, the primer binding free energy ΔG should be in the range of -8 to -13 kcal / mol; primers outside this range will be discarded. In the primer self-complementarity assessment, the maximum number of complementary bases should be 4-15; the GC content should be in the range of 40%-60%; the maximum number of homopolymers should be 4. After primer design, evaluate the dimer and secondary structure formation of all primers, and discard primers with obvious stable secondary structures and severe dimers.
[0029] Preferably, the method further includes a purification step: after cDNA synthesis in step 3), after the first round of nested PCR amplification in step 4), after the second round of nested PCR amplification in step 5), and after sequencing adapter ligation in step 6), the corresponding PCR products are purified respectively; the purification adopts the magnetic bead purification method: 1-2 times the volume of PCR product is added to magnetic beads for mixing and incubation. After the magnetic beads are adsorbed, the supernatant is discarded, the mixture is washed with ethanol, and then the DNA on the magnetic beads is eluted with 0.1×TE buffer.
[0030] Preferably, in step 7), the sequencing data is processed by a bioinformatics analysis process. The fusion gene detection adopts a breakpoint alignment and cross-connection read identification algorithm, the mutation detection adopts a standard process of alignment-variant call-quality filtering, the alternative splicing event detection is based on the exon-connected read count, and the expression level calculation is based on the normalized read count after UMI deduplication.
[0031] Preferably, in step 1), the sample to be tested includes peripheral blood plasma, peripheral blood serum, body fluid, tissue sample, adherent cell sample, or suspended cell sample.
[0032] Preferably, in step 4), the universal forward primer sequence for Outer PCR is TGCGATGCAATGAGAACATTAGAAGTCAGCTGCATGACTGAAGA, and the universal reverse primer sequence for Outer PCR is GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG.
[0033] Preferably, in step 4), the Outer PCR forward primer includes a forward primer with a universal adapter or a forward primer without a universal adapter.
[0034] Preferably, in step 5), the universal reverse primer sequence for inner PCR is CGTGGGCTCGGAGATGTG.
[0035] According to a second aspect of the present invention, a method for detecting an unknown fusion gene is provided, comprising constructing an unknown fusion gene detection library using the method described above, sequencing the unknown fusion gene detection library, and obtaining the sequence of the unknown fusion gene.
[0036] According to the method of the present invention, the key inventive point is that a random primer with adapter and UMI is provided and introduced into the RNA reverse transcription process and used as a reverse transcription primer to generate single-stranded or double-stranded cDNA with UMI and adapter.
[0037] like Figure 3 As shown, the structure of the random primer with adapter and UMI from the 5' end to the 3' end is as follows: adapter, 15-base UMI (H15), 7-base special sequence (X7), and 5-40-base random sequence (Nx).
[0038] The random sequence (Nx) consists of a random nucleotide sequence of 5-40 bases in length, where N is a random base sequence (A, G, C, or T), and the number represents the number of bases. More preferably, the random sequence is 6-20 bases in length. A unique 7-base sequence, GTGCAAG, is attached to the 3' end of the random sequence and is uniquely defined within the random primer. The 15-base UMI (Unique Molecular Identifier) length is designed to significantly reduce the probability of UMI duplication between different molecules compared to a 10-base UMI, while avoiding the accumulation of synthesis and sequencing errors that might occur with a 20-base UMI. According to the research findings of this invention, a 15-base UMI can meet the requirement of unique labeling of the original molecule in practical applications, and its duplication rate is within an acceptable range. The adapter serves as a primer binding site for amplification during subsequent PCR.
[0039] According to the method of the present invention, the targeted genes are analyzed and screened through database data, which includes a series of genes related to the disease, and the design principle is as follows: Figure 4As shown, for the corresponding exons of the target gene, the outer PCR primers, used as the first-round primers for nested PCR, are designed in the 5' end to the middle region of the exon. After the first round of nested PCR amplification, they generate DNA sequences from the exon to the next exon or from the exon to the fusion gene. The inner PCR primers, used as the second-round primers for nested PCR, are designed in the middle region to the 3' end of the exon. The inner PCR primers will bind to the products of the first round of nested PCR amplification during the second round of nested PCR amplification.
[0040] like Figure 4 As shown, nested PCR design requires two pairs of primers. First, the first-round outer PCR primers are used to amplify the target gene region, initially amplifying the DNA sequence containing the target gene. Then, the outer PCR primers, ligating a universal sequence adapter, can be used for further amplification using universal forward primers. Alternatively, outer PCR primers without a universal sequence adapter can be used to amplify the target gene.
[0041] After the first round of nested PCR amplification, a second round of nested PCR amplification was performed using inner PCR primers to amplify the products from the first round. This multi-round amplification of the target gene region increases specificity: the first-round outer PCR primers capture fragments that may contain fusion genes over a wider area, while the second-round inner PCR primers are located closer to the fusion breakpoint, specifically amplifying only when a real fusion event is present. This "broad-to-precise" primer layout effectively reduces non-specific amplification and false positive signals, thus significantly improving the specificity of fusion gene detection. The key difference lies in the design of the outer PCR primers in a broader region outside the breakpoint, while the inner PCR primers are precisely located in the sequence adjacent to the breakpoint. The combination of the two rounds of amplification ensures high specificity and accuracy in detection.
[0042] The procedure of the single-end targeted single-end random molecular tagging method for quantitative analysis of multiple nucleic acid variations provided by the present invention is as follows: Figure 5 As shown, the main steps include:
[0043] S1: Extract total RNA from the sample and bind the RNA sample to random primers;
[0044] S2: Use reverse transcriptase and RRP-UMI to perform one-strand reverse transcription on the RNA sample, converting the RNA into single-stranded cDNA with UMI and adapter;
[0045] S3: Mix the product with the two-strand reverse transcription system and perform two-strand reverse transcription to generate double-stranded cDNA with UMI and adapter from the single-stranded cDNA sample.
[0046] S4: First round of nested PCR amplification: According to the experimental design, the outer PCR primers of the target gene are mixed in a certain proportion to prepare the corresponding PCR system. The cDNA sequence with adapter is added as a template and denaturation, annealing and extension PCR cycles are performed. During the cycle, the target gene primers specifically bind to the target gene region and only the target gene is amplified. Non-target regions are not amplified due to the lack of primers.
[0047] S5: Second round of nested PCR amplification: According to the experimental design, the inner PCR primers of the target gene are mixed in a certain proportion to prepare the corresponding PCR system. In this step, the inner PCR primers specifically bind to the sequence of the target gene amplification product in the previous step and amplify the target gene again to improve the specificity of the target gene. Non-target regions lack primers and are not amplified.
[0048] S6: Sequencing adapter ligation: Prepare the appropriate PCR system and add the adapter primers that are compatible with the NGS sequencer in a certain arrangement and combination order. Amplify the DNA fragment with the adapter sequence to achieve ligation of the sequencing adapter and complete library construction.
[0049] S7: Sequencing and Analysis: The library-constructed products were sequenced using an NGS sequencer, and the FASTQ data after sequencing were compared and screened using bioinformatics analysis software BWA and STAR. The results were then screened using Python and R to obtain the fusion gene detection results.
[0050] Preferably, according to the present invention, after each PCR cycle, a purification step is designed to purify the PCR product, such as... Figure 6 As shown, the magnetic beads are first mixed with the PCR product and incubated. After that, the magnetic beads with DNA are adsorbed onto the tube wall using a magnetic plate. After discarding the solution, ethanol is added for washing. Finally, the DNA on the magnetic beads is eluted with TE buffer, thus completing the purification.
[0051] The purpose of this invention is to provide a method for detecting unknown fusion events in targeted gene detection. It should be understood that, in this invention, an unknown fusion event refers to a type of fusion event that, in targeted gene detection, does not require prior determination of the specific sequence information at both ends of the fusion breakpoint; detection can be performed only if the sequence of one end of the gene is known. This invention successfully overcomes the design limitations of traditional PCR technology, which relies on knowing the sequences at both ends, and can effectively identify and verify novel or rare gene fusions where only one end of the sequence is known.
[0052] The key inventive point of this invention lies in the design of random primers. During reverse transcription, the adapter + UMI is integrated into the reverse transcription random primer (RRP) to obtain single-stranded or double-stranded cDNA with the adapter. This achieves simultaneous adapter addition during the reverse transcription stage, solving the problems of low efficiency and inaccurate quantification after adapter addition following cDNA reverse transcription in existing technologies. Furthermore, the combination of random primers and targeted nested PCR simultaneously achieves unknown fusion capture and highly specific amplification, whereas existing technologies can only achieve one of these. It should be understood that most existing technologies for detecting fusion genes follow the logic of "fixing one end and detecting the other end," and those skilled in the art generally believe that the detection of unknown fusions depends on the knowledge of both ends. This invention is the first to propose that when only one end sequence is known, unbiased capture by random primers combined with UMI quantification breaks the technical path dependence of bidirectional fixation. Furthermore, existing technologies employ the logic of "reverse transcription first, then adapter addition," which reverse transcribes RNA into single-stranded or double-stranded cDNA and then adds adapters using random primers or ligation methods. This invention is the first to propose using RRP-UMI as a primer during the RNA reverse transcription stage to directly reverse transcribe and amplify single-stranded or double-stranded cDNA molecules with adapters, breaking the technological dependence on adding adapters to cDNA or DNA molecules.
[0053] As described in the background section of this invention, next-generation sequencing methods such as WGS and RNA-seq can detect unknown fusion genes, point mutations, and alternative splicing, but these methods are time-consuming and costly. ArcherFUSIONPlex series products and QIAGEN QIAseq RNA Fusion XP series products can achieve targeted sequencing, but their library preparation process includes adapter ligation, and their yield is affected by ligation efficiency. NuProbe's patented products can achieve targeted sequencing without a ligation step, but they lose information about the original RNA molecules, and their experimental steps are numerous, introducing more errors. However, this invention, through the functional integration of random primers, spatial stratification of nested PCR, and innovative molecular markers in UMI, solves the three major technical challenges that have long plagued existing technologies: inaccurate detection of the unknown end, large quantitative errors in low-abundance assays, and high costs of multi-dimensional detection.
[0054] Compared with existing technologies, the single-end targeted single-end random molecular tagging method for quantitative analysis of multiple nucleic acid variations provided by the present invention has the following advantages:
[0055] 1) Compared with methods such as WGS and RNA-seq, this method detects fusion genes targeting the target gene, which has higher sequencing depth and higher sensitivity, and its cost is lower because it only amplifies the target region.
[0056] 2) Compared with PCR technology, this method is based on random primer technology. Random primers can bind to all sequences in the genome without bias, so it can detect fusion events with unknown fusion partners. It has a wider detection range and can detect fusion genes with unknown sequences.
[0057] 3) Compared with Archer FUSIONPlex series products and QIAGEN QIAseq RNA Fusion XP series products, this method combines adapter and UMI design with random primers. Its amplification process replaces the adapter addition step in ordinary sequencing methods, and the random primer adapter addition method has higher ligation efficiency.
[0058] 4) Compared with NuProbe's patented products, this method incorporates the adapter addition step into the RNA reverse transcription process, which is superior to adding adapters at the cDNA or DNA level. It preserves the original quantitative information of the RNA molecule, reduces experimental steps, and lowers experimental biases introduced by unnecessary experimental procedures.
[0059] In summary, the single-end targeted single-end random molecular tagging method for quantitative analysis of multiple nucleic acid variations provided by this invention utilizes specially designed random reverse primers (RRPs) with adapters and UMIs to replace the random primers used in RNA reverse transcription. This adds adapters to the cDNA sequence during RNA reverse transcription, reducing loss during RNA reverse transcription, preserving the original RNA molecular information, and improving adapter ligation efficiency. By using the designed forward primer for the target gene and the RRP reverse primer, fusion sequences of driver genes and unknown genes can be amplified, enabling the detection of fusion genes. This improves the detection rate, accuracy, and ability to capture unknown fusion events. This method is suitable for high-throughput screening and dynamic monitoring of fusion genes in diseases such as tumors, providing reliable technical support for biomarker development and drug efficacy evaluation. Attached Figure Description
[0060] Figure 1 This is an analysis flowchart for existing technology IDT's Archer FUSIONPlex series products;
[0061] Figure 2 This is an analysis flowchart for QIAGEN's QIAseq RNA Fusion XP series products, which are based on existing technology.
[0062] Figure 3 This is a schematic diagram of the design principle of random primers provided by the present invention;
[0063] Figure 4 This is a schematic diagram of the design principle of the targeted primers and nested PCR provided by the present invention;
[0064] Figure 5 This is a flowchart of the single-end targeted single-end random molecular tag method for detecting fusion genes for quantitative analysis of multiple nucleic acid variations provided by the present invention;
[0065] Figure 6 This is a schematic diagram of the magnetic bead purification process provided by the present invention;
[0066] Figure 7 This is a comparison of reverse transcription efficiency using different random primers in Example 1 (quantified using GAPDH gene).
[0067] Figure 8 This is a comparison of the library construction effects of adding connectors at different stages in Example 1;
[0068] Figure 9 This is a schematic diagram comparing the cDNA library construction results of reverse transcription using random primers with and without UMI (RRP) in Example 2.
[0069] Figure 10 This is a quantitative analysis chart of the specific data for the uniformity assessment dimension in Figure 9;
[0070] Figure 11 This is a graph showing the effect of using UMI-RRP to quantify the RNA standard of the fusion gene in Example 3. In this graph, fusion1 refers to the fusion gene FGFR2-COL14A1, and fusion2 refers to the fusion gene TPM3-NTRK1. Detailed Implementation
[0071] The present invention will be further described below with reference to specific embodiments. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Unless otherwise specified, the techniques used in the embodiments are conventional practices in the art, or experimental methods recommended by the reagent kit and instrument manufacturers. Unless otherwise specified, the reagents and materials used in the embodiments are commercially available.
[0072] Example 1
[0073] This embodiment provides a multi-dimensional detection method for tumor-related fusion genes, point mutations, alternative splicing events, and gene expression levels. The detection sites described in this embodiment are obtained through high-throughput sequencing data analysis of normal cervical tissue and cervical cancer samples. During the screening process, breakpoint regions of known tumor-related fusion genes, common mutation hotspots, alternative splicing exon junctions, and target regions of differentially expressed tumor-related genes were located in the samples. Statistical analysis was then used to screen out multiple detection sites with significant differences and clinical relevance. This invention protects, but is not limited to, obtaining high-risk site information through various sequencing methods such as targeted library construction sequencing, whole transcriptome sequencing (RNA-seq), and whole exome sequencing (WES), and can be combined with public databases (including but not limited to TCGA, CCLE, and COSMIC) for site annotation and verification.
[0074] After obtaining the target site set, primers are designed based on the target sites. According to the exon positions, the target primers are designed according to the following requirements: 1) The primer binding free energy ΔG is in the range of -8 to -13 kcal / mol (at 60 ℃ and 0.18 M Na). + (Calculated under the following conditions); 2) The maximum number of complementary bases in the primers is 4-15 bases; 3) The GC content is in the range of 40%-60%; 4) The maximum number of homopolymers is 4 bases; 5) Avoid primers with obvious stable secondary structures and primers that form dimers between primers.
[0075] Following experimental design, functional testing was conducted using tumor cell lines to validate the loci and assess their feasibility and stability in a practical detection system. The library construction process was as follows: First, total RNA was reverse transcribed using primers containing a unique molecular identifier (UMI) to generate UMI-labeled cDNA. The introduction of the UMI was used to remove PCR amplification duplications and reduce quantitative bias in subsequent sequencing data. Subsequently, a detection panel was constructed using specific primers designed for the target loci. Simultaneous detection of fusion gene types and breakpoint locations, point mutation sites and frequencies, alternative splicing patterns and abundance, and gene expression levels was achieved through multiplex PCR amplification combined with a high-throughput sequencing platform. Primer design principles for the detection panel included, but were not limited to: primers covering key functional regions of the target loci; the ability to distinguish specific breakpoints of different fusion genes; coverage of all common variants of mutation hotspots; and cross-exon-exon junctions to identify alternative splicing events, while also capturing reference transcript regions for expression level calculations.
[0076] Sequencing data were processed using a bioinformatics analysis workflow. Fusion gene detection employed breakpoint alignment and junction read identification algorithms, mutation detection followed a standard workflow of alignment-variant recall-quality filtering, alternative splicing event detection was based on exon-linked read counts, and expression level calculation was based on normalized read counts after UMI deduplication. The method described in this embodiment can simultaneously obtain multiple types of molecular-level information in a single experiment. Compared to single detection modes, it can more comprehensively reflect the molecular characteristics of tumors, providing data support for tumor subtyping, prognostic assessment, and personalized treatment.
[0077] The detailed experimental procedure includes the following steps:
[0078] (1) Obtain the nucleic acid of the sample to be tested as a template:
[0079] Nucleic acid was extracted from the sample; a method suitable for the sample type was selected for nucleic acid extraction. In this example, the Yeasen MolPure® Flash Cell / Tissue Total RNA Kit (catalog number 19221) was used. The sample types included peripheral blood leukocytes, peripheral blood plasma / serum, body fluids, cells, and tissues.
[0080] Sample preprocessing:
[0081] 1) Animal tissue: Take less than 30 mg of fresh tissue, add lysis buffer and homogenize (glass homogenizer or electric homogenizer can be used), grind into powder if necessary and then mix with lysis buffer, and shake thoroughly to mix.
[0082] 2) Adherent cells: The cells can be directly added to the culture dish to lyse, or the cells can be collected and then mixed with the lysis buffer, and repeatedly pipetted until there are no cell clusters and the cells are completely lysed.
[0083] 3) Suspension cells: After centrifuging to collect cells, discard the supernatant, add lysis buffer, and repeatedly pipette until the cells are fully broken.
[0084] RNA extraction:
[0085] 1) Transfer the lysis buffer to a DNA removal / RNA adsorption universal column, centrifuge at 13,000 rpm for 1 min to collect the filtrate, add 0.5 volumes of anhydrous ethanol to the filtrate and mix well.
[0086] 2) Transfer the mixture to a DNA removal / RNA adsorption universal column, centrifuge at 13,000 rpm for 30 s, and discard the filtrate.
[0087] 3) Wash once with 700 μL of protein removal solution, centrifuge at 13,000 rpm for 30 s and discard the filtrate.
[0088] 4) Wash once with 500 μL of rinsing solution, centrifuge at 13,000 rpm for 30 s and discard the filtrate.
[0089] 5) Repeat the rinsing step once, and centrifuge to remove the residual filtrate.
[0090] 6) Transfer the RNA binding column to a new collection tube, add 30μL-50μL of RNase-free water to elute the RNA, and centrifuge to collect the eluent.
[0091] 7) The RNA solution can be used immediately or stored at -80°C.
[0092] (2) Reverse transcribe the RNA sample into a cDNA sample:
[0093] Specifically, reverse transcription includes the following steps:
[0094] 1) RNA denaturation
[0095] Mix 14 μL of RNA sample with 1 μL of RRP, pipette to mix thoroughly, and set the PCR program as shown in Table 1 for pre-denaturation:
[0096] Table 1
[0097]
[0098] 2) Synthesis of one-stranded cDNA
[0099] Take 15 μL of the pre-denatured sample, mix it with 8 μL of 1st Reaction Buffer and 2 μL of 1st Strand Enzyme Mix, and set the PCR program as shown in Table 2 for one-strand cDNA synthesis:
[0100] Table 2
[0101]
[0102] 3) Synthesis of second-stranded cDNA
[0103] Take 25 μL of first-strand cDNA sample, mix it with 7 μL of 2nd Reaction Buffer and 3 μL of 2nd Strand Enzyme Mix, and set the PCR program as shown in Table 3 to synthesize second-strand cDNA:
[0104] Table 3
[0105]
[0106] Purify using magnetic beads at a volume 1.8 times that of the PCR product to be purified and elute with 40 μL of 0.1×TE.
[0107] (3) Library construction experiment:
[0108] Specifically, the library construction experiment includes the following steps:
[0109] 1) First round of nested PCR amplification:
[0110] According to the experimental design, the outer primers of the target gene were mixed in a certain proportion, and the corresponding PCR systems were prepared according to the allocation in Table 4. Amplification was then performed using the PCR program shown in Table 5.
[0111] Table 4
[0112]
[0113] Table 5
[0114]
[0115] Purify using magnetic beads at a volume 1.2 times that of the PCR product to be purified and elute with 40 μL of 0.1×TE.
[0116] 2) Second round of nested PCR amplification:
[0117] According to the experimental design, the inner primers of the target gene were mixed in a certain proportion, and the corresponding PCR systems were prepared according to the allocation in Table 6. Amplification was then performed using the PCR program shown in Table 7.
[0118] Table 6
[0119]
[0120] Table 7
[0121]
[0122] 3) Purify using 1.2× magnetic beads and elute with 40 μL 0.1×TE. Dilute the product 10-fold and 1000-fold, respectively. Use the 1000-fold diluted product for qPCR quantification, and use the quantification ct value + 4 as the amplification cycle number for the 10-fold diluted product.
[0123] 4) qPCR quantification:
[0124] According to the experimental design, prepare the corresponding qPCR quantification systems according to the group allocation in Table 8 and perform quantification according to the procedure shown in Table 9:
[0125] Table 8
[0126]
[0127] Table 9
[0128]
[0129] 5) Sequencing adapter ligation: Prepare the corresponding PCR systems according to the group assignments in Table 10, add the adapter primers compatible with the NGS sequencer in a specific arrangement and combination order, and amplify according to the PCR program shown in Table 11:
[0130] Table 10
[0131]
[0132] Table 11
[0133]
[0134] 6) Purify using 1.2× magnetic beads and elute with 40 μL 0.1× TE, quantify using Qubit, mix proportionally and send for sequencing.
[0135] Table 12 Primer Design Table
[0136]
[0137] * N5xx and N7xx can be varied depending on the experimental design, referencing the sequencing sequences from the Illumina NovaSeq 6000 platform.
[0138] like Figure 7 As shown, by using GAPDH as an internal reference gene, the reverse transcription efficiency of different primers was quantified. Based on the standard of defining the ct value as 20% of the highest fluorescence value, the GAPDH ct value of the random primer with adapter was 16, while that of the ordinary random primer was 17. The difference between the two GAPDH ct values was approximately 1, indicating that the reverse transcription efficiency of the random primer with adapter was higher, and the amount of cDNA obtained by reverse transcription was twice that of the ordinary random primer.
[0139] like Figure 8 As shown, compared with using random primers to reverse transcribe cDNA first and then adding adapters (cDNA with adapters), the proportion of correctly amplified reads obtained by directly adding adapters during the RNA reverse transcription step was higher, with an average of 91.1%, while the proportion of correctly amplified reads obtained by adding adapters after reverse transcription into cDNA was 80.7%.
[0140] Example 2
[0141] This embodiment describes a detection experiment using cell line samples. Specifically, to verify the performance of the multidimensional tumor molecular detection panel in detecting fusion genes, point mutations, alternative splicing, and expression levels, this embodiment uses the CaSki cervical cancer cell line stably transfected with standard qualitymids as the detection model. After CaSki cells were expanded to the logarithmic growth phase under standard culture conditions, a total cell volume of approximately 5 × 10⁶ cells was collected. 6 Total RNA was obtained using an RNA extraction kit and then reverse transcribed to obtain cDNA samples.
[0142] During library construction, the experiment was conducted in two parallel groups: the first group used reverse random primers (RRP) with UMI adapters for reverse transcription. UMI is used to label RNA at the molecular level, thereby removing PCR amplification duplications and reducing amplification bias during the analysis stage; the second group used ordinary RRP without UMI as a control group. All reverse transcription products were amplified using specific primers for the detection panel, which covers known breakpoint regions of cervical cancer-related fusion genes, hotspot sites of common oncogene mutations, alternative splicing-sensitive exon linker regions, and regions of characteristic expression marker genes in CaSki cells.
[0143] Sequencing was performed using the Illumina platform in paired-end sequencing mode. The obtained data underwent a unified bioinformatics analysis process, including fusion gene breakpoint identification, variant site retrieval, alternative splicing event analysis, and expression level quantification after UMI deduplication. The results showed that the UMI-RRP group was consistent with the ordinary RRP group in terms of the accuracy of fusion gene breakpoint localization, the sensitivity of low-frequency mutation detection, the consistency of alternative splicing event identification, and the stability of expression level quantification.
[0144] A comparison of the cDNA library construction results using reverse transcription with and without UMI (RRP) random primers is shown below. Figure 9 As shown. Both types of random primers successfully amplified the sequence at the corresponding target location, and their amplification efficiencies were similar. In terms of uniformity assessment, the amplification efficiency differed at each site. This was determined by comparing the difference between the read at each site and the average read across 10 sites (all converted using log10). Figure 10 As shown, the mean difference of the RRP portion is 0.998, while the mean difference of the UMI-RRP portion is 0.936, indicating that the uniformity of the UMI-RRP amplification is better.
[0145] This embodiment demonstrates that combining the UMI strategy with the detection panel can achieve high-precision, multi-type molecular detection in cell line models, providing strong experimental evidence for the widespread application of this technology in clinical sample testing.
[0146] Example 3
[0147] This embodiment describes a detection experiment using fusion gene RNA standards. Specifically, to verify the performance of the quantitative molecular detection panel in accurate quantitative detection, this embodiment uses commercially available fusion gene RNA standards as detection samples. All standards were purchased from Haixing Biotechnology. Fusion gene 1 is FGFR2-COL14A1, catalog number HXSRF1006, with a copy number of 58330 copies / μL as indicated in the instructions; fusion gene 2 is TPM3-NTRK1, catalog number HXSRF1004, with a copy number of 1446 copies / μL as indicated in the instructions. This method was used for detection.
[0148] During library construction, RRP-UMI with UMI adapters was used to capture and quantify the fusion gene in the fusion gene RNA standard. During reverse transcription, adapters were added to each RNA molecule for targeted capture and library construction. Then, the sequencing data were analyzed using Illumina platform in paired-end sequencing mode. Specifically, after routine quality control of the raw sequencing data, UMI sequences (15-base molecular identifiers) were extracted from each read based on the distribution of the 3' adapter and UMI. Reads with the same molecular identifier theoretically originate from the same original RNA molecule. Furthermore, to avoid base errors during PCR and sequencing, a count of UMI molecules was performed, retaining only those with 5 or more reads as true UMI molecules.
[0149] According to the instructions and experimental procedures, the theoretical number of molecules for fusion gene 1 was 93.3 copies, and the theoretical number of molecules for fusion gene 2 was 2.3 copies. After processing the data according to the standard UMI bioinformatics analysis workflow, the results are as follows: Figure 11 As shown, UMI-RRP efficiently captures fusion genes, with the actual number of captured molecules consistent with the theoretical number. Specifically, in the fusion gene FGFR2-COL14A1, the theoretical number of molecules was 93.3, while the average actual number was 11.3, a difference of less than 10 times. Furthermore, in the fusion gene TPM3-NTRK1, the theoretical number of molecules was 2.3, while the average actual number was 4.7, a difference of 2 times. This method can achieve precise quantification at the molecular level.
[0150] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of the invention. Various variations can be made to the above embodiments of the present invention. All simple and equivalent changes and modifications made in accordance with the claims and description of this application fall within the protection scope of the claims of this patent. All aspects not described in detail in this invention are conventional technical content.
Claims
1. A single-end targeted single-end random molecular tagging method for quantitative analysis of multiple nucleic acid variations, characterized in that, Includes the following steps: 1) Sample RNA extraction: Extract total RNA from the sample to be tested; 2) One-stranded cDNA synthesis: The total RNA was reverse transcribed using a one-stranded reverse transcriptase and a random primer RRP-UMI containing a linker and a UMI to generate single-stranded cDNA. The structure of the random primer RRP-UMI containing the linker and UMI from the 5' end to the 3' end is as follows: linker, 15-base UMI, 7-base specific sequence, and 5-40-base random sequence; the 7-base specific sequence is GTGCAAG. 3) Two-stranded cDNA synthesis: The single-stranded cDNA obtained in step 2) is reverse transcribed using a two-stranded reverse transcriptase to generate double-stranded cDNA; 4) First round of nested PCR amplification: Prepare the PCR system, add the outer PCR forward primer designed for the target gene, the outer PCR universal forward primer and the outer PCR universal reverse primer, and perform PCR amplification using the double-stranded cDNA obtained in step 3) as a template; 5) Second round of nested PCR amplification: Prepare the PCR system, add inner PCR forward primers designed for the target gene and universal inner PCR reverse primers, and perform highly specific PCR amplification using the PCR amplification product from step 4) as a template; 6) Sequencing adapter ligation: Prepare the PCR system, add sequencing adapter primers compatible with the NGS sequencer, amplify the PCR amplification product from step 5), ligate the sequencing adapters, and obtain the library preparation product. as well as 7) NGS Sequencing and Analysis: The product was sequenced using an NGS sequencer to obtain FASTQ data, which was then subjected to bioinformatics analysis to screen for fusion gene detection results.
2. The single-end targeted single-end random molecular tagging method according to claim 1, characterized in that, In step 2), the length of the 5-40 base random sequence is 6-20 bases.
3. The single-end targeted single-end random molecular tagging method according to claim 1, characterized in that, In step 2), the modification of the 3' linker portion of RRP-UMI includes at least one of unmodified, complementary nucleotide, blocking probe with secondary structure, phosphorylation modification, and locked nucleic acid modification.
4. The single-end targeted single-end random molecular tagging method according to claim 1, characterized in that, The design of the outer PCR forward primer in step 4) and the inner PCR forward primer in step 5) are both based on one or more of the following parameters: primer binding free energy ΔG, primer self-complementarity, primer dimer formation probability, GC content, homopolymer length, and secondary structure stability; among which, at 60 ℃ and 0.18M Na + Under the given conditions, the primer binding free energy ΔG should be in the range of -8 to -13 kcal / mol; primers outside this range will be discarded. In the primer self-complementarity assessment, the maximum number of complementary bases should be 4-15; the GC content should be in the range of 40%-60%; the maximum number of homopolymers should be 4. After primer design, evaluate the dimer and secondary structure formation of all primers, and discard primers with obvious stable secondary structures and severe dimers.
5. The single-end targeted single-end random molecular tagging method according to claim 1, characterized in that, In step 4), the outer PCR forward primer includes an outer primer with a universal adapter or an outer primer without a universal adapter.
6. The single-end targeted single-end random molecular tagging method according to claim 1, characterized in that, The process also includes purification steps: after cDNA synthesis in step 3), after the first round of nested PCR amplification in step 4), after the second round of nested PCR amplification in step 5), and after sequencing adapter ligation in step 6), the corresponding PCR products are purified respectively; the purification adopts the magnetic bead purification method: 1-2 times the volume of PCR product is added to magnetic beads for mixing and incubation. After the magnetic beads are adsorbed, the supernatant is discarded, the product is washed with ethanol, and then the DNA on the magnetic beads is eluted with 0.1×TE buffer.
7. The single-end targeted single-end random molecular tagging method according to claim 1, characterized in that, In step 7), the sequencing data is processed by a bioinformatics analysis workflow. The fusion gene detection adopts the breakpoint alignment and cross-connection read identification algorithm, the mutation detection adopts the standard workflow of alignment-variant recall-quality filtering, the alternative splicing event detection is based on the exon-connected read count, and the expression level calculation is based on the normalized read count after UMI deduplication.
8. The method according to claim 1, characterized in that, In step 1), the sample to be tested includes peripheral blood plasma, peripheral blood serum, body fluid, tissue sample, adherent cell sample, or suspended cell sample.
9. The single-end targeted single-end random molecular tagging method according to claim 1, characterized in that, In step 4), the universal forward primer sequence for outer PCR is TGCGATGCAATGAGAACATTAGAAGTCAGCTGCATGACTGAAGA, and the universal reverse primer sequence for outer PCR is GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG; in step 5), the universal reverse primer sequence for inner PCR is CGTGGGCTCGGAGATGTG.
10. A method for detecting an unknown fusion gene, characterized in that, The method includes constructing an unknown fusion gene detection library using the single-end targeted single-end random molecular tagging method for quantitative analysis of multiple nucleic acid variations as described in any one of claims 1-9, and sequencing the unknown fusion gene detection library to obtain the sequence of the unknown fusion gene.
Citation Information
Patent Citations
Methods and compositions for sequencing and fusion detection using randomer reverse primers
US20250011833A1
Novel method, primer group and kit for RNA high-throughput sequencing, and application thereof
CN113463202A
Nucleotide sequence, and method for constructing RNA target area sequencing library and application thereof
CN113557300A
High-quality 3 'RNA-seq library building method and application thereof
CN114108103A
Single cell transcriptome sequencing method and application thereof
CN114507711A