Fusion gene enrichment, detection primers and enrichment, detection methods

CN119709944BActive Publication Date: 2026-09-18ZHUHAI LIVZON CYNVENIO DIAGNOSTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410215655.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-27
Publication Date
2026-09-18
Estimated Expiration
2044-02-27

AI Technical Summary

Technical Problem

传统的多重PCR扩增法(mPCR)使用上下游确定性引物建库,这样会漏检碎片化样本中的部分不含有上下游引物的样本;锚定多重PCR扩增法(AMP)是用连接酶将通用接头引物连接到模板上,不管使用T末端粘性连接还是平末端连接方式,连接效率都对建库质量影响巨大;而杂交捕获法建库所使用的起始模板量巨大,这对于组织细针穿刺样本和血液CTC样本都不适合

Benefits of technology

[0039] The beneficial effects achieved by this disclosure include: the enrichment and detection method for fusion genes or splice isoforms provided in this disclosure, and the corresponding enrichment primer set, can be used for NGS library construction under ultra-small starting template conditions. The advantages are as follows: 1. Specific primers with modified groups are used to reduce sequencing noise interference, enrich mutant templates, and improve sensitivity and specificity; 2. No ligase is used in the library construction process; all sequencing primers are ligated to the template using Taq enzyme PCR, which improves library construction efficiency; 3. The combination of specific primers and random primers can be used to detect known and unknown mutations, and the requirements for the fragment size of the initial sample are not high.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004716451910000151
    Figure BDA0004716451910000151
  • Figure BDA0004716451910000161
    Figure BDA0004716451910000161
  • Figure BDA0004716451910000171
    Figure BDA0004716451910000171
Patent Text Reader

Abstract

The present disclosure relates to the field of in vitro diagnosis, in particular to fusion gene enrichment, detection primer and enrichment and detection method. The beneficial effects achieved by the present disclosure include: the fusion gene enrichment and detection method provided by the present disclosure, and the corresponding enrichment random primer group and fusion gene specific primer group can be used for NGS library construction under the condition of ultra-micro initial template amount, and the advantages are as follows: 1, the specific primer and random primer with modification group can be used to detect known and unknown fusion gene mutations, and the fragment size requirement of the initial sample is not high; 2, the modification group is used to enrich the specific primer product, and reduce the sequencing noise interference; 3, the specific Blocker is used to block the amplification of wild type at the fusion gene breakpoint, and the fusion gene mutant fragment is enriched, so that the sensitivity and specificity are improved; 4, no ligase is contained in the library construction process, so that the influence of the ligation efficiency on the library construction efficiency is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of in vitro diagnostic technology, and in particular to primers and methods for enriching and detecting fusion genes. Background Technology

[0002] There are many methods for detecting fusion genes. Currently, the gold standard for clinical detection of fusion genes remains FISH (Fluorescence in situ hybridization) and IHC (Immunohistochemistry). Next-generation sequencing (NGS) methods for detecting fusion genes are constantly being improved and are gradually being used in clinical testing. NGS-based methods for detecting fusion genes are mainly divided into two categories: hybridization capture methods and amplicon sequencing methods. Both methods can detect fusion genes at the DNA and RNA levels.

[0003] Amplicon sequencing can be further divided into two subcategories: traditional multiplex PCR (mPCR) and anchored multiplex PCR (AMP). Both hybridization capture and AMP can detect unknown fusion partner genes. Traditional multiplex PCR (mPCR) uses deterministic upstream and downstream primers for library construction, which may miss fragmented samples that do not contain the upstream and downstream primers. Anchored multiplex PCR (AMP) uses ligase to ligate universal adapter primers to the template. Regardless of whether T-end sticky ligation or blunt end ligation is used, the ligation efficiency has a significant impact on the quality of library construction. Hybridization capture requires a large amount of starting template, which is not suitable for fine-needle aspiration samples and blood CTC samples.

[0004] Therefore, it is necessary to develop an NGS detection kit suitable for ultra-small starting template amounts in tissue fine needle puncture samples and blood CTC samples. Summary of the Invention

[0005] The purpose of this disclosure is to provide a method for enriching and / or detecting unknown fusion mutations from a sample to be tested. This method requires a low initial sample amount, does not require the use of ligases, avoids the influence of ligation efficiency on the enrichment and detection results, and can effectively improve the signal-to-noise ratio. The detection sensitivity and specificity of unknown fusion mutations are also higher.

[0006] To achieve the above objectives, this disclosure provides the following technical solutions:

[0007] In a first aspect, this disclosure provides a method for enriching fusion genes or splice isoforms, comprising amplifying the fusion gene using random primers and specific primers targeting a first fusion gene, wherein the specific primers carry modification groups for purification, and after amplification, recovering the amplification product carrying the modification groups.

[0008] The modified group is used to separate and purify the fusion gene from the amplification product, including but not limited to biotin. The target fragment of the specific primer in the first fusion gene is 1-100 nt away from the fusion breakpoint, preferably 4-79 nt, more preferably 10-79 nt, and even more preferably 10-78 nt. This distance ensures sufficient sensitivity while ensuring that the amplification product mainly contains the fusion gene with the breakpoint.

[0009] In an optional implementation, the first fusion gene includes ALK, ROS1, RET, BRAF, NRG1, FGFR, NTRK, EGFR, ERBB2, MET, ESR1, AML1, BCR, RARα, E2A, TAL1, DEK, AML1, PDGFRA, KMT2A, NPM, NUP98, SET, TEL, TLS, CBFB, VMP1, CLDN18, ARHGAP, CD44, SLC1A2, TSC2, RNF216, AGTARP, TACC2, PPAPDC1A, RAS, RPS6KB1, MAN2A1, FER, DNAJB1, ATP1B1, FIG, PRKACA, PRKACB, IRF2BP2, SEC16A, NOTCH1, or RUNX1. The breakpoint of the first fusion gene is associated with the development of common or serious diseases, or has other important potential value.

[0010] In an optional embodiment, a blocker probe that blocks the amplification of the wild-type first fusion gene is also added to the amplification reaction system to further increase the content of the fusion gene containing the breakpoint in the amplification product.

[0011] In an optional embodiment, the recovery method includes adding a purification reagent for capturing the modifying group to the amplification product; optionally, the purification reagent includes magnetic beads with a modifying group binding molecule attached to them.

[0012] In a second aspect, this disclosure provides an enrichment primer for a fusion gene or splice isoform, comprising a random primer and a specific primer targeting a first fusion gene, wherein the specific primer carries a modification group for purification, the purification modification being used to separate the enriched product formed by the random primer and the specific primer as a primer pair, and the enrichment primer being used to implement the enrichment method described in the first aspect.

[0013] In an optional embodiment, the target fragment of the specific primer in the first fusion gene is 1 to 100 nt away from the fusion breakpoint, preferably 4 to 79 nt, more preferably 10 to 79 nt, and even more preferably 10 to 78 nt; the length of the random primer is 6 to 30 nt, preferably 8 to 25 nt, and even more preferably 16 to 20 nt.

[0014] Preferably, the first fusion gene includes ALK, ROS1, RET, BRAF, NRG1, FGFR, NTRK, EGFR, ERBB2, MET, ESR1, AML1, BCR, RARα, E2A, TAL1, DEK, AML1, PDGFRA, KMT2A, NPM, NUP98, SET, TEL, TLS, CBFB, VMP1, CLDN18, ARHGAP, CD44, SLC1A2, TSC2, RNF216, AGTARP, TACC2, PPAPDC1A, RAS, RPS6KB1, MAN2A1, FER, DNAJB1, ATP1B1, FIG, PRKACA, PRKACB, IRF2BP2, SEC16A, NOTCH1, or RUNX1.

[0015] In an optional embodiment, the modifying group includes biotin.

[0016] Furthermore, the enrichment primers also include a blocker probe that blocks the amplification of the wild-type first fusion gene.

[0017] Thirdly, this disclosure provides an enrichment composition for fusion genes or splice isomers, comprising the enrichment primers described in the second aspect and a purification reagent for capturing modifying groups; optionally, the purification reagent comprises magnetic beads, the magnetic beads being linked to a modifying group binding molecule.

[0018] Fourthly, this disclosure provides the use of the enrichment primers described in the second aspect or the enrichment compositions described in the third aspect in the preparation of gene fusion enrichment reagents or kits.

[0019] Fifthly, this disclosure provides an enrichment reagent or kit for fusion genes or splice isoforms, comprising the enrichment primers described in the second aspect or the enrichment composition described in the third aspect; preferably, it further comprises sampling tools and / or other consumables.

[0020] In a sixth aspect, this disclosure provides a method for detecting fusion genes or splice isoforms, comprising adding detection primer pairs to a sample containing a fusion gene to amplify the fusion gene, using purification reagents to separate the fusion gene enrichment product from the amplification product after amplification, and then using sequencing primer pairs to sequence the fusion gene enrichment product.

[0021] The detection primer pair includes a random primer and a specific primer targeting the first fusion gene. The random primer is connected to a first universal primer, and the specific primer is connected to a second universal primer with a modification group for purification. The purification modification is used to separate the enriched product composed of the random primer and the specific primer as primer pairs.

[0022] The sequencing primer pair includes a first sequencing primer and a second sequencing primer. The first sequencing primer contains a sequence that hybridizes with a first universal primer and a first sequencing primer adapter sequence. The second sequencing primer contains a sequence that hybridizes with a second universal primer and a second sequencing primer adapter sequence.

[0023] In an optional embodiment, the modifying group includes biotin.

[0024] In an optional implementation, a blocker probe that blocks the amplification of the wild-type first fusion gene is also added to the fusion gene amplification reaction system.

[0025] In an optional embodiment, the purification reagent includes magnetic beads with a modified group attached to the molecule.

[0026] In an optional embodiment, the target fragment of the specific primer in the first fusion gene is 1 to 100 nt away from the fusion breakpoint, preferably 4 to 79 nt, more preferably 10 to 79 nt, and even more preferably 10 to 78 nt; the length of the random primer is 6 to 30 nt, preferably 8 to 25 nt, and even more preferably 16 to 20 nt.

[0027] Preferably, the first fusion gene includes ALK, ROS1, RET, BRAF, NRG1, FGFR, NTRK, EGFR, ERBB2, MET, ESR1, AML1, BCR, RARα, E2A, TAL1, DEK, AML1, PDGFRA, KMT2A, NPM, NUP98, SET, TEL, TLS, CBFB, VMP1, CLDN18, ARHGAP, CD44, SLC1A2, TSC2, RNF216, AGTARP, TACC2, PPAPDC1A, RAS, RPS6KB1, MAN2A1, FER, DNAJB1, ATP1B1, FIG, PRKACA, PRKACB, IRF2BP2, SEC16A, NOTCH1, or RUNX1.

[0028] In a seventh aspect, this disclosure provides a detection primer for a fusion gene or splice isoform, comprising a random primer and a specific primer targeting a first fusion gene, wherein the random primer is connected to a first universal primer, and the specific primer is connected to a second universal primer with a modification group for purification, the purification modification being used to separate enriched products composed of the random primer and the specific primer as primer pairs.

[0029] In an optional embodiment, the modifying group includes biotin.

[0030] In an optional embodiment, the detection primers further include a blocker probe that blocks the amplification of the wild-type first fusion gene.

[0031] In an optional embodiment, the target fragment of the specific primer in the first fusion gene is 1 to 100 nt away from the fusion breakpoint, preferably 4 to 79 nt, more preferably 10 to 79 nt, and even more preferably 10 to 78 nt.

[0032] Preferably, the first fusion gene includes ALK, ROS1, RET, BRAF, NRG1, FGFR, NTRK, EGFR, ERBB2, MET, ESR1, AML1, BCR, RARα, E2A, TAL1, DEK, AML1, PDGFRA, KMT2A, NPM, NUP98, SET, TEL, TLS, CBFB, VMP1, CLDN18, ARHGAP, CD44, SLC1A2, TSC2, RNF216, AGTARP, TACC2, PPAPDC1A, RAS, RPS6KB1, MAN2A1, FER, DNAJB1, ATP1B1, FIG, PRKACA, PRKACB, IRF2BP2, SEC16A, NOTCH1, or RUNX1.

[0033] Eighthly, this disclosure provides a detection composition for fusion genes or splice isomers, comprising the detection primers described in the seventh aspect, a purification reagent for capturing modifying groups, and a sequencing primer pair suitable for a sequencing platform;

[0034] The sequencing primer pair includes a first sequencing primer and a second sequencing primer. The first sequencing primer contains a sequence that hybridizes with a first universal primer and a first sequencing primer adapter sequence. The second sequencing primer contains a sequence that hybridizes with a second universal primer and a second sequencing primer adapter sequence.

[0035] Optionally, the purification reagent includes magnetic beads, which are linked to a modifying group-bound molecule.

[0036] Ninthly, this disclosure provides the use of the detection primers described in the seventh aspect or the detection compositions described in the eighth aspect in the preparation of reagents or kits for gene fusion detection or splice isomers.

[0037] In a tenth aspect, this disclosure provides a fusion gene detection or splice isomer detection reagent or kit, comprising the detection primers described in the seventh aspect or the detection composition described in the eighth aspect; preferably, it further comprises sampling tools and / or other consumables.

[0038] In the eleventh aspect, this disclosure provides a diagnostic method for diseases related to gene fusion mutations or abnormal splicing of precursor mRNA, including collecting a sample to be tested, performing sequencing detection on gene fusion mutation-related breakpoints according to the detection method described in the sixth aspect, and estimating the risk of gene fusion mutations in the sample to be tested.

[0039] The beneficial effects achieved by this disclosure include: the enrichment and detection method for fusion genes or splice isoforms provided in this disclosure, and the corresponding enrichment primer set, can be used for NGS library construction under ultra-small starting template conditions. The advantages are as follows: 1. Specific primers with modified groups are used to reduce sequencing noise interference, enrich mutant templates, and improve sensitivity and specificity; 2. No ligase is used in the library construction process; all sequencing primers are ligated to the template using Taq enzyme PCR, which improves library construction efficiency; 3. The combination of specific primers and random primers can be used to detect known and unknown mutations, and the requirements for the fragment size of the initial sample are not high. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0041] Figure 1 This refers to the possible mRNA morphologies present in different samples.

[0042] Figure 2 Reactions that may occur during the first PCR amplification process of the enrichment or detection methods provided in this disclosure;

[0043] Figure 3 The type of amplification product obtained from the first PCR of the enrichment or detection method provided in this disclosure;

[0044] Figure 4 The first PCR reaction process for the enrichment or detection method provided in this disclosure;

[0045] Figure 5 The overall technical approach of the detection method provided in this disclosure;

[0046] Figure 6 This refers to the number of reads of the fusion gene detected under different random primer lengths in Example 4 of this disclosure;

[0047] Figure 7 The test results are for the clinical FFPE samples in Example 5 of this disclosure;

[0048] Figure 8 The results are the detection results of the H2228 sample in the clinical validation of Example 5 of this disclosure. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of this disclosure, but not all embodiments.

[0050] (I) Definitions or terms

[0051] The following abbreviations are used to refer to the definitions and terms used in this disclosure. Unless otherwise defined, all technical terms used herein have the meanings commonly understood by one of ordinary skill in the art. The following terms are provided below.

[0052] 1. Fusion genes

[0053] The "fusion gene" described in this disclosure refers to a chimeric gene formed by two or more genes whose partial or complete coding regions are joined end-to-end and controlled by the same set of regulatory sequences (including promoters, enhancers, ribosome-binding sequences, terminators, etc.). When the first fusion gene is a partial coding region, the site where it connects with other genes or gene fragments is called the breakpoint of the first fusion gene. In existing reports, fusion genes obtained by chimerism of certain genes with other genes or gene fragments at specific breakpoints are closely related to certain specific diseases. Therefore, detecting whether a specific gene contains fusion genes with certain breakpoints can be used to assess the risk of developing certain diseases.

[0054] 2. The "sample" or "sample to be tested" mentioned in this disclosure includes "samples containing fusion genes." These samples typically originate from tumor cells or various bodily fluid samples, such as blood, saliva, urine, pleural effusion, and ascites. Tumor cells can be obtained through surgery, biopsy, or fine-needle aspiration, while bodily fluid samples can be obtained through extraction. These samples contain the genetic information of tumor cells or normal cells, which researchers can detect to identify fusion genes. For example, lung cancer patients can have their tumor tissue surgically removed to obtain information about fusion genes. Researchers can use this information to study the mechanisms of lung cancer occurrence and development, and explore the relationship between fusion genes and lung cancer prognosis. Alternatively, researchers can also detect fusion genes by detecting circulating tumor cells or cell-free DNA in the blood. By obtaining these samples and detecting their genetic information, researchers can gain a deeper understanding of the relationship between fusion genes and tumor occurrence and development, providing strong support for tumor diagnosis, treatment, and prognostic assessment.

[0055] 3. The "splicing isoforms" described in this disclosure refer to the process by which gene transcription produces initial pre-mRNA, which undergoes a series of processing steps, such as capping, tailing, and splicing, ultimately forming mature mRNA. During splicing, introns are removed and exons are joined together. Different splicing methods can produce different mRNAs, thus forming splicing isoforms. Based on the splicing method, splicing isoforms can be divided into three types: Type I splicing isoforms refer to different mRNA molecules produced by different transcripts of the same gene through different splicing methods. These mRNA molecules can jump between exons, thus producing the same or different proteins. For example, some genes can produce two different splicing isoforms, one containing exons A and B, and the other containing exons C and D. Type II splicing isoforms refer to different mRNA molecules produced by different transcripts of the same gene through the same splicing method. These mRNA molecules may have the same exon sequence, but can produce different proteins through mechanisms such as selective tailing and selective splicing. For example, some genes can produce two different splice isoforms, one with a complete open reading frame and the other without. Type III splice isoforms refer to different mRNA molecules produced by the same splicing mechanism from the transcripts of different genes. These mRNA molecules may have the same exon sequence but originate from different genes. For example, some genes can produce different mRNA molecules by splicing the transcript of another gene with the same mechanism. Different mature mRNAs can be considered as fusion genes obtained by splicing different exons. The enrichment or detection method for fusion genes provided in this disclosure can be used to enrich or detect splice isoforms.

[0056] 4. The “ALK” gene described in this disclosure encodes a receptor tyrosine kinase belonging to the insulin receptor superfamily. This protein comprises an extracellular domain, a hydrophobic stretching domain corresponding to a single transmembrane region, and an intracellular kinase domain. It plays a crucial role in brain development and influences specific neurons in the nervous system. This gene has been found to be rearranged, mutated, or amplified in a range of tumors, including anaplastic large cell lymphoma, neuroblastoma, and non-small cell lung cancer. Chromosomal rearrangement is the most common genetic alteration in this gene, leading to the generation of multiple fusion genes during tumorigenesis.

[0057] 5. The “ROS1” gene described in this disclosure encodes a receptor tyrosine kinase that is highly expressed in various tumor cell lines and belongs to the 700 subfamily of tyrosine kinase insulin receptor genes. The protein encoded by this gene is a type I integrated membrane protein with tyrosine kinase activity. This protein can function as a receptor for growth factors or differentiation factors.

[0058] 6. The protein encoded by the “RET” gene described in this disclosure is a transmembrane receptor that participates in multiple cell signaling pathways. These pathways control cell growth, differentiation, and death. When the RET gene is mutated or abnormally expressed, cell growth and differentiation processes may be affected, leading to disease. RET Gene and Cancer: The RET gene plays an important role in various cancers, especially medullary thyroid carcinoma and renal cell carcinoma. Some studies have shown that mutations or abnormal expression of the RET gene may be related to the occurrence and development of cancer. Therefore, studying the role of the RET gene in cancer can provide new insights for cancer diagnosis and treatment.

[0059] 7. The "BRAF" gene described in this disclosure is a proto-oncogene that encodes a serine / threonine protein kinase called the BRAF protein. The BRAF protein plays an important role in cell signal transduction, primarily participating in the MAPK signaling pathway. This pathway involves a cascade of multiple kinases, transmitting signals from the cell surface to the cell nucleus, influencing cell growth, differentiation, and apoptosis.

[0060] 8. The “NRG1” gene described in this disclosure is a member of the brain-derived neurotrophic factor (BDNF) family and plays an important role in the development and function of the nervous system. The NRG1 gene participates in processes such as nerve cell growth, differentiation, synaptic plasticity, and apoptosis by binding to its receptor. The NRG1 gene is expressed in both the central and peripheral nervous systems, particularly in the cerebral cortex, hippocampus, and amygdala. Furthermore, the NRG1 gene is also expressed in tissues such as the heart, kidneys, and skeletal muscle.

[0061] 9. The “FGFR” gene described in this disclosure is an abbreviation for fibroblast growth factor receptor gene. These are transmembrane protein kinase receptors that regulate cell growth, proliferation, and differentiation. The structure of the FGFR gene includes three parts: an extracellular region, a transmembrane region, and an intracellular region. The extracellular region is the ligand-binding region, which can recognize and bind fibroblast growth factors; the transmembrane region is the area that the cell membrane crosses and is composed of hydrophobic amino acids; the intracellular region is the kinase active region, which has the ability to phosphorylate tyrosine residues and can trigger signal transduction.

[0062] 10. The “NTRK” gene described in this disclosure is a gene with important functions. It encodes the NTRK receptor, a protein that plays an important role in the development and function of the nervous system. The NTRK receptor plays a key role in cell signal transduction and can promote cell growth, differentiation, and survival.

[0063] 11. The “EGFR” gene described in this disclosure is a growth factor receptor discovered in the mid-1980s. It was discovered jointly by researchers at the University of Glasgow and the Mondrian Institute. EGFR is a transmembrane protein that is activated by ligands on the cell membrane and can stimulate cell growth, differentiation, and proliferation. EGFR is expressed on many types of cells in the human body, but is mainly concentrated on the surface of epithelial cells in organs such as the skin, digestive tract, and lungs.

[0064] 12. The “ERBB2” gene described in this disclosure, also known as the Her2 gene, is a gene that plays an important role in cell signal transduction. It encodes a transmembrane receptor protein called the Her2 receptor. This receptor is expressed on the cell surface and can bind to specific growth factors, thereby activating a series of signal transduction pathways and promoting cell growth and proliferation.

[0065] 13. The “MET” gene described in this disclosure is short for hepatocyte growth factor receptor gene, which is a proto-oncogene. It encodes a tyrosine kinase transmembrane receptor that can bind to various growth factors and signaling molecules, participating in processes such as cell proliferation, differentiation, migration, and apoptosis.

[0066] 14. The “ESR1” gene described in this disclosure was discovered in 1986 and is located on human chromosome 6. This gene encodes estrogen receptor α protein, a nuclear transcription factor that binds to estrogen and regulates gene expression. The main function of the ESR1 gene is to encode estrogen receptor α protein, which plays an important role in organs and tissues such as the female reproductive system, bones, cardiovascular system, and brain. Estrogen receptor α protein can bind to estrogen and regulate processes such as cell growth, development, and metabolism.

[0067] 15. The "AML1" gene described in this disclosure is a common gene in hematologic malignancies, officially known as the Acute Myeloid Leukemia 1 gene. It is a transcription factor gene that plays an important role in normal hematopoiesis. Abnormalities in the AML1 gene can lead to hematologic malignancies. Acute myeloid leukemia is a common hematologic malignancy characterized by abnormal proliferation of blast cells in the bone marrow, resulting in suppression of normal hematopoietic function.

[0068] 16. The “BCR” gene described in this disclosure is one of the genes in the BCR-ABL complex and is associated with the Philadelphia chromosome. BCR gene testing is a blood test for the early detection and prediction of certain malignant diseases, used to detect aberrations (suspected interstitial crossings) in regions B and C that regulate cell proliferation and apoptosis. BCR gene testing can be used to detect chronic lymphocytic leukemia (CLL) and chronic myeloid leukemia (CML), both of which are associated with aberrations related to BCR.

[0069] 17. The “RARα” gene described in this disclosure is a key gene, also known as retinoic acid receptor α, belonging to the retinoic acid (RA) nuclear receptor family. The main function of the RARα gene is to regulate ligand transcription factors by binding to specific response elements (RAREs) at the promoter sites of target genes. Located on the long arm of chromosome 17, the RARα gene plays a crucial role in hematopoietic cell development and gene expression regulation. In the absence of retinoic acid (RA), the retinoic acid receptor (RAR) forms a heterodimer with the retinol X receptor, while the RARα gene connects to the nuclear co-repressor receptor, forming an aggregate composed of the co-repressor complex, chromatin condensate, and transcriptional repressor.

[0070] 18. The “E2A” gene described in this disclosure is a gene related to leukemia and immune system development. Mutations in the E2A gene can lead to leukemia in acute lymphoblastic leukemia. The E2A gene is the matrix for the production of T cell surface antigens; mutations in this gene in mature T cells lead to abnormal T cell proliferation, thus becoming a major source of leukemia cells. Besides its association with leukemia, mutations in the E2A gene are also associated with intellectual disability. The E2A gene is crucial for the development of a normal immune system and central nervous system; its mutations can lead to a range of developmental disorders and mental illnesses.

[0071] 19. The “TAL1” gene described in this disclosure plays a crucial role in the development of hematopoietic stem cells and the formation of various blood cell lineages. It is located in the 1p32 region of chromosome 1 and belongs to the bHLH transcription factor family. The basic structure of the TAL1 gene includes a DNA-binding region and two α-helices connected by a circular structure. This gene requires the formation of homologous or heterologous dimers to function. Aberrant expression of the TAL1 gene is associated with the development of acute T-lymphoblastic leukemia (T-ALL). Approximately 60% of T-ALL patients exhibit aberrant expression of the TAL1 gene. Of these, 30% are due to the deletion of a DNA segment between SIL1 and TAL1, leading to activation of TAL1 expression by the SIL promoter. Additionally, in 6% of patients, the TAL1 gene recombines with other oncogenes, resulting in TAL1 overexpression. Furthermore, a small number of T-ALL patients exhibit translocations and overexpression of other members of the βHLH family, such as LYL1, TAL2, and BHLHB1.

[0072] 20. The “DEK” gene described in this disclosure is an oncogene that encodes a protein with a SAP domain. This protein binds to cruciate and supercoiled DNA, induces positive supercoils to form closed circular DNA, and participates in splicing site selection during mRNA processing. Chromosomal aberrations involving this region, increased expression of this gene, and the presence of antibodies against this protein are all associated with various diseases. Two transcriptomorphs encoding different subtypes have been identified.

[0073] 21. The "AML1" gene described in this disclosure, short for Acute Myeloid Leukemia Gene, is one of the most frequently mutated genes in human leukemia. The fusion protein AML-1-ETO / MTG8 produced by the AML-1 gene leads to one of the most common chromosomal translocations in acute myeloid leukemia. Furthermore, AML1 gene mutations take many forms, including fusion genes and point mutations. For example, TEL-AML1 is a common genetic mutation in childhood acute lymphoblastic leukemia, formed by the fusion of the TEL gene on chromosome 12 and the AML1 gene on chromosome 21. In addition, AML-1 gene mutations may also be associated with mutations in other genes, such as Flt3 and N-Ras. Mutations in the AML1 gene are also relatively common in hematological malignancies, potentially leading to abnormal cell proliferation and differentiation, thus triggering leukemia and other hematological malignancies. Therefore, detecting AML1 gene mutations is of great significance for the diagnosis and treatment of leukemia and other hematological malignancies.

[0074] 22. The “PDGFRA” gene described in this disclosure is a type III tyrosine kinase receptor, also known as the platelet-derived growth factor receptor. This gene plays an important role in the human body, mainly involved in cell division and proliferation, injury healing, inflammation, and angiogenesis. After PDGFRA gene expression, it binds to platelet-derived growth factor, forming a receptor-ligand complex and becoming activated. This process can activate phosphorylation pathways of phosphatidylinositol, cyclic adenosine monophosphate, and various proteins. These signal transduction pathways further induce the expression of corresponding genes, promote DNA synthesis, and induce cell division and proliferation. The PDGFRA gene plays a crucial role in human development and tissue repair, but in certain situations, such as tumor development and progression, this gene may mutate or be overexpressed.

[0075] 23. The "KMT2A" gene described in this disclosure, short for "lysine-specific methyltransferase2A," is a protein-encoding gene located on chromosome 11 in humans. The KMT2A gene is an epigenetic modification-related gene, primarily involved in the regulation of gene expression. Mutations in the KMT2A gene can lead to several diseases of the blood and skeletal systems, the most significant being childhood acute myeloid leukemia. Furthermore, KMT2A gene mutations can also cause skeletal deformities, such as deformities of the clavicle and radius, as well as developmental delays in motor skills. KMT2A gene mutations may also have some impact on intellectual development, specifically manifesting as intellectual disability and decreased learning ability. These effects may be related to the role of the KMT2A gene in brain development.

[0076] 24. The “NPM” gene described in this disclosure is a gene associated with nuclear phosphoprotein (NPM). NPM genes are widely distributed in various organisms and have multiple biological functions, including regulating the cell cycle, participating in DNA damage repair, and regulating gene expression. Within the cell nucleus, NPMs interact with various proteins to form a complex network structure, playing a crucial regulatory role in intranuclear biological activities. Under certain circumstances, NPM genes may mutate or aberrate, leading to abnormal NPM protein expression or dysfunction. These abnormal NPM proteins may cause abnormalities in cell growth, differentiation, and apoptosis, thereby triggering certain diseases such as cancer. Therefore, research on NPM genes helps to deepen our understanding of their biological functions and regulatory mechanisms, providing new ideas and targets for disease diagnosis and treatment.

[0077] 25. The “NUP98” gene described in this disclosure is a nucleoporin gene located on chromosome 11p15.5. It encodes a 186kDa precursor protein, which undergoes self-protective cleavage to produce a 98kDa nucleoporin and a 96kDa nucleoporin. The 98kDa nucleoporin contains a Gly-Leu-Phe-Gly (GLFG) repeat domain and is involved in many cellular processes, including nuclear importation, nuclear export, mitotic progression, and gene expression regulation. Furthermore, the NUP98 gene is also involved in regulating the transport of proteins and RNA across the nucleus. In acute myeloid leukemia (AML), the NUP98-NSD1 fusion gene can encode a fusion protein with dual acetyltransferase and methyltransferase activities. This fusion protein can be detected in 16.1% of childhood hereditary AML and 2.3% of adult AML, and these patients often have a poor prognosis.

[0078] 26. The “SET” gene described in this disclosure is a gene that plays an important role in organisms. Its full name is “Suppressor of variegation, Enhancer of yellow, Flare”. The SET gene was first discovered in fruit flies and has homologous genes found in many other organisms, such as humans and mice. The protein product of the SET gene, SET protein, is a widely expressed nucleoprotein involved in various biological processes, including cell cycle regulation, apoptosis, and DNA repair. The functions of the SET gene in organisms are multifaceted. Studies have shown that the SET gene, through interactions with other genes, affects the growth, development, and cell differentiation of organisms. Furthermore, the SET gene also participates in the regulation of apoptosis and DNA repair, playing a crucial role in maintaining cell stability and preventing cell carcinogenesis.

[0079] 27. The “TEL” gene described in this disclosure, also known as the ETV6 gene, is a gene located on chromosome 12p13 that encodes a transcription factor belonging to the ETS family. This is a tumor suppressor gene that plays an important role in the development and maintenance of hematopoiesis and vascular networks. The TEL gene is also a member of the ETS transcription factor family. The TEL gene can undergo chromosomal translocations with various genes, among which genes encoding protein kinases are most likely to produce chromosomal translocation effects with the TEL gene. For example, fusions of the TEL gene with ABL kinase, as well as JAK2 and NTRK3 genes, have been discovered. When the TEL gene and ABL gene undergo chromosomal translocation, ABL kinase is continuously activated, cell growth gradually loses its dependence on IL-3, and the expression of c-myc downstream of the signaling pathway is activated.

[0080] 28. The “TLS” gene described in this disclosure, short for TERT promoter, is a gene expressed in many different cell types. It plays a crucial role in the regulation of telomerase, thus influencing cell lifespan. Telomerase is an enzyme that maintains the length of telomeres at the ends of chromosomes during cell division, and the TLS gene regulates cell lifespan by affecting telomerase activity. In normal cells, the length of telomeres at the ends of chromosomes gradually shortens with increasing cell division, leading to cellular aging. However, in certain cell types, the expression of the TLS gene is activated, enhancing telomerase activity and thus maintaining telomere length and extending cell lifespan.

[0081] 29. The “CBFB” gene described in this disclosure is a gene associated with leukemia. It encodes a protein that is a heterodimeric core of the pebp2 / cbf transcription factor family, binding to the β subunit of the transcription factor. This gene locus is a significant factor contributing to leukemia; therefore, detecting the DNA sequence of the CBFB gene locus can determine the presence of pathogenic gene mutations. If such mutations are present, they may lead to leukemia. Early detection of potentially at-risk individuals allows for effective prevention and treatment measures to safeguard health.

[0082] 30. The “VMP1” gene described in this disclosure, also known as the TMEM49 gene, is a stress-inducible gene. It encodes a vacuolar envelope protein closely related to autophagy and is considered an important regulator of this process. Autophagy plays a crucial role in determining cell fate, reducing aging and defunctionalized organelles, recycling cellular components, and responding to cellular stress. Therefore, the relationship between autophagy and tumors has been extensively studied. The protein expressed by the MP1 gene is a transmembrane protein with six hydrophobic regions, allowing it to be localized within the endoplasmic reticulum and Golgi apparatus of the cell membrane. VMP1 has been found essential for maintaining the normal function and integrity of these organelles. In cells lacking VMP1, the secretory pathways of these organelles are disrupted.

[0083] 31. The “CLDN18” gene described in this disclosure is a gene encoding a tight junction protein belonging to the Claudin family. This gene plays a crucial role in maintaining cell polarity and signal transduction, and also plays an important role in the occurrence and development of some diseases. The CLDN18 gene is located in the 3q22.3 region of human chromosome 3. Its first exon has two splicing variants, forming two different protein isoforms, CLDN18.1 and CLDN18.2, each with a different N-terminus sequence of 69 amino acids. Claudin-18 expression is cell-specific; CLDN18.1 is expressed only in alveolar epithelial cells in normal tissues, while CLDN18.2 is expressed only in differentiated gastric mucosal epithelial cells in normal tissues, maintaining the barrier function of the gastric mucosa and preventing the leakage of H+ from gastric acid via the paracellular pathway.

[0084] 32. The “ARHGAP” gene described in this disclosure is a type of gene that plays an important regulatory role in neural activity. Members of this gene family include ARHGAP1, ARHGAP2 (CHN1), ARHGAP3 (CHN2), ARHGAP4, ARHGAP5, ARHGAP6 (STARD8), ARHGAP7 (STARD12 or DLC1), ARHGAP8, ARHGAP9, ARHGAP10, ARHGAP12, ARGGAP13 (SRGAP1), ARHGAP14 (SRGAP2), ARHGAP15, ARHGAP17 (RICH1), ARHGAP18, ARHGAP19, ARHGAP20, ARHGAP21, ARHGAP22, ARHGAP23, ARHGAP24, ARHGAP25, ARHGAP26, STRAD13 (DLC2), HA-1, GMIP, PARG1, PIK3RACGAP1, and FNBP2, etc. This gene family primarily functions in nerve cell development and signal transduction. Studies have shown that mutations in some ARHGAP genes can lead to neurodevelopmental disorders and neurological diseases, such as intellectual disability, autism, and Parkinson's disease. Meanwhile, research has indicated that the ARHGAP11B gene may influence neocortex formation and has been defined as a "smart" gene. Scientists studying the ARHGAP11B gene in marmosets have discovered that this gene forms complex folds and gyri-like structures in the marmoset's cerebral cortex, significantly increasing the number of upper-layer neurons and correspondingly increasing brain volume.

[0085] 33. The "CD44" gene described in this disclosure is a gene located on human chromosome 11, approximately 50 kb in length, composed of 20 highly conserved exons. These exons can be divided into two main categories according to their transcriptional mode: constitutive exons (C) and variant splicing exons (V). The expression product of the CD44 gene is a cell surface glycoprotein that participates in cell-cell interactions, cell adhesion, and cell migration. The transcription product of the CD44 gene exists in various forms, including standard CD44 (CD44S) and CD44V. The CD44 gene plays an important role in many physiological and pathological processes, such as lymphocyte homing, tumorigenesis, and tumor development.

[0086] 34. The proteins encoded by the “SLC1A2” gene described in this disclosure play a crucial role in the pathogenesis of ET. They are key to the removal of glutamate from the synaptic cleft, responsible for approximately 90% of glutamate uptake in most brain regions. Glutamate is an excitatory neurotransmitter in the central nervous system, and a long-term increase in glutamate concentration in the synaptic cleft has neurotoxic effects. This dysregulation of the gene is believed to be associated with several neurological disorders.

[0087] 35. The “TSC2” gene described in this disclosure is an abbreviation for tuberous sclerosis complex 2, a gene that explains approximately 80% of cases of tuberous sclerosis complex. Tuberculous sclerosis complex is a common genetic disorder characterized by wart-like growths on the body surface, as well as malformations or multiple lesions in the brain, heart, kidneys, eyes, and other organs. TSC2 gene testing examines this gene to see if it is affected by substitution or activating mutations. The results of this test are crucial because they help doctors diagnose tuberous sclerosis complex and develop more targeted and effective treatment plans for the future.

[0088] 36. The protein encoded by the “RNF216” gene described in this disclosure is involved in the ubiquitination process. Ubiquitination is a cellular process in which unwanted proteins are labeled as ubiquitin and then broken down. One of the proteins labeled by RNF216 is found in nerve cells (neurons) and plays a role in synaptic plasticity. Synaptic plasticity is the ability of the connections between neurons (synapses) to change and adapt over time in response to experience. This process is crucial for learning and memory.

[0089] 37. The "AGTARP" gene described in this disclosure is an abbreviation for adenylate cyclase-GTPase regulatory protein gene, a gene that plays an important role in the human body. It participates in regulating intracellular signal transduction and metabolic processes, and is crucial for maintaining normal physiological functions. Mutations in the AGTARP gene may lead to a range of diseases, such as cancer and neurodegenerative diseases. Understanding the role and function of the AGTARP gene will help to further explore the pathogenesis of these diseases and provide new ideas and methods for disease diagnosis, prevention, and treatment.

[0090] 38. The “TACC2” gene described in this disclosure is a cancer-related gene. It encodes a protein concentrated in the centrosome throughout the cell cycle, a conserved family of microtubule-interacting proteins. The TACC2 gene is located in a chromosomal region associated with tumorigenesis, its expression may be induced by erythropoietin, and it is thought to influence the progression of breast tumors. Furthermore, several different transcriptomic variants encoding the TACC2 gene have been identified.

[0091] 39. The "PPAPDC1A" gene described in this disclosure plays an important role in the human body, primarily participating in intracellular phosphorylation and glycolysis. The expression of this gene is influenced by various factors, including lifestyle, dietary habits, and environmental factors. By studying the PPADC1A gene, scientists can gain a deeper understanding of its role in cell biology and disease development, providing new ideas and directions for future medical research.

[0092] 40. The “RAS” gene described in this disclosure is a proto-oncogene found in the genomes of humans and many other organisms. It encodes a GTPase, which plays a crucial role in cell signal transduction. The RAS gene participates in regulating cell proliferation and tumorigenesis by modulating processes such as cell growth, differentiation, and apoptosis. The protein encoded by the RAS gene is a GTP-binding protein with a relative molecular mass of 21,000 and possesses GTPase activity. It can bind to and hydrolyze GTP (guanine nucleotides), thereby acting as a switch in signal transduction. When GTP binds, the RAS protein is in an activated state, triggering downstream signal transduction pathways; when GTP is hydrolyzed, the RAS protein is inactivated, and the signal transduction process is terminated.

[0093] 41. The “RPS6KB1” gene described in this disclosure is ribosomal protein S6 kinase B1, a gene in the rat species *Rattus norvegicus*, and its synonyms include p70 S6K-alpha. The gene ID is 83840, and its gene type is protein coding. Furthermore, regarding the RPS6KB1 gene, there exists a product called RPS6KB1 Knockout Lentivirus. This lentivirus, after infecting animal cells, can simultaneously express Cas9, the target gene sgRNA, and the puromycin resistance gene, and can be used to knock out the target gene in animal cells using CRISPR / Cas9 technology.

[0094] 42. The “MAN2A1” gene described in this disclosure, also known as mannosidase α-class 2A member 1, is a protein-encoding gene. Diseases associated with this gene include congenital erythropoiesis-related anemia and systemic autoimmune diseases. Its associated pathways include the translation of structural proteins and the elongation of the N-glycosylated side chain in the middle / trans-Golgi apparatus. Gene ontology (GO) annotations related to this gene include carbohydrate binding and α-mannosidase activity.

[0095] 43. The “FER” gene described in this disclosure is a member of the plant receptor protein kinase gene (RLK) family, specifically belonging to the plant-specific crRLK gene family. The protein-coding region of this gene contains only one short intron in the 5′-UTR (untranslated region), with no introns at other locations. The FER protein is a membrane receptor protein with four distinct domains: an extracellular N-terminal signal peptide, two repeating Malectin-like domains, a central transmembrane domain, and an intracellular serine / threonine kinase domain. The FER protein is polarly expressed on the filamentous organs of ovule synergists and was first identified in the recognition process between pollen tubes and synergists in Arabidopsis thaliana. When a pollen tube enters an ovule with a FER gene mutation, the tip of the pollen tube does not rupture but continues to grow inside the ovule, thus preventing fertilization. As a positive regulator, the FER protein can activate the phosphatase activity of ABI2, thereby negatively regulating the plant's response to ABA.

[0096] 44. The full name of the "DNAJB1" gene described in this disclosure is DnaJ(Hsp40)homolog,subfamily B,member 1, also known as Hdj1, Hsp40, HSPF1, RSPH16B, and Sis1. This gene is located at 19p13.12 on the human chromosome and is a protein coding gene. The protein encoded by the DNAJB1 gene is a molecular chaperone involved in various cellular activities, such as protein folding and oligomeric protein complex assembly. The encoded protein is a highly conserved J-domain protein, and as one of the two major types of molecular chaperones, it participates in a wide range of cellular events.

[0097] 45. The “ATP1B1” gene described in this disclosure is a gene related to Na+. + / K + Genes related to ion transport encode proteins containing Na+. + / K + The β1 subunit of the ion pump. Na + / K + Ion pumps are crucial molecules for maintaining ion balance inside and outside cells, participating in various physiological and pathological processes. Mutations in the ATP1B1 gene can lead to a variety of diseases, such as congenital nephrotic syndrome and pseudoaldosteronism. Furthermore, ATP1B1 gene expression is also associated with the occurrence and development of some cancers.

[0098] 46. ​​The "FIG" gene described in this disclosure is a gene associated with the calcium signaling pathway and plays an important role in the osmotic pressure response in *Aspergillus cristatus*. This gene is 994 nt in length, contains 3 introns, is predicted to encode 271 amino acids, and has four transmembrane regions, classifying it as a hydrophobic protein. Furthermore, the conserved domains and transmembrane structure of the Fig gene are similar to those of other *Aspergillus* Fig genes. The Fig gene likely plays a crucial role in the calcium signaling pathway. It responds to changes in osmotic pressure by transporting calcium ions, which may influence processes such as sexual development, asexual sporulation, and hyphal growth.

[0099] 47. The “PRKACA” gene described in this disclosure is a gene encoding the catalytic subunit of protein kinase A. This gene plays an important role in the human body, particularly in processes such as cell differentiation, proliferation, and apoptosis. The PRKACA gene participates in the holoenzyme synthesis of protein kinase A by encoding the catalytic subunit. This holoenzyme is a tetramer, comprising two regulatory subunits and two catalytic subunits. When cAMP (a second messenger) binds to the holoenzyme, a regulatory subunit dimer can dissociate, while simultaneously binding to the catalytic subunits of four cAMP molecules and two free monomers, thereby activating protein kinase A. Variations in the PRKACA gene are associated with a variety of diseases. Constitutive activation of this gene, caused by somatic mutations or genomic duplications in regions including the gene, is associated with adrenal hyperplasia and adenoma, and with adrenocorticotropic hormone-dependent Cushing's syndrome. Furthermore, variations in the PRKACA gene may also be associated with cancers such as liver cancer.

[0100] 48. The “PRKACB” gene described in this disclosure is a gene encoding the catalytic subunit β of protein kinase A, belonging to the serine / threonine protein kinase family. Located on chromosome 1, this gene participates in the signal transduction of cyclic adenosine monophosphate (cAMP)-dependent protein kinases by encoding cAMP β. CAMP signal transduction plays a crucial role in processes such as cell proliferation and differentiation. There are many alternative splice transcriptomorphs of the PRKACB gene, each with different functions. Mutations in the PRKACB gene are associated with the development of various diseases, including tumors and cardiovascular diseases. In tumors, PRKACB gene mutations can lead to abnormal cell proliferation and apoptosis, thereby promoting tumor development and progression. In cardiovascular diseases, PRKACB gene mutations can affect the contractile and diastolic functions of cardiomyocytes, thus causing heart disease.

[0101] 49. The "IRF2BP2" gene described in this disclosure is an interferon regulator expressed in both humans and animals, primarily functioning in cardiac and skeletal muscle tissues. It comprises two alternative splicing isoforms: IRF2BP2A and IRF2BP2B. The IRF2BP2 gene was first identified in 2003 by Stephen et al. using a yeast two-hybrid system and was identified as a co-acting transcription factor of IRF2. Furthermore, studies have shown that IRF2BP2 is a target gene of p53 and can act as a co-repressor of p53 transcription, inhibiting the expression of downstream genes such as p21 and BAX, thereby playing an important role in biological processes such as cell cycle, cell differentiation, and apoptosis.

[0102] 50. The “SEC16A” gene described in this disclosure is located on human chromosome 9q34.3 and is widely expressed in tissues such as the stomach, ovary, and bone marrow. It encodes a protein that forms part of the Sec16 complex, which plays an important role in protein transport from the endoplasmic reticulum to the Golgi apparatus. Alternative splicing is a post-transcriptional regulatory mechanism that affects protein diversity, participates in cancer development, and is significantly associated with the overall survival of ovarian cancer patients. Alternative splicing of SEC16A has been shown to be associated with the occurrence, invasion, and metastasis of gastric cancer, and its gene expression level is negatively correlated with patient survival.

[0103] 51. The "NOTCH1" gene described in this disclosure is a single-pass transmembrane receptor, initially discovered in the Drosophila nervous system. It is named for the notch (notch) created at the edge of the Drosophila wing due to partial loss-of-function mutations. The protein encoded by the NOTCH1 gene is a highly conserved cell surface receptor belonging to the Notch family, which also includes four other receptors: NOTCH2, NOTCH3, and NOTCH4. In humans, the NOTCH1 gene is located on chromosome 9, with a start site of 139388896 and a stop site of 139440314, on the negative strand. Disorders of Notch signaling not only directly induce tumorigenesis but can also indirectly induce tumorigenesis through multiple other signaling pathways. Abnormalities in the Notch signaling pathway have been found to be closely related to esophageal cancer, gastric cancer, cervical cancer, and colorectal cancer, with Notch1 abnormalities being the most frequently detected in tumor tissues.

[0104] 52. The “RUNX1” gene described in this disclosure, also known as Runt-related transcription factor 1, is a protein-coding gene that plays a crucial role in many biological processes, including cell development and the regulation of specific gene expression. Located on chromosome 21q22, the RUNX1 gene regulates the expression of multiple genes essential for hematopoiesis. Mutations in the RUNX1 gene, particularly germline mutations, are associated with a range of hematologic disorders. Germline mutations can lead to autosomal dominant familial platelet disorders and increase the risk of hematologic malignancies such as acute myeloid leukemia (AML) or myelodysplastic syndromes (MDS). These diseases can progress, for example, from MDS to AML. Among the common gene mutation populations associated with germline RUNX1 mutations, pathogenic mutations include nonsense mutations, frameshift mutations, duplications, deletions, and missense mutations. MDS and AML are the most common hematologic malignancies with germline RUNX1 mutations; they have also been reported in chronic myeloid leukemia (CMML), T-lymphocytic leukemia / lymphoma, and a very small number of B-cell malignancies, including hairy cell leukemia.

[0105] (II) Detailed Technical Solution

[0106] In a specific implementation, in a first aspect, this disclosure provides a method for enriching a fusion gene or splice isoform, comprising amplifying the fusion gene using random primers and specific primers targeting a first fusion gene, wherein the specific primers carry modification groups for purification, and after amplification, recovering the amplification product carrying the modification groups.

[0107] In an optional embodiment, the modifying group includes biotin.

[0108] In an optional embodiment, the target fragment of the specific primer in the first fusion gene is 1 to 100 nt away from the fusion breakpoint, preferably 4 to 79 nt, more preferably 10 to 79 nt, and even more preferably 10 to 78 nt; the length of the random primer is 6 to 30 nt, preferably 8 to 25 nt, and even more preferably 16 to 20 nt.

[0109] In an optional implementation, the first fusion gene includes ALK, ROS1, RET, BRAF, NRG1, FGFR, NTRK, EGFR, ERBB2, MET, ESR1, AML1, BCR, RARα, E2A, TAL1, DEK, AML1, PDGFRA, KMT2A, NPM, NUP98, SET, TEL, TLS, CBFB, VMP1, CLDN18, ARHGAP, CD44, SLC1A2, TSC2, RNF216, AGTARP, TACC2, PPAPDC1A, RAS, RPS6KB1, MAN2A1, FER, DNAJB1, ATP1B1, FIG, PRKACA, PRKACB, IRF2BP2, SEC16A, NOTCH1, or RUNX1.

[0110] In optional embodiments, the breakpoint of the ALK gene is optionally located at the junction of exon 19 and exon 20; the breakpoint of the ROS1 gene is optionally located at the junction of exon 31 and exon 32, the junction of exon 32 and exon 33, or the junction of exon 33 and exon 34; the breakpoint of the RET gene is optionally located at the junction of exon 11 and exon 12; the breakpoint of the NRG1 gene is optionally located at the junction of exon 1 and exon 2; the breakpoint of the BRAF gene is optionally located at the junction of exon 3 and exon 4, or the junction of exon 10 and exon 11; the FGFR gene includes the FGFR1 gene located on chromosome 8, the FGFR2 gene located on chromosome 10, or the FGFR3 gene located on chromosome 4. Optionally, the breakpoint of the ALK gene is optionally located at the junction of exon 31 and exon 32, the junction of exon 32 and exon 33, or the junction of exon 33 and exon 34. The breakpoint is located at the junction of exon 9 and exon 10 of the FGFR1 gene, at the junction of exon 2 and exon 3 of the FGFR2 gene, or at the junction of exon 17 and exon 18 of the FGFR3 gene; the NTRK gene includes the NTRK1 gene located on chromosome 1, the NTRK2 gene located on chromosome 9, or the NTRK3 gene located on chromosome 15. Optionally, the breakpoint is located at the junction of exon 9 and exon 10 of the NTRK1 gene, at the junction of exon 13 and exon 14 of the NTRK2 gene, at the junction of exon 14 and exon 15 of the NTRK3 gene, or at the junction of exon 19 and exon 20 of the NTRK3 gene; the breakpoint of the MET gene is optionally located at the junction of exon 13 and exon 14.

[0111] In an optional embodiment, the specific primers for the ALK gene have a nucleotide sequence that is 80% to 100% identical to the nucleotide sequence shown in any one of SEQ ID No: 1 to 9, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0112] The specific primers for the ROS1 gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID Nos: 10-11, 49-51, and the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0113] The specific primers for the RET gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 12 to 14, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0114] The specific primers for the BRAF gene have a nucleotide sequence that is 80% to 100% identical to the nucleotide sequence shown in any one of SEQ ID No: 15 to 18, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0115] The specific primers for the NRG1 gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 19 to 21, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0116] The specific primers for the FGFR gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 22 to 27, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0117] The specific primers for the NTRK gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 28 to 38, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0118] The specific primers for the MET gene have a nucleotide sequence that is 80% to 100% identical to the nucleotide sequence shown in any one of SEQ ID No: 52 to 53, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0119] In an optional embodiment, a blocker probe that blocks the amplification of the wild-type first fusion gene is also added to the amplification reaction system.

[0120] In an optional embodiment, the recovery method includes adding a purification reagent for capturing the modifying groups to the amplification product.

[0121] Optionally, the purification reagent includes magnetic beads, which are linked to a modifying group-bound molecule.

[0122] Secondly, this disclosure provides enrichment primers for fusion genes or splice isoforms, including random primers and specific primers targeting a first fusion gene, wherein the specific primers have modification groups for purification, the purification modification being used to separate the enriched product composed of random primers and specific primers as primer pairs.

[0123] In an optional embodiment, the target fragment of the specific primer in the first fusion gene is 1 to 100 nt away from the fusion breakpoint, preferably 4 to 79 nt, more preferably 10 to 79 nt, and even more preferably 10 to 78 nt; the length of the random primer is 6 to 30 nt, preferably 8 to 25 nt, and even more preferably 16 to 20 nt.

[0124] Preferably, the first fusion gene includes ALK, ROS1, RET, BRAF, NRG1, FGFR, NTRK, EGFR, ERBB2, MET, ESR1, AML1, BCR, RARα, E2A, TAL1, DEK, AML1, PDGFRA, KMT2A, NPM, NUP98, SET, TEL, TLS, CBFB, VMP1, CLDN18, ARHGAP, CD44, SLC1A2, TSC2, RNF216, AGTARP, TACC2, PPAPDC1A, RAS, RPS6KB1, MAN2A1, FER, DNAJB1, ATP1B1, FIG, PRKACA, PRKACB, IRF2BP2, SEC16A, NOTCH1, or RUNX1.

[0125] The breakpoint of the ALK gene may optionally be located at the junction of exon 19 and exon 20; the breakpoint of the ROS1 gene may optionally be located at the junction of exon 31 and exon 32, the junction of exon 32 and exon 33, or the junction of exon 33 and exon 34; the breakpoint of the RET gene may optionally be located at the junction of exon 11 and exon 12; the breakpoint of the NRG1 gene may optionally be located at the junction of exon 1 and exon 2; the breakpoint of the BRAF gene may optionally be located at the junction of exon 3 and exon 4, or the junction of exon 10 and exon 11; the FGFR gene includes the FGFR1 gene located on chromosome 8, the FGFR2 gene located on chromosome 10, or the FGFR3 gene located on chromosome 4. Optionally, the breakpoint location... The breakpoint is located at the junction of exon 9 and exon 10 of the FGFR1 gene, at the junction of exon 2 and exon 3 of the FGFR2 gene, or at the junction of exon 17 and exon 18 of the FGFR3 gene; the NTRK gene includes the NTRK1 gene located on chromosome 1, the NTRK2 gene located on chromosome 9, or the NTRK3 gene located on chromosome 15. Optionally, the breakpoint is located at the junction of exon 9 and exon 10 of the NTRK1 gene, at the junction of exon 13 and exon 14 of the NTRK2 gene, at the junction of exon 14 and exon 15 of the NTRK3 gene, or at the junction of exon 19 and exon 20 of the NTRK3 gene; the breakpoint of the MET gene is optionally located at the junction of exon 13 and exon 14.

[0126] In an optional embodiment, the specific primers for the ALK gene have a nucleotide sequence that is 80% to 100% identical to the nucleotide sequence shown in any one of SEQ ID No: 1 to 9, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0127] The specific primers for the ROS1 gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID Nos: 10-11, 49-51, and the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0128] The specific primers for the RET gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 12 to 14, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0129] The specific primers for the BRAF gene have a nucleotide sequence that is 80% to 100% identical to the nucleotide sequence shown in any one of SEQ ID No: 15 to 18, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0130] The specific primers for the NRG1 gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 19 to 21, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0131] The specific primers for the FGFR gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 22 to 27, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0132] The specific primers for the NTRK gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 28 to 38, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0133] The specific primers for the MET gene have a nucleotide sequence that is 80% to 100% identical to the nucleotide sequence shown in any one of SEQ ID No: 52 to 53, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0134] In an optional embodiment, the modifying group includes biotin.

[0135] In an optional implementation, a blocker probe that blocks the amplification of the wild-type first fusion gene is also included.

[0136] Thirdly, this disclosure provides an enrichment composition for fusion genes or splice isomers, comprising the enrichment primers described in the second aspect and purification reagents for capturing modifying groups.

[0137] Optionally, the purification reagent includes magnetic beads, which are linked to a modifying group-bound molecule.

[0138] Fourthly, this disclosure provides the use of the enrichment primers described in the second aspect or the enrichment compositions described in the third aspect in the preparation of gene fusion enrichment reagents or kits.

[0139] Fifthly, this disclosure provides an enrichment reagent or kit for fusion genes or splice isoforms, comprising the enrichment primers described in the second aspect or the enrichment composition described in the third aspect; preferably, it further comprises sampling tools and / or other consumables.

[0140] Optionally, the samples collected by the sampling tool include tumor cells and / or body fluids; preferably, the body fluids include blood, saliva, urine, pleural effusion, or ascites.

[0141] In a sixth aspect, this disclosure provides a method for detecting fusion genes or splice isoforms, comprising adding detection primer pairs to a sample containing a fusion gene to amplify the fusion gene, using purification reagents to separate the fusion gene enrichment product from the amplification product after amplification, and then using sequencing primer pairs to sequence the fusion gene enrichment product.

[0142] The detection primer pair includes a random primer and a specific primer targeting the first fusion gene. The random primer is connected to a first universal primer, and the specific primer is connected to a second universal primer with a modification group for purification. The purification modification is used to separate the enriched product composed of the random primer and the specific primer as primer pairs.

[0143] The sequencing primer pair includes a first sequencing primer and a second sequencing primer. The first sequencing primer contains a sequence that hybridizes with a first universal primer and a first sequencing primer adapter sequence. The second sequencing primer contains a sequence that hybridizes with a second universal primer and a second sequencing primer adapter sequence. This detection method can be used to detect whether a fusion gene or splice isoform is present in a suspected case sample, thereby predicting the disease risk of the individual from whom the sample was obtained. Alternatively, this detection method can also be used to detect fusion genes or splice isoforms prepared in vitro, for example, to evaluate the quality of prepared fusion gene products or the splicing efficiency of constructed splice isoform systems.

[0144] In an optional embodiment, the modifying group includes biotin.

[0145] In an optional implementation, a blocker probe that blocks the amplification of the wild-type first fusion gene is also added to the fusion gene amplification reaction system.

[0146] In an optional embodiment, the purification reagent includes magnetic beads with a modified group attached to the molecule.

[0147] In an optional embodiment, the target fragment of the specific primer in the first fusion gene is 1 to 100 nt away from the fusion breakpoint, preferably 4 to 79 nt, more preferably 10 to 79 nt, and even more preferably 10 to 78 nt; the length of the random primer is 6 to 30 nt, preferably 8 to 25 nt, and even more preferably 16 to 20 nt.

[0148] Preferably, the first fusion gene includes ALK, ROS1, RET, BRAF, NRG1, FGFR, NTRK, EGFR, ERBB2, MET, ESR1, AML1, BCR, RARα, E2A, TAL1, DEK, AML1, PDGFRA, KMT2A, NPM, NUP98, SET, TEL, TLS, CBFB, VMP1, CLDN18, ARHGAP, CD44, SLC1A2, TSC2, RNF216, AGTARP, TACC2, PPAPDC1A, RAS, RPS6KB1, MAN2A1, FER, DNAJB1, ATP1B1, FIG, PRKACA, PRKACB, IRF2BP2, SEC16A, NOTCH1, or RUNX1.

[0149] The breakpoint of the ALK gene may optionally be located at the junction of exon 19 and exon 20; the breakpoint of the ROS1 gene may optionally be located at the junction of exon 31 and exon 32, the junction of exon 32 and exon 33, or the junction of exon 33 and exon 34; the breakpoint of the RET gene may optionally be located at the junction of exon 11 and exon 12; the breakpoint of the NRG1 gene may optionally be located at the junction of exon 1 and exon 2; the breakpoint of the BRAF gene may optionally be located at the junction of exon 3 and exon 4, or the junction of exon 10 and exon 11; the FGFR gene includes the FGFR1 gene located on chromosome 8, the FGFR2 gene located on chromosome 10, or the FGFR3 gene located on chromosome 4. Optionally, the breakpoint location... The breakpoint is located at the junction of exon 9 and exon 10 of the FGFR1 gene, at the junction of exon 2 and exon 3 of the FGFR2 gene, or at the junction of exon 17 and exon 18 of the FGFR3 gene; the NTRK gene includes the NTRK1 gene located on chromosome 1, the NTRK2 gene located on chromosome 9, or the NTRK3 gene located on chromosome 15. Optionally, the breakpoint is located at the junction of exon 9 and exon 10 of the NTRK1 gene, at the junction of exon 13 and exon 14 of the NTRK2 gene, at the junction of exon 14 and exon 15 of the NTRK3 gene, or at the junction of exon 19 and exon 20 of the NTRK3 gene; the breakpoint of the MET gene is optionally located at the junction of exon 13 and exon 14.

[0150] In an optional embodiment, the specific primers for the ALK gene have a nucleotide sequence that is 80% to 100% identical to the nucleotide sequence shown in any one of SEQ ID No: 1 to 9, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0151] The specific primers for the ROS1 gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID Nos: 10-11, 49-51, and the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0152] The specific primers for the RET gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 12 to 14, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0153] The specific primers for the BRAF gene have a nucleotide sequence that is 80% to 100% identical to the nucleotide sequence shown in any one of SEQ ID No: 15 to 18, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0154] The specific primers for the NRG1 gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 19 to 21, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0155] The specific primers for the FGFR gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 22 to 27, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0156] The specific primers for the NTRK gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 28 to 38, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0157] The specific primers for the MET gene have a nucleotide sequence that is 80% to 100% identical to the nucleotide sequence shown in any one of SEQ ID No: 52 to 53, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0158] In an optional implementation, the sample includes tumor cells and / or body fluids; preferably, the body fluids include blood, saliva, urine, pleural effusion, or ascites.

[0159] In a seventh aspect, this disclosure provides a detection primer for a fusion gene or splice isoform, comprising a random primer and a specific primer targeting a first fusion gene, wherein the random primer is connected to a first universal primer, and the specific primer is connected to a second universal primer with a modification group for purification, the purification modification being used to separate enriched products composed of the random primer and the specific primer as primer pairs.

[0160] In an optional embodiment, the modifying group includes biotin.

[0161] In an optional embodiment, the detection primers further include a blocker probe that blocks the amplification of the wild-type first fusion gene.

[0162] In an optional embodiment, the target fragment of the specific primer in the first fusion gene is 1 to 100 nt away from the fusion breakpoint, preferably 4 to 79 nt, more preferably 10 to 79 nt, and even more preferably 10 to 78 nt; the length of the random primer is 6 to 30 nt, preferably 8 to 25 nt, and even more preferably 16 to 20 nt.

[0163] Preferably, the first fusion gene includes ALK, ROS1, RET, BRAF, NRG1, FGFR, NTRK, EGFR, ERBB2, MET, ESR1, AML1, BCR, RARα, E2A, TAL1, DEK, AML1, PDGFRA, KMT2A, NPM, NUP98, SET, TEL, TLS, CBFB, VMP1, CLDN18, ARHGAP, CD44, SLC1A2, TSC2, RNF216, AGTARP, TACC2, PPAPDC1A, RAS, RPS6KB1, MAN2A1, FER, DNAJB1, ATP1B1, FIG, PRKACA, PRKACB, IRF2BP2, SEC16A, NOTCH1, or RUNX1.

[0164] The breakpoint of the ALK gene may optionally be located at the junction of exon 19 and exon 20; the breakpoint of the ROS1 gene may optionally be located at the junction of exon 31 and exon 32, the junction of exon 32 and exon 33, or the junction of exon 33 and exon 34; the breakpoint of the RET gene may optionally be located at the junction of exon 11 and exon 12; the breakpoint of the NRG1 gene may optionally be located at the junction of exon 1 and exon 2; the breakpoint of the BRAF gene may optionally be located at the junction of exon 3 and exon 4, or the junction of exon 10 and exon 11; the FGFR gene includes the FGFR1 gene located on chromosome 8, the FGFR2 gene located on chromosome 10, or the FGFR3 gene located on chromosome 4. Optionally, the breakpoint location... The breakpoint is located at the junction of exon 9 and exon 10 of the FGFR1 gene, at the junction of exon 2 and exon 3 of the FGFR2 gene, or at the junction of exon 17 and exon 18 of the FGFR3 gene; the NTRK gene includes the NTRK1 gene located on chromosome 1, the NTRK2 gene located on chromosome 9, or the NTRK3 gene located on chromosome 15. Optionally, the breakpoint is located at the junction of exon 9 and exon 10 of the NTRK1 gene, at the junction of exon 13 and exon 14 of the NTRK2 gene, at the junction of exon 14 and exon 15 of the NTRK3 gene, or at the junction of exon 19 and exon 20 of the NTRK3 gene; the breakpoint of the MET gene is optionally located at the junction of exon 13 and exon 14.

[0165] In an optional embodiment, the specific primers for the ALK gene have a nucleotide sequence that is 80% to 100% identical to the nucleotide sequence shown in any one of SEQ ID No: 1 to 9, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0166] The specific primers for the ROS1 gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID Nos: 10-11, 49-51, and the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0167] The specific primers for the RET gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 12 to 14, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0168] The specific primers for the BRAF gene have a nucleotide sequence that is 80% to 100% identical to the nucleotide sequence shown in any one of SEQ ID No: 15 to 18, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0169] The specific primers for the NRG1 gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 19 to 21, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0170] The specific primers for the FGFR gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 22 to 27, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0171] The specific primers for the NTRK gene have nucleotide sequences that are 80% to 100% identical to the nucleotide sequences shown in any one of SEQ ID No: 28 to 38, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0172] The specific primers for the MET gene have a nucleotide sequence that is 80% to 100% identical to the nucleotide sequence shown in any one of SEQ ID No: 52 to 53, wherein the identity is preferably 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, 98%, or 100%.

[0173] Eighthly, this disclosure provides a detection composition for fusion genes or splice isomers, comprising the detection primers described in the seventh aspect, a purification reagent for capturing modifying groups, and a sequencing primer pair suitable for a sequencing platform;

[0174] The sequencing primer pair includes a first sequencing primer and a second sequencing primer. The first sequencing primer contains a sequence that hybridizes with a first universal primer and a first sequencing primer adapter sequence. The second sequencing primer contains a sequence that hybridizes with a second universal primer and a second sequencing primer adapter sequence.

[0175] Optionally, the purification reagent includes magnetic beads, which are linked to a modifying group-bound molecule.

[0176] Ninthly, this disclosure provides the use of the detection primers described in the seventh aspect or the detection compositions described in the eighth aspect in the preparation of reagents or kits for gene fusion detection or splice isomers.

[0177] In a tenth aspect, this disclosure provides a fusion gene detection reagent or kit, comprising the detection primers described in the seventh aspect or the detection composition described in the eighth aspect; preferably, it further comprises sampling tools and / or other consumables.

[0178] In an optional implementation, the sample includes tumor cells and / or body fluids; preferably, the body fluids include blood, saliva, urine, pleural effusion, or ascites.

[0179] In the eleventh aspect, this disclosure provides a diagnostic method for diseases related to gene fusion mutations or abnormal splicing of precursor mRNA, including collecting a sample to be tested, performing sequencing detection on gene fusion mutation-related breakpoints according to the detection method described in the sixth aspect, and estimating the risk of gene fusion mutations in the sample to be tested.

[0180] For example, the method for enriching unknown fusion genes used in a specific embodiment of this disclosure includes:

[0181] (1) Extract total RNA from the sample to be tested and reverse transcribe it to obtain cDNA; (2) PCR amplification: use random primers and specific primers to amplify the cDNA obtained in step (1) for the first time to obtain the fusion gene amplification product; the specific primers are modified with biotin; (3) Secondary PCR amplification: use streptavidin magnetic beads to purify the product, and perform secondary amplification on the purified material to obtain the unknown fusion gene library.

[0182] For example, a method for detecting unknown fusion genes used in a specific embodiment of this disclosure includes:

[0183] (1) Extract total RNA from the sample to be tested and reverse transcribe it to obtain cDNA; (2) PCR amplification: use random primers and specific primer pairs to amplify the cDNA obtained in step (1) for the first time to obtain the fusion gene amplification product; the random primers are connected to the 5' end of the first universal primer (read-1) sequence; the specific primers are connected to the 5' end of the second universal primer (read-2) sequence, and the second universal primer is also modified with biotin; (3) Secondary PCR amplification: use streptavidin magnetic beads to purify the product, and perform a second PCR amplification on the purified product. In the CR reaction, in order to use the Illumina sequencing platform to sequence the amplified products, the upstream primer used in the second PCR was P5+i5+read1 (5' end to 3' end, where P5 is the upstream sequencing primer adapted to the Illumina sequencing platform and i5 is the corresponding upstream tag sequence); the downstream primer was P7+i7+read2 (5' end to 3' end, where P7 is the downstream sequencing primer adapted to the Illumina sequencing platform and i7 is the corresponding downstream tag sequence), and finally the target region library was obtained; (4) Library quality control and sequencing.

[0184] The optional nucleotide sequence of the upstream primer P5+i5+read1 includes 5'-(P5)AATGATACGGCGACCACCGAGAT CTACAC(i5)NNNNNNNN(read1)TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' (SEQ ID No: 39).

[0185] The optional nucleotide sequence of the downstream primer P7+i7+read2 includes 5'-(P7)CAAGCAGAAGACGGCATACGAGAT(i7)NNNNNNNN(read2)GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-3' (SEQ ID No: 40).

[0186] In actual clinical samples, FFPE (Formalin-fixed paraffin-embedding, tissue samples) have low mRNA quality and severe fragmentation; fresh tissue and CTCs (circulating tumor cells) have high-quality mRNA with intact Poly A tails. Possible mRNA morphologies in different samples are as follows: Figure 1 As shown. When using enrichment primers or detection primers to perform the first PCR amplification of clinical samples, the possible reactions are as follows. Figure 2 As shown, the specific primers are those described in this disclosure.

[0187] The types of amplification products obtained, such as Figure 3 As shown, where:

[0188] 1. Product 1 is the product of random primers. The product is not biotinylated and can be purified and filtered out using streptavidin magnetic beads.

[0189] 2. Product 2 is the target product;

[0190] 3. Product 3 is a wild-type product. The amplification of most wild-type products can be blocked by the Blocker, thus enriching the generation of the target product.

[0191] 4. Product 4 is the product of biotinylated specific primers and random primers, and the product does not pass through the fusion mutation breakpoint. The generation of such products can be reduced by adjusting the distance between the specific primers and the fusion mutation breakpoint, as well as by the competitive amplification effect of the Blocker.

[0192] It should be noted that some clinical tissue samples contain very low levels of wild-type mRNA, meaning the content of product 3 in the first PCR product is low. Furthermore, the fragment in the first fusion gene targeted by the specific primers provided in this disclosure is 4–79 nt away from the breakpoint. This effectively reduces the amount of fusion genes without breakpoints in the amplification product. Therefore, for mRNAs from such sources, a blocker is not required during the first PCR amplification process.

[0193] Furthermore, the fusion gene product containing the breakpoint obtained from the first PCR was subjected to a second PCR using sequencing primers, as follows: Figure 4 As shown, the final product can be directly used for gene sequencing on the Illumina sequencing platform. The overall technical route is as follows: Figure 5 As shown.

[0194] When the total RNA contains splice isomers obtained by alternative splicing of pre-mRNA, the above enrichment and detection methods can be used to detect the splice isomers.

[0195] The following detailed description of some embodiments of this disclosure is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0196] The cell lines or plasmids used in the following examples are as follows:

[0197] NCI-H2228 Zhejiang Meisen Cell Technology Co., Ltd. Product Number CTCC-001-0340 - SW780 Saibaikang Bio icell-h278 - TPC-1 Saibaikang Bio icell-h309 - HCC78 Zhejiang Meisen Cell Technology CTCC-001-0872 - MDA MB 175Ⅶ Wuhan Pronosei Biotechnology CL-0381B - KG-1a Saibaikang Bio icell-h308 - plasmid ABM28599-6 Shanghai Sangon Biotech Synthetic plasmids SEQ ID No: 41 plasmid ABM28599-7 Shanghai Sangon Biotech Synthetic plasmids SEQ ID No: 42 plasmid ABM28599-8 Shanghai Sangon Biotech Synthetic plasmids SEQ ID No: 43 KM12-SM Zhejiang Meisen Cell Technology CTCC-001-0315 - plasmid ABM28599-1 Shanghai Sangon Biotech Synthetic plasmids SEQ ID No: 44 plasmid ABM28599-2 Shanghai Sangon Biotech Synthetic plasmids SEQ ID No: 45 plasmid ABM28599-3 Shanghai Sangon Biotech Synthetic plasmids SEQ ID No: 46 plasmid ABM28599-4 Shanghai Sangon Biotech Synthetic plasmids SEQ ID No: 47 plasmid ABM28599-5 Shanghai Sangon Biotech Synthetic plasmids SEQ ID No: 48 Hs746T ATCC HTB-135 -

[0198] Example 1: Detection of Fusion Genes

[0199] This embodiment describes an exemplary method for detecting fusion genes. Using different starting amounts of fusion cell line templates as examples, the accuracy of the detection method provided in this disclosure is verified.

[0200] (1) RNA extraction and reverse transcription

[0201] RNA extraction steps: RNA samples were extracted using the OMEGA Total RNA Kit I.

[0202] RNA reverse transcription steps: Add 1 μg of RNA sample and random primer 6N, and use the Novozymes HiScript III 1st Strand cDNA Synthesis Kit (+gDNA wiper) to reverse transcribe the above sample into cDNA for later use.

[0203] (2) PCR-1 reaction steps

[0204] The primer extension master mixture was prepared by mixing 5 μL cDNA, 5 μL PCR buffer (4×VAHTS Multi-PCR Mix), 8 μL nucleic acid-free water, 1 μL random primers (final concentration 2.5 μM), and 1 μL biotin-labeled specific primers (final concentration 2.4 μM). The reaction procedure is shown in the table below:

[0205]

[0206] After the PCR-1 reaction, the extension product was collected and analyzed using BeaverBeads (streptavidin magnetic beads) from Beaver Biotechnology. TM Streptavidin (300 nm) was used to purify the extended product, yielding streptavidin magnetic beads-bound biotinylated nucleic acid.

[0207] The random primers contain a read1 sequence at their 5' end. The length of the random primers can be 6–30 nt, or a mixture of random primers of different nt lengths. In this embodiment, 16 nt random primers are used. The specific primers contain the sequence of the known fusion site and contain a read2 sequence at their 5' end, and are labeled with biotin. The specific primers can be one or more sequences targeting a single site, or a combination of one or more sequences targeting a single site for multi-site detection. The design principle for read-1 and read-2 is 16–30 nt; this scheme selects 23 nt.

[0208] For example, the specific primers used in this disclosure are shown in the table below: where the italicized fragment is the read-2 sequence and the regular fragment is the specific primer sequence.

[0209]

[0210]

[0211]

[0212]

[0213]

[0214] Note: (fusion site exonA-RP / FP) indicates the fusion breakpoint location detected by this primer, where exonA-RP indicates the breakpoint is between exon A-1 and exon A; exonA-FP indicates the breakpoint is between exon A and exon A+1; A is a positive integer.

[0215] "*" indicates that the base to the left of "*" has a thiolated modification, "+" indicates that the base to the left of "+" has a locked nucleic acid modification, and " / 3ddC / " indicates a dideoxyribonucleotide modification.

[0216] Specifically, the PCR-1 reaction step also includes an internal reference primer, which can be one or more primers; more specifically, the internal reference primer is a primer combination composed of multiple internal reference primers. Specifically, the concentration of the internal reference primer is selected from 50 to 200 nM. In some specific embodiments, the internal reference primer combination used is a combination of internal reference primers N029, N030, N174, N064, and N065.

[0217] (3) PCR-2 amplification steps

[0218] The sequencing product was prepared by mixing 23 μL of nucleic acid-free water, 25 μL of PCR reaction buffer (VAHTS HiFi Universal Amplification Mix), biotinylated nucleic acid bound to magnetic beads obtained in step (2), 1 μL of the first upstream primer, and 1 μL of the first downstream primer. The reaction procedure is shown in the table below:

[0219]

[0220] After the PCR-2 reaction, the PCR product was collected and eluted with nucleic acid-free water and magnetic beads. DNA Selection Beads were used to purify the DNA, and the purified product was then quantified using Qubit.

[0221] The first upstream primer contains a sequence that hybridizes with the read1 base sequence and a first sequencing adapter sequence; the first downstream primer contains a sequence that hybridizes with the read2 base sequence and a second sequencing adapter sequence; the number of base overlaps between the first upstream primer and the read-1 sequence is 16 to 30, preferably 23; the number of base overlaps between the first downstream primer and the read-2 sequence is 16 to 30, preferably 23.

[0222] (4) Sequencing

[0223] Sequencing was performed using the Illumina sequencing system according to standard protocols.

[0224] (5) Data Analysis

[0225] S1: Perform targeted sequencing on the library preparation samples to obtain the original fastq file;

[0226] S2: Perform quality control on the original fastq file, removing adapter sequences, low-quality data, and excessively short reads (Q30 > 80%); S3: Count the number of target sequences (specific primers and target sequences) read from 1G of data (if the number of reads for the internal reference primers is greater than 200, the library construction is considered successful; if the number of reads for the target sequence is greater than 200, the fusion is considered to be expressed); Specifically, the target sequence is combined with sequence alignment, and the target sequence must contain at least 10 nt of bases on both sides of the template fusion gene mutation breakpoint.

[0227] Example 2: Experimental Investigation of the Distance Between Specific Primers and Breakpoints

[0228] Using NCI-H2228 cell line cDNA as a template, the cDNA information is shown in the table below:

[0229] NCI-H2228 EML4-ALK.E6bA20.COSF411

[0230]

[0231] Using N143, N144, N032, N146, N092, N148, and N149 as specific primers and N029, N030, N174, N064, and N065 as internal reference primers, experiments were conducted according to steps 1 to 5 of Example 1. The results are shown in the table above. The results indicate that N032 is the most effective primer for detecting EML4-ALK.E6aA20.COSF411.

[0232]

[0233] Specific primer combination 1, consisting of N143, N144, and N032, is characterized by its closest proximity to the breakpoint. Specific primer combination 2, consisting of N143, N032, N092, and N148, is characterized by the best detection performance of a single primer. Specific primer combination 3, consisting of N143, N144, N032, N146, N092, N148, and N149, is characterized by containing all specific primers. The above primer combinations were tested according to steps 1-5 of Example 1. The results are shown in the table above. The results indicate that using the specific primer combination with the best detection performance results in better performance than other combinations, and also better than using single primers.

[0234] Example 3: Validation of the detection system based on other fusion sites

[0235] The following are listed: N143, N032, N092, N148, N102, N209, N213, N214, N078, N184, N185, N187, N188, N215, N216, N217, N181, N182, N276, N277, N204, N205, N104, N105, N180, N266, N267, NTRK2, N108, N273, N2 74, N271, and N110 are detection panels. N029, N030, N174, N064, and N056 are used as internal reference primer sets. Following steps 1 to 5 in Example 1, the presence of fusion genes related to ALK, RET, BRAF, NRG1, FGFR1, FGFR2, FGFR3, NTRK1, NTRK2, NTRK3, and ROS1 and their read counts in the test samples are detected respectively.

[0236] Cell line cDNA or plasmid DNA with different mutation information was used as a mutation template, and WBC cDNA was used as a wild-type template to form a detection template. The detection template was then tested using a panel according to the steps in Example 1. The detection results are as follows:

[0237] NCI-H2228 ALK Fusion EML4-ALK.E6aA20.COSF411 11445 NCI-H2228 ALK Fusion EML4-ALK.E6bA20.COSF474 2160 SW780 RET Fusion CCDC6-RET.C1R12.COSF1271 12,917 TPC-1 RET Fusion CCDC6-RET.C1R12.COSF1271 5,006 HCC78 ROS1 Fusion SLC34A2-ROS1.S4R32.COSF1196 3,004 MDA MB 175Ⅶ NRG1 Fusion TENM4-NRG1.T12N2 41,678 KG-1a FGFR1 fusion FGFR1OP2-FGFR1.F4F10 11,373 plasmid ABM28599-6 FGFR2 Fusion FGFR2-ATE1.F2A9 7,466 SW780 FGFR3 Fusion FGFR3-BAIAP2L1.F17B2 6,757 plasmid ABM28599-7 BRAF fusion BRAF-SLC26A4.B3S6 10,350 plasmid ABM28599-8 BRAF fusion MKRN1-BRAF.M4B11 3,909 KM12-SM NTRK1 Fusion TPM3-NTRK1.T7N10.COSF1329 11,315 plasmid ABM28599-1 NTRK2 Fusion NACC2-NTRK2.N4N13.COSF1449 75,439 plasmid ABM28599-2 NTRK2 Fusion QKI-NTRK2.Q6N16.COSF1447 18,942 plasmid ABM28599-3 NTRK3 Fusion NTRK3-ETV6.N14E6.COSF825 65,637 plasmid ABM28599-4 NTRK3 Fusion ETV6-NTRK3.E4N15.COSF823 81,786 plasmid ABM28599-5 NTRK3 Fusion ETV6-NTRK3.E5N15.COSF571 18,202

[0238] Experimental conclusion:

[0239] For the different samples mentioned above, the Panel in this embodiment can accurately detect the corresponding fusion gene target sequence, and the number of reads for the target sequence can meet the library construction requirements for sequencing to determine fusion sequences. Therefore, the method provided in this scheme can use specific primers designed based on known fusion gene sites, combined with random primers, to achieve the enrichment of fusion genes and sequencing library construction, and covers known fusion genes and unknown fusion genes in the "unknown gene - known fusion gene site" range.

[0240] Example 4: Experiment on random primer length

[0241] The specific primer set is based on N143(ALK exon20-RP), N144(ALK exon20-RP), N145(ALK exon20-RP), N032(ALK exon20-RP), N146(ALK exon20-RP), N092(ALK exon20-RP), N147(ALK exon20-RP), N148(ALK exon20-RP), and N149(ALK exon20-RP). The detection template is shown in the table below.

[0242]

[0243] The experiment was conducted according to steps 1-5 of Example 1, using random primers of different lengths. The number of reads for EML4-ALK.E6A20.COSF411 and EML4-ALK.E6A20.COSF474 with different primer lengths was measured. The results are shown in the table below. Figure 6 As shown, when the length of the random primer is 16-20 nt, more reads are obtained.

[0244] EML4-ALK.E6A20.COSF411 1,270 1,907 3,710 7,889 4,698 271 EML4-ALK.E6A20.COSF474 5,534 13,867 37,026 33,588 10,100 1,453

[0245] Example 5 Detection Limit

[0246] (1) Cell line validation

[0247] Total RNA of 1 μg from cell lines H2228, Hs746T, TPC-1, and HCC78 was reverse transcribed using a reverse transcription kit. The reverse transcription volume was 20 μl, resulting in a cDNA sample concentration of 50 ng / μl. Equal volumes of cDNA from the four cell lines were mixed to prepare a positive control, FR-02, at a concentration of 50 ng / μl. FR-02 was quantified using digital PCR to determine the copy number concentration of each fusion-positive site in the sample. The results are as follows:

[0248] Preparation of positive control material: FR-02 (cDNA concentration: 50 ng / μl)

[0249]

[0250] 10 μL of FR-02 sample was taken and diluted 10-fold and 100-fold with enzyme-free water to obtain FR-02-10X and FR-02-100X, respectively. The copy number concentrations of each fusion-positive site in FR-02-10X and FR-02-100X were determined by digital PCR as follows:

[0251]

[0252]

[0253] 5 μL of FR-02-100X sample was taken and mixed with 5 μL of HEK293 cell line cNDA sample at a concentration of 19.5 ng / μL to obtain sample FR-03. The copy number concentrations of each fusion positive site in FR-03 were as follows by digital PCR quantification:

[0254]

[0255] Using FR-02-10X (2 μL), FR-02-100X (2 μL), and FR-03 (1 μL) as starting templates, and N143, N032, N092, N148, N102, N265, N213, N214, N316, and N078 as detection panels, and N029, N030, N174, N064, and N056 as internal control primers, the ALK, RET, MET, and ROS1-related fusion genes and their read counts in the test samples were detected according to steps 1-5 of Example 1. The results are as follows:

[0256]

[0257]

[0258] As shown in the table above, the detection method provided in this disclosure can detect ALK, RET, MET, and ROS1-related fusion genes in all three test groups: FR-02-10X, FR-02-100X, and FR-03. Specifically, in the FR-02-100X test with a loading amount of 1 ng, it can support the detection of fusion-positive mutations with a total copy count as low as 50; in the FR-03 test with a loading amount of 10 ng, it can support the detection of fusion-positive mutations with a total copy count as low as 12. Therefore, the technical solution provided in this disclosure can achieve a method for enriching and detecting unknown fusion genes with low starting amounts and low loading amounts.

[0259] (2) Validation of clinical FFPE samples

[0260] This embodiment uses the clinical FFPE sample preserved by the applicant as the test sample. This sample has been confirmed as an EML4-ALK fusion case by the Festo 10-gene lung cancer mutation detection kit. At the same time, the H2228 cell line is used as a control. The primer set and reaction system of Example 4 are used to detect the number of reads of EML4-ALK.E6A20.COSF411 and EML4-ALK.E6A20.COSF474 in the two samples, respectively.

[0261] The qsep (high-performance capillary electrophoresis) data of FFPE samples and H2228 cell line are as follows: Figure 7 and Figure 8 As shown, the RNA length of the FFPE sample is distributed across multiple lengths, while the RNA length of the H2228 cell sample is mainly distributed around 2k and between 4k and 6k. The fusion gene detection results of the FFPE sample and the H2228 cell line are shown in the table below:

[0262] Sample loading amount (ng) 10 10 ALK-ex20-N143-EML4-ALK.E6A20.COSF411 53,092 20,794 ALK-ex20-N143-EML4-ALK.E6A20.COSF474 33,452 11,992

[0263] As can be seen, based on the detection method of this scheme, when using low-quality FFPE samples as detection samples, the number of reads for EML4-ALK.E6A20.COSF411 is 53092, and the number of reads for EML4-ALK.E6A20.COSF474 is 33452, which is higher than the number of reads for high-quality RNA samples of H2228 cell samples. This indicates that the fusion gene enrichment method and detection method provided by this invention can be applied to both high-quality and low-quality RNA samples.

[0264] Example 6: Detection system using Blocker

[0265] Hs746T-cD was used as the template for the MET 14 exon skipping mutant, and HCC78-cD was used as the template for the MET 14 exon skipping wild-type. N120 was used as the specific primer, and N029, N30, N174, and N064 were used as internal control primers. N099 was used as the MET14 exon skipping wild-type blocker. Experimental groups 1 to 4 were set up as shown in the table below. The number of reads containing the MET13 exon region and the number of reads containing the MET14 exon skipping mutant gene were detected in experimental groups 1 to 4 according to steps 1 to 5 in Example 1.

[0266] Experimental groups 1-4

[0267]

[0268]

[0269] The experimental results are shown in the table below:

[0270]

[0271] The experimental results show that the use of MET Blocker inhibits the amplification of MET wild-type in wild-type samples, but has virtually no effect on the amplification of MET14 exon skipping in mutant samples. Therefore, it is evident that adding a blocker probe to the fusion gene enrichment and detection method provided in this invention can increase the concentration of fusion genes in samples after library construction, thereby improving enrichment and detection efficiency.

[0272] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this disclosure.

Claims

1. A method for enriching fusion genes or splice isoforms for purposes other than disease diagnosis, characterized in that, When the enrichment target is a fusion gene, the method includes amplifying the fusion gene using random primers and a specific primer set targeting the first fusion gene, wherein the specific primers have modification groups for purification, and after amplification, the amplification product with modification groups is recovered. The specific primer has a target fragment in the first fusion gene that is 10–78 nt away from the fusion breakpoint; the random primer has a length of 16–20 nt; the first fusion gene is ALK, ROS1, RET, BRAF, NRG1, FGFR, or NTRK. When the enrichment target is a splice isoform, the splice isoform is a splice isoform of the MET gene. The method includes amplification using random primers and a specific primer set targeting the MET gene splice isoform. The specific primers have modification groups for purification. After amplification, the amplification product with modification groups is recovered. The specific primers have a target fragment distance of 10–78 nt from the breakpoint in the MET gene splice isoform; the random primers have a length of 16–20 nt.

2. The enrichment method according to claim 1, wherein, The modifying group includes biotin.

3. The enrichment method according to claim 1 or 2, wherein, The breakpoint of the ALK gene is located at the junction of exon 19 and exon 20; The breakpoint of the ROS1 gene is located at the junction of exon 31 and exon 32, the junction of exon 32 and exon 33, or the junction of exon 33 and exon 34. The breakpoint of the RET gene is located at the junction of exon 11 and exon 12; The breakpoint of the NRG1 gene is located at the junction of exon 1 and exon 2; The breakpoint of the BRAF gene is located at the junction of exon 3 and exon 4, or at the junction of exon 10 and exon 11. The FGFR gene includes the FGFR1 gene located on chromosome 8, the FGFR2 gene located on chromosome 10, or the FGFR3 gene located on chromosome 4. The breakpoint is located at the junction of exon 9 and exon 10 of the FGFR1 gene, at the junction of exon 2 and exon 3 of the FGFR2 gene, or at the junction of exon 17 and exon 18 of the FGFR3 gene. The NTRK gene includes the NTRK1 gene located on chromosome 1, the NTRK2 gene located on chromosome 9, or the NTRK3 gene located on chromosome 15. The breakpoint is located at the junction of exon 9 and exon 10 of the NTRK1 gene, at the junction of exon 13 and exon 14 of the NTRK2 gene, at the junction of exon 14 and exon 15 of the NTRK3 gene, or at the junction of exon 19 and exon 20 of the NTRK3 gene. The breakpoint of the MET is located at the junction of exon 13 and exon 14.

4. The enrichment method according to claim 3, wherein, The nucleotide sequences of the specific primers for the ALK gene are shown in SEQ ID No: 1-8; The nucleotide sequences of the specific primers for the ROS1 gene are shown in SEQ ID No: 10-11; The nucleotide sequences of the specific primers for the RET gene are shown in SEQ ID No: 13-14; The nucleotide sequences of the specific primers for the BRAF gene are shown in SEQ ID No: 15-18; The nucleotide sequences of the specific primers for the NRG1 gene are shown in SEQ ID No: 20-21; The nucleotide sequences of the specific primers for the FGFR gene are shown in SEQ ID No: 22-27; The nucleotide sequences of the specific primers for the NTRK gene are shown in SEQ ID No: 28-29, 32, 34, 36 or 38; The nucleotide sequences of the specific primers for MET are shown in SEQ ID No: 52-53.

5. The enrichment method according to any one of claims 1-4, wherein, When the enrichment target is a fusion gene, the amplification reaction system further includes a blocker probe that blocks the amplification of the wild-type first fusion gene; when the enrichment target is a MET gene splice isoform, the amplification reaction system further includes a blocker probe that blocks the amplification of the wild-type MET gene.

6. The enrichment method according to any one of claims 1-4, wherein, The recovery method includes adding a purification reagent to the amplification product to capture the modifying groups.

7. The enrichment method according to claim 6, wherein, The purification reagent includes magnetic beads, which are linked to a modifying group-bound molecule.

8. The use of enrichment primers in the preparation of reagents for enriching fusion genes or splice isoforms, characterized in that, When the enrichment target is a fusion gene, the enrichment primers include random primers and specific primers targeting the first fusion gene. The specific primers have modification groups for purification. The purification modification is used to separate the enrichment product using random primers and specific primers as primer pairs. The specific primer has a target fragment in the first fusion gene that is 10–78 nt away from the fusion breakpoint; the random primer has a length of 16–20 nt; the first fusion gene is ALK, ROS1, RET, BRAF, NRG1, FGFR, or NTRK. When the enrichment target is a splice isomer, the splice isomer is a splice isomer of the MET gene, and the enrichment primers include random primers and specific primers targeting the MET gene splice isomer. The specific primers have modification groups for purification, and the purification modification is used to separate the enriched product using random primers and specific primers as primer pairs. The specific primers have a target fragment distance of 10–78 nt from the breakpoint in the MET gene splice isoform; the random primers have a length of 16–20 nt.

9. The use of enrichment primers in the preparation of kits for enriching fusion genes or splice isoforms, characterized in that, When the enrichment target is a fusion gene, the enrichment primers include random primers and specific primers targeting the first fusion gene. The specific primers have modification groups for purification. The purification modification is used to separate the enrichment product using random primers and specific primers as primer pairs. The specific primer has a target fragment in the first fusion gene that is 10–78 nt away from the fusion breakpoint; the random primer has a length of 16–20 nt; the first fusion gene is ALK, ROS1, RET, BRAF, NRG1, FGFR, or NTRK. When the enrichment target is a splice isomer, the splice isomer is a splice isomer of the MET gene, and the enrichment primers include random primers and specific primers targeting the MET gene splice isomer. The specific primers have modification groups for purification, and the purification modification is used to separate the enriched product using random primers and specific primers as primer pairs. The specific primers have a target fragment distance of 10–78 nt from the breakpoint in the MET gene splice isoform; the random primers have a length of 16–20 nt.

10. The use according to claim 8 or 9, wherein, The breakpoint of the ALK gene is located at the junction of exon 19 and exon 20; The breakpoint of the ROS1 gene is located at the junction of exon 31 and exon 32, the junction of exon 32 and exon 33, or the junction of exon 33 and exon 34. The breakpoint of the RET gene is located at the junction of exon 11 and exon 12; The breakpoint of the NRG1 gene is located at the junction of exon 1 and exon 2; The breakpoint of the BRAF gene is located at the junction of exon 3 and exon 4, or at the junction of exon 10 and exon 11. The FGFR gene includes the FGFR1 gene located on chromosome 8, the FGFR2 gene located on chromosome 10, or the FGFR3 gene located on chromosome 4. The breakpoint is located at the junction of exon 9 and exon 10 of the FGFR1 gene, at the junction of exon 2 and exon 3 of the FGFR2 gene, or at the junction of exon 17 and exon 18 of the FGFR3 gene. The NTRK gene includes the NTRK1 gene located on chromosome 1, the NTRK2 gene located on chromosome 9, or the NTRK3 gene located on chromosome 15. The breakpoint is located at the junction of exon 9 and exon 10 of the NTRK1 gene, at the junction of exon 13 and exon 14 of the NTRK2 gene, at the junction of exon 14 and exon 15 of the NTRK3 gene, or at the junction of exon 19 and exon 20 of the NTRK3 gene. The breakpoint of the MET is located at the junction of exon 13 and exon 14.

11. The use according to claim 10, wherein, The nucleotide sequences of the specific primers for the ALK gene are shown in SEQ ID No: 1-8; The nucleotide sequences of the specific primers for the ROS1 gene are shown in SEQ ID No: 10-11; The nucleotide sequences of the specific primers for the RET gene are shown in SEQ ID No: 13-14; The nucleotide sequences of the specific primers for the BRAF gene are shown in SEQ ID No: 15-18; The nucleotide sequences of the specific primers for the NRG1 gene are shown in SEQ ID No: 20-21; The nucleotide sequences of the specific primers for the FGFR gene are shown in SEQ ID No: 22-27; The nucleotide sequences of the specific primers for the NTRK gene are shown in SEQ ID No: 28-29, 32, 34, 36 or 38; The nucleotide sequences of the specific primers for MET are shown in SEQ ID No: 52~53.

12. The use according to any one of claims 8-11, wherein, The modifying group includes biotin.

13. The use according to any one of claims 8-11, wherein, When the enrichment target is a fusion gene, the amplification reaction system further includes a blocker probe that blocks the amplification of the wild-type first fusion gene; when the enrichment target is a MET gene splice isoform, the amplification reaction system further includes a blocker probe that blocks the amplification of the wild-type MET gene.

14. An enrichment composition for fusion genes or splice isomers, comprising the enrichment primers of any one of claims 8-12 and a purification reagent for capturing modifying groups, said purification reagent comprising magnetic beads linked to modifying group binding molecules.

15. The enrichment composition according to claim 14, wherein, When the enrichment target is a fusion gene, the enrichment composition further includes a blocker probe that blocks the amplification of the wild-type first fusion gene; when the enrichment target is a MET gene splice isoform, the enrichment composition further includes a blocker probe that blocks the amplification of the wild-type MET gene.

16. Use of the enrichment composition of claim 14 or 15 in the preparation of an enrichment agent for fusion genes or splice isomers, wherein the fusion gene is ALK, ROS1, RET, BRAF, NRG1, FGFR, or NTRK; and the splice isomer is a splice isomer of the MET gene.

17. Use of the enrichment composition of claim 14 or 15 in the preparation of a kit for enriching fusion genes or splice isomers, wherein the fusion gene is ALK, ROS1, RET, BRAF, NRG1, FGFR, or NTRK; and the splice isomer is a splice isomer of the MET gene.

18. A fusion gene or splice isoform enrichment reagent, comprising the enrichment primers of any one of claims 8-13.

19. A fusion gene or splice isoform enrichment agent, comprising the enrichment composition of claim 14 or 15.

20. A kit for enriching fusion genes or splice isoforms, comprising the enrichment primers of any one of claims 8-13.

21. A kit for enriching fusion genes or splice isoforms, comprising the enrichment composition of claim 14 or 15.

22. The enrichment kit according to claim 20 or 21, wherein, It also includes sampling tools.

23. The enrichment kit according to claim 22, wherein, The samples collected by the sampling tool include tissue samples, tumor cells, and / or body fluids.

24. The enrichment kit according to claim 23, wherein, The bodily fluids include blood, saliva, urine, pleural effusion, or peritoneal effusion.

25. A method for detecting fusion genes or splice isoforms for non-disease diagnostic purposes, characterized in that, When the target of detection is a fusion gene, the method includes adding a detection primer pair to a sample containing the fusion gene to amplify the fusion gene, using a purification reagent to separate the fusion gene enrichment product from the amplification product after amplification, and then using a sequencing primer pair to sequence the fusion gene enrichment product. The detection primer pair includes a random primer and a specific primer targeting the first fusion gene. The random primer is connected to a first universal primer, and the specific primer is connected to a second universal primer with a modification group for purification. The purification modification is used to separate the enriched product composed of the random primer and the specific primer as primer pairs. The sequencing primer pair includes a first sequencing primer and a second sequencing primer. The first sequencing primer contains a sequence that hybridizes with a first universal primer and a first sequencing primer adapter sequence. The second sequencing primer contains a sequence that hybridizes with a second universal primer and a second sequencing primer adapter sequence. The specific primer has a target fragment in the first fusion gene that is 10–78 nt away from the fusion breakpoint; the random primer has a length of 16–20 nt; the first fusion gene is ALK, ROS1, RET, BRAF, NRG1, FGFR, or NTRK. When the target of detection is a splice isoform, the splice isoform is a splice isoform of the MET gene. The method includes adding detection primer pairs to a sample containing a splice isoform of the MET gene for amplification, using purification reagents to separate enriched products from the amplification products after amplification, and then using sequencing primer pairs to sequence the enriched products. The detection primer pair includes a random primer and a specific primer targeting the MET gene splice isoform. The random primer is connected to a first universal primer, and the specific primer is connected to a second universal primer with a modification group for purification. The purification modification is used to separate the enriched product composed of the random primer and the specific primer as primer pairs. The sequencing primer pair includes a first sequencing primer and a second sequencing primer. The first sequencing primer contains a sequence that hybridizes with a first universal primer and a first sequencing primer adapter sequence. The second sequencing primer contains a sequence that hybridizes with a second universal primer and a second sequencing primer adapter sequence. The specific primers have a target fragment distance of 10–78 nt from the breakpoint in the MET gene splice isoform; the random primers have a length of 16–20 nt.

26. The detection method according to claim 25, wherein, The modifying group includes biotin.

27. The detection method according to claim 25 or 26, wherein, When the target gene is a fusion gene, the amplification reaction system further includes a blocker probe that blocks the amplification of the wild-type first fusion gene; when the target gene is a MET gene splice isoform, the amplification reaction system further includes a blocker probe that blocks the amplification of the wild-type MET gene.

28. The detection method according to any one of claims 25-27, wherein, The purification reagent includes magnetic beads, which are linked to a modifying group-bound molecule.

29. The detection method according to any one of claims 25-28, wherein, The breakpoint of the ALK gene is located at the junction of exon 19 and exon 20; The breakpoint of the ROS1 gene is located at the junction of exon 31 and exon 32, the junction of exon 32 and exon 33, or the junction of exon 33 and exon 34. The breakpoint of the RET gene is located at the junction of exon 11 and exon 12; The breakpoint of the NRG1 gene is located at the junction of exon 1 and exon 2; The breakpoint of the BRAF gene is located at the junction of exon 3 and exon 4, or at the junction of exon 10 and exon 11. The FGFR gene includes the FGFR1 gene located on chromosome 8, the FGFR2 gene located on chromosome 10, or the FGFR3 gene located on chromosome 4. The breakpoint is located at the junction of exon 9 and exon 10 of the FGFR1 gene, at the junction of exon 2 and exon 3 of the FGFR2 gene, or at the junction of exon 17 and exon 18 of the FGFR3 gene. The NTRK gene includes the NTRK1 gene located on chromosome 1, the NTRK2 gene located on chromosome 9, or the NTRK3 gene located on chromosome 15. The breakpoint is located at the junction of exon 9 and exon 10 of the NTRK1 gene, at the junction of exon 13 and exon 14 of the NTRK2 gene, at the junction of exon 14 and exon 15 of the NTRK3 gene, or at the junction of exon 19 and exon 20 of the NTRK3 gene. The breakpoint of the MET is located at the junction of exon 13 and exon 14.

30. The detection method according to claim 29, wherein, The nucleotide sequences of the specific primers for the ALK gene are shown in SEQ ID No: 1-8; The nucleotide sequences of the specific primers for the ROS1 gene are shown in SEQ ID No: 10-11; The nucleotide sequences of the specific primers for the RET gene are shown in SEQ ID No: 13-14; The nucleotide sequences of the specific primers for the BRAF gene are shown in SEQ ID No: 15-18; The nucleotide sequences of the specific primers for the NRG1 gene are shown in SEQ ID No: 20-21; The nucleotide sequences of the specific primers for the FGFR gene are shown in SEQ ID No: 22-27; The nucleotide sequences of the specific primers for the NTRK gene are shown in SEQ ID No: 28-29, 32, 34, 36 or 38; The nucleotide sequences of the specific primers for MET are shown in SEQ ID No: 52~53.

31. The detection method according to any one of claims 25-30, wherein, The samples include tissue samples and / or body fluids.

32. The detection method according to claim 31, wherein, The bodily fluids include blood, saliva, urine, pleural effusion, or peritoneal effusion.

33. A detection composition for fusion genes or splice isoforms, comprising detection primers, purification reagents for capturing modifying groups, and sequencing primer pairs suitable for sequencing platforms; The sequencing primer pair includes a first sequencing primer and a second sequencing primer. The first sequencing primer contains a sequence that hybridizes with a first universal primer and a first sequencing primer adapter sequence. The second sequencing primer contains a sequence that hybridizes with a second universal primer and a second sequencing primer adapter sequence. The purification reagent includes magnetic beads, which are linked to a modifying group bound to a molecule. When the target gene to be detected is a fusion gene, the detection primers include random primers and specific primers targeting the first fusion gene. The random primers are connected to a first universal primer, and the specific primers are connected to a second universal primer with a modification group for purification. The purification modification is used to separate the enriched product using the random primers and specific primers as primer pairs. The target fragment of the specific primer in the first fusion gene is 10–78 nt away from the fusion breakpoint. The random primers are 16–20 nt in length. The first fusion gene is ALK, ROS1, RET, BRAF, NRG1, FGFR, or NTRK. When the target for detection is a splice isoform, the splice isoform is a splice isoform of the MET gene. The detection primers include random primers and specific primers targeting the splice isoform of the MET gene. The random primers are connected to a first universal primer, and the specific primers are connected to a second universal primer with a modification group for purification. The purification modification is used to separate the enriched product using the random primers and specific primers as primer pairs. The target fragment of the specific primer in the MET gene splice isoform is 10–78 nt from the breakpoint; the random primers are 16–20 nt in length.

34. The detection composition according to claim 33, wherein, The modifying group includes biotin.

35. The detection composition according to claim 33 or 34, wherein, When the target gene is a fusion gene, the detection composition further includes a blocker probe that blocks the amplification of the wild-type first fusion gene; when the target gene is a splice isoform, the detection composition further includes a blocker probe that blocks the amplification of the wild-type MET gene.

36. The detection composition according to any one of claims 33-35, wherein, The breakpoint of the ALK gene is located at the junction of exon 19 and exon 20; The breakpoint of the ROS1 gene is located at the junction of exon 31 and exon 32, the junction of exon 32 and exon 33, or the junction of exon 33 and exon 34. The breakpoint of the RET gene is located at the junction of exon 11 and exon 12; The breakpoint of the NRG1 gene is located at the junction of exon 1 and exon 2; The breakpoint of the BRAF gene is located at the junction of exon 3 and exon 4, or at the junction of exon 10 and exon 11. The FGFR gene includes the FGFR1 gene located on chromosome 8, the FGFR2 gene located on chromosome 10, or the FGFR3 gene located on chromosome 4. The breakpoint is located at the junction of exon 9 and exon 10 of the FGFR1 gene, at the junction of exon 2 and exon 3 of the FGFR2 gene, or at the junction of exon 17 and exon 18 of the FGFR3 gene. The NTRK gene includes the NTRK1 gene located on chromosome 1, the NTRK2 gene located on chromosome 9, or the NTRK3 gene located on chromosome 15. The breakpoint is located at the junction of exon 9 and exon 10 of the NTRK1 gene, at the junction of exon 13 and exon 14 of the NTRK2 gene, at the junction of exon 14 and exon 15 of the NTRK3 gene, or at the junction of exon 19 and exon 20 of the NTRK3 gene. The breakpoint of the MET is located at the junction of exon 13 and exon 14.

37. The detection composition according to claim 36, wherein, The nucleotide sequences of the specific primers for the ALK gene are shown in SEQ ID No: 1-8; The nucleotide sequences of the specific primers for the ROS1 gene are shown in SEQ ID No: 10-11; The nucleotide sequences of the specific primers for the RET gene are shown in SEQ ID No: 13-14; The nucleotide sequences of the specific primers for the BRAF gene are shown in SEQ ID No: 15-18; The nucleotide sequences of the specific primers for the NRG1 gene are shown in SEQ ID No: 20-21; The nucleotide sequences of the specific primers for the FGFR gene are shown in SEQ ID No: 22-27; The nucleotide sequences of the specific primers for the NTRK gene are shown in SEQ ID No: 28-29, 32, 34, 36 or 38; The nucleotide sequences of the specific primers for MET are shown in SEQ ID No: 52-53.

38. Use of the detection composition of any one of claims 33-37 in the preparation of a detection reagent for a fusion gene or splice isomer, wherein the fusion gene is ALK, ROS1, RET, BRAF, NRG1, FGFR or NTRK; and the splice isomer is a splice isomer of the MET gene.

39. Use of the detection composition of any one of claims 33-37 in the preparation of a detection kit for a fusion gene or splice isoform, wherein the fusion gene is ALK, ROS1, RET, BRAF, NRG1, FGFR or NTRK; and the splice isoform is a splice isoform of the MET gene.

40. A detection reagent for fusion genes or splice isoforms, comprising the detection composition of any one of claims 33-37.

41. A detection kit for fusion genes or splice isoforms, comprising the detection composition of any one of claims 33-37 or the detection reagent of claim 40.

42. The detection kit according to claim 41, wherein, It also includes sampling tools.

43. The detection kit according to claim 42, wherein, The samples collected by the sampling tool include tissue samples, tumor cells, and / or body fluids.

44. The test kit according to claim 43, wherein, The bodily fluids include blood, saliva, urine, pleural effusion, or peritoneal effusion.

Citation Information

Patent Citations

  • Targeted sequencing method and kit for detecting gene variation

    CN116745432A

  • Multiple digital PCR detection kit and detection method thereof

    CN117025765A