A targeted high-throughput sequencing method for detecting splicing isoforms

By constructing a sequencing library through reverse transcription and click chemistry reactions and adopting a targeted high-throughput sequencing method, the problem of difficulty in efficiently detecting and analyzing splicing isomers in existing technologies is solved, and accurate quantification and low-cost high-throughput sequencing are achieved.

CN118302538BActive Publication Date: 2025-09-19SHANGHAI INTRONCURE BIOTECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480000814.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-04-27
Filing Date
2024-02-20
Publication Date
2025-09-19
Estimated Expiration
2044-02-20

AI Technical Summary

Technical Problem

Existing technologies make it difficult to efficiently and accurately detect and analyze splicing isomers, especially unknown splicing isomers. In addition, existing methods are costly and complex to operate, making it difficult to meet the needs of real-time detection.

Method used

A targeted high-throughput sequencing method is used to reverse transcribe RNA using a reverse transcription primer. After adding 3'-modified dNTP, a click chemistry reaction is carried out with an oligonucleotide fragment modified with an alkyne group. Subsequently, PCR amplification is performed to construct a sequencing library, and finally high-throughput sequencing is performed.

Benefits of technology

It realizes the analysis of multiple targeted splicing isoforms, accurately quantifies splicing isoforms, has high detection sensitivity, can identify unknown splicing isoforms, is simple to operate, low cost, and is suitable for real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118302538B_ABST
    Figure CN118302538B_ABST
Patent Text Reader

Abstract

A targeted high-throughput sequencing method for detecting splice isoforms, including a method for establishing a sequencing library for high-throughput sequencing, comprising the following steps: 1) reverse transcription of sample RNA using a reverse transcription primer, adding common dNTPs and 3'-modified dNTPs to generate a first-strand cDNA; 2) ligating an oligonucleotide fragment with an alkyne modification at the 5' end to the cDNA fragment obtained in step 1) via a click chemistry reaction; 3) performing PCR amplification using the reaction product of step 2) as a template; and 4) obtaining a sequencing library. Targeted enrichment primers can be introduced during the reverse transcription or PCR amplification steps. The enrichment step reaction system contains multiple gene-specific primers, each designed based on the downstream exon segment of the alternative splicing event in the transcript, ensuring that the random-length fragments generated by reverse transcription cover the splice site.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to PCT international application No. PCT / CN2023 / 091367 filed on April 27, 2023, and this application cites the full text of the above-mentioned PCT international application. Technical Field

[0003] The present invention relates to the field of biotechnology, and in particular to a targeted high-throughput sequencing method for detecting splicing isoforms. Background Art

[0004] Alternative splicing (AS) refers to the process by which different mRNA splice isoforms are generated from a single mRNA precursor through different splicing modes (selecting different splice site combinations). Alternative splicing allows for a degree of flexibility, allowing the final mRNA structure to selectively include portions of exons and / or introns from the precursor RNA, resulting in variations in the protein's sequence composition. Therefore, alternative splicing is a key mechanism of proteome diversity.

[0005] During gene expression, gene transcription first synthesizes pre-mRNA, which contains all introns and exons as well as non-coding regions at both ends (5' and 3' Untranslated regions). RNA splicing is an essential step for pre-mRNA to mature into mRNA. Under the action of the spliceosome, introns in pre-mRNA are removed and all exons and UTR sequences are connected. When alternative splicing occurs, the same pre-mRNA can generate two or more splice isomers. Alternative splicing can be summarized into five forms, the two most common of which are intron retention and exon skipping. As follows Figure 1 As shown in the figure, intron retention is when the intron is not removed during the pre-mRNA splicing process and is still retained on the mRNA; exon skipping is when the exon is cut and removed by the spliceosome during the pre-mRNA splicing process, resulting in exon deletion in the subsequently generated mRNA.

[0006] Currently available commercial RNA library construction kits primarily utilize different fragmentation methods to construct total mRNA sequencing libraries. None target splicing isoforms, particularly those with unknown splicing isoforms. This is primarily due to the fact that classic RNA-seq involves random fragmentation of total RNA or messenger RNA (mRNA), followed by reverse transcription using random hexamer and / or oligo(dT) primers to form double-stranded DNA, followed by adapter ligation for library construction. These methods target only the entire transcriptome and cannot specifically analyze a small subset of interest. Furthermore, they are expensive. Because RNA-seq fragments are randomly located, typically not at splice junctions, extremely high sequencing depth is required to accurately detect different splicing isoforms, especially when considering the numerous potential alternative splicing sites to be validated. Consequently, the sequencing costs are extremely high. Currently available commercial targeted RNA sequencing kits are all designed and developed for known targets, and therefore cannot meet the requirements for sequencing unknown splicing isoforms. At the same time, the targeted RNA sequencing methods reported in the literature are difficult to apply to commercial and non-research applications. This is primarily due to the reliance on commercial companies such as Agilent Technologies and IDT to design and synthesize hybridization probes for target enrichment. This process is often time-consuming and requires continuous optimization, without the flexibility to modify (add or delete) existing probes. New potential targets discovered in real time cannot be validated promptly and efficiently, and the cost is substantial. Even with optimized target enrichment systems, library construction typically requires at least two days and involves complex hands-on procedures. Furthermore, a key requirement for splice isoform validation is to include as many isoforms as possible, but currently available methods may not be able to fully address this. Another major factor is that existing methods (such as those from Agilent) can only analyze known splice isoforms, necessitating primer design and the potential for competition between amplification primers for different splice isoforms, even in single droplets with highly diluted templates. The classic method for quantitative analysis of low-throughput splicing isoforms is qRT-PCR. However, different amplification efficiencies between different splicing isoforms can cause systematic deviations in the fluorescence signal, and even with tedious calibration steps, only relative quantification can be achieved. In addition, the detection of low-abundance splicing isoforms requires high standards for primer design, experimental operations, and equipment, making quantitative results difficult to replicate.

[0007] Therefore, it is particularly important and urgent to develop a more efficient and accurate RNA multiple targeted enrichment library construction method. Summary of the Invention

[0008] In order to solve the problems existing in the prior art, the purpose of the present disclosure is to provide a targeted high-throughput sequencing method for detecting splicing isoforms.

[0009] In order to achieve the above objectives, the present disclosure adopts the following specific technical solutions:

[0010] In one aspect, the present disclosure provides a method for establishing a sequencing library for high-throughput sequencing, comprising the following steps:

[0011] (1) Reverse transcription of the sample RNA is performed using a reverse transcription primer, and ordinary dNTPs and 3'-modified dNTPs are added to the reaction to obtain the first strand of cDNA, wherein the 3'-modified dNTPs are selected from one or more of the following: AzNTP, AmNTP, propargyl-NTP, and HalNTP;

[0012] (2) connecting the oligonucleotide fragment with an alkyne modification at the 5′ end to the cDNA fragment obtained in step (1) by a click chemistry reaction, wherein the oligonucleotide fragment comprises a random sequence and a complementary segment sequence of universal sequencing primer 1 (seq1);

[0013] (3) performing PCR amplification using the reaction product of step (2) as a template;

[0014] (4) Obtain a sequencing library.

[0015] In another aspect, the present disclosure provides a high-throughput sequencing library constructed according to the above method.

[0016] In another aspect, the present disclosure provides a targeted high-throughput sequencing method for detecting splicing isoforms, comprising the following steps:

[0017] (1) Extracting target cell RNA and constructing a sequencing library according to the above method;

[0018] (2) Based on the above sequencing library, high-throughput sequencing is performed to obtain sequencing information of the target gene in the cell sample.

[0019] In another aspect, the present disclosure provides a method for establishing a sequencing library for high-throughput sequencing, a use of the high-throughput sequencing library and / or the targeted high-throughput sequencing method in the assessment of off-target events;

[0020] Preferably, the off-target event assessment includes the following:

[0021] (1) Accurately determine the genomic location where off-target trans-splicing occurs in trans-splicing factors;

[0022] (2) Quantitative analysis of on-target trans-splicing and off-target trans-splicing.

[0023] The beneficial effects achieved by the present disclosure are at least as follows:

[0024] (1) Enable analysis of multiple targeted splicing isoforms. The throughput of traditional qRT-PCR methods is limited by the number of fluorophores in the detection instrument, bleed-through between fluorophores, and mutual interference between primers in the reaction system. Currently, the throughput of commonly used qRT-PCR is quadruple, while the HTAS (High-throughput Targeted Alternative Splicing) analysis platform disclosed in this paper can simultaneously analyze 5-100 (or more) pre-mRNA splicing events;

[0025] (2) Achieve accurate quantification of splicing isoforms: High-throughput sequencing performs quantitative analysis of splicing isoforms at the single-molecule level (digital signal), while traditional qRT-PCR quantification relies on fluorescence intensity (analog signal). Analog signals only support relative quantification, while digital signals can achieve absolute quantification of sample splicing isoforms. In addition, the unique advantage of HTAS is that the addition of a certain proportion of 3' modified dNTPs during the reverse transcription process can randomly terminate the extension of the first-chain cDNA. Therefore, the same type of splicing isoforms can generate cDNA products of different lengths, and the amplification efficiency is no longer affected by the different lengths of splicing isoform PCR products during library construction. Finally, the introduction of random sequences (N5; up to 12-16 nucleotides) in the design of the 5' linker can eliminate systematic bias (PCR skewing) in the PCR amplification process; mixed sample experiments have shown that the correlation r>0.99 between the actual detection value of HTAS and the theoretical value of the mixed sample;

[0026] (2) Detection sensitivity: The minimum intron-retained (IR) splicing isoform ratio detected in the examples is 0.1%. As the sequencing depth increases, the IR detection sensitivity is expected to reach 1 / 10. 5 or lower, its detection sensitivity is much higher than traditional RNA-seq, especially for low-abundance transcripts;

[0027] (4) Detection of unknown splicing isoforms: Due to the diversity of tissues and cells, a large number of splicing isoforms are still not annotated in the transcriptome in the current detection system. Traditional qRT-PCR protocols are mainly designed for known splicing isoforms, while the HTAS platform can detect both known and unknown splicing isoforms simultaneously, eliminating the interference of missing annotations in isoform analysis and improving the specificity of quantitative analysis;

[0028] (5) Achieve the evaluation of trans-splicing off-target events: Trans-splicing is a new technology for the targeted editing of mRNA, but off-target phenomena are one of the main bottlenecks in its clinical application. Since off-target events are expected to occur in any pre-mRNA in the transcriptome, there is currently no method to systematically evaluate off-target phenomena. HTAS only needs to know the sequence of the trans-splicing molecule (Pre-mRNA Trans-splicing Molecule) to simultaneously perform qualitative and quantitative analysis of on-target trans-splicing and off-target trans-splicing;

[0029] (6) Simple operation and low cost: No tedious library construction and enrichment steps are required, and the entire library construction time is approximately 4-5 hours. The cost per sample is significantly lower than existing products or methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 Schematic diagram of the principle of splicing isomerization.

[0031] Figure 2 Schematic diagram for high-throughput sequencing library construction (Scheme 1).

[0032] Figure 3 Schematic diagram for high-throughput sequencing library construction (Scheme 2).

[0033] Figure 4 Schematic diagram of the construction principle of high-throughput sequencing library (Scheme 3).

[0034] Figure 5 Schematic diagram of the correlation between the actual detection values ​​of splicing isoform IVT products with different mixing ratios and the theoretical values ​​of mixed samples. DETAILED DESCRIPTION

[0035] I. Definition

[0036] Unless otherwise indicated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, terms and laboratory procedures related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, and immunology used herein are those widely used in the respective fields and are standard procedures. To facilitate a better understanding of the present invention, definitions and explanations of relevant terms are provided below.

[0037] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as those commonly understood by those skilled in the art in the art. The terms used in the specification of this disclosure are only for the purpose of describing specific embodiments and are not intended to limit this disclosure.

[0038] The terms "including" and "having" and any variations thereof in this disclosure are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps is not limited to the listed steps or modules, but may optionally include steps not listed, or may optionally include other steps inherent to the process, method, product, or device.

[0039] The "plurality" mentioned in the present disclosure refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. At the same time, in order to better understand the present disclosure, the definitions and explanations of relevant terms are provided below. As used herein and unless otherwise specified, the term "about" or "approximately" means within plus or minus 10% of a given value or range. Where an integer is required, the term means within plus or minus 10% of a given value or range, rounded up or down to the nearest integer.

[0040] With respect to polypeptide sequences, the phrase "substantially identical" is understood to mean exhibiting at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference polypeptide sequence. With respect to nucleic acid sequences, the term is understood to mean nucleotide sequences that exhibit at least greater than 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference nucleic acid sequence.

[0041] The term "Oligo(dT) primer" used in this disclosure refers to a repetitive oligonucleotide sequence composed of approximately 12-25 polythymine Ts, which can specifically anneal to the ploy(A) tail of eukaryotic mRNA. Therefore, it is not suitable for RNA that lacks a ploy(A) tail structure, such as prokaryotic RNA or miRNA, nor is it suitable for degraded RNA, such as RNA in FFPE samples. Oligo(dT)n VN is composed of a fixed and specific sequence + (T)20 or so + degenerate base VN. The presence of a fixed and specific sequence is to facilitate the design of universal PCR downstream primers. The VN in the Oligo(dT)n VN sequence refers to the presence of an anchor base at the 3' end. ("V" represents dATP, dGTP or dCTP; "N" represents any one of dATP, dTTP, dGTP, and dCTP). The role of the anchor base is to be able to specifically bind to the 5' end of Poly(A) to prevent excessive T bases from being reverse transcribed.

[0042] The term "click chemistry" as used in this disclosure is well known in the art and generally refers to a fast reaction that is easy to purify and region-specific. Click chemistry is a class of reactions that allows a selected substrate to be linked to a specific molecule. Click chemistry is not a single specific reaction, but rather describes a way of producing products according to examples in nature, which also produces substances by connecting small modular units. In many applications, click reactions connect biomolecules and reporter molecules. Click chemistry is not limited to biological conditions: the concept of "click" reactions has been used in pharmacology and various biomimetic applications. However, the application is particularly useful in the detection, localization and identification of biomolecules. A typical click reaction is the classic click reaction is the copper-catalyzed reaction of azides and alkynes to form a 5-membered heteroatom ring: Cu(I)-catalyzed azide-alkyne cycloaddition (CuAAC).

[0043] As used in this disclosure, the terms "contacting," "adding," "reacting," "treating," and the like mean bringing one reactant, reagent, solvent, catalyst, reactive group, etc., into contact with another reactant, reagent, solvent, catalyst, reactive group, etc. The reactants, reagents, solvents, catalysts, reactive groups, etc., may be added individually, simultaneously, or separately and in any order that achieves the desired result. The reactants, reagents, solvents, catalysts, reactive groups, etc. may be added with or without heating or cooling equipment and may optionally be added under an inert atmosphere.

[0044] The term "complementary" as used in this disclosure refers to the broad concept of sequence complementarity between regions of two polynucleotide chains or between two nucleotides by base pairing. It is known that adenine nucleotides can form specific hydrogen bonds ("base pairing") with thymine or uracil nucleotides. Similarly, it is known that cytosine nucleotides can base pair with guanine nucleotides.

[0045] The term "library" used in the present disclosure is intended to mean a collection of nucleic acids with different chemical compositions (e.g., different sequences, different lengths, etc.) when used with respect to nucleic acids. Typically, the nucleic acids in a library will be different species having a common feature or characteristic of a certain genus or class, but are otherwise different to some extent. For example, a library can comprise nucleic acid species that are different in nucleotide sequence but similar in having a sugar-phosphate backbone. Libraries can be created using techniques known in the art. The nucleic acids exemplified herein can include nucleic acids obtained from any source, including, for example, a genome (e.g., a human genome) or a digestion of a genomic mixture. In another example, nucleic acids can be those obtained from a metagenomic study of a specific environment or ecosystem. The term also includes artificially created nucleic acid libraries, such as DNA libraries.

[0046] The terms "random primer" or "random hexamer primer" or "Random hexamer" or "Random hexamer primer" as used in this disclosure are well known in the art and generally refer to short oligodeoxyribonucleotides (d(N)6) of random sequence that anneal to random complementary sites on target DNA or RNA and serve as primers for DNA synthesis by DNA polymerase or reverse transcriptase.

[0047] The term "AzNTP (3'-azido-2',3'dNTP)" as used in the present disclosure is an azide deoxynucleotide, wherein the base is selected from adenine, guanine, cytosine and thymine.

[0048] The term "Add on PCR" as used in this disclosure means that in addition to sequences complementary to the template, the primers involved in PCR also have some other sequences. These other sequences do not participate in the current round of PCR reaction, but the generated PCR products will serve as templates for the next PCR reaction because of the additional sequences.

[0049] II. Detailed description of specific implementation methods

[0050] In some embodiments, the present disclosure provides a method for establishing a sequencing library for high-throughput sequencing, comprising the following steps:

[0051] (1) Reverse transcription of the sample RNA is performed using a reverse transcription primer, and ordinary dNTPs and 3'-modified dNTPs are added to the reaction to obtain the first strand of cDNA, wherein the 3'-modified dNTPs are selected from one or more of the following: AzNTP, AmNTP, propargyl-NTP, and HalNTP;

[0052] (2) connecting the oligonucleotide fragment with an alkyne modification at the 5′ end to the cDNA fragment obtained in step (1) by a click chemistry reaction, wherein the oligonucleotide fragment comprises a random sequence and a complementary segment sequence of universal sequencing primer 1 (seq1);

[0053] (3) performing PCR amplification using the reaction product of step (2) as a template;

[0054] (4) Obtain a sequencing library.

[0055] In some embodiments, the click chemistry reaction refers to a CuAAC click reaction, i.e., a copper ion-catalyzed azide-alkyne cycloaddition reaction.

[0056] In some embodiments, the reverse transcription PCR reaction in step (1) includes five reaction stages: the specific reaction conditions are 25°C for 10 min, 37°C for 10 min, 50°C for 45 min, 85°C for 2 min, and 12°C hold.

[0057] In some embodiments, the enzymes used in the reverse transcription PCR reaction in step (1) include HiScriptIII Reverse Transcriptase (R302-01, Nanjing Novozymes Biotechnology Co., Ltd.), SuperScript TM III Reverse Transcriptase (18080093, ThermoFisher SCIENTIFIC), HiFi II M-MLV(H-) Reverse Transcriptase (CW0743, Kangwei Century Biotechnology Co., Ltd.), Reverse Transcriptase [M-MLV, RNaseH-] (AE101-02, Beijing Quanshijin Biotechnology Co., Ltd.), MutiScript II Reverse Transcriptase (MD311, Feipeng Biotechnology Co., Ltd.).

[0058] In some embodiments, the reverse transcription primers described in step (1) are random primers or gene-specific primer group 1 designed based on the downstream exons of the retained introns of the targeted gene, and the 5' ends of the primers in the gene-specific primer group 1 carry the universal sequencing primer 2 sequence (seq2).

[0059] In some embodiments, the click chemistry reaction system in step (2) further includes vitamin C, a copper (II)-TBTA composition, and DMSO.

[0060] In some embodiments, the PCR amplification reaction in step (3) includes four reaction stages: the first PCR amplification reaction is 1 cycle, and the specific reaction conditions are 94°C reaction for 1 min, 60°C reaction for 30 s, and 68°C reaction for 10 min; the second PCR amplification reaction includes 12 cycles, and the specific reaction conditions are 94°C reaction for 30 s, 60°C reaction for 30 s, and 68°C reaction for 2 min; the third PCR amplification reaction includes 1 cycle, and the specific reaction conditions are 68°C reaction for 5 min; the fourth PCR amplification reaction includes 1 cycle, and the specific reaction conditions are 12°C.

[0061] In some embodiments, the PCR reaction system in step (3) further includes PCR buffer solutions such as MgCl2, DMSO, Tris-HCl, EDTA, NaCl, and KCl.

[0062] In some embodiments, the enzyme used in the PCR reaction in step (3) is Taq DNA polymerase.

[0063] In some embodiments, the PCR amplification primers used in step (3) include universal sequencing primer 1 and gene-specific primer group 2, wherein the 5' end of the primer in the gene-specific primer group 2 carries the universal sequencing primer 2 sequence.

[0064] In some embodiments, both gene-specific primer groups 1 and 2 are designed based on exons downstream of alternative splicing events of retained introns of specific targeted genes, wherein the targeting site of gene-specific primer 2 is shifted 5-100 bases upstream of the site of gene-specific primer 1.

[0065] In some embodiments, the targeting site of gene-specific primer 2 is shifted upstream by 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 bases compared to the site of gene-specific primer 1.

[0066] In some embodiments, the targeting site of gene-specific primer group 2 is shifted 20-50 bases upstream from the position of gene-specific primer group 1.

[0067] In some embodiments, the number of target genes targeted by the gene-specific primer group 1 and the gene-specific primer group 2 is greater than or equal to 1.

[0068] In some embodiments, both the gene-specific primer groups 1 and 2 are designed based on the exons downstream of the retained introns of the specific targeted gene.

[0069] In some embodiments, the target positions of the gene-specific primer groups 1 and 2 may partially overlap but not completely coincide.

[0070] In some embodiments, the molar concentration ratio of the common dNTPs and the 3'-modified dNTPs added in step (1) is 1:1-1:100, and the molar concentration ratio can be 1:100, 1:90, 1:80, 1:75, 1:70, 1:65, 1:60, 1:55, 1:50, 1:45, 1:40, 1:35, 1:30, 1:25, 1:20, 1:19, 1:18, 1:17, 1:16, 1:15, 1:14, 1:13, 1:12, 1:11 or 1:10; in a preferred embodiment, the molar concentration ratio of the common dNTPs and the 3'-modified dNTPs added in step (1) is 1:50.

[0071] In a preferred embodiment, the molar concentration ratio of the common dNTP and the 3'-modified dNTP added in step (1) is 1:20.

[0072] In some embodiments, the random sequence in step (2) comprises 4-16 nucleotides.

[0073] In some embodiments, the random sequence in step (2) comprises 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or 16 nucleotides.

[0074] In some embodiments, the random sequence in step (2) comprises 5 nucleotides.

[0075] In some embodiments, the PCR amplification in step (3) includes the step of performing PCR amplification on the reaction product of step (2) using a universal sequencing primer 1 and a gene-specific primer group 2, and then adding a linker structure to the amplified product, wherein the linker structure includes a P5 / P7 linker and a nucleic acid barcode, and the nucleic acid barcode is connected to the primer P5 and / or P7 linker end.

[0076] In some embodiments, the step (3) is: using the reaction product of step (2) as a template, adding P5 and P7 linkers to perform a PCR reaction.

[0077] In some embodiments, the P5 adapter primer is shown as SEQ ID NO:54.

[0078] In some embodiments, the P7 adapter primers are shown as SEQ ID NOs: 55-57.

[0079] In some embodiments, the nucleic acid barcode is attached to a single end of a primer P7 adapter.

[0080] In some embodiments, the nucleic acid barcode is attached to both ends of primer P5 and P7 adapters.

[0081] In some embodiments, the nucleic acid barcode is divided into nucleic acid barcode 5 and nucleic acid barcode 7.

[0082] In some embodiments, the nucleotide sequence of the nucleic acid barcode 5 is selected from the sequence shown in any one of SEQ ID NOs. 10-35.

[0083] In some embodiments, the nucleotide sequence of the nucleic acid barcode 7 is selected from the sequence shown in any one of SEQ ID NOs. 36-53.

[0084] In some embodiments, the PCR amplification reaction in step (3) includes four reaction stages: the first PCR amplification reaction is 1 cycle, and the specific reaction conditions are 94°C reaction for 30 seconds; the second PCR amplification reaction includes 18 cycles, and the specific reaction conditions are 94°C reaction for 30 seconds, 68°C reaction for 30 seconds, and 72°C reaction for 30 seconds; the third PCR amplification reaction includes 1 cycle, and the specific reaction conditions are 72°C reaction for 5 minutes; the fourth PCR amplification reaction includes 1 cycle, and the specific reaction conditions are 12°C.

[0085] In some embodiments, the PCR reaction system in step (3) further includes PCR buffer solutions such as MgCl2, DMSO, Tris-HCl, EDTA, NaCl, and KCl.

[0086] In some embodiments, the enzyme used in the PCR reaction in step (3) is Taq DNA polymerase.

[0087] In some embodiments, the universal sequencing primer 1 sequence is selected from any one of SEQ ID NO: 3, SEQ ID NO: 5 or SEQ ID NO: 6.

[0088] In some embodiments, the complementary segment sequence of the universal sequencing primer 1 is shown as SEQ ID NO:4.

[0089] In some embodiments, the universal sequencing primer 2 sequence (seq2) is selected from any one of SEQ ID NO:7 or SEQ ID NO:8.

[0090] In some embodiments, the present disclosure provides a high-throughput sequencing library constructed according to the above method.

[0091] In some embodiments, the present disclosure provides a targeted high-throughput sequencing method for detecting splicing isoforms, comprising the following steps:

[0092] (1) Extracting sample RNA and constructing a sequencing library according to the above method;

[0093] (2) Based on the above sequencing library, high-throughput sequencing is performed to obtain sequencing information of the target gene in the cell sample.

[0094] In some embodiments, the sample is a tissue, cell, and / or body fluid sample.

[0095] In some embodiments, the body fluid sample includes one or more of blood, saliva, urine, breast milk, cerebrospinal fluid, amniotic fluid, ascites, bile, and pleural effusion.

[0096] In some embodiments, the sample is a cell sample.

[0097] In some embodiments, the cell samples include, but are not limited to, MCF10A, MCF7, HeLa, HEK293T, and / or MDA-MB-231.

[0098] In some embodiments, the target genes include but are not limited to ATP13A1, CXXC1, ECHDC2, FGFRL1, HMGN3, KLHL17, NAXD, LZTR1, SELENBP1, JMJD8, PSMB1, HIGD2A, HNRNPAB, SMARCC1, ATP5IF1, HIGD2B, RPS21, UQCC5, NFATC3, PCNP and / or OSGEP.

[0099] In some embodiments, the present disclosure provides a method for establishing a sequencing library for high-throughput sequencing, the use of the high-throughput sequencing library and / or the targeted high-throughput sequencing method in the assessment of off-target events.

[0100] In some embodiments, the off-target event assessment comprises the following:

[0101] (1) Accurately determine the genomic location where off-target trans-splicing occurs in trans-splicing factors;

[0102] (2) Quantitative analysis of on-target trans-splicing and off-target trans-splicing.

[0103] In some embodiments, the high-throughput sequencing method can be used to detect any form of alternative splicing and any combination of transcripts in the transcriptome.

[0104] In some embodiments, the construction process of the sequencing library for high-throughput sequencing is as follows, and its principle diagram is shown in the attached Figure 2 :

[0105] 1) Reverse transcription: The first strand of cDNA is synthesized by reverse transcription using RNA as a template. A random hexamer primer is used as a reverse transcription primer, and ordinary dNTPs and a certain concentration of 3' modified dNTPs are added for reverse transcription. The 3' modified dNTP is AzNTP (3'-azido-2', 3'dNTP), which can prevent the addition of the next dNTP during chain extension. At the same time, the 3' modification group should be available for chemical connection to the downstream linker. Random Hexamer is a single-stranded DNA with a random sequence of 6 bases. The characteristics of the random sequence can help it randomly bind to different fragments of RNA. During the reverse transcription process, under the action of reverse transcriptase, dNTPs are added to the 3' end of the primer using RNA as a template to synthesize single-stranded cDNA. When the 3' modified dNTP replaces dNTPs and is added to the cDNA single strand, chain extension is terminated, and the synthesis of the first strand of cDNA is completed.

[0106] 2) Click chemistry ligation: The cDNA obtained above is subjected to a click chemistry reaction (i.e., CuAAC click reaction—monovalent copper ion-catalyzed azide-alkyne cycloaddition reaction) with an oligonucleotide modified with an alkyne group at the 5' end. The oligonucleotide adapter sequence contains a complementary segment sequence of 5'N5+NGS universal sequencing primer (seq1), which serves as a PCR primer binding site in subsequent sequencing library construction. N5 is a sequence composed of 5 random bases, which is used to ensure the molecular diversity of the sequencing library and to remove systematic biases and sequencing errors during PCR amplification in the result analysis. The length of the random base sequence can be adjusted to N4, N5, N6, N7, N8, N9, N10, N11, N12, N13, N14, N15, and N16. Wherein N is selected from any base in ATCG.

[0107] 3) Targeted enrichment: Using the click chemistry reaction product as a template, the target gene fragment is targeted and enriched through PCR reaction. The primers in the PCR system include a universal sequencing primer 1 (seq1, whose 3' end sequence is complementary to the 3' end of the above-mentioned alkyne primer) and a gene-specific primer group. The primer group is designed based on the exons downstream of the retained intron of each targeted gene, and a universal sequencing primer 2 sequence (seq2) is added to the 5' end of the primer. The CLICK product contains an azide-alkyne linker, so the DNA polymerase needs to use Taq or other polymerases that can amplify this template.

[0108] 4) Sequencing Adapter Ligation and Enrichment: Using the enriched fragments as templates, a PCR reaction is performed to add intact P5 and P7 adapters to the enriched fragments for subsequent NGS sequencing. A nucleic acid barcode (sample barcode) is also added. Nucleic acid barcodes facilitate sample mixing during high-throughput sequencing, increasing detection throughput and reducing instrument error during sample sequencing.

[0109] In some embodiments, the reverse transcription primer used in the high-throughput sequencing library construction process is a gene-specific primer group 1 designed based on the downstream exons of the retained introns of the targeted gene, and a universal sequencing primer sequence (seq2) is added to the 5' end of the specific primer; in some embodiments, in the actual construction process, sequencing adapter ligation and enrichment can be performed directly after the click chemistry ligation product is purified; in some embodiments, the construction process of the sequencing library for high-throughput sequencing is as follows, and its principle diagram is shown in the attached figure. Figure 3 :

[0110] 1) Reverse transcription: RNA is used as a template to synthesize the first strand of cDNA by reverse transcription. A gene-specific primer group 1 designed based on the downstream exons of the retained introns of the targeted gene is used as a reverse transcription primer, and a universal sequencing primer sequence (seq2) is added to the 5' end. Ordinary dNTPs and a certain concentration of 3' modified dNTPs are added to carry out the reverse transcription reaction, wherein the 3' modified dNTP is AzNTP (3'-azido-2', 3'dNTP), which can prevent the addition of the next dNTP from binding during chain extension, and the 3' modified group should be available for chemical connection to the downstream linker. During the reverse transcription process, under the action of reverse transcriptase, dNTPs are added to the 3' end of the primer using RNA as a template to synthesize single-stranded cDNA. When the 3' modified dNTP replaces dNTPs and is added to the cDNA single strand, chain extension is terminated, and the synthesis of the first strand of cDNA is completed.

[0111] 2) Click chemistry ligation: The cDNA obtained above is subjected to a click chemistry reaction (i.e., CuAAC click reaction—monovalent copper ion-catalyzed azide-alkyne cycloaddition reaction) with an oligonucleotide modified with an alkyne group at the 5' end. The oligonucleotide adapter sequence contains a complementary segment sequence of 5'N5+NGS universal sequencing primer (seq1), which serves as a PCR primer binding site in subsequent sequencing library construction. N5 is a sequence composed of 5 random bases, which is used to ensure the molecular diversity of the sequencing library and to remove systematic biases and sequencing errors during PCR amplification in the result analysis. The length of the random base sequence can be adjusted to N4, N5, N6, N7, N8, N9, N10, N11, N12, N13, N14, N15, and N16.

[0112] 3) Sequencing adapter connection and enrichment: Using the click chemistry reaction product as a template, add complete P5 and P7 adapters for subsequent PCR amplification. The P5 adapter can complement the Seq 1 complementary sequence in the click chemistry reaction product, and the P7 adapter can complement the Seq 2 complementary sequence generated during the PCR reaction. And add a nucleic acid barcode (samplebarcode). The nucleic acid barcode can achieve sample mixing during high-throughput sequencing, improve detection throughput, and reduce instrument errors during sample sequencing. After several rounds of PCR reactions, the P5 and P7 adapters are connected and the fragments to be sequenced are enriched for subsequent NGS sequencing. The CLICK product contains an azide-alkyne linker, so the DNA polymerase needs to use Taq or other polymerases that can amplify this template.

[0113] In some embodiments, the reverse transcription primer used in the high-throughput sequencing library construction process is a gene-specific primer group 1 designed based on the downstream exons of the retained introns of the targeted gene; in some embodiments, a specific gene PCR is performed after the click chemistry ligation reaction to remove the influence of ribosomal RNA in the template on the sequencing data; in some embodiments, the construction process of the sequencing library for high-throughput sequencing is as follows, and its principle diagram is shown in the attached figure. Figure 4 :

[0114] 1) Reverse transcription: RNA is used as a template to synthesize the first strand of cDNA by reverse transcription. A gene-specific primer group 1 designed based on the downstream exons of the retained introns of the targeted gene is used as a reverse transcription primer, and a universal sequencing primer sequence (seq2) is added to the 5' end. Ordinary dNTPs and a certain concentration of 3' modified dNTPs are added to carry out the reverse transcription reaction, wherein the 3' modified dNTP is AzNTP (3'-azido-2', 3'dNTP), which can prevent the addition of the next dNTP from binding during chain extension, and the 3' modified group should be available for chemical connection to the downstream linker. During the reverse transcription process, under the action of reverse transcriptase, dNTPs are added to the 3' end of the primer using RNA as a template to synthesize single-stranded cDNA. When the 3' modified dNTP replaces dNTPs and is added to the cDNA single strand, chain extension is terminated, and the synthesis of the first strand of cDNA is completed.

[0115] 2) Click chemistry ligation: The cDNA obtained above is subjected to a click chemistry reaction (i.e., CuAAC click reaction—monovalent copper ion-catalyzed azide-alkyne cycloaddition reaction) with an oligonucleotide modified with an alkyne group at the 5' end. The oligonucleotide adapter sequence contains a complementary segment sequence of 5'N5+NGS universal sequencing primer (seq1), which serves as a PCR primer binding site in subsequent sequencing library construction. N5 is a sequence composed of 5 random bases, which is used to ensure the molecular diversity of the sequencing library and to remove systematic biases and sequencing errors during PCR amplification in the result analysis. The length of the random base sequence can be adjusted to N4, N5, N6, N7, N8, N9, N10, N11, N12, N13, N14, N15, and N16.

[0116] 3) Targeted enrichment: Using the click chemistry reaction product as a template, the target gene fragment is targeted and enriched through PCR reaction. The primers in the PCR system include a universal sequencing primer 1 (seq1, whose 3' end sequence is complementary to the 3' end of the above-mentioned alkyne primer) and a gene-specific primer group 2. Primer group 2 is designed based on the exon downstream of the retained intron of each targeted gene, but should be shifted 5-100 bases upstream than the position of the targeted reverse transcription primer group 1, and can even partially overlap, but not be completely consistent, and a universal sequencing primer 2 sequence (seq2) is added to the 5' end of the primer. The CLICK product contains an azide-alkyne linker, so the DNA polymerase needs to use Taq or other polymerases that can amplify this template.

[0117] 4) Sequencing Adapter Ligation and Enrichment: Using the enriched fragments as templates, a PCR reaction is performed to add intact P5 and P7 adapters to the enriched fragments for subsequent NGS sequencing. A nucleic acid barcode (sample barcode) is also added. Nucleic acid barcodes facilitate sample mixing during high-throughput sequencing, increasing detection throughput and reducing instrument error during sample sequencing.

[0118] In some embodiments, the fragments in the constructed high-throughput sequencing library all include the following elements: P5 sequence, sample tag 5, universal sequencing primer 1 (seq1), inserted DNA fragment, universal sequencing primer 2 (seq2), sample tag 7, and P7 adapter.

[0119] In some embodiments, the nucleotide sequences of the P5 sequence and P7 are shown as SEQ ID NO: 1 (AATGATACGGCGACCACCGAGATCTACAC) and SEQ ID NO: 2 (CAAGCAGAAGACGGCATACGAGA T), respectively.

[0120] In some embodiments, the universal sequencing primer includes two types, PE adapter and Nextera adapter.

[0121] In some embodiments, the universal sequencing primers 1 and 2 are selected from any combination of the sequences in Table 1 below.

[0122] Table 1 Sequence information of universal sequencing primers

[0123]

[0124] In some embodiments, the nucleotide sequence of the oligonucleotide Hex_N5_Seq1rc is as shown in SEQ ID NO:9 (NNNNNAGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT).

[0125] In some embodiments, the sample tag 5 and the sample tag 7 are selected from any combination of the sequences in Table 2 below.

[0126] Table 2 Sequence information

[0127]

[0128]

[0129] In some embodiments, the sequence information of the P5 and P7 linkers of the complete sequence is shown in Table 3.

[0130] Table 3 Sequence information

[0131]

[0132] Example

[0133] The technical solutions of the present disclosure are further described below by way of specific implementation. It should be understood by those skilled in the art that the embodiments are only for the purpose of helping to understand the present disclosure and should not be regarded as specific limitations of the present disclosure.

[0134] Example 1: Construction of a targeted high-throughput sequencing platform for detecting splicing isoforms

[0135] 1. RNA Extraction

[0136] Extract RNA from target cells using a commercial kit, such as the FineProtect Universal RNA Extraction Kit from Jifan Biotechnology (Beijing) Co., Ltd. Ensure the quality of RNA extraction during the extraction process to facilitate subsequent experiments.

[0137] 2. RNA Quantification

[0138] RNA is quantified using commercial kits and the specific detection of RNA concentration using fluorescent dye technology, such as Invitrogen's Qubit TM RNA High Sensitivity (HS) Quantitation Kit.

[0139] 3. Reverse transcription

[0140] Based on the quantitative concentration determined above, RNA of the same mass was selected as the template for synthesizing the first strand of cDNA, and a random hexamer primer was used as the reverse transcription primer. Common dNTPs and 3'-modified dNTPs at a 1 / 20 molar ratio were added to the reverse transcription system for reverse transcription reaction.

[0141] The components of the reverse transcription reaction system are shown in Table 4, and the reverse transcription reaction conditions are shown in Table 5.

[0142] Table 4: Transcription reaction system (20 μL / reaction)

[0143]

[0144] Table 5: Reverse transcription reaction conditions

[0145] temperature time 25℃ 10min 37℃ 10min 50℃ 45min 85℃ 2min 12℃ Hold

[0146] 4. cDNA Purification

[0147] After first-strand cDNA synthesis, purify the product using a commercial kit, such as the Zymo DNA Clean & Concentrator-5 kit. Elute the purified product with purified water to a volume of 10 μL.

[0148] 5. Click ligation

[0149] The purified cDNA and oligonucleotides containing alkyne modifications at the 5' end were subjected to a click reaction (i.e., CuAAC click reaction - copper ion-catalyzed azide-alkyne cycloaddition reaction) at room temperature for 1 hour in the presence of a copper(II)-TBTA complex catalyst.

[0150] The components of the click chemistry reaction system are shown in Table 6, and the nucleic acid sequence of the oligonucleotides is shown in SEQ ID NO: 9.

[0151] Table 6: Click chemistry reaction system (37.8 μL / reaction)

[0152]

[0153] 6. Click ligation product purification

[0154] Click chemistry reaction products were purified using a commercial kit, such as the Zymo DNA Clean & Concentrator-5 kit. The purified product was eluted with pure water to a volume of 10 μL.

[0155] 7. Targeted Enrichment

[0156] The purified product was used as a template, and primers designed for the downstream exons of each retained intron of the targeted gene were used as the targeted gene-specific primer combination and a universal sequencing primer (PE1_p26) for PCR reaction.

[0157] The components of the PCR reaction system are shown in Table 7, the preparation of the primer combination mixture is shown in Table 8, the nucleic acid sequence of the PE1_p26 primer is shown in SEQ ID NO: 6, and the qPCR reaction conditions are shown in Table 9.

[0158] Table 7: PCR reaction system (25 μL / reaction)

[0159]

[0160] Table 8: Primer combination mixture components

[0161] Components Addition volume (μL) Each targeted gene primer (100 μM each) 10 each purified water Replenish to 100

[0162] Table 9: PCR reaction conditions

[0163]

[0164] 8. Purification of targeted enrichment products

[0165] Purify the targeted enrichment PCR product using a commercial kit, such as the Zymo DNA Clean & Concentrator-5 kit. Elute the purified product with purified water to a volume of 10 μL.

[0166] 9. Sequencing adapter ligation and enrichment

[0167] Using the target-enriched fragments obtained above as templates and intact P5 and P7 adapters as primers, a PCR reaction is performed for subsequent NGS sequencing. The P5 adapter complements the sequence containing the complementary segment of Seq 1 in the click reaction product, and the P7 adapter complements the sequence containing the complementary segment of Seq 2 generated during the PCR reaction. Including a nucleic acid barcode (sample barcode) in a single or dual Add-on PCR primer allows for sample mixing during high-throughput sequencing.

[0168] The components of the PCR reaction system are shown in Table 10, the nucleic acid sequences of the complete P5 and P7 adapter primers are shown in Table 3, and the qPCR reaction conditions are shown in Table 11.

[0169] Table 10: PCR reaction system (25 μL / reaction)

[0170]

[0171] Table 11: PCR reaction conditions

[0172]

[0173] 10. High-throughput sequencing and result analysis

[0174] Based on the sequencing library amplified in the above steps, high-throughput sequencing is performed to obtain sequencing information of the target gene in the cell sample, including but not limited to analysis of multiple targeted splicing isoforms, precise quantification of splicing isoforms, and assessment of unknown splicing isoforms and trans-splicing off-target events.

[0175] Example 2: Detection of splicing isoforms in MCF10A, MCF7, and MDA-MB-231 cell lines

[0176] The high-throughput sequencing platform described in Example 1 was used to detect splicing isoforms in MCF10A, MCF7, and MDA-MB-231 cell lines.

[0177] Wherein, steps 1-6 are the same as those in Example 1.

[0178] During the targeted enrichment process in step 7, PCR reactions were performed using primers designed for the downstream exons of the retained introns of the ATP13A1, CXXC1, ECHDC2, FGFRL1, HMGN3, KLHL17, and OSGEP genes as a targeted enrichment primer combination and the universal sequencing primer PE1_p26. The composition of the primer combination mixture is shown in Table 12, and the primer sequences for each targeted gene are shown in Table 13.

[0179] Table 12: Primer combination mixture components

[0180] Components Addition volume (μL) ATPE7R_PE2 (100 μM) 10 KLHE11R_PE2 (100 μM) 10 CXXE12R_PE2 (100 μM) 10 ECHE6R_PE2 (100 μM) 10 FGFE8R_PE2 (100 μM) 10 HMGE6R_PE2 (100 μM) 10 OSGE5R_PE2 (100 μM) 10 purified water Replenish to 100

[0181] Table 13: Nucleic acid sequences of primers

[0182]

[0183]

[0184] Steps 8-10 are the same as those in Example 1.

[0185] High-throughput sequencing results showed that both normal spliceosomes and splice isoforms (intron retention) were observed for the target genes in MCF10A, MCF7, and MDA-MB-231 cell lines. The ratio of splice isoforms in the targeted genes was successfully quantitatively detected. Specific test results are shown in Table 14. These data demonstrate that the targeted high-throughput sequencing platform constructed in this disclosure can effectively detect splice isoforms.

[0186] Table 14: Detection results of target genes in MCF10A, MCF7, and MDA-MB-231 cell lines

[0187]

[0188] Example 3: Detection of splice isoforms at different ratios using in vitro transcribed RNA (IVT) simulation

[0189] In vitro RNA synthesis samples (IVT products) were prepared, including two splice isoform sequences: the normal splice isoform and the splice isoform (intron-retained). The IVT products were mixed with simulated splice isoforms at different ratios, and detected using the targeted high-throughput sequencing platform constructed in this disclosure. The correlation between the actual detection value of the mixed sample and the theoretical value of the mixed sample was calculated. The specific experimental details are as follows:

[0190] Step 1: Obtain target RNA synthesis sample by in vitro transcription

[0191] The synthesized normal splice form and splice isoform (intron retention) DNA were subjected to in vitro transcription reaction using commercial kits, such as Novagen T7 High Yield RNA Transcription Kit (TR101).

[0192] The in vitro transcription reaction system is shown in Table 15.

[0193] Table 15: In vitro transcription system (20 μL system)

[0194]

[0195] After incubating the above system at 37°C for 2 hours, 1 μL of DNase I was added and incubated at 37°C for another 15 minutes to remove the DNA template.

[0196] Step 2: RNA recovery and purification

[0197] In vitro transcribed RNA was recovered using phenol-ethanol precipitation. The final precipitated RNA product was dissolved in double-distilled water, and the RNA concentration was determined using a Nano-Drop. The RNA was then diluted to 2 μM, 0.2 μM, 0.02 μM, and 0.002 μM for later use.

[0198] Step 3: Mixing IVT products of different splicing isoforms

[0199] The IVT products were mixed in a 1:1 volume ratio as shown in Table 16.

[0200] Table 16: IVT Product Mix

[0201]

[0202] Step 4: Reverse transcription

[0203] The mixed IVT products were diluted 100-fold and used as templates for synthesizing the first strand of cDNA. The reverse primer on the downstream exon of the target intron was used as the reverse transcription primer. Ordinary dNTPs and 3'-modified dNTPs at a 1 / 15 molar concentration were added to the reverse transcription system for reverse transcription reaction.

[0204] The components of the reverse transcription reaction system are shown in Table 17, and the reverse transcription reaction conditions are shown in Table 18.

[0205] Table 17: Transcription reaction system (20 μL / reaction)

[0206]

[0207]

[0208] Table 18: Reverse transcription reaction conditions

[0209] temperature time 25℃ 10min 37℃ 10min 50℃ 45min 85℃ 2min 12℃ Hold

[0210] Among them, steps 5-6 are performed the same as steps 4-6 in Example 1.

[0211] Step 8: Targeted Enrichment

[0212] The purified product was used as a template, and a PCR reaction was performed using primers designed for the downstream exons of the target intron as a target gene-specific primer combination and a universal sequencing primer (PE1_p26).

[0213] The PCR reaction system components are shown in Table 19, the nucleic acid sequence of the PE1_p26 primer is shown in SEQ ID NO: 6, and the qPCR reaction conditions are shown in Table 20. The primer sequences are shown in Table 21.

[0214] Table 19: PCR reaction system (25 μL / reaction)

[0215]

[0216] Table 20: PCR reaction conditions

[0217]

[0218]

[0219] Table 21: Nucleic acid sequences of primers

[0220]

[0221] Among them, steps 9-10 are performed the same as steps 8-9 in Example 1.

[0222] The high-throughput sequencing results showed that the correlation between the actual detection value of splicing isoform IVT products with different mixing ratios and the theoretical value of mixed samples was r>0.99. The specific test results are shown in Table 22 and Figure 5 .

[0223] Table 22: Detection results of RNA simulations with different ratios of splice isoforms

[0224]

[0225] Example 4: Detection of splicing isoforms in HeLa cell lines

[0226] The high-throughput sequencing platform described in Example 1 was used to detect splicing isoforms in the HeLa cell line.

[0227] Wherein, steps 1-2 are the same as those in Example 1.

[0228] During the targeted enrichment reverse transcription process in step 3, gene-specific primers are used as reverse transcription primers, and common dNTPs and 3'-modified dNTPs at a 1 / 15 molar concentration ratio are added to the reverse transcription system to perform a reverse transcription reaction.

[0229] The gene-specific primer sequences are shown in Table 23, the gene-specific primer combination mixture preparation is shown in Table 24, the reverse transcription reaction system components are shown in Table 25, and the reverse transcription reaction conditions are shown in Table 26.

[0230] Table 23: Reverse transcription reaction gene-specific primer sequences

[0231]

[0232]

[0233] Table 24: Gene-specific primer mix preparation

[0234] Components Addition amount Single primer (200 μM) 1 μL Double distilled water Make up to 100 μL

[0235] Table 25: Transcription reaction system (20 μL / reaction)

[0236]

[0237] Table 26: Reverse transcription reaction conditions

[0238] temperature time 25℃ 10min 37℃ 10min 50℃ 45min 85℃ 2min 12℃ Hold

[0239] Wherein, steps 4-6 are the same as those in Example 1.

[0240] During the targeted enrichment process of step 7, primers designed to retain the downstream exons of the introns of ATP13A1 gene, CXXC1 gene, ECHDC2 gene, FGFRL1 gene, HMGN3 gene, KLHL17 gene, OSGEP gene, NAXD gene, LZTR1 gene, SELENBP1 gene, and JMJD8 gene were used as a targeted enrichment primer combination and a universal sequencing primer PE1_p26 for PCR reaction. Compared to the gene-specific primers in reverse transcription, the primers for targeted enrichment PCR are closer to the introns by about 20-50bp. The primer sequences of the targeted genes are shown in Table 27, the primer mixture preparation is shown in Table 28, the PCR reaction system components are shown in Table 29, the nucleic acid sequence of the PE1_p26 primer is shown in SEQ ID NO: 6, and the PCR reaction conditions are shown in Table 30.

[0241] Table 27: Primer combination mixture preparation

[0242]

[0243] Table 28: Primer mix preparation

[0244] Components Addition amount Single primer (200 μM) 1 μL Double distilled water Make up to 20 μL

[0245] Table 29: PCR reaction system (25 μL / reaction)

[0246]

[0247] Table 30: PCR reaction conditions

[0248]

[0249] Steps 8-10 are the same as those in Example 1.

[0250] The final product was subjected to high-throughput sequencing. The results showed that both normal spliceosomes and splice isoforms (intron retention) were observed in the target gene of the HeLa cell line. The ratio of splice isoforms in the targeted gene was successfully quantitatively detected. The specific test results are shown in Table 31. The above data show that the targeted high-throughput sequencing platform constructed in the present disclosure can effectively detect splice isoforms.

[0251] Table 31: Detection results of target genes in HeLa cell lines

[0252]

[0253]

[0254] Example 5: Assessment of trans-splicing off-target events

[0255] The high-throughput sequencing platform described in Example 1 was used to detect on-target trans-splicing and off-target trans-splicing events that occurred after HEK293T cells were transfected with a mini-gene and a trans-splicing factor (Pre-mRNA Trans-splicing Molecule).

[0256] Wherein, steps 1-2 are the same as those in Example 1.

[0257] During the targeted enrichment reverse transcription process in step 3, the trans-splicing molecule and the mini-gene specific primer are used as reverse transcription primers, and ordinary dNTPs and 3'-modified dNTPs at a 1 / 15 molar concentration ratio are added to the reverse transcription system for reverse transcription reaction.

[0258] The gene-specific primer sequences are shown in Table 32, the gene-specific primer combination mixture preparation is shown in Table 33, the reverse transcription reaction system components are shown in Table 34, and the reverse transcription reaction conditions are shown in Table 35.

[0259] Table 32: Reverse transcription reaction gene-specific primer sequences

[0260] name Primer sequence 5'→3' Contained nucleic acid sequence number HTAS_SP_ts_R1 GTTGCCATCCTCCTTGAAAT SEQ ID NO: 93 HTAS_SP_cis_R1 CCACCTAGTGCTTTAGCAATAT SEQ ID NO: 94

[0261] Table 33: Gene-specific primer mix preparation

[0262] Components Addition amount Single primer (200 μM) 1 μL Double distilled water Make up to 100 μL

[0263] Table 34: Transcription reaction system (20 μL / reaction)

[0264]

[0265] Table 35: Reverse transcription reaction conditions

[0266] temperature time 25℃ 10min 37℃ 10min 50℃ 45min 85℃ 2min 12℃ Hold

[0267] Wherein, steps 4-6 are the same as those in Example 1.

[0268] During the targeted enrichment process in step 7, a PCR reaction was performed using a trans-splicing molecule and a mini-gene-specific primer as a targeted enrichment primer combination and a universal sequencing primer PE1_p26. Compared to the specific primers used in reverse transcription, the primers for the trans-splicing molecule and mini-gene targeted enrichment PCR were approximately 20-50 bp closer to the 3' splicing site and the 5' splicing site, respectively. The primer sequences for each targeted gene are shown in Table 36, the primer mix preparation is shown in Table 37, the PCR reaction system components are shown in Table 38, the nucleic acid sequence of the PE1_p26 primer is shown in SEQ ID NO: 6, and the PCR reaction conditions are shown in Table 39.

[0269] Table 36: Primer combination mixture preparation

[0270]

[0271] Table 37: Primer mix preparation

[0272] Components Addition amount Single primer (200 μM) 1 μL Double distilled water Make up to 20 μL

[0273] Table 38: PCR reaction system (25 μL / reaction)

[0274]

[0275] Table 39: PCR reaction conditions

[0276]

[0277] Steps 8-10 are the same as those in Example 1.

[0278] The final product was subjected to high-throughput sequencing, and the results of high-throughput sequencing showed that after HEK293T cells were transfected with mini-genes and trans-splicing factors (Pre-mRNA Trans-splicing Molecule), target trans-splicing products and off-target trans-splicing products were visible. At the same time, the genomic location where off-target trans-splicing of trans-splicing factors occurred was accurately determined (there were many events, not all of which were listed). The specific test results are shown in Table 40. The above data show that the targeted high-throughput sequencing platform constructed in the present disclosure can effectively realize the evaluation of trans-splicing off-target events, and simultaneously perform qualitative and quantitative analysis of target trans-splicing and off-target trans-splicing.

[0279] Table 40: Detection results of trans-splicing and off-target trans-splicing in HEK293T cells

[0280]

[0281] The above-described embodiments merely represent several implementation methods of the present disclosure. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person of ordinary skill in the art could make various modifications and improvements without departing from the spirit of the present disclosure, all of which fall within the scope of protection of the present disclosure. Therefore, the scope of protection of the present patent shall be determined by the appended claims.

Claims

1. A method for establishing a sequencing library for high-throughput sequencing, comprising the following steps: (1) Reverse transcription of the sample RNA is performed using a reverse transcription primer, and ordinary dNTPs and 3'-modified dNTPs are added to the reaction to obtain the first strand of cDNA, wherein the 3'-modified dNTPs are selected from one or more of the following: AzNTP, AmNTP, propargyl-NTP, and HalNTP; (2) connecting the oligonucleotide fragment with an alkyne modification at the 5' end to the cDNA fragment obtained in step (1) by a click chemistry reaction, wherein the oligonucleotide fragment contains a complementary segment sequence of universal sequencing primer 1; (3) Using the reaction product of step (2) as a template, universal sequencing primer 1 and gene-specific primer group 2 are used to perform PCR amplification and targeted enrichment; (4) Using the enriched fragments from step (3) as templates, add complete P5 / P7 adapters and nucleic acid barcodes to obtain a sequencing library; The reverse transcription primers in step (1) are gene-specific primer group 1 designed based on the downstream exons of the retained introns of the target gene, and the 5' ends of the primers in the gene-specific primer group 1 have a universal sequencing primer 2 sequence; The 5' end of the primer in the gene-specific primer group 2 carries the universal sequencing primer 2 sequence; The targeting site of the gene-specific primer group 2 is shifted 5-100 bases upstream of the site of the gene-specific primer group 1.

2. The method according to claim 1, wherein both gene-specific primer groups 1 and 2 are designed based on exons downstream of alternative splicing events of retained introns of the targeted gene. 3 . The method according to claim 1 , wherein the number of target genes targeted by gene-specific primer group 1 and gene-specific primer group 2 is greater than or equal to 1.

4. The method according to claim 1, wherein The molar concentration ratio of the normal dNTP and the 3' modified dNTP added in step (1) is 1:1-1:

100.

5. The method according to claim 1, wherein The molar concentration ratio of the normal dNTP and the 3' modified dNTP added in step (1) is 1:

20. The method according to claim 1 , wherein the nucleic acid barcode is connected to the primer P5 and / or P7 adapter end.

7. The method according to claim 1, wherein The universal sequencing primer 1 sequence is selected from any one of SEQ ID NO: 3, SEQ ID NO: 5 or SEQ ID NO:

6.

8. The method according to claim 1, wherein The complementary segment sequence of the universal sequencing primer 1 is shown in SEQ ID NO:

4.

9. The method according to any one of claims 1 to 8, wherein: The universal sequencing primer 2 sequence is selected from any one of SEQ ID NO: 7 or SEQ ID NO:

8.

10. A high-throughput sequencing library constructed according to the method according to any one of claims 1 to 9.

11. A targeted high-throughput sequencing method for detecting splicing isoforms, comprising the following steps: (1) Extracting sample RNA and constructing a sequencing library according to any one of claims 1 to 9; (2) Based on the above sequencing library, high-throughput sequencing is performed to obtain sequencing information of the target gene in the sample.

12. The method according to claim 11, wherein The sample is selected from the following cell samples: MCF10A, MCF7, HeLa, HEK293T and / or MDA-MB-231.

13. The method according to claim 11 or 12, wherein: The target genes include ATP13A1, CXXC1, ECHDC2, FGFRL1, HMGN3, KLHL17, NAXD, LZTR1, SELENBP1, JMJD8, PSMB1, HIGD2A, HNRNPAB, SMARCC1, ATP5IF1, HIGD2B, RPS21, UQCC5, NFATC3, PCNP and / or OSGEP.

14. Use of the method for establishing a sequencing library for high-throughput sequencing according to any one of claims 1 to 9, the high-throughput sequencing library according to claim 10, and / or the targeted high-throughput sequencing method according to claims 11 to 13 in the assessment of off-target events, The off-target event assessment includes the following: (1) Accurately determine the genomic location where off-target trans-splicing occurs in trans-splicing factors; (2) Quantitative analysis of on-target trans-splicing and off-target trans-splicing.

Citation Information

Patent Citations

  • Poly(A)-ClickSeq Click-Chemistry for Next Generation 3-End Sequencing Without RNA Enrichment or Fragmentation

    US20190256547A1