Method for acquiring full-length transcript mRNA sequence

By introducing universal bases into reverse transcription and two-strand synthesis processes, and combining them with endonuclease and polymerase chain reaction, the cost and accuracy problems of obtaining full-length transcript sequences from single cells in existing technologies have been solved, achieving low-cost, high-throughput, and high-accuracy full-length transcript sequence acquisition.

WO2025189419A9PCT designated stage Publication Date: 2026-04-23MGI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
MGI TECH CO LTD
Filing Date
2024-03-14
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing technologies are insufficient for obtaining full-length transcript sequences from single cells at low cost and high throughput, and traditional methods suffer from problems with sequencing accuracy and high cost.

Method used

By introducing universal bases in reverse transcription and two-strand synthesis, and using endonucleases to recognize and digest specific nucleic acid fragments, combined with polymerase chain reaction, permanent nucleotide markers are formed, ensuring the accurate location of sequencing reads and enabling the acquisition of full-length transcript sequences.

Benefits of technology

This technology enables low-cost and high-accuracy acquisition of full-length transcript sequences from single cells, improving the precision and assembly integrity of sequencing data analysis and allowing for accurate localization of cell origin.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024081653_23042026_PF_FP_ABST
    Figure CN2024081653_23042026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method for acquiring a full-length transcript sequence. The method comprises: performing reverse transcription treatment on the mRNA to be tested; performing second-strand synthesis treatment on a product after the reverse transcription treatment; performing digestion treatment on a product after the second-strand synthesis treatment in the presence of endonuclease, wherein the endonuclease specifically recognizes a nucleic acid fragment with universal bases; performing polymerase chain reaction treatment on a product after the digestion treatment; and performing sequencing treatment on a product after the polymerase chain reaction treatment, so as to obtain a full-length sequence of the mRNA to be tested. The reverse transcription treatment or the second-chain synthesis treatment is performed in the presence of universal bases.
Need to check novelty before this filing date? Find Prior Art

Description

Methods for obtaining full-length transcript mRNA sequences Technical Field

[0001] This application relates to the field of biotechnology, specifically to a method for obtaining a full-length transcript mRNA sequence. Background Technology

[0002] In recent years, with the advancement of high-throughput sequencing technology, transcriptome sequencing (RNA-seq) has been increasingly widely used in cells and tissues. Traditional short-read sequencing methods analyze gene expression differences, alternative splicing, and gene fusions by comparing reads (sequencing reads). However, due to the complexity of the transcriptome, traditional methods struggle to assemble complete transcripts, affecting the accuracy of subsequent analyses. Single-molecule long-read sequencing technology can obtain full-length transcript information, but its sequencing accuracy and cost remain challenges. Furthermore, high-throughput single-cell transcriptome sequencing technology plays a crucial role in studying the transcriptome of individual cells, but currently, commonly used methods can only obtain information from the 3' end of mRNA, limiting alternative splicing and isoform studies.

[0003] To address this issue, researchers proposed a method combining high-throughput single-cell RNA library preparation and single-molecule sequencing. The specific steps involve obtaining a full-length cDNA library from a single cell using a single-cell RNA library preparation kit, and then using a single-molecule sequencing platform to obtain the full-length transcript from the single cell. This method overcomes the challenge of obtaining full-length transcripts from single cells, but its high cost limits its large-scale application.

[0004] Therefore, methods for obtaining full-length transcript sequences from single cells at low cost and high throughput still need improvement.

[0005] Summary of the Invention

[0006] This application was made by the inventor based on the discovery of the following problems and facts:

[0007] Through analysis of existing technologies (Figure 1), the inventors discovered that although traditional transcriptome sequencing technology can assemble transcripts, it is difficult to accurately distinguish the original mRNA molecule source of the sequencing data, thus making it difficult to accurately assemble complete single transcripts. It also faces the challenge of not being able to detect more transcript types.

[0008] Although the Smart-seq method can obtain full-length cDNA during library preparation, it does not fully utilize the information from this full-length cDNA during subsequent library preparation and sequencing. Because short-read sequencing data cannot accurately distinguish the original mRNA molecules from which they originate, the problem of not being able to accurately assemble complete individual transcripts persists. Furthermore, the Smart-seq method cannot be combined with high-throughput single-cell sequencing technologies, limiting the high-throughput sequencing capability of full-length single-cell transcripts.

[0009] While single-molecule long-read sequencing methods can obtain complete full-length transcripts, they are very expensive; furthermore, long-read sequencing has lower accuracy. High-throughput single-cell RNA library preparation combined with single-molecule long-read sequencing can sequence complete full-length transcripts from single cells at high throughput, but it is also very expensive. Furthermore, due to the lower accuracy of single-molecule sequencing, high-throughput single-cell sequencing can interfere with the sequencing of cell tags, causing transcript loss and affecting the accuracy of transcript analysis. If high-accuracy sequencing data is required, the sequencing throughput needs to be increased, further increasing sequencing costs.

[0010] This application aims to address at least one of the technical problems existing in the prior art. To this end, the inventors propose a means for efficiently obtaining full-length transcript mRNA.

[0011] In view of this, in one aspect of this application, a method for obtaining a full-length transcript sequence is proposed. According to an embodiment of this application, the method includes: reverse transcription of the mRNA to be tested; two-strand synthesis of the reverse transcription product; digestion of the two-strand synthesis product in the presence of a nuclease, wherein the nuclease specifically recognizes a nucleic acid fragment with a universal base; polymerase chain reaction (PCR) of the digestion product to obtain a sequencing library; and sequencing of the sequencing library to obtain the full-length sequence of the mRNA to be tested; wherein the reverse transcription or the two-strand synthesis is performed in the presence of a universal base. This method is the first to employ the incorporation of a universal base into the cDNA first strand or the cDNA second strand amplified from the cDNA first strand, thereby generating a permanent nucleotide marker in the DNA strand produced by the subsequent PCR. This marker method can accurately determine its origin in sequencing reads, providing a reliable means for the precise assembly of single mRNA molecules and realizing the acquisition of a full-length transcript sequence.

[0012] According to embodiments of this application, the method for obtaining the full-length transcript sequence described above may further include at least one of the following technical features:

[0013] According to embodiments of this application, the full-length transcript sequence is derived from multiple single-cell samples.

[0014] In one example of this application, the labeling method of the mRNA molecule is correlated with the cell tag in RNA sequencing, which can accurately determine the cellular origin of the transcript and provide an effective approach for high-throughput single-cell full-length transcript sequencing.

[0015] According to embodiments of this application, cell tags are added to the mRNA of single cells based on microdroplet, micropore technology, or in-situ chip.

[0016] According to embodiments of this application, RNA is reverse transcribed using microdroplets, micropores, or in-situ chips to obtain cDNA samples with cell tags, and cDNA samples from different single cells are mixed, wherein the cDNA from different cells carries different cell tags.

[0017] According to embodiments of this application, the universal base includes at least one selected from dITP, 8-Oxo-dGTP, dPTP, dKTP, and 2OH-dATP. In some preferred examples of this application, the universal base is selected from dITP (2'-deoxy-inosine-5'-triphosphate). According to embodiments of this application, the above-mentioned universal base can arbitrarily pair with various nucleotides, such as dITP pairing with A, T, C, or G. For example, 8-Oxo-dGTP (8-Oxo-2'-deoxyguanosine 5'-triphosphate) can pair with A or C.

[0018] According to embodiments of this application, the reverse transcription process is performed in the presence of universal bases. The reverse transcription, amplification, and digestion processes are carried out as follows: cleaving single cells to release mRNA; reverse transcription of the mRNA to be tested in the presence of dNTPs and universal bases; performing two-strand synthesis on the reverse transcription product in the presence of dNTPs and the absence of universal bases; and digesting the two-strand synthesis product in the presence of a nuclease, which specifically recognizes nucleic acid fragments containing universal bases. In the reverse transcription step, a universal base is introduced to mark the cDNA molecule, thereby precisely locating the source of sequencing reads and improving the accuracy of sequencing data analysis.

[0019] According to embodiments of this application, the reverse transcription process is performed in the presence of universal bases. The reverse transcription, amplification, and digestion processes are carried out as follows: single cells are immobilized and digested on a chip; the mRNA to be tested is reverse transcribed in the presence of dNTPs and universal bases; the reverse transcription product is subjected to double-strand synthesis in the presence of dNTPs and the absence of universal bases; the double-strand synthesis product is digested in the presence of a nuclease, which specifically recognizes nucleic acid fragments containing universal bases. In the reverse transcription step, a marker is introduced into the cDNA molecule by introducing universal bases, thereby accurately locating the source of sequencing reads and improving the accuracy of sequencing data analysis. The chip surface is immobilized with capture sequences.

[0020] According to the above embodiments, the chip surface has discrete capture regions, the capture sequences of different capture regions have different tag sequences, and the capture sequences within the same capture region have the same tag sequence.

[0021] According to the above embodiments, the capture sequence is used to capture the nucleic acid to be tested in the cell, preferably, the capture sequence is used to capture mRNA or cDNA sequences.

[0022] According to the above embodiments, the capture sequence is an oligonucleotide primer with poly T, which may include at least one selected from universal primers, fixed sequences, cell tags, and unique molecular tags (UMI).

[0023] According to the above embodiments, the capture sequence is a template conversion primer (TSO) that may include at least one selected from universal primers, fixed sequences, cell tags, and unique molecular tags (UMI).

[0024] According to the embodiments of this application, referring to Figure 2, in the reverse transcription process, a universal base or a combination of a universal base and a dU base is randomly introduced into the cDNA first strand obtained by reverse transcription of mRNA; in the two-strand synthesis process, the cDNA first strand with the universal base is extended or amplified, thereby forming a nucleotide tag in the DNA product (cDNA second strand) obtained by amplification of the cDNA first strand, where nucleotides that pair with the universal base form a nucleotide tag; in the digestion process, endonuclease specifically recognizes the universal base or the dU base and digests the cDNA first strand, leaving only the cDNA second strand. Then, in the subsequent polymerase chain reaction, only the cDNA second strand is used as a template for further amplification, forming a permanent nucleic acid tag in the amplification product, providing a tag for the subsequent assembly of sequencing reads, and obtaining the sequencing results of the full-length transcript mRNA.

[0025] According to embodiments of this application, the reverse transcription process is performed in a reverse transcription system, which further includes a reverse transcriptase, the reverse transcriptase comprising at least one selected from MMLV reverse transcriptase and AMV reverse transcriptase. In some examples of this application, enzymes with similar functions to MMLV reverse transcriptase and AMV reverse transcriptase may also be used as alternatives.

[0026] According to embodiments of this application, the molar ratio of dNTPs to universal bases in the reverse transcription processing system is 9:1 to 1:9. The molar ratio of dNTPs to universal bases can optionally be 8:1, 7:1, 6:1, 5:1, 4:1, 3:1, 2:1, 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, or 1:8. According to a preferred embodiment of this application, the molar ratio of dNTPs to universal bases is 4:1. The molar ratio of dNTPs to universal bases in the reverse transcription processing system according to embodiments of this application allows universal bases to be inserted into the reverse transcription product—cDNA—at a suitable ratio and position, effectively avoiding base bias problems that may be caused by excessively high universal base content, thereby preventing the loss or overexpression of certain specific sequences. Simultaneously, it also eliminates the influence of excessively low universal base content, which could prevent complete labeling of the fragmented product. The molar ratio of dNTPs to universal bases in the reverse transcription processing system according to embodiments of the present invention further improves the accuracy and comprehensiveness of subsequent sequencing assembly.

[0027] According to embodiments of this application, the dNTPs include at least one of dATP, dTTP, dCTP, and dGTP. In some examples of this application, the molar ratio of dATP, dTTP, dCTP, dGTP, and the universal base is 2.25:2.25:2.25:2.25:1 to 0.25:0.25:0.25:0.25:9. The molar ratio of dATP, dTTP, dCTP, dGTP, and the universal base can optionally be 2:2:2:2:1, 1.5:1.5:1.5:1.5:1, 1:1:1:1:1, 0.25:0.25:0.25:0.25:1, 0.25:0.25:0.25:0.25:3, 0.25:0.25:0.25:0.25:6, or 0.25:0.25:0.25:0.25:9. According to a preferred embodiment of this application, the molar ratio of dATP, dTTP, dCTP, dGTP, and the universal base is 1:1:1:1:1. The molar ratio of dNTPs to universal bases in the reverse transcription processing system according to embodiments of the present invention enables the universal bases to be inserted into the first strand of the reverse transcription product—cDNA—at an appropriate ratio and position. This effectively avoids base bias problems that may be caused by excessive universal base content, thereby preventing the loss or overexpression of certain specific sequences. Simultaneously, it also eliminates the impact of insufficient universal base content leading to incomplete labeling of the fragmented product. The molar ratio of dNTPs to universal bases in the reverse transcription processing system according to embodiments of the present invention further improves the accuracy and comprehensiveness of subsequent sequencing assembly.

[0028] According to embodiments of this application, the two-strand synthesis processing system is carried out within a two-strand synthesis processing system, which further includes a first DNA polymerase having the activity of amplifying nucleic acid templates containing the universal bases. In some examples of this application, any DNA polymerase having the activity of amplifying nucleic acid templates containing the universal bases can be used.

[0029] According to embodiments of this application, the first DNA polymerase is selected from Phanta Uc ultra-fidelity DNA polymerase, BGI Platinum ultra-fidelity DNA polymerase, BGI Mars ultra-fidelity DNA polymerase, BGI Saturn ultra-fidelity DNA polymerase, and Beyotime Q6U. TM At least one of the high-fidelity DNA polymerases.

[0030] According to embodiments of this application, the reverse transcription processing system further includes template-converting oligonucleotides. In this step, the template-converting oligonucleotides are used to perform template conversion in the reverse transcription processing system, guiding the synthesis of the 3' end of cDNA.

[0031] According to embodiments of this application, in the reverse transcription system, the concentration of the dNTP is 0.1–0.9 mmol / L, preferably 0.8 mmol / L; the concentration of the reverse transcriptase is 2–20 U / μL, preferably 10 U / μL; the concentration of the universal base is 0.1–0.9 mmol / L, preferably 0.2 mmol / L; and the concentration of the template-converting oligonucleotide is 0.2–2 μmol / L, preferably 1 μmol / L. After extensive experimental optimization, the inventors determined the optimal concentration of the universal base in the reverse transcription system to be 0.2 mmol / L. Under the reverse transcription system described in this embodiment of the invention, the base bias problem that may be caused by excessively high universal base concentration can be effectively avoided, thereby preventing the loss or overexpression of certain specific sequences. Simultaneously, the effect of excessively low universal base concentration leading to incomplete labeling of the fragmentation product is also eliminated.

[0032] According to an embodiment of this application, in the two-strand synthesis treatment system, the concentration of the first DNA polymerase is 0.002 U / μL to 0.2 U / μL, preferably 0.02 U / μL; the concentration of the dNTP is 0.2 to 2 mmol / L, preferably 1 mmol / L.

[0033] In some specific examples of this application, the above implementation method can be referred to Figure 2.

[0034] According to embodiments of this application, the two-strand synthesis process is performed in the presence of universal bases. The reverse transcription process, the two-strand synthesis process, and the digestion process are performed as follows: the mRNA to be tested is reverse transcribed in the presence of dNTPs and in the absence of universal bases, wherein the reverse transcription primers include dU bases (deoxyuracil); the reverse transcription product is subjected to two-strand synthesis in the presence of dNTPs and universal bases; the product after two-strand synthesis is digested in the presence of a nuclease, wherein the nuclease specifically recognizes nucleic acid fragments containing dU bases; the digested product is subjected to three-strand synthesis in the presence of dNTPs and in the absence of universal bases; the product after three-strand synthesis is digested in the presence of a nuclease, wherein the nuclease specifically recognizes nucleic acid fragments containing universal bases, thereby obtaining a three-stranded RNA polymerase with nucleotides paired with universal bases forming a nucleotide tag, which is then used to form a permanent nucleotide tag in a polymerase chain reaction.

[0035] According to embodiments of this application, referring to Figure 3, in the reverse transcription process, a reverse transcription product (cDNA first strand) is obtained by reverse transcription of mRNA, wherein the reverse transcription primer (oligo dT primer) contains dU bases, and the reverse transcription system contains dU bases; in the two-strand synthesis process, the cDNA first strand is amplified under the condition of the presence of universal bases, thereby introducing universal bases into the DNA product (cDNA second strand) obtained by amplifying the cDNA first strand; in the digestion process, a nuclease specifically recognizes the cDNA first strand containing dU bases, digests the dU bases in the cDNA first strand, leaving the cDNA second strand; in the three-strand synthesis process, the digested cDNA second strand containing universal bases is used as a template to amplify the cDNA third strand, wherein the nucleosides in the cDNA third strand that pair with universal bases... Acids form nucleotide tags, which are then permanently formed in polymerase chain reactions (PCRs). During digestion, endonucleases specifically recognize cDNA double strands with universal bases, digesting the cDNA double strand and leaving the DNA product (cDNA triple strand) obtained from the amplification of the cDNA double strand. In subsequent PCRs, only the DNA product (cDNA triple strand) obtained from the amplification of the cDNA double strand is used as a template for further amplification. Permanent nucleic acid tags are formed in the amplification products, providing tags for the assembly of subsequent sequencing reads, and obtaining the sequencing results of full-length transcript mRNA.

[0036] According to embodiments of this application, the reverse transcription process is performed in a reverse transcription system, which further includes a reverse transcriptase, the reverse transcriptase comprising at least one selected from MMLV reverse transcriptase and AMV reverse transcriptase. In some examples of this application, enzymes with similar functions to MMLV reverse transcriptase and AMV reverse transcriptase may also be used as alternatives.

[0037] According to embodiments of this application, the two-strand synthesis process is performed in a two-strand synthesis system, wherein the molar ratio of dNTPs to universal bases in the two-strand synthesis system is 9:1 to 1:9. The molar ratio of dNTPs to universal bases can optionally be 8:1, 7:1, 6:1, 5:1, 4:1, 3:1, 2:1, 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, or 1:8. According to a preferred embodiment of this application, the molar ratio of dNTPs to universal bases is 4:1. According to embodiments of the present invention, the molar ratio of dNTPs to universal bases in the two-strand synthesis system enables universal bases to be inserted into the first strand of the reverse transcription product—cDNA—at a suitable ratio and position, effectively avoiding base bias problems that may be caused by excessive universal base content, thereby preventing the loss or overexpression of certain specific sequences. Simultaneously, it also eliminates the influence of insufficient universal base content leading to incomplete labeling of the fragmented product. The molar ratio of dNTPs to universal bases in the two-strand synthesis processing system according to embodiments of the present invention further improves the accuracy and comprehensiveness of subsequent sequencing assembly.

[0038] According to embodiments of this application, the dNTPs include at least one of dATP, dTTP, dCTP, and dGTP. In some examples of this application, the molar ratio of dATP, dTTP, dCTP, dGTP, and the universal base is 2.25:2.25:2.25:2.25:1 to 0.25:0.25:0.25:0.25:9. The molar ratio of dATP, dTTP, dCTP, dGTP, and the universal base can optionally be 2:2:2:2:1, 1.5:1.5:1.5:1.5:1, 1:1:1:1:1, 0.25:0.25:0.25:0.25:1, 0.25:0.25:0.25:0.25:3, 0.25:0.25:0.25:0.25:6, or 0.25:0.25:0.25:0.25:9. According to a preferred embodiment of this application, the molar ratio of dATP, dTTP, dCTP, dGTP, and the universal base is 1:1:1:1:1. The molar ratio of dNTPs to universal bases in the two-strand synthesis processing system according to embodiments of the present invention enables the universal bases to be inserted into the first strand of the reverse transcription product—cDNA—at an appropriate ratio and position. This effectively avoids base bias problems that may be caused by excessive universal base content, thereby preventing the loss or overexpression of certain specific sequences. Simultaneously, it also eliminates the impact of insufficient universal base content leading to incomplete labeling of the fragmented product. Furthermore, the molar ratio of dNTPs to universal bases in the two-strand synthesis processing system according to embodiments of the present invention further improves the accuracy and comprehensiveness of subsequent sequencing assembly.

[0039] According to an embodiment of this application, the two-strand synthesis process is carried out in a two-strand synthesis process system, which further includes a second DNA polymerase.

[0040] According to embodiments of this application, the second DNA polymerase is selected from Phanta Uc ultra-fidelity DNA polymerase, BGI Platinum ultra-fidelity DNA polymerase, BGI Mars ultra-fidelity DNA polymerase, BGI Saturn ultra-fidelity DNA polymerase, and Beyotime Q6U. TM At least one of the high-fidelity DNA polymerases.

[0041] According to embodiments of this application, the triple-strand synthesis process is performed in a triple-strand synthesis system, which further includes a third DNA polymerase having the activity of amplifying nucleic acid templates containing the universal bases. In some examples of this application, any DNA polymerase having the activity of amplifying nucleic acid templates containing the universal bases can be used. In some specific examples of this application, the third DNA polymerase is selected from Phanta Uc ultra-fidelity DNA polymerase, BGI Platinum ultra-fidelity DNA polymerase, BGI Mars ultra-fidelity DNA polymerase, BGI Saturn ultra-fidelity DNA polymerase, and Beyotime Q6U. TM At least one of the high-fidelity DNA polymerases.

[0042] According to an embodiment of this application, in the reverse transcription treatment system, the concentration of the dNTP is 0.2–2 mmol / L, preferably 1 mmol / L; the concentration of the reverse transcriptase is 2–20 U / μL, preferably 10 U / μL; and the concentration of the template-converting oligonucleotide is 0.2–2 μmol / L, preferably 1 μmol / L.

[0043] According to embodiments of this application, in the two-strand synthesis treatment system, the concentration of the second DNA polymerase is 0.002 U / μL to 0.2 U / μL, preferably 0.02 U / μL; the concentration of the dNTP is 0.2 to 2 mmol / L, preferably 0.8 mmol / L; and the concentration of the universal base is 0.2 to 2 mmol / L, preferably 0.2 mmol / L. After extensive experimental optimization, the inventors determined that the optimal concentration of the universal base in the two-strand synthesis treatment system is 0.2 mmol / L. This avoids the base bias problem that may be caused by excessively high universal base concentrations, thereby preventing the loss or overexpression of certain specific sequences. Simultaneously, it also eliminates the influence of excessively low universal base concentrations, which could prevent complete labeling of the fragmentation product.

[0044] According to an embodiment of this application, in the triple-strand synthesis treatment system, the concentration of the third DNA polymerase is 0.002 U / μL to 0.2 U / μL, preferably 0.02 U / μL, and the concentration of the dNTP is 0.2 to 2 mmol / L, preferably 0.8 mmol / L.

[0045] In some specific examples of this application, the above implementation method can be referred to Figure 3.

[0046] According to embodiments of this application, the endonuclease includes at least one selected from USER enzyme, UDG enzyme, endonuclease VIII, hSMUG1 enzyme, hAAG enzyme, and Fpg enzyme. In a preferred example of this application, the endonuclease is selected from hAAG enzyme.

[0047] According to embodiments of this application, the polymerase chain reaction (PCR) treatment is performed under conditions of a fourth DNA polymerase, the presence of dNTPs, and the absence of the universal base. In some examples of this application, the polymerases selected are Phanta Uc ultra-fidelity DNA polymerase, BGI Platinum ultra-fidelity DNA polymerase, BGI Mars ultra-fidelity DNA polymerase, BGI Saturn ultra-fidelity DNA polymerase, and Beyotime Q6U. TM At least one of the high-fidelity DNA polymerases.

[0048] According to an embodiment of this application, the full-length sequence of the mRNA to be tested is obtained by: filtering and aligning the sequencing data obtained by the sequencing process to identify unique molecular tags composed of single nucleotide variants; and assembling the full-length sequence of the mRNA to be tested based on the overlap information of reads containing the unique molecular tags.

[0049] According to embodiments of this application, the sequencing process includes: fragmenting the polymerase chain reaction product (using conventional methods such as sonication or fragmented enzymes), constructing a library from the fragmented product, and sequencing the library-constructed product. The fragmentation and library construction processes described in this step are conventional steps and can be performed based on the instructions of a standard reagent kit.

[0050] According to embodiments of this application, the above alignment process is performed by comparing the sequencing results with a reference genome. Based on the alignment results, single nucleotides that differ from the reference genome in the first sequencing data are identified to determine single nucleotide variant combinations; and sequencing reads having combinations of said single nucleotide variants are assembled to obtain the full-length sequence of the mRNA to be tested. As described above, regardless of whether universal bases are introduced during reverse transcription or referenced in two-strand synthesis, nucleotides that pair with universal bases will remain in the amplified DNA product, thus being permanently retained in subsequent polymerase chain reactions. During the alignment process, the single nucleotides that differ from the reference genome in the first sequencing data are the positions where universal bases were introduced in the amplified product. Different sequencing reads have different combinations of these single nucleotide variants, and these combinations of single nucleotide variants can serve as unique molecular tags for the sequencing reads. Sequencing reads with molecular tags are assembled to obtain the full-length sequence of the mRNA to be tested.

[0051] According to embodiments of this application, prior to the reverse transcription treatment, the method further includes pre-hybridization of the mRNA to be tested.

[0052] According to embodiments of this application, the pre-hybridization process includes hybridizing the mRNA to be tested with oligonucleotide primers containing polyT.

[0053] According to embodiments of this application, the oligonucleotide primer with poly T may include at least one selected from poly T, universal primers, fixed sequences, cell tags, and unique molecular tags (UMI).

[0054] According to embodiments of this application, the template conversion primer (TSO) may include at least one selected from universal primers, fixed sequences, cell tags, and unique molecular tags (UMI).

[0055] In one example of this application, the universal primer refers to a primer used for amplification or guiding a sequencing reaction; the fixed sequence refers to one or more known sequences, 3-10 bases in length, added to the oligonucleotide primer, used as a spacer sequence between different functional sequences; the cell tag refers to a known tag sequence added to the oligonucleotide primer, generally selected from 8-12 bases, which can be used to distinguish different cells; the UMI is selected from randomly synthesized oligonucleotides and used to label each molecule in the sample, facilitating the counting of each molecule. In one example of this application, the UMI is used to count the copy number of each gene.

[0056] In a second aspect of this application, a method for detecting the full-length transcriptome of a single cell is proposed. According to embodiments of this application, based on the method described in the first aspect or any embodiment of the first aspect, a full-length transcriptome sequence of a single cell to be tested is obtained; based on the cell tag sequence in the full-length transcriptome sequence of the single cell to be tested, full-length transcriptome sequences of single cells from different sources are determined. In some examples of this application, this method can obtain the full-length transcriptome sequences of all single cells to be tested at low cost and high accuracy, and further distinguishes all single cells to be tested based on the cell tag sequence in the full-length transcriptome sequence.

[0057] According to embodiments of this application, the above-described single-cell full-length transcriptome detection method may further include at least one of the following technical features:

[0058] According to an embodiment of this application, the method further includes: obtaining the copy number of different genes in a single cell based on the UMI sequence in the full-length transcript sequence.

[0059] According to an embodiment of this application, the method further includes: obtaining single-cell transcriptome information based on a comprehensive analysis of the UMI sequence and cell tag sequence in the full-length transcript sequence.

[0060] In a third aspect, this application provides a full-length transcriptome detection kit. According to embodiments of this application, the kit comprises at least one of: dNTPs, universal bases, dU bases, endonucleases, reverse transcriptases, DNA polymerases, reverse transcription primers, TSO primers, and polyT-containing oligonucleotide primers. In some examples of this application, single-cell full-length transcriptome sequencing can be performed cost-effectively, easily, and rapidly using the aforementioned kit.

[0061] According to embodiments of this application, the kit further comprises: endonuclease buffer, reverse transcriptase buffer, DNA polymerase buffer, deionized water, and at least one of the kit instructions.

[0062] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be described in detail here. Attached Figure Description

[0063] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0064] Figure 1 is a schematic diagram of the principle of short-read full-length transcript sequencing based on intramolecular markers as described in this application;

[0065] Figure 2 is a schematic diagram of the full-length transcript library construction and sequencing process described in this application (incorporation of universal bases in the reverse transcription step);

[0066] Figure 3 is a schematic diagram of the full-length transcript library construction and sequencing process described in this application (universal bases are incorporated during the second-strand synthesis process);

[0067] Figure 4 is a schematic diagram of the high-throughput single-cell full-length transcript sequencing workflow described in this application;

[0068] Figure 5 is a schematic diagram of the types of mRNA isoforms of each gene described in one embodiment of this application;

[0069] Figure 6 is a schematic diagram of the length distribution of transcripts according to an embodiment of this application. Detailed Implementation

[0070] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0071] In this application, unless otherwise stated, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0072] In this application, unless otherwise stated, the term "reverse transcription" refers to the transcription of the mRNA to be tested into the corresponding cDNA, typically through the action of reverse transcriptase, which converts the mRNA molecule into complementary cDNA. In one example of this application, a U base or a universal base is introduced into the cDNA in the presence of a universal base, thereby forming a specific molecular tag in the DNA product formed by the subsequent amplification reaction.

[0073] In this application, unless otherwise stated, the term "amplification process" refers to an extension or amplification reaction using DNA or RNA as a template under the action of a polymerase to obtain a DNA product. In one example of this application, a universal base label is introduced through the amplification process, thereby forming a specific molecular tag in the DNA product formed by the subsequent amplification reaction.

[0074] In this application, unless otherwise stated, the term "digestion treatment" refers to the use of endonucleases to cleave nucleotide chains containing universal or U bases, thereby achieving the purpose of digesting nucleotide chains containing universal or U bases. The term "digestion treatment" in this application can also refer to permeabilization or lysis treatment of cells. Permeabilization refers to creating pores on the cell membrane surface using enzymes, electroporation, or other digestion reagents to allow reverse transcription reagents or intracellular nucleic acids to enter or exit the cell membrane.

[0075] In this application, unless otherwise stated, the term "sequencing process" refers to nucleic acid sequence determination, the same as "nucleic acid sequencing" or "gene sequencing", which refers to determining the base sequence of the primary structure of nucleic acid molecules, which can be achieved by sequencing by synthesis (SBS) and / or sequencing by ligation (SBL).

[0076] The term "sequencing" can be performed using a sequencing platform. According to embodiments of this application, selectable sequencing platforms include, but are not limited to, Illumina's HiSeq, MiSeq, Nextseq, and Novaseq sequencing platforms; Thermo Fisher / Life Technologies' Ion Torrent platform; BGI Genomics' BGISEQ and MGISEQ / DNBSEQ platforms; and single-molecule sequencing platforms. In some preferred embodiments of this application, the sequencing platform is selected from BGI Genomics' MGISEQ platform. In some embodiments of this application, the sequencing method can be selected as single-end sequencing, paired-end sequencing, or any sequencing method supported by the chosen automated sequencing platform. The sequence obtained from sequencing is called a sequencing sequence, also known as a read, and the length of the sequencing sequence or read is also called the read length.

[0077] In this application, the term "universal base" refers to a non-natural base (A, T, C, G) or a base capable of complementary pairing with a variety of natural bases (A, T, C, G). Examples include dITP, 8-Oxo-dGTP, dPTP, dKTP, 2OH-dATP, and dUTP. For instance, dITP can pair with A, T, C, or G; 8-Oxo-dGTP (8-Oxo-2'-deoxyguanosine 5'-triphosphate) can pair with A, T, C, or C; and dPTP (6H,8H-3,4-Dihydropyrimido(4,5-c)(1,2)oxazin-7-one-8-β-D-2'-deoxy-ribofuranoside-5'-triphosphate, 6H,8H-3,4-di ... '-Deoxyfuranoside-5'-triphosphate) can pair with A or G; dKTP (N6-Methoxy-2,6-diaminopurine-2'-deoxyriboside-5'-O-triphosphate) can pair with T or C; 2OH-dATP (2-Hydroxy-2'-deoxyadenosine-5'-triphosphate, Triethylammonium salt) can pair with A, T, C or G; dUTP (deoxyuridine triphosphate) can pair with A.

[0078] In this application, unless otherwise stated, the term "reference genome" refers to the genome sequence of the species corresponding to the known sample. It can be a genome sequence obtained through public channels or obtained by sequencing assembly. It can be the entire genome sequence or a part of the genome of interest. For example, when analyzing human samples, multiple versions of human genome sequences provided by public databases can be used, such as hg19.

[0079] Methods for obtaining full-length transcript mRNA sequences

[0080] In one aspect of this application, a method for obtaining full-length transcript mRNA sequences is proposed, which is particularly suitable for high-throughput short-read sequencing. Using this method to obtain full-length transcript mRNA sequences solves the problems of high cost and low accuracy associated with single-molecule long-read transcriptome sequencing. Simultaneously, it also addresses the challenges of traditional short-read transcriptome sequencing, which cannot accurately detect individual transcript molecules and cannot fully determine transcript length.

[0081] In some examples of this application, the label of a single mRNA can also be matched one-to-one with the label of a single cell in high-throughput single-cell RNA sequencing to determine the cell origin of the transcript and achieve high-throughput single-cell full-length transcript sequencing.

[0082] For ease of understanding, the technical solution of this application will be described in detail below with reference to the accompanying drawings.

[0083] Referring to Figure 2, in one example of this application, the mRNA to be tested is first reverse transcribed using a mixture of primers containing Poly(A), template-changing oligonucleotides (TSO), reverse transcriptase (such as MMLV or AMV-type reverse transcriptase), and dNTPs (containing universal bases) to obtain a cDNA one-strand containing universal bases. During this process, the universal bases replace the natural bases A, T, C, and G incorporation into the cDNA one-strand during reverse transcription.

[0084] In one example of this application, the TSO primer includes a universal primer sequence and a UMI sequence; the primer with Poly(A) includes a universal sequence and a UMI sequence.

[0085] Then, using DNA polymerase and universal TSO primers, the cDNA first strand is extended as a template. By pairing with various natural bases, the base types at multiple positions in the mRNA sequence are changed, forming a uniquely labeled cDNA second strand. For example, the universal base I can pair with four bases: A, T, C, and G. If the original mRNA position is C, the corresponding cDNA second strand will change to A, T, C, or G. If the cDNA second strand changes to C, it matches the original mRNA position C and is not used as a molecular marker. If the corresponding cDNA second strand changes to A, T, or G, and does not match the original mRNA position C, it can serve as a unique molecular marker.

[0086] Furthermore, the fragment containing universal bases was removed using endonuclease, leaving only the cDNA double strand with molecular markers, and the retained cDNA double strand was then amplified by PCR.

[0087] Finally, the PCR amplification products were fragmented for library construction, and the assembly and acquisition of full-length transcript mRNA molecules were achieved through filtering, alignment, and analysis of single nucleotide variant (SNV) combinations of sequencing data.

[0088] In some examples of this application, the library construction method is selected from conventional laboratory library construction methods or the Tn5 library construction method.

[0089] In some examples of this application, the sequencing data filtering and alignment methods are selected from conventional methods.

[0090] In one example of this application, the assembly of the full-length transcript mRNA molecule can be performed as follows: First, the sequencing data is quality controlled and filtered to remove low-quality reads and fragments containing technical errors. Next, high-quality reads are aligned with a reference genome to determine their genomic location. Unique molecular tags are formed by identifying bases that do not match the reference genome (single nucleotide variants, SNVs). These tags can represent different variant positions and are important indicators for distinguishing different molecules. Utilizing these SNV combinations and the overlap information of the reads, an overlap assembly algorithm (such as De Bruijn graphs or graph-based assembly methods) is used to sequentially assemble the reads into a full-length transcript mRNA molecule. Further correction and splicing are performed to resolve potential mismatches or base errors, ensuring the assembly result is as accurate and complete as possible.

[0091] In some examples, the assembly of full-length transcripts may also be evaluated and validated, typically using other experimental methods (such as RT-PCR) or comparisons with other known databases to verify the accuracy and integrity of the resulting full-length transcripts.

[0092] In a specific example of this application, referring to Figure 4, taking high-throughput single-cell full-length transcript sequencing as an example, the assembly process of the full-length transcript mRNA molecule is as follows:

[0093] First, individual cells are isolated (e.g., using microdroplet or micropore techniques) and labeled. Single-cell lysis (using standard laboratory methods) is then performed in the droplets or micropores to release mRNA. The polyA of the mRNA hybridizes with oligonucleotide primers containing polyT (containing a specific universal primer, an 8–12 nt fixed sequence, a cell barcode, and a unique molecular tag (UMI)). Next, reverse transcription and template conversion occur, during which universal bases, such as dITP, 8-Oxo-dGTP, dPTP, dKTP, 2OH-dATP, and dUTP, are incorporated in specific proportions. These universal bases replace the A, T, C, and G in the original mRNA and are incorporated into the resulting cDNA first strand.

[0094] Subsequently, specific endonucleases, such as USER and UDG, were used to digest the first strand of cDNA, removing fragments containing universal bases and leaving only the second strand of cDNA containing molecular markers. The second strand of cDNA was then amplified by PCR, and a library was constructed. After library sequencing, the data was filtered and aligned. Molecular tags were formed using SNVs, and these tags were used to assemble reads to reconstruct the full-length mRNA molecule.

[0095] The assembly of the full-length transcript mRNA molecule can be performed as follows: First, the sequencing data undergoes quality control and filtering to remove low-quality reads and fragments containing technical errors. Next, high-quality reads are aligned with a reference genome to determine their genomic location. Unique molecular tags are formed by identifying bases that do not match the reference genome (single nucleotide variants, SNVs). These tags represent different variant locations and are important indicators for distinguishing different molecules. Utilizing these SNV combinations and read overlap information, an overlap assembly algorithm (such as De Bruijn graphs or graph-based assembly methods) is used to sequentially assemble the reads into a full-length transcript mRNA molecule. Further correction and splicing are performed to resolve potential mismatches or base errors, ensuring the assembly result is as accurate and complete as possible.

[0096] Finally, reads with cell tags were screened using known 8-12 base fixed sequences, and their corresponding complete mRNA molecules were determined based on the molecular tags in these reads.

[0097] In some examples of this application, cell tags and molecular tags can also be used to distribute reads to individual cells, confirm the expression level of mRNA molecules in each cell and the information of specific mRNAs, so as to achieve full-length transcript sequencing and expression analysis of all mRNA molecules in a single cell.

[0098] In one example of this application, full-length transcript sequencing combining single-cell and spatial transcriptomes can also be achieved based on the above method.

[0099] Referring to Figure 3, in one example of this application, firstly, the mRNA to be tested is transcribed into a single strand of cDNA via reverse transcription. This step does not involve the addition of universal bases, and the reverse transcription primers contain dU bases. Next, in the amplification process, dNTPs containing universal bases are used. These universal bases replace the native bases A, T, C, and G in the cDNA during the amplification process.

[0100] In one example of this application, the universal bases include dITP, 8-Oxo-dGTP, dPTP, dKTP, 2OH-dATP, dUTP, etc.

[0101] Subsequently, using the first strand of cDNA as a template, a second amplification was performed under the condition of the presence of universal bases, thereby introducing universal bases into the second strand of cDNA.

[0102] Furthermore, the cDNA one strand containing universal bases was identified and excised using the User enzyme for digestion.

[0103] Further, using the second strand of cDNA as a template, a third amplification was performed under conditions that did not contain universal bases in order to obtain the third strand of cDNA.

[0104] Furthermore, endonuclease was used for digestion to identify and remove the second strand of cDNA containing universal bases, leaving the third strand of cDNA with molecular markers. The remaining third strand of cDNA was then subjected to PCR amplification to obtain the amplification product.

[0105] Finally, the PCR amplification products were fragmented for library construction, and the assembly and acquisition of full-length transcript mRNA molecules were achieved through filtering, alignment, and analysis of single nucleotide variant (SNV) combinations of sequencing data.

[0106] In some examples of this application, the library construction method is selected from conventional laboratory library construction methods or the Tn5 library construction method.

[0107] In some examples of this application, the sequencing data filtering and alignment methods are selected from conventional methods.

[0108] In one example of this application, the assembly of the full-length transcript mRNA molecule can be performed as follows: First, the sequencing data is quality controlled and filtered to remove low-quality reads and fragments containing technical errors. Next, high-quality reads are aligned with a reference genome to determine their genomic location. Unique molecular tags are formed by identifying bases that do not match the reference genome (single nucleotide variants, SNVs). These tags can represent different variant positions and are important indicators for distinguishing different molecules. Utilizing these SNV combinations and the overlap information of the reads, an overlap assembly algorithm (such as De Bruijn graphs or graph-based assembly methods) is used to sequentially assemble the reads into a full-length transcript mRNA molecule. Further correction and splicing are performed to resolve potential mismatches or base errors, ensuring the assembly result is as accurate and complete as possible.

[0109] In some examples, the assembly of full-length transcripts may also be evaluated and validated, typically using other experimental methods (such as RT-PCR) or comparisons with other known databases to verify the accuracy and integrity of the resulting full-length transcripts.

[0110] In summary, this application involves labeling full-length mRNA molecules before library construction and sequencing, followed by sequencing using a next-generation sequencing platform. This reduces sequencing costs and increases sequencing accuracy compared to single-molecule sequencing platforms. Furthermore, because the entire molecule, including the intermediate regions, is labeled, data from the sequenced RNA intermediate regions can be used to assemble transcripts, reducing the difficulty of assembling individual transcript molecules, increasing the length and integrity of assembled transcripts, and reducing data waste. It can also be combined with current high-throughput single-cell library construction technologies to achieve high-throughput single-cell full-length transcriptome sequencing, expanding the scope of applications.

[0111] The present invention will now be described with reference to examples. It should be noted that these embodiments are merely descriptive and do not limit the invention in any way. Where specific techniques or conditions are not specified in the embodiments, they shall be performed in accordance with the techniques or conditions described in the literature in the art or according to the product instructions. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products.

[0112] Example 1: Library construction and sequencing using universal human RNA standards

[0113] 1. The commercially available universal human RNA standard (UHRR, 740000, Agilent) was used as the sample for the experiment. UHRR was dissolved to a final concentration of 2 μg / μL according to the instructions.

[0114] 2. Take 0.5 μL of UHRR sample and dilute it to 100 ng / μL, then take 1 μL of the diluted RNA sample into a PCR tube.

[0115] 3. Add 1 μL of 10 mM oligo dT primer to the PCR tube, and then add 1 μL of nuclease-free water.

[0116] 4. After mixing and centrifuging the PCR tube, place it in a PCR instrument and incubate at 72°C for 3 minutes. Immediately after incubation, place the PCR tube in an ice box to cool, allowing the oligo dT primers to hybridize with the polyA tail of the mRNA.

[0117] oligo dT primer: AAGGAGTGGTATGAATGGAGAGTAGTTTTTTTTTTTTTTTTTTTVN (SEQ ID NO: 1).

[0118] 5. Prepare the RT reaction mixture according to the formula in Table 1;

[0119] Table 1

[0120] Note: In the 25mM dATP / dTTP / dCTP / dGTP / dITP mixture: dATP, dTTP, dCTP, dGTP, dITP, and dITP each account for 20% of the total molar composition, and are prepared according to the ratio of 100mM dATP:100mM dTTP:100mM dCTP:100mM dGTP:100mM dGTP:100mM dITP = 0.8:0.8:0.8:0.8:0.8.

[0121] The dNTP Set Solution (dATP, dCTP, dTTP, dGTP, 100mM each) is from Yisheng, catalog number 10122ES74; the dITP solution (100mM) is from Thermo Fisher, catalog number R1191.

[0122] The reverse transcriptase Hiscript III is from Novozymes, catalog number R302-01, and contains 5x Hiscript III Buffer;

[0123] The RNase inhibitor Murine is from Novozymes, catalog number R301-01;

[0124] Betaine is from Sigma, catalog number B0300; 1M DTT is from Sigma, 646563, diluted with water to 100mM.

[0125] Magnesium chloride 1M MgCl2 is from Thermo Fisher, item number AM9530G.

[0126] TSO: AAGGAGTGGTATGAATGGAGAGTAGAT / rG / rG / iXNA_G (SEQ ID NO: 2);

[0127] rG represents ribonucleic acid modification; iXNA_G represents locked nucleic acid modification.

[0128] 6. Add the RT reaction mixture to the PCR tube, mix briefly, and then centrifuge.

[0129] 7. Place the PCR tube on the PCR instrument and perform the reaction according to the procedure in Table 2.

[0130] Table 2

[0131] 8. After the reaction is complete, briefly centrifuge to collect the liquid in the PCR tube to the bottom of the tube.

[0132] 9. Prepare the cDNA amplification reaction solution according to Table 3.

[0133] Table 3

[0134] Note: Phanta Uc ultra-fidelity DNA polymerase is from Novozymes, catalog number P507-01, and contains 5x Uc buffer;

[0135] 25mM dNTPs are from ENZYMATICS, item number N2050-25;

[0136] TSO primer: AAGGAGTGGTATGAATGGAGAGTAG (SEQ ID NO: 3).

[0137] 10. Add the cDNA amplification reaction solution to the PCR tube, mix briefly, and then centrifuge.

[0138] 11. Place the PCR tube on the PCR instrument and perform the extension reaction according to the procedure in Table 4.

[0139] Table 4

[0140] 12. After the extension reaction is complete, briefly centrifuge to collect the liquid in the PCR tube to the bottom of the tube.

[0141] 13. Add 1 μL of hAAG enzyme (NEB, M0313S) to the PCR tube, mix briefly, and then centrifuge.

[0142] 14. Place the PCR tubes on the PCR instrument and perform digestion and PCR reaction according to the procedure in Table 5 (PCR reaction system is the same as in Table 3).

[0143] Table 5

[0144] 15. After the reaction is complete, briefly centrifuge to collect the liquid in the PCR tube to the bottom of the tube.

[0145] 16. Add 30 μL of DNA purification magnetic beads to the PCR tube, pipette and mix thoroughly 10 times, then incubate at room temperature for 10 min.

[0146] 17. After a brief centrifugation, place the PCR tube on a magnetic rack and let it stand for 5 minutes until the liquid is clear. Use a pipette to aspirate and discard the supernatant.

[0147] 18. Add 150 μL of 75% ethanol to the tube, let stand for 30 seconds, and wash the magnetic beads.

[0148] 19. Repeat the previous step, trying to aspirate as much liquid as possible from the tube. If a small amount of liquid remains on the tube wall, the PCR tube can be centrifuged briefly. After separation on a magnetic rack, use a small-capacity pipette to aspirate the liquid from the bottom of the tube.

[0149] 20. Keep the PCR tube fixed on the magnetic rack, open the tube cap, and let it dry at room temperature until the surface of the magnetic beads is no longer reflective and cracked.

[0150] 21. Remove the PCR tube from the magnetic rack, add 32 μL of molecular-grade water to elute the DNA, and gently pipette 10 times until completely mixed.

[0151] 22. Incubate at room temperature for 5 minutes.

[0152] 23. Centrifuge briefly, place the PCR tube on a magnetic rack, let stand for 5 minutes until the liquid is clear, and transfer 30 μL of supernatant to a new PCR tube.

[0153] 24. Use The dsDNA HS Assay Kit was used to quantify the PCR product. 100 ng of the PCR product was transferred to a new PCR tube and the volume was increased to 45 μL with TE buffer.

[0154] 25. Use the reagents in the MGI Easiy Fast DNA Enzyme Digestion Library Preparation Kit (940-001193-00, MGI) to construct the library.

[0155] 26. Prepare the end-of-life repair and A reaction solution on ice according to Table 6.

[0156] Table 6

[0157] 27. Add the prepared end-fracture repair reaction solution to the PCR tube, vortex 3 times for 3 seconds each time, and then centrifuge briefly to collect the reaction solution to the bottom of the tube.

[0158] 28. When the temperature of the PCR instrument drops to 4℃, place the PCR tube on the PCR instrument and perform the break-up and end-repair A reaction according to the conditions in Table 7.

[0159] Table 7

[0160] 29. After the reaction is complete, briefly centrifuge to collect the reaction solution to the bottom of the tube.

[0161] 30. Remove the DNA purification beads in advance and allow them to equilibrate at room temperature for at least 30 minutes. Shake well before use.

[0162] 31. Pipette 48 μL of magnetic beads into a PCR tube and vortex to mix.

[0163] 32. Incubate at room temperature for 5 minutes.

[0164] 33. Briefly centrifuge the PCR tube, place it on a magnetic rack and let it stand for 2 minutes until the liquid is clear. Carefully aspirate the supernatant and discard it. Aspirate as much liquid as possible from the tube. If a small amount of liquid remains on the tube wall, briefly centrifuge the tube again, separate it on a magnetic rack, and then use a small-range pipette to aspirate the liquid from the bottom of the tube.

[0165] 34. Remove the PCR tube from the magnetic rack, add 45 μL of TE buffer to elute the DNA, vortex to mix, and then centrifuge briefly.

[0166] 35. Add 5 μL of adapter (from the MGIEAsy Fast DNA Enzyme Digestion Library Preparation Kit) to the PCR tube, vortex 3 times for 3 seconds each time, and then briefly centrifuge to collect the reaction solution to the bottom of the tube.

[0167] 36. Prepare the connector connection reaction solution on ice according to the formula in Table 8.

[0168] Table 8

[0169] 37. Add the prepared adapter ligation reaction solution to the PCR tube, vortex 6 times for 3 seconds each time, and then centrifuge briefly to collect the reaction solution to the bottom of the tube.

[0170] 38. Place the PCR tubes above on the PCR instrument and perform the adapter ligation reaction according to the conditions in Table 9.

[0171] Table 9

[0172] 39. After the reaction is complete, the reaction solution is collected to the bottom of the tube by instantaneous centrifugation.

[0173] 40. Remove the DNA purification beads in advance and allow them to equilibrate at room temperature for at least 30 minutes. Shake thoroughly before use.

[0174] 41. After aspirating 22 μL of TE buffer into the PCR tube, aspirate 20 μL of DNA purification magnetic beads into the PCR tube and vortex to mix.

[0175] 42. Incubate at room temperature for 5 minutes.

[0176] 43. Centrifuge the PCR tube briefly, place it on a magnetic rack and let it stand for 2 minutes until the liquid is clear. Carefully aspirate the supernatant and discard it.

[0177] 44. Keep the PCR tube fixed on the magnetic rack, add 150 μL of 80% ethanol to rinse the magnetic beads and tube walls, let stand for 30 seconds, carefully aspirate the supernatant and discard it.

[0178] 45. Repeat the previous step, and try to dry the liquid in the tube. If there is a small amount of liquid remaining on the tube wall, you can centrifuge the PCR tube briefly, separate it on a magnetic rack, and then use a small-range pipette to dry the liquid at the bottom of the tube.

[0179] 46. ​​Keep the PCR tube fixed on the magnetic rack, open the tube cap, and let it dry at room temperature until the surface of the magnetic beads is no longer reflective and cracked.

[0180] 47. Remove the PCR tube from the magnetic rack, add 20 μL of TE buffer for elution, and vortex to mix.

[0181] 48. Incubate at room temperature for 5 minutes.

[0182] 49. Centrifuge the PCR tube briefly, place it on a magnetic rack and let it stand for 2 minutes until the liquid is clear. Carefully aspirate 19 μL of the supernatant into a new PCR tube.

[0183] 50. Prepare the PCR reaction solution on ice according to Table 10.

[0184] Table 10

[0185] 51. Add 31 μL of the prepared PCR reaction solution to the PCR tube, vortex 3 times for 3 seconds each time, and then centrifuge briefly to collect the reaction solution to the bottom of the tube.

[0186] 52. Place the PCR tube on the PCR instrument and perform the PCR reaction according to the conditions in Table 11.

[0187] Table 11

[0188] 53. After the reaction is complete, the reaction solution is collected to the bottom of the tube by instantaneous centrifugation.

[0189] 54. Remove the DNA purification beads in advance and allow them to equilibrate at room temperature for at least 30 minutes. Shake thoroughly before use.

[0190] 55. Pipette 38 μL of DNA purification magnetic beads into a PCR tube and vortex to mix.

[0191] 56. Incubate at room temperature for 5 minutes.

[0192] 57. Centrifuge the PCR tube briefly, place it on a magnetic rack and let it stand for 2 minutes until the liquid is clear. Carefully aspirate the supernatant and discard it.

[0193] 58. Keep the PCR tube fixed on the magnetic rack, add 150 μL of 80% ethanol to rinse the magnetic beads and tube walls, let stand for 30 seconds, carefully aspirate the supernatant and discard it.

[0194] 59. Repeat the previous step, and try to aspirate as much liquid as possible from the tube. If a small amount of liquid remains on the tube wall, you can centrifuge the PCR tube briefly, separate it on a magnetic rack, and then use a small-range pipette to aspirate the liquid from the bottom of the tube.

[0195] 60. Keep the PCR tube fixed on the magnetic rack, open the tube cap, and let it dry at room temperature until the surface of the magnetic beads is no longer reflective and cracked.

[0196] 61. Remove the PCR tube from the magnetic rack, add 32 μL of TE buffer for elution, and vortex to mix.

[0197] 62. Incubate at room temperature for 5 minutes.

[0198] 63. Briefly centrifuge the PCR tube, place it on a magnetic rack and let it stand for 2 minutes until the liquid is clear. Carefully aspirate 30 μL of the supernatant into a new PCR tube.

[0199] 64. Use The dsDNA HS Assay Kit was used to quantify the PCR product. 360 ng of the PCR product was transferred to a new 0.2 mL PCR tube and the volume was increased to 48 μL with TE Buffer.

[0200] 65. Place the PCR tube on the PCR instrument and perform the reaction according to the conditions in Table 12.

[0201] Table 12

[0202] 66. After the reaction is complete, immediately place the PCR tube on ice and let it stand for 2 minutes before adding the single-stranded cyclization reaction solution.

[0203] 67. Based on the reaction number, prepare the single-chain cyclization reaction solution on ice according to the formula in Table 13.

[0204] Table 13

[0205] 68. Use a pipette to add 12.1 μL of the prepared single-stranded circularization reaction solution to a PCR tube, vortex 3 times for 3 seconds each time, and then centrifuge briefly to collect the reaction solution to the bottom of the tube.

[0206] 69. Place the PCR tube on the PCR instrument and perform the reaction according to the conditions in Table 14.

[0207] Table 14

[0208] 70. After the reaction is complete, centrifuge the PCR tube briefly and place it on ice, then immediately proceed to the next step of the reaction.

[0209] 71. Prepare the enzyme digestion reaction solution on ice in advance according to the formula in Table 15.

[0210] Table 15

[0211] 72. Use a pipette to add 4 μL of the prepared enzyme digestion reaction solution to the PCR tube, vortex 3 times for 3 seconds each time, and then centrifuge briefly to collect the reaction solution to the bottom of the tube.

[0212] 73. Place the PCR tube on the PCR instrument and perform the reaction according to the conditions in Table 16.

[0213] Table 16

[0214] 74. After the reaction is complete, centrifuge briefly to collect the reaction solution to the bottom of the tube.

[0215] 75. Immediately add 7.5 μL of Digestion Stop Buffer to the PCR tube, vortex 3 times for 3 seconds each time, and then briefly centrifuge to collect the reaction solution to the bottom of the tube. Transfer all the reaction solution to a new 1.5 mL centrifuge tube.

[0216] 76. Remove the DNA for purification in advance, equilibrate at room temperature for at least 30 minutes, and shake well before use.

[0217] 77. Pipette 170 μL of DNA purification magnetic beads into a centrifuge tube. Gently pipette at least 10 times until all the magnetic beads are suspended. On the last pipette, make sure that all liquid in the pipette tip and the magnetic beads are transferred into the centrifuge tube.

[0218] 78. Incubate at room temperature for 5 minutes.

[0219] 79. Centrifuge the centrifuge tube briefly, place it on a magnetic rack, let it stand for 2 minutes until the liquid is clear, carefully aspirate the supernatant with a pipette and discard it.

[0220] 80. Keep the centrifuge tubes on the magnetic rack, add 200 μL of freshly prepared 80% ethanol to rinse the magnetic beads and tube walls, let stand for 30 seconds, then carefully aspirate and discard the supernatant.

[0221] 81. Repeat the previous step, trying to remove as much liquid as possible from the tube. If a small amount remains on the tube wall, centrifuge the tube briefly. After separating the tubes on a magnetic rack, use a small-capacity pipette to remove the liquid from the bottom of the tube.

[0222] 82. Keep the centrifuge tubes on the magnetic rack, open the centrifuge tube caps, and allow them to dry at room temperature until the surface of the magnetic beads is no longer reflective.

[0223] 83. Remove the centrifuge tube from the magnetic rack, add 22 μL of TE buffer for elution, and gently pipette at least 10 times until all magnetic beads are suspended.

[0224] 84. Incubate at room temperature for 5 minutes.

[0225] 85. Centrifuge the centrifuge tube briefly, place it on a magnetic rack, and let it stand for 2 minutes until the liquid is clear. Use a pipette to transfer 20 μL of the supernatant to a new 1.5 mL centrifuge tube.

[0226] 86. Use The ssDNA Assay Kit is a real-time fluorescence reagent kit used to quantify the product after enzyme digestion and purification, following the instructions of the kit.

[0227] 87. Use the MGISEQ-2000 sequencer for sequencing. Use the MGISEQ-2000 high-throughput sequencing reagent kit (PE100) (MGI, catalog number 1000012536) to prepare and sequence DNB according to the instructions.

[0228] 88. After sequencing is completed, the sequencing data is filtered and aligned. Base combinations in the reads that do not align with the reference genome are used as molecular tags. Reads with the same combination are repeatedly overlapped and assembled to eventually produce a full-length transcript. Different read combinations indicate that they do not originate from the same mRNA molecule. The number of different mRNA combinations (representing the number of transcripts) is counted, and then the gene and transcript types, lengths, and expression levels are calculated.

[0229] Experimental results:

[0230] 1) The database construction results are shown in Table 17.

[0231] Table 17

[0232] 2) The analysis results are shown in Table 18.

[0233] Table 18

[0234] The results showed that more than 20,000 genes could be detected in the UHRR samples, with an average of more than 3 mRNA isoforms detected for each gene (Figure 5). Furthermore, the average length of the detected transcripts was over 3kb, with the majority concentrated around 1500bp, consistent with the distribution of mRNA (Figure 6).

[0235] Example 2: Single-cell library construction and sequencing of human 293T cell line

[0236] 1. Prepare cell lysis buffer according to Table 19 below.

[0237] Table 19

[0238] Note: The RNase inhibitor Murine is from Novozymes, catalog number R301-01; the Triton X-100 solution is from Sigma, catalog number 93443.

[0239] 2. Prepare a 200μL PCR tube, add 2μL of cell lysis buffer and 1μL of 10mM oligo dT primer to the PCR tube.

[0240] oligo dT primer: AAGGAGTGGTATGAATGGAGAGTAGTTTTTTTTTTTTTTTTTTTVN (SEQ ID NO: 1).

[0241] 3. Use a BD flow cytometer to sort human 293T cell line (Pronosai, CL-0005) single cells into the 200μL PCR tube prepared in the previous step.

[0242] 4. After centrifuging the PCR tube, place it in a PCR instrument and incubate at 72°C for 3 minutes. Immediately after incubation, place the PCR tube in an ice box to cool, allowing the oligo dT primers to hybridize with the polyA tail of the mRNA.

[0243] 5. Prepare the RT (reverse transcription) reaction mixture according to the formula in Table 20;

[0244] Table 20

[0245] Note: In the 25mM dATP / dTTP / dCTP / dGTP / dITP mixture: dATP, dTTP, dCTP, dGTP, dITP, and dITP each account for 20% of the total molar composition, and are prepared according to the ratio of 100mM dATP:100mM dTTP:100mM dCTP:100mM dGTP:100mM dGTP:100mM dITP = 0.8:0.8:0.8:0.8:0.8.

[0246] The dNTP Set Solution (dATP, dCTP, dTTP, dGTP, 100mM each) is from Yisheng, catalog number 10122ES74; the dITP solution (100mM) is from Thermo Fisher, catalog number R1191.

[0247] The reverse transcriptase Hiscript III is from Novozymes, catalog number R302-01, and contains 5x Hiscript III Buffer;

[0248] The RNase inhibitor Murine is from Novozymes, catalog number R301-01;

[0249] Betaine is from Sigma, catalog number B0300; 1M DTT is from Sigma, 646563, diluted with water to 100mM.

[0250] Magnesium chloride 1M MgCl2 is from Thermo Fisher, item number AM9530G.

[0251] TSO: AAGGAGTGGTATGAATGGAGAGTAGAT / rG / rG / iXNA_G (SEQ ID NO: 2);

[0252] rG represents ribonucleic acid modification; iXNA_G represents locked nucleic acid modification.

[0253] 6. Add the RT reaction mixture to the PCR tube and centrifuge briefly.

[0254] 7. Place the PCR tube on the PCR instrument and perform the reaction according to the procedure in Table 21.

[0255] Table 21

[0256] 8. After the reaction is complete, briefly centrifuge to collect the liquid in the PCR tube to the bottom of the tube.

[0257] 9. Prepare the cDNA amplification reaction solution according to Table 22.

[0258] Table 22

[0259] Note: Phanta Uc ultra-fidelity DNA polymerase is from Novozymes, catalog number P507-01, and contains 5x Uc buffer;

[0260] 25mM dNTPs are from ENZYMATICS, item number N2050-25;

[0261] TSO primer: AAGGAGTGGTATGAATGGAGAGTAG (SEQ ID NO: 3).

[0262] 10. Add the cDNA amplification reaction solution to the PCR tube and centrifuge briefly.

[0263] 11. Place the PCR tube on the PCR instrument and perform the extension reaction according to the procedure in Table 23.

[0264] Table 23

[0265] 12. After the extension reaction is complete, briefly centrifuge to collect the liquid in the PCR tube to the bottom of the tube.

[0266] 13. Add 1 μL of hAAG enzyme (NEB, M0313S) to the PCR tube and centrifuge briefly.

[0267] 14. Place the PCR tubes on the PCR instrument and perform digestion and PCR reaction according to the procedure in Table 24 (PCR reaction system is the same as in Table 3).

[0268] Table 24

[0269] 15. After the reaction is complete, briefly centrifuge to collect the liquid in the PCR tube to the bottom of the tube.

[0270] 16. Add 30 μL of DNA purification magnetic beads to the PCR tube, pipette and mix thoroughly 10 times, then incubate at room temperature for 10 min.

[0271] 17. After a brief centrifugation, place the PCR tube on a magnetic rack and let it stand for 5 minutes until the liquid is clear. Use a pipette to aspirate and discard the supernatant.

[0272] 18. Add 150 μL of 75% ethanol to the tube, let stand for 30 seconds, and wash the magnetic beads.

[0273] 19. Repeat the previous step, trying to aspirate as much liquid as possible from the tube. If a small amount of liquid remains on the tube wall, the PCR tube can be centrifuged briefly. After separation on a magnetic rack, use a small-capacity pipette to aspirate the liquid from the bottom of the tube.

[0274] 20. Keep the PCR tube fixed on the magnetic rack, open the tube cap, and let it dry at room temperature until the surface of the magnetic beads is no longer reflective and cracked.

[0275] 21. Remove the PCR tube from the magnetic rack, add 32 μL of molecular-grade water to elute the DNA, and gently pipette 10 times until completely mixed.

[0276] 22. Incubate at room temperature for 5 minutes.

[0277] 23. Centrifuge briefly, place the PCR tube on a magnetic rack, let stand for 5 minutes until the liquid is clear, and transfer 30 μL of supernatant to a new PCR tube.

[0278] 24. Use The dsDNA HS Assay Kit was used to quantify the PCR product. 100 ng of the PCR product was transferred to a new PCR tube and the volume was increased to 45 μL with TE buffer.

[0279] 25. Use the reagents in the MGI Easiy Fast DNA Enzyme Digestion Library Preparation Kit (940-001193-00, MGI) to construct the library.

[0280] 26. Prepare the end-of-life repair and A reaction solution on ice according to Table 25.

[0281] Table 25

[0282] 27. Add the prepared end-fracture repair reaction solution to the PCR tube, vortex 3 times for 3 seconds each time, and then centrifuge briefly to collect the reaction solution to the bottom of the tube.

[0283] 28. When the temperature of the PCR instrument drops to 4℃, place the PCR tube on the PCR instrument and perform the break-up and end-repair A reaction according to the conditions in Table 26.

[0284] Table 26

[0285] 29. After the reaction is complete, briefly centrifuge to collect the reaction solution to the bottom of the tube.

[0286] 30. Remove the DNA purification beads in advance and allow them to equilibrate at room temperature for at least 30 minutes. Shake well before use.

[0287] 31. Pipette 48 μL of magnetic beads into a PCR tube and vortex to mix.

[0288] 32. Incubate at room temperature for 5 minutes.

[0289] 33. Briefly centrifuge the PCR tube, place it on a magnetic rack and let it stand for 2 minutes until the liquid is clear. Carefully aspirate the supernatant and discard it. Aspirate as much liquid as possible from the tube. If a small amount of liquid remains on the tube wall, briefly centrifuge the tube again, separate it on a magnetic rack, and then use a small-range pipette to aspirate the liquid from the bottom of the tube.

[0290] 34. Remove the PCR tube from the magnetic rack, add 45 μL of TE buffer to elute the DNA, vortex to mix, and then centrifuge briefly.

[0291] 35. Add 5 μL of adapter (from the MGIEAsy Fast DNA Enzyme Digestion Library Preparation Kit) to the PCR tube, vortex 3 times for 3 seconds each time, and then briefly centrifuge to collect the reaction solution to the bottom of the tube.

[0292] 36. Prepare the connector connection reaction solution on ice according to the formula in Table 27.

[0293] Table 27

[0294] 37. Add the prepared adapter ligation reaction solution to the PCR tube, vortex 6 times for 3 seconds each time, and then centrifuge briefly to collect the reaction solution to the bottom of the tube.

[0295] 38. Place the PCR tubes above on the PCR instrument and perform the adapter ligation reaction according to the conditions in Table 28.

[0296] Table 28

[0297] 39. After the reaction is complete, the reaction solution is collected to the bottom of the tube by instantaneous centrifugation.

[0298] 40. Remove the DNA purification beads in advance and allow them to equilibrate at room temperature for at least 30 minutes. Shake well before use.

[0299] 41. After aspirating 22 μL of TE buffer into the PCR tube, aspirate 20 μL of DNA purification magnetic beads into the PCR tube and vortex to mix.

[0300] 42. Incubate at room temperature for 5 minutes.

[0301] 43. Centrifuge the PCR tube briefly, place it on a magnetic rack and let it stand for 2 minutes until the liquid is clear. Carefully aspirate the supernatant and discard it.

[0302] 44. Keep the PCR tube fixed on the magnetic rack, add 150 μL of 80% ethanol to rinse the magnetic beads and tube walls, let stand for 30 seconds, carefully aspirate the supernatant and discard it.

[0303] 45. Repeat the previous step, and try to dry the liquid in the tube. If there is a small amount of liquid remaining on the tube wall, you can centrifuge the PCR tube briefly, separate it on a magnetic rack, and then use a small-range pipette to dry the liquid at the bottom of the tube.

[0304] 46. ​​Keep the PCR tube fixed on the magnetic rack, open the tube cap, and let it dry at room temperature until the surface of the magnetic beads is no longer reflective and cracked.

[0305] 47. Remove the PCR tube from the magnetic rack, add 20 μL of TE buffer for elution, and vortex to mix.

[0306] 48. Incubate at room temperature for 5 minutes.

[0307] 49. Centrifuge the PCR tube briefly, place it on a magnetic rack and let it stand for 2 minutes until the liquid is clear. Carefully aspirate 19 μL of the supernatant into a new PCR tube.

[0308] 50. Prepare the PCR reaction solution on ice according to Table 29.

[0309] Table 29

[0310] 51. Add 31 μL of the prepared PCR reaction solution to the PCR tube, vortex 3 times for 3 seconds each time, and then centrifuge briefly to collect the reaction solution to the bottom of the tube.

[0311] 52. Place the PCR tube on the PCR instrument and perform the PCR reaction according to the conditions in Table 30.

[0312] Table 30

[0313] 53. After the reaction is complete, the reaction solution is collected to the bottom of the tube by instantaneous centrifugation.

[0314] 54. Remove the DNA purification beads in advance and allow them to equilibrate at room temperature for at least 30 minutes. Shake thoroughly before use.

[0315] 55. Pipette 38 μL of DNA purification magnetic beads into a PCR tube and vortex to mix.

[0316] 56. Incubate at room temperature for 5 minutes.

[0317] 57. Centrifuge the PCR tube briefly, place it on a magnetic rack and let it stand for 2 minutes until the liquid is clear. Carefully aspirate the supernatant and discard it.

[0318] 58. Keep the PCR tube fixed on the magnetic rack, add 150 μL of 80% ethanol to rinse the magnetic beads and tube walls, let stand for 30 seconds, carefully aspirate the supernatant and discard it.

[0319] 59. Repeat the previous step, and try to aspirate as much liquid as possible from the tube. If a small amount of liquid remains on the tube wall, you can centrifuge the PCR tube briefly, separate it on a magnetic rack, and then use a small-range pipette to aspirate the liquid from the bottom of the tube.

[0320] 60. Keep the PCR tube fixed on the magnetic rack, open the tube cap, and let it dry at room temperature until the surface of the magnetic beads is no longer reflective and cracked.

[0321] 61. Remove the PCR tube from the magnetic rack, add 32 μL of TE buffer for elution, and vortex to mix.

[0322] 62. Incubate at room temperature for 5 minutes.

[0323] 63. Briefly centrifuge the PCR tube, place it on a magnetic rack and let it stand for 2 minutes until the liquid is clear. Carefully aspirate 30 μL of the supernatant into a new PCR tube.

[0324] 64. Use The dsDNA HS Assay Kit was used to quantify the PCR product. 360 ng of the PCR product was transferred to a new 0.2 mL PCR tube and the volume was increased to 48 μL with TE Buffer.

[0325] 65. Place the PCR tube on the PCR instrument and perform the reaction according to the conditions in Table 31.

[0326] Table 31

[0327] 66. After the reaction is complete, immediately place the PCR tube on ice and let it stand for 2 minutes before adding the single-stranded cyclization reaction solution.

[0328] 67. Based on the reaction number, prepare the single-chain cyclization reaction solution on ice according to the formula in Table 32.

[0329] Table 32

[0330] 68. Use a pipette to add 12.1 μL of the prepared single-stranded circularization reaction solution to a PCR tube, vortex 3 times for 3 seconds each time, and then centrifuge briefly to collect the reaction solution to the bottom of the tube.

[0331] 69. Place the PCR tube on the PCR instrument and perform the reaction according to the conditions in Table 33.

[0332] Table 33

[0333] 70. After the reaction is complete, centrifuge the PCR tube briefly and place it on ice, then immediately proceed to the next step of the reaction.

[0334] 71. Prepare the enzyme digestion reaction solution on ice in advance according to the formula in Table 34.

[0335] Table 34

[0336] 72. Use a pipette to add 4 μL of the prepared enzyme digestion reaction solution to the PCR tube, vortex 3 times for 3 seconds each time, and then centrifuge briefly to collect the reaction solution to the bottom of the tube.

[0337] 73. Place the PCR tube on the PCR instrument and perform the reaction according to the conditions in Table 35.

[0338] Table 35

[0339] 74. After the reaction is complete, centrifuge briefly to collect the reaction solution to the bottom of the tube.

[0340] 75. Immediately add 7.5 μL of Digestion Stop Buffer to the PCR tube, vortex 3 times for 3 seconds each time, and then briefly centrifuge to collect the reaction solution to the bottom of the tube. Transfer all the reaction solution to a new 1.5 mL centrifuge tube.

[0341] 76. Remove the DNA for purification in advance, equilibrate at room temperature for at least 30 minutes, and shake well before use.

[0342] 77. Pipette 170 μL of DNA purification magnetic beads into a centrifuge tube. Gently pipette at least 10 times until all the magnetic beads are suspended. On the last pipette, make sure that all liquid in the pipette tip and the magnetic beads are transferred into the centrifuge tube.

[0343] 78. Incubate at room temperature for 5 minutes.

[0344] 79. Centrifuge the centrifuge tube briefly, place it on a magnetic rack, let it stand for 2 minutes until the liquid is clear, carefully aspirate the supernatant with a pipette and discard it.

[0345] 80. Keep the centrifuge tubes on the magnetic rack, add 200 μL of freshly prepared 80% ethanol to rinse the magnetic beads and tube walls, let stand for 30 seconds, then carefully aspirate and discard the supernatant.

[0346] 81. Repeat the previous step, trying to remove as much liquid as possible from the tube. If a small amount remains on the tube wall, centrifuge the tube briefly. After separating the tubes on a magnetic rack, use a small-capacity pipette to remove the liquid from the bottom of the tube.

[0347] 82. Keep the centrifuge tubes on the magnetic rack, open the centrifuge tube caps, and allow them to dry at room temperature until the surface of the magnetic beads is no longer reflective.

[0348] 83. Remove the centrifuge tube from the magnetic rack, add 22 μL of TE buffer for elution, and gently pipette at least 10 times until all magnetic beads are suspended.

[0349] 84. Incubate at room temperature for 5 minutes.

[0350] 85. Centrifuge the centrifuge tube briefly, place it on a magnetic rack, and let it stand for 2 minutes until the liquid is clear. Use a pipette to transfer 20 μL of the supernatant to a new 1.5 mL centrifuge tube.

[0351] 86. Use The ssDNA Assay Kit is a real-time fluorescence reagent kit used to quantify the product after enzyme digestion and purification, following the instructions of the kit.

[0352] 87. Use the MGISEQ-2000 sequencer for sequencing. Use the MGISEQ-2000 high-throughput sequencing reagent kit (PE100) (MGI, catalog number 1000012536) to prepare and sequence DNB according to the instructions.

[0353] 88. After sequencing is completed, the sequencing data is filtered and aligned. Base combinations in the reads that do not align with the reference genome are used as molecular tags. Reads with the same combination are repeatedly overlapped and assembled to eventually produce a full-length transcript. Different read combinations indicate that they do not originate from the same mRNA molecule. The number of different mRNA combinations (representing the number of transcripts) is counted, and then the gene and transcript types, lengths, and expression levels are calculated.

[0354] Experimental results:

[0355] 1) The database construction results are shown in Table 36.

[0356] Table 36

[0357] 2) The analysis results are shown in Table 37.

[0358] Table 37

[0359] The results showed that more than 10,000 genes could be detected in the 293T-1 sample, with an average of 2.6 mRNA isoforms detected per gene and an average length of more than 1,500.

[0360] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0361] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for obtaining a full-length transcript sequence, characterized in that, include: The mRNA to be tested was reverse transcribed. The reverse transcription product was subjected to two-strand synthesis. The product of the two-strand synthesis process is digested in the presence of a nuclease, which specifically recognizes nucleic acid fragments with universal bases. The digestion products were processed by polymerase chain reaction to obtain sequencing libraries. as well as The sequencing library is sequenced to obtain the full-length sequence of the mRNA to be tested; The reverse transcription process or the two-strand synthesis process is carried out under conditions where universal bases are present.

2. The method of claim 1, wherein, The full-length transcript sequence was derived from multiple single-cell samples.

3. The method of claim 2, wherein, Cellular tags are added to mRNA in single cells using microdroplet, micropore, or chip technologies.

4. The method of claim 1, wherein, The universal base includes at least one selected from dITP, 8-Oxo-dGTP, dPTP, dKTP, and 2OH-dATP; preferably dITP.

5. The method of claim 1, wherein, The reverse transcription process is performed in the presence of universal bases, and the reverse transcription process, the double-strand synthesis process, and the digestion process are performed in the following manner: The mRNA to be tested was reverse transcribed in the presence of dNTPs and universal bases. The reverse transcription product was subjected to two-strand synthesis in the presence of dNTPs and the absence of universal bases. The product of the two-strand synthesis process is digested in the presence of a nuclease that specifically recognizes nucleic acid fragments with universal bases.

6. The method of claim 1, wherein, The double-strand synthesis process is performed under conditions where universal bases are present, and the reverse transcription process, the double-strand synthesis process, and the digestion process are performed in the following manner: The mRNA to be tested was reverse transcribed in the presence of dNTPs and in the absence of universal bases. The reverse transcription primers included dU bases. The reverse transcription product was subjected to second-strand synthesis in the presence of dNTPs and universal bases. The product after the two-strand synthesis was digested in the presence of a nuclease, which specifically recognizes nucleic acid fragments containing dU bases. The digestion product was subjected to triplet synthesis in the presence of dNTPs and the absence of universal bases. The product after triple-strand synthesis is digested in the presence of a nuclease that specifically recognizes nucleic acid fragments with universal bases.

7. The method according to claim 5 or 6, characterized in that, The reverse transcription processing system further includes a reverse transcriptase, which includes at least one selected from MMLV reverse transcriptase and AMV reverse transcriptase.

8. The method of claim 5, wherein, In the reverse transcription treatment system, the molar ratio of dNTPs to universal bases is 9:1 to 1:9, preferably 4:1; Optionally, the dNTPs include at least one of dATP, dTTP, dCTP, and dGTP.

9. The method of claim 8, wherein, The molar ratio of dATP, dTTP, dCTP, dGTP and universal bases is 2.25:2.25:2.25:2.25:1 to 0.25:0.25:0.25:0.25:9; preferably 1:1:1:1:

1.

10. The method of claim 6, wherein, The two-chain synthesis treatment is carried out in a two-chain synthesis treatment system, wherein the molar ratio of the dNTP to the universal base in the two-chain synthesis treatment system is 9:1 to 1:9, preferably 4:1; Optionally, the dNTPs include at least one of dATP, dTTP, dCTP, and dGTP.

11. The method of claim 10, wherein, The molar ratio of dATP, dTTP, dCTP, dGTP and the universal base is 2.25:2.25:2.25:2.25:1 to 0.25:0.25:0.25:0.25:

9.

12. The method of claim 11, wherein, The molar ratio of dATP, dTTP, dCTP, dGTP and the universal base is 1:1:1:1:

1.

13. The method of claim 5, wherein, The two-strand synthesis process is carried out in a two-strand synthesis process system, which further includes a first DNA polymerase, which has the activity of amplifying nucleic acid templates containing the universal bases.

14. The method of claim 6, wherein, The two-strand synthesis process is carried out in a two-strand synthesis process system, which further includes a second DNA polymerase.

15. The method of claim 1, wherein, The full-length sequence of the mRNA to be tested was obtained in the following manner: The sequencing data obtained from the sequencing process is filtered and compared to identify unique molecular tags composed of single nucleotide variants; The full-length sequence of the mRNA to be tested is assembled based on overlapping information containing the unique molecular tag reads.

16. The method of claim 6, wherein, The triple-strand synthesis process is carried out in a triple-strand synthesis system, which further includes a third DNA polymerase that has the activity of amplifying nucleic acid templates containing the universal bases.

17. The method of claim 1, wherein, The endonuclease includes at least one selected from USER enzyme, UDG enzyme, endonuclease VIII, hSMUG1 enzyme, hAAG enzyme and Fpg enzyme, preferably hAAG enzyme.

18. The method of claim 1, wherein, The polymerase chain reaction treatment was carried out under conditions in which the fourth DNA polymerase, dNTPs, and the universal bases were absent.

19. The method of claim 1 or 5 or 6, wherein, The reverse transcription processing system further includes template-converting oligonucleotides.

20. A method of single-cell full-length transcriptome detection, comprising: include: Based on the method described in any one of claims 1 to 19, obtain the full-length transcript sequence of the single cell to be tested; Based on the cell tag sequences in the full-length transcriptome sequences of the single cells to be tested, the full-length transcriptome sequences of single cells from different sources are determined.

21. The method of claim 20, wherein, Further includes: Based on the UMI sequence in the full-length transcript sequence, the copy number of different genes in a single cell is obtained.

22. The method of claim 20, wherein, Further includes: Based on the comprehensive analysis of the UMI sequence and cell tag sequence in the full-length transcriptome, single-cell full-length transcriptome information is obtained.

23. A full-length transcriptome detection kit, characterized in that, include: At least one of the following: dNTP, universal base, dU base, endonuclease, reverse transcriptase, DNA polymerase, reverse transcription primer, TSO primer, and oligonucleotide primer containing poly T.