Template conversion oligonucleotide, kit and application thereof

The library was constructed by template conversion of oligonucleotide reverse transcription and barcode probes, which solved the sequencing problem caused by non-specific binding of TSO, achieved efficient single-cell and spatio-temporal and spatial-computing research, and improved data utilization and sequencing accuracy.

CN120400136AInactive Publication Date: 2025-08-01BGI RESEARCH HANGZHOU
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510898779.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing single-cell and spatio-temporal omics studies, the ineffective sequence problems caused by nonspecific binding of TSO affect sequencing accuracy and flux, and traditional biotin labeling methods are cumbersome and prone to amplification preference problems.

Method used

Template conversion oligonucleotides were used for reverse transcription, and non-specific binding was prevented by modification at the 3' end, simplifying the library construction process, avoiding amplification preferences, and using barcode probes and amplification linker regions to build the library.

Benefits of technology

It improves the data utilization rate of single-molecule sequencing and target gene capture sensitivity, reduces background interference, is compatible with long-term, short-read, long-term sequencing platforms, and improves the accuracy and throughput of sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120400136A_ABST
    Figure CN120400136A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of molecular biology, and particularly relates to template conversion oligonucleotide, a kit and application of the template conversion oligonucleotide. In the first aspect of the invention, a template conversion oligonucleotide is provided, the template conversion oligonucleotide comprises a template conversion region at the 3'end, the template conversion region comprises a plurality of nucleotides, the nucleotides comprise sugars and basic groups, and the basic groups are guanine; wherein in the plurality of nucleotides, the 3 '-hydroxyl site of the glycosyl of the nucleotide at the 3'-terminal is provided with a modification for preventing the 3 '-hydroxyl from forming a phosphodiester bond. TSO of conventional reverse transcription is specifically modified, then conventional reverse transcription processes of single cells, space-time omics and the like are carried out, additional enrichment operation is not needed, and the operation process is simple. Not only is the data utilization rate of single molecule sequencing improved, but also the sensitivity of target gene capture is improved, and background interference is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of molecular biology, and particularly relates to a template-switching oligonucleotide, a kit and its application. Background Art

[0002] The rapid development of single-cell and spatio-temporal omics technologies has provided powerful tools for in-depth analysis of cell heterogeneity and its spatial distribution characteristics in tissues. Nevertheless, most current large-scale omics studies mainly rely on gene expression data. Although next-generation sequencing (NGS) technology is favored for its low error rate and economy, its short read length limits the ability to obtain full-length gene transcript information. Although the micro-well sequencing technology of the Smart-seq series can achieve full-length transcriptome sequencing at the single-cell level, its throughput limitation and dependence on short-fragment splicing may lead to inaccurate identification of duplicate copy variations in the gene coding region. The progress of single-molecule sequencing technology has brought new opportunities for capturing full-length transcript information.

[0003] However, single-molecule sequencing technology also faces certain challenges in terms of accuracy and sequencing throughput, and these challenges can be traced back to the invalid sequences in existing single-cell and spatio-temporal omics libraries. A large part of these invalid sequences comes from the non-specific binding of TSO to non-target nucleic acid molecules. Refer to Figure 1 , and the combined TSO serves as a primer to initiate the transcription of non-target fragments, resulting in non-target products with TSO sequences at both ends. On NGS platforms where the throughput is not restricted, this problem is often overlooked, but on single-molecule sequencing platforms with limited throughput, the impact of this problem is amplified. On the other hand, due to the existence of TSO non-specific sequences, it may cause amplification bias, resulting in deviation or masking of target fragment capture, affecting the interpretation of data by researchers. Therefore, although the library construction process based on short reads can remove this part of the read length to a certain extent, the bias introduced by cDNA amplification in the early stage is still inevitable.

[0004] In order to alleviate the bottleneck problem of long-read throughput, in current single-cell library construction processes based on PacBio and ONT, biotin labeling of the target product is generally used, and then streptavidin is used to capture the target product to solve the influence of TSO non-specific sequences. However, this scheme requires biotin labeling of DNA products, magnetic bead capture enrichment, PCR amplification or enzymatic digestion to enrich the target product, making the whole process cumbersome. In addition, releasing the target product by PCR amplification is likely to cause the problem of amplification bias brought by over-amplification, thus affecting the amplification of the target fragment or causing the loss of true signals. Summary of the Invention

[0005] The present invention provides a template-switching oligonucleotide, a kit and its application. Using this template-switching oligonucleotide, no labeling is required, the whole process is simpler, and the problem of missing true signals caused by amplification bias will not occur.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows: The first aspect of the present invention provides a template-switching oligonucleotide, which includes: a template-switching region located at the 3'-end of the template-switching oligonucleotide, and an amplification adapter region located in the 5'-direction of the template-switching region; The length of the template-switching region is 2 to 5 nucleotides; the 3'-terminal nucleotide of the template-switching region has a 3'-terminal modification, and the 3'-terminal modification is used to prevent the 3'-terminal nucleotide of the template-switching region from forming a phosphodiester bond with the free 5'-phosphate group of deoxynucleotides and / or deoxynucleotide sequences.

[0007] In some embodiments of the present invention, the 3'-terminal nucleotide with a 3'-terminal modification refers to a locked nucleic acid, or a 3'-terminal nucleotide whose 3'-terminal hydroxyl group is replaced by an -O-R or -R' group.

[0008] Wherein, R is any one of an alkyl group with 1 to 18 carbon atoms, an alkenyl group with 2 to 18 carbon atoms, an alkynyl group with 2 to 18 carbon atoms, an aryl group with less than 18 carbon atoms, an aralkyl group with less than 18 carbon atoms, a heteroaryl group with less than 18 carbon atoms, a phosphate group, an azidomethyl group, an allyl group, a nitrobenzyl group, a carbamate group, a thiophosphate group; R' is any one of H, an alkyl group with 1 to 18 carbon atoms, an alkenyl group with 2 to 18 carbon atoms, an alkynyl group with 2 to 18 carbon atoms, an aryl group with less than 18 carbon atoms, an aralkyl group with less than 18 carbon atoms, a heteroaryl group with less than 18 carbon atoms.

[0009] In some embodiments of the present invention, the template-switching region consists of 3 ribonucleotides.

[0010] In some embodiments of the present invention, the sequence of the template-switching region is rGrGrG.

[0011] The second aspect of the present invention provides a kit, which includes the aforementioned template-switching oligonucleotide.

[0012] In some embodiments of the present invention, the kit further includes at least one of reverse transcriptase, reverse transcription primer, DNA polymerase, RNase inhibitor, polyethylene glycol, betaine, dNTP and solid-phase carrier.

[0013] Among them, a barcode probe is directly or indirectly immobilized on the solid-phase carrier; the barcode probe includes, from the 5'-end to the 3'-end: an amplification adapter region, a barcode sequence, and a capture sequence; the barcode sequence is a spatial barcode or a cellular barcode, and the capture sequence is complementary to a partial sequence of the RNA from the sample to be tested.

[0014] In a third aspect of the present invention, there is provided a method for constructing a spatial transcriptome sequencing library or a single-cell transcriptome sequencing library by using the aforementioned template-switching oligonucleotide.

[0015] In some embodiments of the present invention, the method includes the following steps: (1) Using the barcode probe as a primer, performing a reverse transcription reaction with the RNA in the tissue or cell as a template to obtain cDNA; among them, the barcode probe is directly or indirectly immobilized on the solid-phase carrier, and includes, from the 5'-end to the 3'-end: an amplification adapter region', a barcode sequence, and a capture sequence; the barcode sequence is a spatial barcode or a cellular barcode, and the capture sequence is complementary to a partial sequence of the RNA from the sample to be tested; (2) Introducing a non-template polymerization sequence with a length of 2 to 5 nucleotides at the 3'-end of the cDNA to obtain an extended product; (3) Hybridizing the non-template polymerization sequence of the extended product with the template-switching region of the aforementioned template-switching oligonucleotide, and continuing to extend with the template-switching oligonucleotide as a template to obtain a library starting molecule; the library starting molecule includes, from the 5'-end to the 3'-end in sequence: an amplification adapter region', a barcode sequence, a capture sequence, a cDNA sequence, and a complementary sequence of the amplification adapter region; (4) Amplifying or enriching the library starting molecule by using a first library amplification primer and a second library amplification primer to obtain the spatial transcriptome sequencing library or the single-cell transcriptome sequencing library; among them, the 3'-sequence of the first library amplification primer is complementary to the complementary sequence of the amplification adapter region, and the 3'-sequence of the second library amplification primer is complementary to the complementary sequence of the amplification adapter region'.

[0016] In a fourth aspect of the present invention, there is provided a sequencing method, including the step of sequencing the spatial transcriptome library or the single-cell transcriptome library obtained by the aforementioned method.

[0017] In some embodiments of the present invention, the sequencing is full-length sequencing performed by using a single-molecule sequencing method.

[0018] In some embodiments, for the library constructed by selective capture at the 3' end, the utilization rate of valid data in the library constructed by the above method is more significantly improved. Specifically, taking single-cell transcriptome sequencing as an example, for the method of constructing a library by selective capture from the 3' end, the construction of the library can refer to CN112005115A.

[0019] The beneficial effects of the present invention are as follows: By specifically modifying the TSO of conventional reverse transcription, and then performing the reverse transcription process of conventional single-cell transcriptome, spatial-temporal transcriptomics, and transcriptome (bulk RNA-seq), no additional enrichment operation is required, and the operation process is simple. The potential value lies in not only helping to improve the data utilization rate of single-molecule sequencing, but also helping to improve the sensitivity of target gene capture and reduce background interference.

[0020] At the same time, it is compatible with long and short read length sequencing platforms, enabling accurate analysis of complex transcripts and single-base level sequencing. In the future, it is expected to become a conventional library construction kit in the fields of NGS and single-molecule sequencing. The present invention uses a blocked TSO sequence, which can not only reduce the influence of amplification bias existing in both single-molecule long read length sequencing (i.e., third-generation sequencing) and short read length sequencing (i.e., second-generation sequencing), but also improve the utilization rate of valid data in long read length sequencing. It greatly improves the utilization rate of single-cell omics data, provides high-quality data support for in-depth understanding of cell states and functions, and promotes the development of precision medicine. In addition, transcriptome analysis combined with spatial-temporal information can also reveal the dynamic change rules of transcripts in different cell types and their microenvironments, opening up a new technical path for spatial-temporal omics research. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram of the non-specific initiation of TSO.

[0022] Figure 2 It is the comparison result of the data utilization rate after constructing a single-molecule sequencing library for different cell lines using different template-switching nucleotides in the embodiments of the present invention. Among them, C represents the existing conventional TSO (control group), P represents that the 3'-OH of the ribose of the last nucleotide at the 3' end of the conventional TSO is replaced with 3'-P (phosphate group modification), H represents that 3'-OH is replaced with 3'-H (deoxy treatment), and C3 represents that 3'-OH is replaced with 3'-C3 Spacer (alkyl chain C3 modification). DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] In the description of the present invention, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0024] In the first aspect of the present invention, there is provided a template-switching oligonucleotide, which comprises: a template-switching region located at the 3'-end of the template-switching oligonucleotide, and an amplification adapter region located in the 5'-direction of the template-switching region.

[0025] Among them, the template-switching region comprises 2 to 5 nucleotides, for example, it can be 2, 3, 4, or 5 nucleotides. The 3'-terminal nucleotide of the template-switching region has a 3'-terminal modification, and the 3'-terminal modification is used to prevent the 3'-terminal nucleotide of the template-switching region from forming a phosphodiester bond with the free 5'-phosphate group of a deoxynucleotide and / or a deoxynucleotide sequence. It can be understood that the formation of a phosphodiester bond between the 3'-terminal nucleotide of the template-switching region and the free 5'-phosphate group can be an extension reaction mediated by a polymerase, a ligation reaction mediated by a ligase, etc. Thus, preventing the formation of the phosphodiester bond also prevents the polymerase-mediated extension reaction, the ligase-mediated ligation reaction, etc. Therefore, when the template-switching oligonucleotide non-specifically binds to a non-target nucleic acid molecule, it cannot serve as a primer to initiate the transcription of a non-target fragment, reducing the appearance of non-target products.

[0026] Nucleotides include sugars and bases. In some embodiments, the bases include adenine (A), cytosine (C), guanine (G), and thymine (T). In some embodiments, the bases also include one or more other types of bases such as uracil (U), hypoxanthine (I), etc. Among the multiple nucleotides in the above-mentioned template-switching region of the template-switching oligonucleotide, the base is guanine. Thus, the template-switching oligonucleotide can bind to cDNA by complementary pairing of the guanine in the template-switching region with the cytosine added by reverse transcriptase at the 3'-end of cDNA. In some embodiments, the sugar is ribose or deoxyribose. Thus, the multiple nucleotides in the template-switching region can be ribonucleotides or deoxyribonucleotides.

[0027] In some embodiments, the 3'-terminal nucleotide with a 3'-terminal modification refers to a locked nucleic acid. In some embodiments, the locked nucleic acid is a bicyclic structure formed by the 2'-O atom of the nucleotide glycosyl and the 4'-C through a methylene bridge.

[0028] In some embodiments, a 3'-terminal nucleotide with a 3'-terminal modification refers to a 3'-terminal nucleotide in which the 3'-terminal hydroxyl group is replaced by an -O-R or -R' group. Among them, the 3'-terminal hydroxyl group refers to the hydroxyl group connected to the 3'-C of the sugar of the nucleotide, and after replacement, 3'-C-O-R or 3'-C-R' is formed.

[0029] Among them, R is any one of an alkyl group having 1 to 18 carbon atoms, an alkenyl group having 2 to 18 carbon atoms, an alkynyl group having 2 to 18 carbon atoms, an aryl group having less than 18 carbon atoms, an aralkyl group having less than 18 carbon atoms, a heteroaryl group having less than 18 carbon atoms, a phosphate group, an azidomethyl group, an allyl group, a nitrobenzyl group, a carbamate group, and a thiophosphate group.

[0030] R' is any one of H, an alkyl group having 1 to 18 carbon atoms, an alkenyl group having 2 to 18 carbon atoms, an alkynyl group having 2 to 18 carbon atoms, an aryl group having less than 18 carbon atoms, and an aralkyl group having less than 18 carbon atoms.

[0031] Among the above groups, some groups (such as azidomethyl and carbamate) or the 3'-deoxynucleotide structure (such as 3'-CH) are used to eliminate the hydroxyl activity, and some groups (such as a spacer having 1 to 18 carbon atoms) block the catalytic activity of DNA polymerase through the steric hindrance of the spacer at the 3'-end, inhibiting the primer extension reaction.

[0032] Among the above groups, the alkyl group having 1 to 18 carbon atoms, the alkenyl group having 2 to 18 carbon atoms, and the alkynyl group having 2 to 18 carbon atoms include any one of a straight-chain hydrocarbon group, a branched-chain hydrocarbon group, and a cycloalkyl group. Taking the alkyl group as an example, the alkyl group having 1 to 18 carbon atoms includes any one of a straight-chain alkyl group having 1 to 18 carbon atoms, a branched-chain alkyl group having 3 to 18 carbon atoms, and a cycloalkyl group having 3 to 18 carbon atoms. In some embodiments, the alkyl group having 1 to 18 carbon atoms includes any one of methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, sec-butyl, tert-butyl, n-pentyl, isopentyl, neopentyl, tert-pentyl, n-hexyl, isohexyl, n-heptyl, 2-methylhexyl, n-octyl, 2-ethylhexyl, n-nonyl, n-decyl, etc.

[0033] In some embodiments, the heteroaryl group having less than 18 carbon atoms may be such that at least one C atom in the aromatic ring is replaced by a heteroatom (such as N, O, S). In some embodiments, the heteroaryl group having less than 18 carbon atoms includes 1 to 3 heteroatoms.

[0034] In some embodiments, R or R' may each independently be optionally any one of a C3 spacer, a C6 spacer, a C12 spacer, spacer 9, spacer 18, etc.

[0035] In some embodiments, the sugar of at least one nucleotide in the template conversion region is ribose. For example, the sugars of 1, 2, 3, 4, or 5 nucleotides can be ribose. This can enhance the binding strength between the nucleotides in the template conversion region and the cytosine nucleotides added to the cDNA ends, improving the stability and specificity of the pairing. Correspondingly, in the template conversion region, 0, 1, 2, 3, or 4 nucleotides can be deoxynucleotides, as long as at least one nucleotide is a ribonucleotide. In some embodiments, the template conversion region consists of ribonucleotides. In some embodiments, the base of at least one nucleotide in the template conversion region is guanine. For example, the bases of 1, 2, 3, 4, or 5 nucleotides can be guanine. In some embodiments, the template conversion region consists of guanine ribonucleotides. In some embodiments, the sequence of the template conversion region is rGrGrG, that is, three consecutive guanine ribonucleotides. In some embodiments of the present invention, the sequence of the template conversion region is 5'-rGrGrG-3', where the rG at 3' is a locked nucleic acid.

[0036] In some embodiments, the number of nucleotides in the template conversion oligonucleotide is 10 - 30, for example, it can be 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30.

[0037] In some embodiments, a functional sequence can also be included between the template conversion region and the amplification adapter region, such as a unique molecular identifier (UMI) sequence.

[0038] In some embodiments, after template conversion is completed, the complementary sequence of the amplification adapter region is integrated into the cDNA, facilitating subsequent library construction and sequencing.

[0039] In a second aspect of the present invention, a kit is provided, which includes the aforementioned template conversion oligonucleotide.

[0040] In some embodiments, the kit further includes at least one of reverse transcriptase, reverse transcription primer, DNA polymerase, RNase inhibitor, polyethylene glycol, betaine, dNTP, and solid phase carrier.

[0041] In some embodiments, the reverse transcriptase includes but is not limited to at least one of MMLV reverse transcriptase, HIV reverse transcriptase, and AMV reverse transcriptase. The reverse transcriptase can be a wild-type reverse transcriptase or a product obtained by genetic engineering modification based on the wild-type reverse transcriptase. For example, it is a mutant reverse transcriptase obtained by random mutation, site-directed mutation, DNA shuffling, etc. for one or more purposes such as reducing RNase H activity, improving synthesis ability, and improving thermal stability.

[0042] In some embodiments, the DNA polymerase includes, but is not limited to, at least one of Taq DNA polymerase, Klenow DNA polymerase, Bst DNA polymerase, Pfu DNA polymerase, Tfi DNA polymerase, Tfl DNA polymerase, Vent DNA polymerase, KOD DNA polymerase, Phi29 DNA polymerase, etc. It can be understood that the DNA polymerase can be the wild type of the above-mentioned polymerase or mutants, modified forms (such as antibody modification, chemical modification, etc.) for improving specificity, fidelity, heat resistance, amplification rate, etc.

[0043] In some embodiments, the solid-phase carrier can be made of at least one material selected from inorganic substances, natural polymers, and synthetic polymers, including but not limited to cellulose and its derivatives (such as nitrocellulose), resins, glass, silica gel, polystyrene, agarose, gelatin, polyvinylpyrrolidone, vinyl-acrylamide copolymer, polyacrylamide, latex, dextran, rubber, silicon, plastics, natural sponges, metal plastics, hydrogels, etc. In some embodiments, the solid-phase carrier has a planar structure, such as a slide, a chip, a microchip, an array. In some embodiments, the solid-phase carrier or its surface is non-planar, such as the inner (outer) surface of a tube or a container. In some embodiments, the solid-phase carrier includes microspheres or beads. In some embodiments, the solid-phase carrier includes an array of beads or pores. In some embodiments, the solid-phase carrier includes any one of beads, chips, glass, sensors, electrodes, silicon wafers.

[0044] In some embodiments, a barcode probe is directly or indirectly immobilized on the solid-phase carrier.

[0045] Among them, the barcode probe includes, from the 5' end to the 3' end: an amplification adapter region, a barcode sequence, and a capture sequence.

[0046] The barcode sequence is a spatial barcode or a cellular barcode, which are respectively used to identify the spatial location of transcriptome data in spatial transcriptome sequencing and the cellular origin of transcriptome data in single-cell transcriptome sequencing. In some embodiments, the length of the barcode sequence is 3 to 30 nucleotides, for example, it can be 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 15, 16, 18, 20, 22, 24, 25, 26, 28, 30 nucleotides. When the barcode sequence is a spatial barcode, different barcode probes ( "different" means different spatial barcode sequences, and the capture sequences can be the same or different) are usually fixed at different positions on the same solid-phase carrier, that is, each position (i.e., a spot) can fix a barcode probe or a cluster of barcode probes, and among the barcode probes at different positions or the clusters of barcode probes at different positions, the barcode sequences of the barcode probes are different. When the barcode sequence is a cellular barcode, different barcode probes (different means different cellular barcode sequences, and the capture sequences can be the same or different) are usually fixed on different solid-phase carriers, that is, each solid-phase carrier has a unique barcode probe or a cluster of barcode probes, and the barcode sequences are different from those of the barcode probes or the clusters of barcode probes on other solid-phase carriers.

[0047] The capture sequence is complementary to a partial sequence of the RNA from the sample to be tested, so that the RNA from the sample to be tested can be bound to the solid-phase carrier. In some embodiments, the capture sequence can be a polyT sequence, which can be complementary base-paired with the polyA tail of the RNA molecule. In some embodiments, the capture sequence can be a target-specific sequence, whereby targeted capture sequencing and the like can be performed. In some embodiments, the capture sequence can also be a random sequence.

[0048] In some embodiments, a cleavage site can also be included in the 5' direction of the amplification adapter region, such as a USER cleavage site or an exonuclease cleavage site, for releasing the product (such as a labeled cDNA molecule); of course, when there is no cleavage site, alkaline solution elution release can also be used, or double-strand release can be performed after double-strand synthesis.

[0049] In some embodiments, functional sequences such as UMI (unique molecular identifier) can also be included between the barcode sequence and the capture sequence.

[0050] In the third aspect of the present invention, a method for constructing a spatial transcriptome sequencing library or a single-cell transcriptome sequencing library using the aforementioned template-switching oligonucleotide is provided.

[0051] In some embodiments, the method includes the following steps: (1) Using the barcode probe as a primer and the RNA in the tissue or cell as a template, a reverse transcription reaction is carried out to obtain cDNA; wherein, the barcode probe is directly or indirectly immobilized on a solid-phase carrier and includes, from the 5'-end to the 3'-end: an amplification adapter region, a barcode sequence, and a capture sequence; the barcode sequence is a spatial barcode or a cellular barcode, and the capture sequence is complementary to a partial sequence of the RNA from which the test sample is derived; (2) A non-template polymerization sequence with a length of 2 to 5 nucleotides is introduced at the 3'-end of the cDNA to obtain an extended product; (3) The non-template polymerization sequence of the extended product hybridizes with the template-switching region of the aforementioned template-switching oligonucleotide, and continues to extend using the template-switching oligonucleotide as a template to obtain a library starting molecule; the library starting molecule includes, in sequence from the 5'-end to the 3'-end: an amplification adapter region, a barcode sequence, a capture sequence, a cDNA sequence, and a complementary sequence of the amplification adapter region; (4) The library starting molecule is amplified or enriched using a first library amplification primer and a second library amplification primer to obtain a spatial transcriptome sequencing library or a single-cell transcriptome sequencing library; wherein, the 3'-sequence of the first library amplification primer is complementary to the complementary sequence of the amplification adapter region, and the 3'-sequence of the second library amplification primer is the same as the amplification adapter region.

[0052] In some embodiments, steps (1) to (3) can be completed in the same reaction system. At this time, the reverse transcription reaction in step (1) and the non-template sequence polymerization and continued extension in steps (2) and (3) can be completed by the same reverse transcriptase. In some other embodiments, they can also be completed in different reaction systems.

[0053] In some embodiments, the non-template polymerization sequence includes a polymerization sequence of a single base (such as 5'-CCC-3') or a polymerization sequence of multiple different bases (such as 5'-CGC-3').

[0054] In some embodiments, the invalid fragments that can be reduced by the above library construction method include at least one of the following: non-specific TSO sequences, tandem sequences of barcode probes and TSOs, etc.

[0055] Among them, the non-specific TSO sequence includes a TSO sequence lacking a barcode sequence, and the sequence structure can be, for example, TSO-cDNA-TSO complementary sequence; the tandem sequence of barcode probe and TSO includes a sequence lacking cDNA, and the sequence structure can be, for example, barcode probe-TSO complementary sequence.

[0056] In some embodiments, the amplification method includes but is not limited to at least one of polymerase chain reaction (PCR), isothermal amplification (such as loop-mediated isothermal amplification LAMP, recombinase polymerase amplification RPA, rolling circle amplification RCA, cross primer amplification CPA, strand displacement amplification SDA, helicase-dependent amplification HDA), etc.

[0057] In some embodiments, different types of cDNA libraries, such as single-ended or double-ended, chained or circular, can be constructed based on different sequencing platforms.

[0058] In a fifth aspect, the present invention provides a sequencing method, comprising the step of sequencing the spatial transcriptome library or the single-cell transcriptome library obtained by the aforementioned method.

[0059] In some embodiments, sequencing is full-length sequencing without interruption. In some embodiments, sequencing is full-length sequencing performed using a single molecule sequencing method.

[0060] In some embodiments, specific sequencing methods include any one of first-generation sequencing, second-generation sequencing, and third-generation sequencing. First-generation sequencing includes Maxam-Gilbert sequencing technology, Sanger dideoxy sequencing technology, pyrosequencing technology, fluorescent automated sequencing technology, and hybridization sequencing technology, etc.; second-generation sequencing includes 454 Roche GS FLX, Illumina Solexa, SOLiD, Ion Torrent, BGISEQ, etc.; and third-generation sequencing includes HeliScope, PacBio HiFi, PacBio CLR, ONT, Cyclone WT, etc.

[0061] In some embodiments, sequencing methods include at least one of single-cell sequencing, spatiotemporal omics sequencing, bulk RNA sequencing, and direct RNA sequencing. In some embodiments, single-cell sequencing requires pre-capture of the cells. Capture methods include limiting dilution, flow cytometry, laser cutting, microscopy, and novel capture methods leveraging microfluidics. In some embodiments, tissue samples are frozen or paraffin-embedded and then sectioned to achieve spatial sequencing. The tissue sections are then attached to specific supports (e.g., solid-phase supports such as microarrays and magnetic beads) for nucleic acid capture and labeling.

[0062] Another aspect of the present invention provides the use of the aforementioned template switching oligonucleotide, the aforementioned kit, the aforementioned method for constructing a spatial transcriptome sequencing library or a single-cell transcriptome sequencing library, or the aforementioned sequencing method in omics detection or preparation of omics detection products.

[0063] In some embodiments, the omics detection includes the omics detection of the transcriptome. In some embodiments, the omics detection of the transcriptome includes the omics detection of single-cell transcriptome (scRNA-seq), the omics detection of spatial transcriptome, and the omics detection of non-single-cell transcriptome (bulk RNA sequencing), specifically, it can be Smart-seq, Smart-seq2, 10X Genomics, MGI (DNBelab C) transcriptome sequencing, etc. In some embodiments, the omics detection also includes the integration of transcriptome detection with other omics (such as genome, proteome, metabolome), for example, it can be single-cell multi-omics detection, spatial omics detection, etc.

[0064] In some embodiments, the foregoing template conversion oligonucleotide, kit, spatial transcriptome sequencing library or single-cell transcriptome sequencing library construction method is applied to the following fields: (a) the preparation of a sequencing library before omics analysis; or (b) the preparation of a sequencing library product for omics analysis. The sequencing platform includes next-generation sequencing or single-molecule long-read (third-generation) sequencing, such as any one of single-molecule real-time sequencing, nanopore sequencing, etc.

[0065] In some embodiments, the sequencing library is a library constructed by selectively capturing the 3'-ends of mRNA molecules of the transcriptome. In this case, the above method can more significantly improve the utilization rate of effective data.

[0066] Among them, taking the single-cell transcriptome sequencing as an example, for the method of selectively capturing and constructing a library from the 3'-end, the construction of the library can refer to CN112005115A.

[0067] The content of the present invention will be further described in detail below through specific examples.

[0068] It should be understood that these examples are only used to illustrate the present invention and not to limit the scope of the present invention.

[0069] For the experimental methods without specific conditions noted in the following examples, they are usually carried out under conventional conditions or according to the conditions recommended by the manufacturers. The materials, reagents, etc. used in this example are, unless otherwise specified, reagents and materials obtained from commercial sources.

[0070] Example 1 This example provides the construction and sequencing of a DNBelab C4 single-cell transcriptome library, and the specific process is as follows: I. Single-cell library optimization 1. Preparation of droplet generation system Prepare the reaction system according to the standard operating procedure of the DNBelab C Series High-Throughput Single-Cell RNA Library Preparation Kit Set V3.0 (MGI, 940-001819-00).

[0071] 1.1 Prepare the sample-phase suspension The cell suspensions of different samples are prepared according to the kit's recommended instructions. Among them, the human and mouse cell lines are a 1:1 mixture of HEK293T and NIH3T3. Human peripheral blood mononuclear cells (PBMCs) are from the peripheral blood mononuclear cells of healthy volunteers, and mouse brain cell nuclei are from the nuclei extracted by dissociating mouse brains. The single-cell suspensions from the above-prepared different sample sources are loaded onto the machine at 20,000 cells per single chip. Gently pipette and aspirate to mix the prepared cell suspension evenly, and prepare the sample-phase suspension by taking each component in the kit set according to Table 1: Table 1. Sample-phase suspension system

[0072] Among them, replace RT Primer-V3 with 100 μmol / L TSO primer. After preparation, gently pipette and mix evenly, centrifuge briefly, and place on ice for later use.

[0073] The sequence of the conventional TSO primer (RT Primer-V3) is: 5'-AAGCAGTGGTATCAACGCAGAG / rG / / rG / / rG / -3'; The sequence of the TSO primer for the experimental group is: 1) 5'-AAGCAGTGGTATCAACGCAGAG / rG / / rG / / XNA_G / -3'; 2) 5'-AAGCAGTGGTATCAACGCAGAG / rG / / rG / / rG / -3 Among them, rG is guanosine ribonucleotide; in 1), XNA_G is locked nucleic acid-modified guanosine ribonucleotide, and the 3'-OH of the sugar group is replaced with 3'-P (phosphate group modification) respectively. In 2), the last rG at the 3'-end is 3'-H (deoxy treatment) of guanosine ribonucleotide or is modified with 3'-C3 Spacer (alkyl chain C3 modification).

[0074] (Note: This TSO modification is not limited to the single-cell C4 commercial kit and can be flexibly modified according to the experimental platform. Similarly, the above PCR amplification primers, amplification enzymes, and purification magnetic beads can be replaced similarly, not limited to the components of this product) 1.2 Prepare the magnetic bead-phase suspension Take out Cell Beads-V3 and Index Carrier from the kit set, invert or pipette up and down until completely mixed. Pipette the single-sample dosage of Cell Beads-V3 and Index Carrier into a 0.2 mL low-binding PCR tube according to Table 2. Place the PCR tube on a magnetic rack and let it stand for 3 - 5 minutes, then slowly discard the supernatant, avoiding loss of magnetic beads. Remove the PCR tube from the magnetic rack and sequentially add Beads Buffer and Lysis Buffer-V3 from the kit set. After preparation, gently pipette up and down until completely mixed, centrifuge briefly, and place on ice for later use.

[0075] Table 2. Magnetic bead phase suspension system

[0076] 2. Droplet generation Prepare the slide according to the operation requirements of the kit set instructions and place it in the droplet generator. Sequentially add 80 μL of the sample phase suspension, 950 μL of droplet generation oil (P100 Oil), and 100 μL of the magnetic bead phase suspension in order, and start droplet generation.

[0077] 3. Obtain cDNA products 1) After droplet generation, collect the droplets into a PCR tube and perform reverse transcription reaction according to the operation of the kit set instructions (42 °C, 90 min, 10 cycles (50 °C, 2 min, 42 °C, 2 min), 85 °C, 5 min, 4 °C hold).

[0078] 2) After the reverse transcription reaction, add the demulsifying agent (Breakage Reagent, MGI, 940 - 001820 - 00), take the aqueous phase (middle layer) after demulsification for 0.6x (magnetic bead volume: sample volume) magnetic bead purification (DNA Clean Beads, MGI, 940 - 001820 - 00), and then perform PCR amplification using cDNA Amp Primer-V3 (MGI, 940 - 001819 - 00) (the upstream primer is the adapter sequence of TSO, and the downstream primer is the 5' end adapter sequence of the probe) and cDNA Amp Enzyme (MGI, 940 - 001819 - 00) (95 °C, 3 min, 10 cycles (98 °C, 20 s, 65 °C, 30 s, 72 °C, 3 min), extend at 72 °C for 10 min, 12 °C, hold).

[0079] 3) After the amplification reaction, perform magnetic bead purification at 0.6x (magnetic bead volume: sample volume) (DNA Clean Beads, MGI, 940-001820-00). Further, perform concentration detection and fragment distribution analysis on the purified full-length cDNA product.

[0080] 4. Short-read library construction and sequencing 1) From the full-length cDNA product obtained in the previous step, take 1 / 3 of the product for fragmentation, end repair, adapter ligation, and amplification reaction of the fragmented product (MGI, 940-001821-00); 2) Perform DNBSEQ sequencing on the amplified product after library construction.

[0081] II. Single-molecule library construction and sequencing 1) Take the full-length cDNA product obtained in step 3 of the single-cell library optimization process and construct and sequence the single-molecule library according to the instruction manual of the CycloneSEQ library preparation reagent kit (MGI, H940-000001-00), including processes such as end repair with A addition and adapter ligation.

[0082] 2) Perform single-molecule sequencing on the library product according to the input amount specified in the kit instruction manual.

[0083] Among them, if the full-length cDNA product obtained in step 3 is not enough for the starting amount of single-molecule library construction, it can be appropriately amplified. Similarly, the cDNA obtained by this process is compatible with other currently commercial single-molecule sequencing platforms. Just follow the commercial instruction manuals of each platform.

[0084] III. Result analysis Calculate the data utilization rate of the sequencing results according to the following formula: Data utilization rate = number of target read lengths / total number of read lengths × 100%.

[0085] The results are as Figure 2 , and the long-read sequencing data results of the single-molecule library show that, whether it is human and mouse cell lines, PBMC, mouse brain cell nuclei, etc., various TSO structural variants obtained by modifying the terminal nucleotide at the 3'-hydroxyl site of TSO or adding a molecular spacer can systematically improve the effective data utilization rate. This confirms that the TSO three-end capping strategy can effectively inhibit non-specific amplification mediated by TSO by blocking polymerase non-template extension through multiple paths, providing a standardized solution for the process optimization of single-cell and spatial transcriptome technologies.

[0086] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. Template-converting oligonucleotide, characterized in that, The template-switching oligonucleotide includes: a template-switching region located at the 3'-end of the template-switching oligonucleotide, and an amplification adapter region located in the 5'-direction of the template-switching region; The length of the template-switching region is 2 to 5 nucleotides; the 3'-terminal nucleotide of the template-switching region has a 3'-terminal modification, and the 3'-terminal modification is used to prevent the 3'-terminal nucleotide of the template-switching region from forming a phosphodiester bond with the free 5'-phosphate group of a deoxynucleotide and / or a deoxynucleotide sequence.

2. The template-converted oligonucleotide according to claim 1, wherein The 3'-terminal nucleotide having a 3'-terminal modification refers to a locked nucleic acid, or a 3'-terminal nucleotide whose 3'-terminal hydroxyl group is replaced by an -O-R or -R' group; wherein, R is any one of an alkyl group having 1 to 18 carbon atoms, an alkenyl group having 2 to 18 carbon atoms, an alkynyl group having 2 to 18 carbon atoms, an aryl group having less than 18 carbon atoms, an aralkyl group having less than 18 carbon atoms, a heteroaryl group having less than 18 carbon atoms, a phosphate group, an azidomethyl group, an allyl group, a nitrobenzyl group, a carbamate group, and a thiophosphate group; R' is any one of H, an alkyl group having 1 to 18 carbon atoms, an alkenyl group having 2 to 18 carbon atoms, an alkynyl group having 2 to 18 carbon atoms, an aryl group having less than 18 carbon atoms, and an aralkyl group having less than 18 carbon atoms.

3. The template-converted oligonucleotide according to claim 1, characterized in that, The template-switching region consists of 3 ribonucleotides.

4. The template-converted oligonucleotide according to claim 3, wherein The sequence of the template-switching region is rGrGrG.

5. Kit, characterized in that, Comprising the template-switching oligonucleotide according to any one of claims 1 to 4.

6. The kit according to claim 5, wherein, The kit further includes at least one of reverse transcriptase, reverse transcription primer, DNA polymerase, RNase inhibitor, polyethylene glycol, betaine, dNTP, and a solid-phase carrier; wherein, a barcode probe is directly or indirectly immobilized on the solid-phase carrier; the barcode probe includes, from the 5'-end to the 3'-end: an amplification adapter region', a barcode sequence, and a capture sequence; the barcode sequence is a spatial barcode or a cellular barcode, and the capture sequence is complementary to a partial sequence of the RNA from the sample to be tested.

7. A method for constructing a spatial transcriptome sequencing library or a single-cell transcriptome sequencing library using the template-switching oligonucleotide according to any one of claims 1 to 4.

8. The method according to claim 7, wherein Comprising the following steps: (1) Using the barcode probe as a primer, performing a reverse transcription reaction with the RNA in the tissue or cell as a template to obtain cDNA; wherein, the barcode probe is directly or indirectly immobilized on the solid-phase carrier, and includes, from the 5'-end to the 3'-end: an amplification adapter region', a barcode sequence, and a capture sequence; the barcode sequence is a spatial barcode or a cellular barcode, and the capture sequence is complementary to a partial sequence of the RNA from the sample to be tested; (2) Introducing a non-template polymerization sequence with a length of 2 to 5 nucleotides at the 3'-end of the cDNA to obtain an extended product; (3) Hybridizing the non-template polymerization sequence of the extended product with the template-switching region of the template-switching oligonucleotide according to any one of claims 1 to 4, and continuing to extend using the template-switching oligonucleotide as a template to obtain a library starting molecule; the library starting molecule includes, from the 5'-end to the 3'-end in sequence: an amplification adapter region', a barcode sequence, a capture sequence, a cDNA sequence, and a complementary sequence of the amplification adapter region. (4) Amplify or enrich the library starting molecules using the first library amplification primer and the second library amplification primer to obtain the spatial transcriptome sequencing library or the single-cell transcriptome sequencing library; wherein, the 3'-end sequence of the first library amplification primer is complementary to the complementary sequence of the amplification adaptor region, and the 3'-end sequence of the second library amplification primer is complementary to the complementary sequence of the amplification adaptor region'.

9. A sequencing method, characterized in that, It includes the step of sequencing the spatial transcriptome library or the single-cell transcriptome library obtained by the method according to claim 7 or 8.

10. The sequencing method according to claim 9, wherein The sequencing is full-length sequencing using a single-molecule sequencing method.

Citation Information

Patent Citations

  • Methods characterizing multiple analytes from individual cells or cell populations

    CN112005115A

  • Self-folding amplification of target nucleic acid

    CN102725424A

  • Synthesis of double-stranded nucleic acids

    CN106460052A

  • High-throughput polynucleotide library sequencing and transcriptome analysis

    CN111148849A

  • Compositions and methods for improved cdna synthesis

    CN113811610A